.LOG For the EFD tier, choose RAID5 (3+1) for the FC tier, choose RAID1 — this is explained in further detail in its own section below. For the SATA tier, choose RAID6 (6+2) — avoid RAID5 for resiliency reasons, and avoid RAID6 (14+2) because it is not capable of performing optimized/coalesced full stripe writes. ============ symcfg discover symcfg sync -local symcfg sync -vpdata symcfg sync -fast ======== symcfg list -v detailed info on the arrays in the DB symcfg list –thin –pool –detail –gb Info about thin pool capacity (optional –v for verbose) symdisk list –dskgrp_summary Physical disk info symcfg –dir all list –v Info about Symm directors (optional –v for verbose) symsg list List storage groups (optional –v for verbose) symtier –sid xxx list –offline List tier info (optional –v) symfast –sid xxx list –vp List fast policies (optional –v) symfast –sid xxx list –demand –assoc FAST Demand report by SG Association — shows the potential for each SG to tier up based on assigned policies ========================== More balance way Consider a "Building block" approach DA IOPS Small random I/Os are : 15k fc/sas disk can do 150iops 10k fc/sas disk can do 120iops 7.2k SAS disk can do 50iops EFD disk can do 1000 IOPS DA configurations are : multiple of 8 drives per Vmax 10k ( 959) or 20K engine multiple of 16 drives per Vmax 10k ( 987) or 40K engine EFDs - Balance Calculating number of Disks needed: Fast VP system needing 50,000 IOPS Raid protection :- Raid1 : 1host write = 2 write Raid 5 : 1 host write = 2 reads+2 writes Raid 6 : 1 host write = 3 reads + 3 write Frontend connections :- Use all the zero ports on director first and then the one ports planning to use the average I/O rate may not be good enough writes tend to come in bursts you need enough front-end cpus to absorb the bursts Virtual provisioning : - TDAT size - 8 physical hypers per disk dirive per pool Raid 1 TDAT count should be even Raid 5 (3+1) for EFD Raid 1 for FC tier ( 300gb drives preferred, 900 gb drive may be too large Raid 6 for SATA FAST VP and SATA drives: FAST VP will not make SATE dries any faster cache partition for SATA TDATs ( 5876 ) Monitoring the symmetrix over Time Watch for growth trends in your workload - SPA ( symmetrix performance Analyzer) Look out for significant increases in response time - Host based tools ( iostat, sar, RMF) Monitor utilization metrics in WLA/STP/SPAS - better to be proactive than waiting to hit the wall Any utilizations well over 50% should be considered a possible source for future issues with growth Dont over-complicate the configuration 10:56 PM 1/6/2015 http://vjswami.com/category/storage/ http://sanartisan.wordpress.com/category/performance/ http://www.rajeshvu.com/portal/tag/vmax http://storageboy.com/category/symmetrix/ http://www.sandirect.com/ http://www.scottbrightwell.org/ https://community.emc.com/community/products/vipr https://github.com/emcvipr/controller-openstack-cinder http://www.manageengine.com/products/opstor/ http://searchstorage.techtarget.com/ 3:25 PM 1/15/2015 3:25 PM 1/15/2015 Storage operatuions challenges:- Application Complexity Wasted Storage uncertain recoverablity opaque costs SAL failres performance problems wasted effot and incidents uncontrolled growth performance problems utilization and optimization data protection complaince performance Trending & reporting storage configuration managemnet storage and application capacity trends SLA Achievement reporting application chargeback understand host to storage relationships chargeback or showback reports multi-tenancy/Dashboards Dashboards block/file/object Automate capacity reporting analyze capacity usage & trends plan & justify new purchases match to EMC support matrix All - operations - complaince - storage compliance SLA Achievement reporting automate report delivery via role-based access or email Data protection complaince utilizationi optimations ( move to lower cost storage without violation SLAs service assurance suite Visualize, analyze & optimize performance trending & reporting storage configuration management - define polices and best practices, validate compliance to EMC support matrix, ensure proper configuration to meet service levels & track configuration changes New in SRM 3.5:- Vplex enhancements: topology views end-2-end data path reporting new capacity reports chargeback reporting configuration validation Vipr enhancements: virtual pool --->service level mappint chargeback reporting deeper visibility into physical resources 12:58 AM 1/17/2015 monitoring measuring -iops -throughput -queue depth -read/write -block side -random/sequential 4kiops * 256 block side = 1000mb/s 50k iops * 4k block size = 200MB/s iops * blocksize = throughput performance monitor add disk read/writ Bytes/sec adddisk read/wrotes/sec key metrics for efective storage capacity storage architecture dis vistualization 12:41 AM 1/20/2015 can you be more specific on what happened? 1. What were the service times? Is there any performance analysis available from the host side? 2. Which disks were reported to have been affected (array/LUN IDs)? 3. Was there any increase in data volume being processed at that time? 4. Is there a MIM raised for this? 5. Are there any findings on the OS/DB side yet? Also, please provide emcreports output from the host. 10:42 PM 1/20/2015 The following questions are important to understand the nature of the performance issue and should be answered before a more thorough and targeted investigation might occur: 1.What is the performance currently (quantify the performance issue)? How is the performance being measured? 2.What is the expected performance? How is the performance measured? 3.Where is the performance degradation being seen (specific client or volume group)? 4.Is the performance issue tied to a specific time period or workload? 5.Has the system ever performed at the expected level? 6.Have any changes been made that pre-empted the performance issue (firmware changes, file system layout changes, and/or software upgrades)? Performance can be tuned from the file system/application perspective as well as from the E-Series storage perspective. The following are some basic factors to investigate from the storage perspective: The following are some basic factors to investigate from the storage perspective: Storage Factors Is cache enabled? Is the state currently suspended? Are any volume groups degraded? Are there any exclusive operations in progress (drive reconstruction and dynamic volume expansion)? Are there more partial writes versus full? See 'write algorithms' in evfShowVol in the stateCapture. Are there any drive-side issues reported (Destination driver events or check conditions in MEL, high SAS phy errors, and drive-side timeouts)? Are there any indications of a slow drive (iditnall queue depth)? Are there any degraded SAS wide ports or FC drive side channels? Are there any controller reboots or failovers during the time frame that performance impact is reported? Is the RAID level appropriate for the workload? Is the RAID segment size in harmony with the host application (important to avoid read-modify-write penalties)? What is the typical I/O size? Is the I/O random/sequential/mixed? what I would check: 1 .Configuration of all storage arrays. ip’s, passwords, raid groups, physical disks, storage processors, luns, connected hosts, support contact and information needed to contact spport. 2. SAN switch, configuration, # of ports, ports used, connected HBA’s or storage controllers, model, support information. 3. End to end storage layout diagram. 4. Metrics pertaining to iops per port, aggregated to iops per switch. Disk space utilization, RAID group utilization, Storage processor iops, %bandwidth used. You get the idea. Start small, and you can keep adding to it. For IP-based storage systems, including SAN and NAS, the ExtraHop platform analyzes all transactions traversing the network in real time, extracting health and performance metrics for clients and servers, including application-level details contained at L7 such as methods and errors. Proactively fix potential problems with trend-based, early-warning alerts for storage latency Determine the root cause of application slowdowns with cross-tier, correlated visibility across storage, database, web, and network tiers Audit storage usage with metrics for reads, writes, and metadata requests for each user Pinpoint problem files by analyzing transaction details for individual servers and server groups Select the right WAN optimization candidates for SAN and NAS systems 12:16 AM 1/21/2015 Examples of metrics and measurements for storage efficiency and optimization include the following: •Macro (e.g., facilities such as power usage effectiveness) and micro (device or component level) •Time (performance or activity) vs. availability vs. space (capacity) •Performance metrics, including IOPS, bandwidth, and response time or latency •Additional performance metrics, including reads, writes, random, sequential or IO size •Storage capacity metrics, including percent utilization as well as reduction ratios •Other capacity metrics, including raw, formatted, free, allocated or allocated not used Metrics can be obtained from in-house, third-party, or operating system and application-specific tools. Other metrics can be estimated or simulated; for example, benchmarks running specific workloads such as those from the Transaction Processing Performance Council (TPC), Storage Performance Council (SPC), Standard Performance Evaluation Corporation (SPEC) or Microsoft Exchange Solution Reviewed Program (ESRP). Compound metrics, those made up of multiple metrics, include cost per GB and cost per IOP, along with capacity per watt or activity per watt, such as IOPS or bandwidth per watt of energy used. ========================== Here is a list of common storage performance metrics: •IOPS: I/O operations per second where the I/O can be of various size •Latency: The response time where lower is better for time-sensitive applications •MTBF: Mean time between failures indicates reliability or availability •MTTR: Mean time to repair or replace a failed component or storage device •Quality of Service (QoS): Refers to performance, availability or general service experience •Recovery point objective (RPO): To what point in time is data saved or lost •Recovery time objective (RTO): How quickly data or applications can be made available •SPC: Storage Performance Council workload (IOP, bandwidth and others) •TPC: Transaction Processing Council workload comparisons Other metrics include uptime, planned or unplanned downtime, errors or defects, and missed windows for data protection or other infrastructure resource management tasks. Remember to keep idle and active modes of operation in perspective when comparing tiered storage. Applications that rely on performance or data access need to be compared on an activity basis, while applications and data that are focused more on data retention should be compared on a cost per-capacity basis. For example, active, online and primary data that needs to provide performance should be looked at in terms of activity per-watt per-footprint cost, while inactive or idle data should be looked at on a capacity per-watt per-footprint cost basis. Given that productivity is also a tenet of storage efficiency, metrics that shed light on how effectively resources are being used are important. For example, QoS, performance, transactions, IOPS, files serviced or other activity-based metrics should be looked at to determine how effective and productive storage resources are. =============== Virtual Instruments VirtualWisdom4 VirtualWisdom software is a hardware-agnostic SAN monitoring tool that provides a real-time view of utilization and performance. This newest version added at-a-glance performance views, upgraded analytics capabilities and case-based alarms. A NAS monitoring probe offers increased visibility. ============== Storage Array Monitoring •Monitor Physical components - Controllers, ports,drives •Monitor logical components - LUNs, Volumes, Storage Groups •Monitor health,availability and utilization of resources •Monitor sensor faults, battery, Power supply status ====================== EMC asked to provide answer on listed below questions: 1. Specify the operating system, type, and version. 2. What is the exact issue being seen... i.e technical description, When did the event occur, ongoing? 3. Are there any error messages? Provide outputs or screenshots. 4. What devices are affected? 5. Are there multiple hosts seeing the same issue? 6. Is any EMC software involved (PowerPath, Solutions Enabler, etc.)? 11:02 PM 1/21/2015 EMC have analyzed the provided outputs and came up with the following: - The application logs collected in the grabs have no information logged before 01/20/15 @ 8:05 AM - The system event file is missing the logs from 01/16/15 @ 1:09PM through 01/16/15 @ 7:31 PM - PowerPath version currently not supported, need to update to minimum support level of 5.5 SP1 - There was no dead path / failover issue that would have caused latency with the devices EMC also performed a checkout on the array and it came clean. We filed a request for performance analysis regarding the time the issue was observed and awaiting EMC feedback. However, I think it might be difficult to establish the root cause without any input from DB side, so EMC will conclude with some recommendations but no guarantee these will actually address or indicate a potential root cause. I’ll be out of the office tomorrow, Dawid Orczyk will be monitoring the case and provide updates if necessary. 2:17 AM 1/22/2015 4:23 AM 1/29/2015 uncontrolled growth application complexity performance problems wasted storage opaque cost SLA failures wasted efforts & incidnents uncertain recoverablity Business demond application performance TCO speed to service flexibility 11:43 PM 2/23/2015 Please feel free to forward this invite to anyone that may have been missed and has information regarding this incident. The purpose of the meeting is to confirm ownership, understand the sequence of events, and ensure full scope of the impact is captured. Please come prepared to discuss the following: • Root Cause Analysis and Determination: How and why did the incident happen? • Identification of true nature of the incident: Internal/External • Known Defects: problems occurring on a platform, application, or environment • Provide architecture diagram to visualize the components involved causing the incident • Were there any contributing factors that elongated Mean Time to Repair & Restoration of Services? • How were we notified of the issue? Alert Monitoring? Customer? Vendor? • Confirm businesses impacted, application impacted, what transactions, by whom, and where? • Confirm Timeline (Impact start/end times) • What was done to fix problem? Quick Fix? Resolution? Workaround? • Was COB or a workaround of any type invoked or considered during this outage • Was it successful? Failover as designed? • If not why and what is being put in place going forward? • Last successful test? • What actions will be taken to prevent a reoccurrence? • Resiliency or Redundancy: Adequate / sufficient tolerance built in • Were Change activities a trigger or a cause in this outage? • Approved RFC? Risk Level? • Vendor? • BAU or Marketplace Request? Few notes to remember: • This is a “NO BLAME” meeting and you are encouraged to openly discuss the associated activities that took place before, during and after the incident. • Please come prepared to discuss your involvement, as well as providing ideas for prevention /process improvement. • Forward this invite to other support folks whom you think are needed for this meeting [Ex – In country Support, GNCC, Engineering, Architecture, Data center Operations etc.] • Please perform internal investigations and log collections before the meeting and not during the meeting • Please dial in to the conference bridge or come to the meeting room 5 minutes early, so that we can get started on time and spend less time in conferencing people in. • See bottom of meeting request for all non-UK/US/Singapore access and toll free numbers. 11:01 PM 2/26/2015 emerging technology team meeting - 26/2/2015 Bryan/Sal/Shibu/Peter/uttham/mike/Kery/Raja/Mat Gary/Al - no longer with citi introduction to everyone lens staff - call 10:35 PM 3/9/2015 SRM Vipr training ViPR SRM - software to visualize, analyze vipr srm product documentation - documentation for Vipr configured usable : Luns, pools, any other usable storage RAID overhead : cost if RAID protection of configured usable Hot spare : hot spare drives unconfigured : empty space you can use to configure unusable : fragmentation and short-stroked drives Free : LUNs either not mapped or not masked or both Designated replicas not in a replica relationship Pool Free : free spacec in the pool Used for block : mapped and masked to a block port on the array used for file : mapped and masked to a port used by the file array Used for virtual : mapped and masked to a VPLEX Data enrichment is needed to change the loc or DC Service levels are applied to CONFIG/URED Luns service level capacity is sum of LUN measures 2.1 do lab ex 2 part 1 now chargeback report service levels : configured LUN proerties Group:Group of hosts who has access to LUNs primary used : primary presented Data enrichmentf /opt/apg/doc/apg-Frontend-user-guide.pdf • Disk I/O response time R1-R2-BCV mapping for Symmetrix devices can you show 'EMC configuration validation Report' and the results EMC configuration validation Report - SRM ViPR continuously validates compliance with design best practices and the EMC Support Matrix to ensure the environment is always configured right to meet service level requirements. 12:04 AM 3/31/2015 Poll - Devices, SNMP, SMI, WMI, SSH, XML backend process the metric name, value & time-stamp DB : Mysql & alert on a reported value -email -trap -log entry -run a script - sms - auto provision - self-healing ( up a dn interface) 6:22 PM 4/16/2015 8:58 PM 6/17/2015 1-2-1 with Shibu 5:05 AM 9/7/2015 6:00 PM 10/1/2015 2:13 AM 1/15/2016 alerts on xbar issues in Cisco problems are in ISL ( port flopping, capacity issue ) - monitoring on switch side run books - diagnize slow drain issue - perform basic checks in DCNM ( Kevin and tom ) - better health checks on isl Host locks and reserve locks, hba settings not setup properties - AIX / SA community to control the repeated problems Scott, Mike, Ashok, Sal, Sameer, John, Raja, Kerry Hines, james,lokesh, Denis, susan, steve, eswara