What is Enable RF flow control ---- SRDF/A prevent RF directors from flooding the link or the R2 Symmetrix with work. It works by limiting the number of jobs an RF director can place on the SRDF link at one time, thereby balancing the work across all available RFs and across the link availability Page Data Sets: Your EMC CE needs to set Enable Page Date Set Mode to YES in the IMPL.bin file to ensure synchronous replication of all page data sets Configuring Delta Set Extension (DSE) disable dev 451:4a6 in pool DEFAULT_POOL , type = SAVEDEV; create pool rdfa_dse_pool , type = rdfa_dse; add dev 451:4a6 to pool rdfa_dse_pool , type = rdfa_dse , member_state = enable; set ra group 1 rdfa_transmit_idle = enable; rdfa_dse_pool = rdfa_dse_pool; emulation = fba; rdfa_dse_autostart = enable; Next, use symconfigure to commit this change: symconfigure -sid xx -file cmd preview -v symconfigure -sid xx -file cmd commit -v -cons_exempt flag where to set http://sanengineers.wordpress.com/srdfa-best-practices/ - DCX Symmetrix monitoring Switches - Cisco, DCX Unisphere for VMAX Symmetrix Performance monitoring pool balance varience zero space reclaim push or pull? SRDF from thick to thin and larger size R2 Each SRDF base solution operates in one of the following modes of operation: ◆ synchronous ◆ semi-synchronous ◆ adaptive copy ◆ asynchronous Synchronous and semi-synchronous modes are the primary modes of operation while adaptive copy modes are the secondary modes of operation asynchronous mode mirrors R1 devices by maintaining a dependent-write consistent copy of the data on the secondary (R2) site at all times. SRDF/A session data is transferred from the primary to the secondary site in cycles Semi-synchronous mode is supported with Enginuity versions prior to 5773. Semi-synchronous mode allows the R1 and R2 devices to be out of synchronization by one write I/O operation. Adaptive copy modes write pending and adaptive copy disk. Adaptive copy modes do not guarantee a dependent-write consistent copy of data on R2 devices. Number of tracks out of synchronization between the R1 and the R2 devices at any given time is determined by the maximum skew value SRDF/A will capture a delta set of writes and send them in cycles across the link. In addition to the new writes, SRDF/A will include up to 30,000 invalid tracks per cycle. This is a design feature and the 30,000 track value was chosen to prevent cache from being flooded by the invalid tracks Therefore, EMC generally recommends as a best practice to synchronize the boxes in Adaptive Copy Disk mode to below 30,000 invalid tracks before activating SRDF/A SRDF/A works in cycle of four steps ( recieve the data, gather the data for the till the predefined cycle time, send the data, and wait for acknowledgments). in SRDF/A writes are destaged to disk only after they have been copied over the RDF links. In a given point of time there are one one "cycle switc "which is send over the rdf link and waiting for the ACK and the SAME time another cycle is doing "recieve the data and gathering data " . So if a failure ocurs during this stage you lose maximum two "cycle switches of data" one being waiting for the ACK and the other one being the "receive/gather" state. That is how the "twice the cycle time" .. SRDF/A default cycle switch time is 30 second, and most recents code(5875) you can see the default is 15 seconds. symcfg list -rdfg all` symdev show on the R1 volume shows the R1 R2 time lag srdf A will give you RPO of 2cycles.. DR is always 2 cycles behind the production To understand how delta set pushes the data to R2, you can look at the SRDF product guide. In a nutshell, SRDF/A works in Cycle. 1) Capture 2) Transmit 3) Recieve 4) Apply Lets assume your minimum cycle time is 30 sec, so R1 symm will start capturing the data for 30 sec and then do a cycle switching (Capture cycle becomes transmit and tranmit becomes Capture). Once the data reaches to transmit, it starts sending it over the RDF link to R2 Symm where it is received in Receive cycle, again after 30 sec, cycle switching happens (receive cycle becomes apply and apply becomes receive) and whatever data in Apply cycle is destaged to the disk. SyncInProg A synchronization is currently in progress between the R1 and the R2. There are existing invalid tracks between the two pairs and the logical links between both sides of an SRDF pair are up. Synchronized The R1 and the R2 are currently in a synchronized state. The same content exists on the R2 as the R1. There are no invalid tracks between the two pairs. Split The R1 and the R2 are currently ready to their hosts, but the links are not ready or write disabled. Failed Over The R1 is currently not ready or write disabled and operations have been failed over to the R2. R1 Updated The R1 is currently not ready or write disabled to the host, there are no local invalid tracks on the R1 side, and the links are ready or write disabled. R1 UpdInProg The R1 is currently not ready or write disabled to the host, there are invalid local (R1) tracks on the source side, data is being copied from the R2 to the R1 device, and the links are ready. Suspended The SRDF links have been suspended and are not ready or write disabled. If the R1 is ready while the links are suspended, any I/O will accumulate as invalid tracks owed to the R2. Consistent The R2 SRDF/A capable devices are in a consistent state. Consistent state signifies the normal state of operation for device pairs operating in asynchronous mode. Transmit Idle The SRDF/A session cannot push data in the transmit cycle across the link because the link is down. Symmetrix array keeps an account of the tracks that are "owed" to the other side. The owed tracks are known as remote invalids feature allows you to set a specific length of time for Enginuity to wait when a down link is detected before updating the link status. If the link status is still Not Ready after the link limbo time expires, devices are marked Not Ready to the link. http://richgoldstein.net/content/emc/srdf_intro --- Device states symrdf -rdf -sid 123 ping data from R1 to a larger R2 device You can copy data from an R1 device to a larger R2 device but the following restrictions apply: All swap and SRDF/Star operations are blocked. If SYMAPI_RDF_CREATEPAIR_LARGER_R2 is set to DISABLE in the options file, all createpair operations are blocked. Data mirrored to a larger R2 device cannot be restored back to its R1 device. Concatenated metadevices are not supported but striped metadevices are supported. Consistency exempt feature in Enginuity 5874 provides the ability to dynamically add and remove volumes from an active SRDF/A session without affecting the state of the SRDF/A session or the reporting of the SRDF pair state for each of the volumes in the active session that are not the target of the add or remove operation. This is achieved by marking the volumes being added or removed as “exempt” from being considered when calculating the consistency state of the volumes in the SRDF/A session or when deciding if the SRDF/A session should be dropped to maintain dependent write consistency on the R2 side. Setting the consistency exempt flag on a volume allows the volume to be added or removed from an active SRDF/A SRDF group using either a create, delete, or move operation without requiring the other volumes in the SRDF group to be suspended prior to the operation An R1 SRDF mirror indicates that data stored on the R1 device is also remotely mirrored to the R2 device. Likewise, an R2 SRDF mirror indicates that data stored on the R2 device is remotely mirrored to the R1 device Static Devices need to be converted to RDF before the pair created. Static requires Bin changes however dynamic can be changed online Enginuity version 5875, SRDF supports zero space reclamation. Zero space reclamation is an Enginuity feature that allows you to remotely mirror a thick SRDF device to a thin SRDF device while avoiding mirroring pre-allocated zero data chunks that may be associated with a thick SRDF device RAID 6 A RAID 6 group consists of 8 or 16 RAID members. Data and horizontal and diagonal parity blocks are distributed across all RAID members. RAID 6 protects data in the event of one or two drive failures. Fabric fan out ratio is number of host HBA connected to storage port. The standard is 10:1 and is defined by the Storage vendors Symmetrix director flags (bits) The following FA director flag settings need to be configured to support Windows Server 2003 and 2008 Common Serial Number (C) Host SCSI Compliance 2007 (OS2007) SCSI-3 SPC-2 Compliance (SPC2) SCSI-3 compliance (SC3) For FC Switch Base Topology (FC−SW), Enable Auto Negotiation (EAN), Point−to−Point (PP) , Unique WWN (UWN). Additionally for Windows 2008 Failover Cluster, the Persistent Reservation attribute SCSI3_persist_reserv must be enabled on each Symmetrix DMX device used. This should NOT be done to devices for Windows 2003 clusters. EMC recommends that the SCSI3_persist_reserv attribute be only set on devices that require it. To display what director flags have been set per host initiator, run the following Solutions Enabler command. symmaskdb -sid XXXX list database -v (without the -v you will not see the flags) To enable the necessary director flags for an initiator on director 4a port 0 using symmask: symmask -sid xxxx -wwn xxxxxxxxxxxxx -dir 4a -p 0 set hba_flags on C,OS2007,SC3,SPC2 -enable symmask refresh (this command is required after performing the above command) To enable the necessary director flags for an initiator group using Symaccess: symaccess -sid xxxx -type init -name myig1 set ig_flags on C,OS2007,SC3,SPC2 -enable To display the flags from the previous command using symmaccess: symaccess -sid xxxx -type init show myig1 -detail To enable the SC3, SPC2 & OS2007 flags globally on the FA 1D port 0 via symconfigure, create a text file similar to below.. set port 1d:0 SCSI_3=enable, SPC2_Protocol_Version=enable, SCSI_Support1=enable; the SC-3, SPC-2, and OS2007 Edit Director flags can be enabled with the Symmetrix online. These flags can be enabled via a EMC CE applied bin file change. The flags may be changed one port at a time via multiple bin file changes or they may all be enabled simultaneously in one bin file change. This will invoke the online Change_Director_Flags (CdfOnl) script on the Symmetrix Service Processor. The SymmWin Change_Director_Flags (CdfOnl) script will not set the affected ports offline when changing these flags. However if enabling the SPC-2 Edit Director flag in the bin file via an ECC or Solutions Enabler configuration change (symconfigure set port command) then the EMC software will require that the affected FA ports be set offline before the activity can be performed In all cases, the changed state of these Edit Director flags will not be detected until the host HBA is logged out and logs back into the Symmetrix fibre channel (FA) Can these flags be changed with the hosts online? Yes, however a host reboot is required. For example (Symmetrix V-Max at 587x with with VMware 3.5, VMware ESX 4.0): Common serial number (C) Auto negotiation (EAN) enabled Fibrepath enabled on this port (ACLX) SCSI 3 (SC3) (Optional) OS2007 is optional (it can be enabled if required by other hosts in a port sharing heterogeneous host environment) 5875 1.FAST VP/Sub-LUN auto-tiering 2.VLUN Mobility for VP 3.VMware VAAI support 4.Federated Live Migration 5.Rename thin pools 6.Native 10GB Ehernet support 7.DARE 8.Virtual Storage Integrator 9.Zero Space Reclaim during Migration and Replication 5876 1. Host IO limits - either on Bandwidth or IOPS per storage groups 2. Cascading of Storage groups 3. Windows 2012 support - Thin provisioning awareness/Storage Space Reclamation/ODX 4. FTS - Encapsulation/External Provisioning(4th FAST tier) 5. FST VP SRDF aware Performance metrics - Prosphere FE Port - Throughput in KBps - % busy Device - Sampled avg read time - Smapled avg write time - KBs read/s - KBs write/s System - % busy - Reads/s - writes/s - KBs read/write /s Gate Keepers - A Gatekeeper device may not respond to an application or server I/O request until a Symmetrix command request is completed. If a Symmetrix is executing a complex or time-consuming command, such as a query of the Remote R2 Symmetrix, line latency may cause a delay and the device can, in effect, be locked for some time. If the time is too long, the application using that Gatekeeper as a data device may experience I/O timeouts or perform poorly. Why does the Symmetrix use Gatekeepers and not some other means of control? - The Symmetrix is a high performance storage system that scales up to the largest storage capacities available in the industry. It also has very comprehensive monitoring, performance and diagnostic capabilities. As a result, command dialogs with the system can be extensive and can involve a good deal of data. Symmetrix uses Gatekeepers as a scalable, high-performance means of command and control, suitable to the capabilities of the array. The deterministic nature of disk I/O is particularly powerful when used for array-wide or multi-array operations such as consistent splits, in which conventional networking may encounter packet collisions or retries, which could prevent the timely operation of disaster recovery commands. - Devices that absolutely should not be used as Gatekeepers can be put into a gkavoid file as part of the Solutions Enabler setup - Devices intended for Gatekeeper use only can be put into the gkselect file. Devices in this file will take priority in the selection process. Solutions Enabler may choose data devices as a Gatekeeper for a particular Symmetrix if none its Gatekeepers in the gkselect file exist, or are offline. - Meta Devices cannot be configured as Gatekeepers - Six Gatekeepers should prove sufficient for any Open Systems control server - Gatekeepers should be RAID protected, they are not SRDF, or virtual or snap devices, or devices in a save pool or thin pool. - Gatekeeper devices must be mapped and masked to single hosts only and should not be shared for concurrent I/O across hosts License management •Symlmf add -type emclm -sid xxxx file (for adding Symmetrix Array Based Licenses) •Symlmf add -type emclm -file (for adding Symmetrix Host Based Licenses) •Symlmf show -type emclm -sid xxxx (This command displays the license file that was last used for installation of licenses) •Symlmf query -type emclm -sid xxxx (This command displays the state of all licenses for a Symmetrix that were activated by the license file or were deemed "in-use" at the time of upgrade from Enginuity 5874. In addition, it displays the current usage by license by the Symmetrix. •Symlmf list -type host (This command displays the Host Based Licenses) SRDF/A Target devices should have same configuration layout at R2 side The default device write pending limit (amount of cache slots per volume) should be the same or higher in the R2 as in the R1. This may require more physical cache in the R2 than in the R1. SRDF/A can sometimes reduce the overall bandwidth by 20% over Synchronous SRDF If you model on 15 minute data, you must configure enough cache and bandwidth to keep SRDF/A active for 15 minutes at a minimum. Cycle times may elongate past the minimum during this period. Never guarantee minimum cycle times Do not mix Synchronous and SRDF/A on the same adapters. Directors supporting SRDF/A should not be shared with any other SRDF SRDF/A should be monitored during the initial roll-out to ensure that all components were properly sized and configured. Data needs to be collected via STP or WLA and then run through the tools again to verify the initial projections were correct. STP at 5x71 microcode includes SRDF/A statistics, which can be very beneficial SRDF/A activation is considerate of cache utilization. SRDF/A will capture a delta set of writes and send them in cycles across the link. In addition to the new writesSRDF/A will include up to 30,000 invalid tracks per cycle. This is a design feature and the 30,000 track value was chosen to prevent cache from being flooded by the invalid tracks SRDF/A will drop when 94% of System WP limit is reached. There is a parameter called "Snow Cache Use" or "Max Cache Usage" limit that controls this. If DSE is configured in the Frame, Engineering recommends lowering the SRDF/A "Snow Cache Use" percentage to 74%. Performance Run IOstat -xk 2 Service time (storage wait) is the time it takes to actually send the I/O request to the storage and get an answer back – this is the time the storage system (EMC in this case) needs to handle the I/O. It varies between 3.8 and 7 ms on average Average Wait (host wait) is the time the I/O’s wait in the host I/O queue. Average Queue Size is the average amount of I/O’s waiting in the queue. A very simple example: Assume a storage system handles each I/O in exactly 10 milliseconds and there are always 10 I/O’s in the host queue. Then the Average Wait will be 10 x 10 = 100 milliseconds. In this specific test you can verify that storage wait * queue size ≈ host wait average service time in wait queue, in milliseconds. If the wait queue is crossed more than 10 then we need to check with SAN team regarding this A database operating system process will see the Average Wait and not the service time. So Oracle AWR, for example, will report the higher number and storage performance tools will see the lower number. There is a big difference and I’ve experienced miscommunication on this between database, server and storage administrators on several occasions. Make sure you talk about the same statistics when discussing performance problems! Meta Expansion ----------------- Concatenated Meta: For concatenated meta devices, the 'form meta from dev 123' will NOT overwrite the data on dev 123. An existing data device can safely be used as a metahead. However, any LUNs added as metamembers (via the 'add dev' option) will lose any existing data. Striped Meta: For striped meta devices, the 'form meta from dev 123' WILL CAUSE DATA LOSS on dev 123. DO NOT form a striped meta from a device with pertinent data on it. Also, the same rule applies for metamembers as with concatenated meta devices - namely, that data will be erased when they are added to a meta device. Workaround for striped meta devices: If it becomes necessary to expand a single LUN into a striped metaLUN while preserving data. use one of the following workarounds: 1.Form a concatenated meta device with the LUN as a metahead. Then use SymConfigure to convert the meta device to striped. Be sure to use the 'protect data' option. Refer to the EMC Solutions Enabler Symmetrix Configuration Change CLI Guide available on Powerlink for more information. 2.Create a striped meta device out of separate devices and use replication software to copy the data from the original LUN to the new striped metaLUN. Notes Note on forming Thin Metas: For forming thin metas the same rules apply, forming a concatenated meta will not cause the data on the lun to be lost, but forming a striped thin meta WILL CAUSE THE DATA ON THE LUN TO BE DESTROYED. Expanding thin metas: With Solutions Enabler 7.2 and Enginuity 5875 we support expanding striped thin meta's using the protect_data option. A BCV must be used to backup during the expansion. The BCV needs to be a regular device (not thin) and be of the same size and # of meta members as the thin meta to be expanded. ========================================== Vault Drives: All Clariions have Vault Drives. They are the first five (5) disks in all Clariions. Disks 0_0_0 through 0_0_4. The Vault drives on the Clariion are going to contain some internal information that is pre-configured before you start putting data on the Clariion. Vault Drives contains Vault area, PSM Lun, Flare database Lun and Operating System. The Vault: The vault is a ‘save area’ across the first five disks to store write cache from the Storage Processors in the event of a Power Failure to the Clariion, or a Storage Processor Failure. The PSM Lun: The Persistent Storage Manager Lun stores the configuration of the Clariion. Such as Disks, Raid Groups, Luns, Access Logix information, SnapView configuration, MirrorView and SanCopy configuration as well. Flare Database LUN: The Flare Database LUN will contain the Flare Code that is running on the Clariion. I like to say that it is the application that runs on the Storage Processors that allows the SPs to create the Raid Groups, Bind the LUNs, setup Access Logix, SnapView, MirrorView, SanCopy, etc… Operating System: The Operating System of the Storage Processors is stored to the first five drives of the Clariion. =========================================== Total SRDF Groups supported 250 64 Groups on Single Port for SRDF