| Unique Visitors |
A little while ago I wrote an article titled 5 Tips to Prevent 80% of Virtualization Problems. This article was all about storage and how to configure your storage and the dangers to watch out for. This is because problems in virtualized environments are predominantly caused by or related to storage in one way or another. In that article I explained the impact of queue depths on performance and also some of the dangers of making the HBA device queue depths too high. What I didn’t know at the time I wrote the previous article was that the default queue depth for QLogic HBA’s was changed between vSphere 4.1 and 5.x. This article will being you up to date on the changes and the impacts of the change in default values between vSphere 4.x and 5.x.
The HBA device queue depths are important as I outlined in 5 Tips to Prevent 80% of Virtualization Problems because it has a big impact on the number of parallel IO’s that a VM can issue and that can be serviced. It also has an impact on the number of LUNs that you need to support to achieve the same performance. If your queue depth is too low you can have very high latency when your VM’s are trying to issue a high number of parallel IO’s as they’ll all queue up inside the hypervisor. If your HBA device queue depth is too large you could have lots of IO’s either queueing up in the HBA itself, or you could overload your storage array. So you need to strike the right balance.
This article is a follow up to Cormac Hogans article titled Heads Up! Device Queue Depth on QLogic HBA’s that was published in response to some queries VMware had received from one of their Technical Account Managers. I would recommend you read Cormac’s article for the reasons why the change to the defaults was made, as I won’t cover that here.
The following are the default HBA device queue depths when using QLogic HBA’s for FibreChannel or FCoE SAN connectivity:
Note: This change does not effect Emulex HBA’s, only QLogic.
The VMware KB 1267 – Changing the queue depth for QLogic and Emulex HBAs, which documents the process for changing device queue depths for QLogic and Emulex HBA’s has been updated to include the default queue depths for the adapters per vSphere version.
So is this change to the default really significant? I think it’s significant in that it wasn’t documented anywhere and in fact the QLogic HBA documentation still lists the default as 32. It’s also significant due to the impact of an overload condition can be quite a dramatic negative storage performance hit, which could take a while to troubleshoot. But for a very long time it had been a common best practice for VMware to recommend changing the HBA device queue depth on QLogic HBA’s to 64 from the default of 32. In most cases this had a positive impact on performance with reduced IO latencies. If you are using Storage I/O Control it will dynamically adjust queue slots between different VM’s on a shared datastore and you don’t need to worry about the device queue depths.
Storage I/O Control takes away the worry and will adjust performance to ensure the latency thresholds are met (by default 30ms). If you have vSphere Enterprise Plus (4.1 and above) and you have Multiple VM’s per Datastore you should be making use of Storage I/O Control. The device queue depth is used when there is only one VM per datastore and Disk.SchedNumReqOutstanding is used when there are multiple VM’s per datastore, in which case the per device queue depth is ignored. As Paudie O’Riordan, one of VMware’s Senior Staff Technical Support Engineers says “let the computer (SIOC) make the decision, not the finger and the wind”.
However there are a few cases where the queue depth of 64 had a detrimental impact and that was largely when non-virtual systems were sharing the same storage array as the vSphere hosts. In this case the vSphere hosts got a far larger proportion of the array’s IO resources and this could impact the performance of the non-virtual systems. I would recommend that where possible you don’t share storage arrays between your virtual and non-virtual environments, which would avoid these types of impacts. In cases where that is not possible you will need to carefully consider the quality of service and storage IO isolation requirements and impacts that high performance vSphere hosts could have on the overall storage array.
The Queue Depth for all devices on the QLogic HBA is a total of 4096. So if you have a per device (per LUN) queue depth of 32 you can support 128 LUN’s at full queue depth, without queueing in the HBA. If you increase the queue depth to 64 (as is the new default in 5.x) then you can support only 64 LUN’s at full queue depth. You can still have more LUN’s configured based on the assumption that not all LUNs will be using all the queue depth all at the same time, so you can effectively overcommit queues in essence. But it would pay to consider the impact of a large queue depth if all VM’s do start issuing IO’s. As Cormac says in his article “If you hit the adapter queue limit, then you won’t be able to reach the device queue depth, and may possibly have I/Os retried due to queue full conditions.”
Now this is only for the HBA queue depths. What about the target ports or storage processor ports on the array? A lot of array storage processor ports will have a queue depth of 2048. You should check with your storage vendor what the Target Port Queue Depth is for your array, if any. As you can see if the HBA is only configured to issue IO’s to one target port a single HBA could easily overwhelm the storage processor and this could cause a QFULL. Fortunately your design should have LUN’s configured across multiple target storage processor ports and multiple storage processors to reduce the risk of overloading. So what happens in a QFULL scenario? Well you can read the QLogic Document titled Execution Trottle and Queue Depth with VMware and QLogic HBA’s. In essence the vSphere Host will set the queue depth to the minimum, which is 1. You can just imagine what this would do to your performance.
Final Word
I recommend you read Cormac’s article Heads Up! Device Queue Depth on QLogic HBA’s and read the QLogic document Execution Trottle and Queue Depth with VMware and QLogic HBA’s. Overall the change in default queue depth for QLogic HBA’s should be positive for performance in most environments. In some environments however you may need to adjust the settings to reduce the risks that I have outlined here. It’s far better to be armed with this knowledge than suddenly have your storage performance fall off a cliff and not know what might have caused it. If you have vSphere Enterprise Plus (4.1 or above) and you have Multiple VM’s per Datastore you should be making use of Storage I/O Control.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
I have had a number of customers who are running HP server hardware report that their hosts are constantly getting disconnected from the network, including their management NIC’s (sometimes causing isolation events), and also they sometimes are getting Purple Screens of Death (PSOD). As you can probably guess this is causing them some major pain. HP has issued an advisory regarding these problems that you need to review if you have any of the affected NIC’s – NC522SFP, NC523SFP, NC375T, NC375i, NC522m, CN1000Q.
I have previously written about the problems I experienced with a few customers running the NC522 and NC523 NIC’s in my article HP Critical Advisory – NC522 and NC523 10Gb/s Server Adapters. The customer that I was working with when I came across this problem originally (well before the advisory went out) had a particularly serious problem as the NIC’s were also used for storage access and management. This lead partially to me writing When Management NIC’s Go Down. Fortunately for my customer they now have a stable environment, but they went through dozens of firmware and driver updates, and eventually had to get the cards replaced.
Now there is a new advisory as of December 2012 regarding a broader set of NIC’s and systems that are having some serious problems and causing VMware vSphere hosts to become disconnected from the network and causing PSOD’s. You can find the HP Advisory Here – HP ProLiant and HP StorageWorks Systems: HP NC375i, NC375T, NC522m, NC522SFP, NC523SFP, CN1000Q Network Adapters – FIRMWARE UPGRADE REQUIRED to Avoid the Loss and Automatic Recovery of Ethernet Connectivity or Adapter Unresponsiveness. The title of the advisory really says it all. VMware has issued KB 2012455 regarding this problem. Note that this is not a VMware issue, it’s a hardware issue, and you should upgrade to the firmware / driver combination that resolves the problem as soon as possible.
I hope that once you upgrade the firmware / drivers your environment will become stable as you would normally expect. When working with HP on these types of issues I have found them to be generally responsive when you get to the right people. I would encourage you to work with your account manager and the HP technical support teams to get these issues resolved. If the problems persist after upgrading the firmware as advised then I would strongly recommend you consider replacing the NIC’s with an alternative model after discussions with HP.
Final Word
NIC disconnections and PSOD’s of this type should be extremely rare in the overall scheme of things. I have not come across many of these types of situations in the 10 years I’ve been working with VMware solutions. But when you come across these types of problems they need to be resolved as soon as possible. The best way to approach it is to log support requests with both VMware and your hardware vendors. Hopefully you strike these types of hardware problems during QA testing before your infrastructure goes live into production, but that is not always the case. If you don’t have a QA process for your hardware that includes burn in then I would recommend you consider it.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
Etherchannel or Load Based Teaming has been a popular topic of conversation ever since Load Based Teaming was introduced in vSphere 4.1. Generally the consideration for Etherchannel starts because people are not aware that Load Based Teaming exists as an option, they are not familiar with how virtual networking in vSphere works, or they’ve just always used it. It is quite common for non VMware Admins to think the virtual networking in vSphere acts just like a normal server in which one uplink is active and the others are strictly for failover with no load balancing capability. This is not the case with vSphere and of the five available teaming options only one provides failover only without any load balancing, the four other options all provide load balancing of multiple host uplinks. If you want to know if you should use Etherchannel or Load Based Teaming, and why, keep reading.
vSphere Network Teaming Options
This article assumes vSphere 4.1 or above, but even in previous versions Etherchannel may not be a good choice. The first thing you should know about vSphere Networking is that the out of the box vNetwork Distributed Switch (vDS) support 5 different teaming options, but does not support LACP (802.3ad) or Dynamic Etherchannel. Static Etherchannel is the only form of Etherchannel currently supported (and static link aggregation 802.3ad). You can utilize LACP only if you deploy the Cisco Nexus 1000v or another add on vDS to you environment. This article will not discuss LACP in any detail for this reason.
The five teaming options are:
Route based on originating virtual port
Route based on IP Hash (only one supported with Static Etherchannel and Static 802.3ad)
Route based on Source MAC address
Route based on physical NIC load (Load Based Teaming or LBT)
Use explicit failover order (Not a load balancing algorithm)
All of the choices except “Use explicit failover order” offers uplink load balancing for the vSphere hosts. So you have four options if you are primarily concerned with load balancing the vSphere host uplinks. This however is not the same as a single virtual machine with a single IP address load balancing multiple uplinks and in most cases even this has very little real benefit. I won’t explain all of the various options here as they are covered in the VMware Product Documentation and the purpose of this article is to discuss Etherchannel and Load Based Teaming.
Etherchannel and IP Hash Load Balancing
IP Hash Load Balancing, which requires Static Etherchannel or static 802.3ad be configured on the switching infrastructure, uses a hashing algorithm based on source and destination IP address to determine which host uplink egress traffic should be routed through. VMware’s support and configuration of Etherchannel is explained in VMware KB 1004048. Frank Denneman explains the mechanics of how this works in his article IP-Hash versus LBT, and Ken Cline wrote a good explanation in his article The Great vSwitch Debate – Part 3.
It is possible for some workloads to load balance multiple host uplinks, but in reality the use cases for this are few. To be able to load balance multiple links a workload would have to be sending large amounts of traffic to many destinations (each unlikely to be the same pattern). Each traffic flow will only ever go over a single host uplink, and therefore an individual flow is limited to a single host uplink. The load balancing is calculated on egress traffic only and is not utilization aware.
Etherchannel and IP Hash Load Balancing is technically very complex to implement and has a number of prerequisites and limitations such as:
Configuring Etherchannel and IP Hash Load balancing is a very technically complex process that can be error prone if the correct process is not followed. It is easy to knock a hosts management interfaces off the network during configuration (Refer to VMware KB 1022751). It is very hard to achieve an even balance and very easy that one or more uplink become congested while others are more lightly loaded due to the nature of the IP hashing. In many cases there may be no performance gains, even through there are additional overheads to calculate the IP Hashes for every conversation.
The best use case I can think of for IP Hash Load Balancing is for a download server or very high traffic single web server where no other load balancers are involved and it is not possible to deploy multiple VM’s and load balancers for the purpose (which presents a single point of failure). In this cases you may achieve good load balance and utilization of multiple links, even if this load balancing is not dynamic, and there is little control of congestion. But is the additional technical complexity for such a small use case really worth it? Do you truly need to be able to achieve more throughput from a single VM than a single uplink can sustain? In an environment with many VM’s consolidated onto the host do you want a single VM to be able to monopolize host networking to the detriment of others?
If your only reason for wanting to use Etherchannel and IP Hash Load Balancing is to distribute load over multiple host uplinks and provide redundancy and failover then it is most likely not the best choice. In fact if this is your only objective any other of the teaming methods would be a better choice (excluding explicit failover order). The complexity and limitations, when in most cases there will be no performance gain, doesn’t seem to make it worthwhile. This brings us nicely onto Load Based Teaming.
Load Based Teaming
Load Based Teaming, or Route based on Physical NIC Load is an option on the vDS that has been available since vSphere 4.1. It is a more dynamic load teaming algorithm that evaluates uplink utilization every 30 seconds and if utilization exceeds 75% will move VM’s between host uplink ports. LBT will work on standard access or trunk port, and does not require any special switch configuration, such as stacking or Virtual Port Channels. Each VM will be limited to the bandwidth of a single host uplink. LBT works on both ingress and egress traffic flows. It is incredibly simple to set up and use. Frank Denneman has wrote about LBT when it was first released and then followed up with IP-Hash versus LBT as previously mentioned.
The advantages of LBT are:
The only downside is a single VM vNIC is limited to the bandwidth of a single host uplink. For a VM to effectively utilize multiple host uplinks it would need to be multi-homed by configuring it with multiple vNIC’s. This is a very small downside when the benefits are so great for the majority of workloads and situations and the sheer simplicity of the configuration and use.
What about LACP?
If you have the Nexus 1000v vDS in your environment (or vSphere 5.1 vDS) and you have switching infrastructure capable of supporting Virtual Port Channels then you may benefit from using LACP. LACP with the Nexus 1000v has 19 different hashing algorithms (vSphere 5.1 vDS has only one algorithm). LACP still suffers from the technical complexity as Etherchannel and some of the same limitations, such as not being able to span switches without special configuration and support. However if you are using Nexus 1000v you have chosen a somewhat more complex configuration in addition to the other features and benefits provided. The additional load balancing methods offer a much greater chance to gain even load balance from many more situations than with static Etherchannel, even though a single conversation is still limited to a single host uplink. If you have the Nexus 1000v and infrastructure capable of supporting LACP across multiple switches so the switch is not a single point of failure then this would be a superior choice than using mac pinning.
Final Word
Use Load Based Teaming unless you have no other option, and even then you should seriously consider not using Etherchannel and IP Hash Load Balancing. The complexity, increased overheads and lack of probable real world benefits make IP Hash a poor choice for most use cases. Remember LACP is not currently supported on the native VMware vDS and I think even if VMware decided to support LACP in future the case for LBT in preference to LACP would still be strong. I would be interested to hear your thoughts on this topic.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2012 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
I’ve started to see reports recently of I/O errors when running very high I/O workloads on Windows 2008 and Windows 2008 R2 VM’s. Mostly this was during artificial benchmark tests run against MS SQL and Exchange 2010 with Jetstress. However it could impact production workloads. Upon further investigation it appears these I/O errors are a known defect with a certain version of the PVSCSI driver that comes with VMware Tools and can affect vSphere 4.0 U1, 4.1 and 5.0. Here I’ll cover more about this potentially serious issue and how to fix it.
This problem is described in VMware KB 2004578 – Windows 2008 R2 virtual machine using a Paravirtual SCSI adapter reports the error: Operating system error 1117 encountered along with the versions of the PVSCSI driver that are impacted and a link to the fix. Microsoft has also included a knowledge base article on their site with regard to this, refer to MS KB 2519834 – SQL Server reports “Operating system error 1117 (I/O Device Error)” on VMware ESX environments that are configured to use PVSCSI adapters.
Although the referenced KB articles describe a situation with SQL it is possible for this to happen under any high I/O workload on the impacted versions of Windows, including for example Exchange. The information available right now doesn’t mention Windows 7 VM’s. But Win7 VM’s are generally less susceptible to the same high I/O workloads as Exchange and SQL servers. Even though Win7 VM’s are less susceptible to the same load conditions that would cause this issue the PVSCSI driver in Win7 is still affected by this problem and should be updated. In the case of VDI desktops could be re-provisioned if they experienced this issue.
What makes this issue potentially serious is that in the worst case (rare) scenario this problem could lead to data corruption. This makes it very important that you upgrade or patch your vSphere environment to address this defect. With ESXi 5.0 the patch is included with Update 1. For ESXi 4.1 you should deploy Patch 04 described in VMware KB 2009144 – VMware ESXi 4.1 Patch ESXi410-201201402-BG: Updates VMware Tools.
On completion of the ESXi Patch process, you will be required to update VMware Tools (System Restart Required) on both required VMs and Templates. You should then confirm the PVSCSI StorPort device driver version has been updated to 1.1.3.0 or later. For more information about the VMware Paravirtual SCSI adapter and supported VMs, refer to VMware KB 1010398 – Configuring disks to use VMware Paravirtual SCSI (PVSCSI) adapters.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2012 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.