| Unique Visitors |
I have had a number of customers who are running HP server hardware report that their hosts are constantly getting disconnected from the network, including their management NIC’s (sometimes causing isolation events), and also they sometimes are getting Purple Screens of Death (PSOD). As you can probably guess this is causing them some major pain. HP has issued an advisory regarding these problems that you need to review if you have any of the affected NIC’s – NC522SFP, NC523SFP, NC375T, NC375i, NC522m, CN1000Q.
I have previously written about the problems I experienced with a few customers running the NC522 and NC523 NIC’s in my article HP Critical Advisory – NC522 and NC523 10Gb/s Server Adapters. The customer that I was working with when I came across this problem originally (well before the advisory went out) had a particularly serious problem as the NIC’s were also used for storage access and management. This lead partially to me writing When Management NIC’s Go Down. Fortunately for my customer they now have a stable environment, but they went through dozens of firmware and driver updates, and eventually had to get the cards replaced.
Now there is a new advisory as of December 2012 regarding a broader set of NIC’s and systems that are having some serious problems and causing VMware vSphere hosts to become disconnected from the network and causing PSOD’s. You can find the HP Advisory Here – HP ProLiant and HP StorageWorks Systems: HP NC375i, NC375T, NC522m, NC522SFP, NC523SFP, CN1000Q Network Adapters – FIRMWARE UPGRADE REQUIRED to Avoid the Loss and Automatic Recovery of Ethernet Connectivity or Adapter Unresponsiveness. The title of the advisory really says it all. VMware has issued KB 2012455 regarding this problem. Note that this is not a VMware issue, it’s a hardware issue, and you should upgrade to the firmware / driver combination that resolves the problem as soon as possible.
I hope that once you upgrade the firmware / drivers your environment will become stable as you would normally expect. When working with HP on these types of issues I have found them to be generally responsive when you get to the right people. I would encourage you to work with your account manager and the HP technical support teams to get these issues resolved. If the problems persist after upgrading the firmware as advised then I would strongly recommend you consider replacing the NIC’s with an alternative model after discussions with HP.
Final Word
NIC disconnections and PSOD’s of this type should be extremely rare in the overall scheme of things. I have not come across many of these types of situations in the 10 years I’ve been working with VMware solutions. But when you come across these types of problems they need to be resolved as soon as possible. The best way to approach it is to log support requests with both VMware and your hardware vendors. Hopefully you strike these types of hardware problems during QA testing before your infrastructure goes live into production, but that is not always the case. If you don’t have a QA process for your hardware that includes burn in then I would recommend you consider it.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
I recently had a situation at a customer site where a physical NIC failure (Read the Critical Advisory I blogged about) caused an All Paths Down (APD) and management network failure. This required a replacement of the NIC, and in our case to a different NIC type, which didn’t have a compatible driver already installed on the host. See if you can guess where this is heading. New NIC installed, Host BIOS, NIC Firmware all updated, and then the real trouble started!
This forced us to move the management vmkernel port over to a different NIC temporarily on a standard vSwitch by force through the DCUI (hence blowing away the network config on the host). Because the NIC that failed (dual port 10Gb/s NIC) had the management and storage access configured on it (iSCSI) the host was not able to read it’s vDS configuration from the SAN. When we finally got the host reconnected to vCenter by relocating it’s vmkernel port this caused all of the vDS configuration to be unavailable. This was compounded by the fact that the failed NIC was replaced being a new NIC of a different type, and the driver didn’t support update manager. So we had to install it manually.
So now we have the host back into vCenter running on another NIC, which fortunately had the correct VLAN’s trunked to it. We’ve updated the NIC driver on the host and the new 10Gb/s NIC is visible. It’s got the same vmnic numbers for both uplinks due to it being installed in the same PCI slot location. Software iSCSI initiator is configured correctly as per the previous state. The host now thinks it’s attached to the vDS. But for some reason the vDS will not sync with the host no matter what we try. We were not even able to remove the host from the vDS, as vCenter thought the port previously occupied by the vmkernel port for management was still operational, which of course it wasn’t.
So the question at this point was how were we to get the host back and configured correctly on the network without completely rebuilding it or configuring all the networking from command line by force? The answer was by using Host Profiles! When we had originally configured this environment we had taken a host profile baseline and had all the hosts in compliance.
Host Profiles to the Rescue! I quickly checked the compliance of the other hosts that were available just to ensure that there was nothing wrong with our baseline host profile. Everything was good. I applied the baseline back to our problem host by puting it into maintenance mode, and then applying the profile. After entering all the relevant vmkernel port IP details for all the vmk ports and waiting a couple of minutes for the configuration everything was looking good.
Finally we rebooted the host to ensure that it started up correctly and could see all the datastores and paths. The host didn’t start! It got caught in an endless loop trying to boot from it’s CDROM and NIC’s. By this stage I was pulling out my hair to figure out why a host that was previously running after the NIC change would suddenly not boot off the internal SD card.
After considering the options for a couple of minutes I decided to enter the BIOS and check all the start order of the devices. Even though the SD card was in the list of start up devices it wasn’t the first in the order. I updated the order so the SD card was first, which in theory should have no baring on it’s ability to start up, as it was previously working. Then I exited the BIOS and with fingers crossed hoped that the host would boot. After it took a couple of minutes to go through post there was a huge relief when the hypervisor started to boot. Once the host had booted we found all the NIC’s and vDS configuration was correct, all the datastores were visible, and finally the host was back up and running.
Lessons learned from this experience:
I hope you never have to experience a situation quite the same as this. Hopefully you’ll be able to address these types of scenarios in your designs before they become problems in production. But that is not always possible due to customer constraints and requirements. Hopefully this will give you some ideas of what can be done to address these sorts of problems when they arise and also demonstrate the value of Enterprise Plus licenses. Let me know some of the more problematic troubleshooting and failure scenarios you’ve come across and what you had to do to get them fixed in a timely manner.
Prior to this problem all hosts within this environment had experienced NIC failures with the same type of NIC, but not to the same extent as this host. All the hosts had been configured with static power at maximum performance, and the fans configured to enhanced cooling. This had made the problems less frequent, but didn’t really address the root cause.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
HP has issued a critical customer advisory regarding some of their high performance server adapters. I’ve had a number of customers impacted by the NC522 and NC523 10Gb/s server adapters losing connectivity. It’s good to see there is a firmware update now that may solve this.
Customer Advisory c02964542, ProLiant and StorageWorks Systems: Certain NC-Series and CN1000Q Network Adapters – FIRMWARE UPGRADE REQUIRED to Avoid the Loss and Automatic Recovery of Ethernet Connectivity or Adapter Unresponsiveness Requiring a Server Reboot to Recover, has been released/revised.
[Updated 22/03/2021] This article is now quite dated and HPE should have long fixed this firmware issue in their NIC’s. This should be viewed as historic only.
Affected Item(s):
HP NC375i Integrated Quad Port Multifunction Gigabit Server Adapter, HP ProLiant DL370 G6 Server series, HP NC522SFP Dual Port 10GbE Gigabit Server Adapter, HP NC522m Dual Port 10GbE Multifunction BL-c Adapter, HP NC375T PCI Express Quad Port Gigabit Server Adapter, HP ProLiant DL580 G7 Server series, HP NC523SFP 10Gb 2-port Server Adapter, HP ProLiant DL980 G7 Server series, HP ProLiant DL585 G7 Server series, HP CN1000Q Dual Port Converged Network Adapter, HP Business Data Warehouse Appliance, HP D2D4312 Backup System, HP D2D4324 Backup System
Description:
IMPORTANT: The network adapter firmware and driver upgrades provided in the Resolution are required to prevent the loss and recovery of Ethernet connectivity, or adapter unresponsiveness requiring a reboot to recover, from occurring. HP recommends performing these upgrades at the customer’s earliest possible convenience. Neglecting to perform the recommended action and not performing the recommended resolution could result in the potential for subsequent errors to occur.
The network adapters listed in the Scope section (below) may encounter either of the following:
* The adapter may temporarily lose Ethernet connectivity, and then automatically recover.
OR
* The adapter may stop responding, requiring a server reboot to recover the operation of the adapter.
Note: There is a low probability of this occurring when operating under a normal network workload.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +.