| Unique Visitors |
Hi Allen, I assume you've upgraded the on board and add on NIC cards to the latest firmware as well? You might well be right, or it could just be a manufacturing fault with the batch of cards. I've seen it before where a series of components manufactured around the same time all had faults. I hope you're able to get this resolved to your satisfaction.
]]>Well, we've got those very same BIOS settings implemented already (though I need to verify "enhanced cooling" I suppose). Unfortunately, they are still not stable for us and considering the integrated NICs have the same issues, I am just struggling to believe it is a hardware issue instead of a poorly written driver.
]]>Hi Allen, I have a customer that did manage to get stability out of their NC522SFP's after the latest firmware and driver update. But in the end they still got the cards replaced. Other things I've learned to improve the situation include changing the server BIOS to Static High Performance or Maximum Performance, Enhanced Cooling, and turning off C States. These are normally recommended settings for vSphere servers, but the enhanced cooling really seemed to make a difference to the NC522SFP's as they were prone to getting very hot and more problems would then occur.
]]>The QLogic CNA's have also had the same type of problems. It's a big concern when these two big manufacturers both have stability and reliability issues at the same time. My Broadcom and Intel NIC's have been flawless though.
]]>Everyone who picked Emulex as supplier has been burnt by this. IBM, Dell and HP all use them. At least with HP, their G8 lines no longer force you to take on Emulex and you can actually choose to go back to Broadcomm.
The problem is worse if you use these Emulex NICs for IP-base storage.
]]>Hi Jonas, VMware relies heavily on the OEM vendors of the hardware to qualify their products for use with VMware vSphere. The qualification tests may well include burn in testing. But this does not mean every server gets burned in before it is shipped. Also customer burn in tests don't need to take months. 48 hours would be normal.
It's a reality that not all servers that come out of the factor are defect free or have the most up to date firmware or drivers. If you want to ensure your environment is reliable, some measure of QA testing of your hardware is important prior to putting it into production. Also hardware these days is largely software and anyone that's been in the software business or is a user of software knows that it has bugs. So why would modern hardware be any different?
I've been burned too many times by hardware bugs not to do the necessary brief testing prior to production use. Even today (literally today) a batch of new servers for a customer would not perform vMotion due to firmware bugs with their CNA's. The bugs were only fixed in a fairly recent combination of firmware and drivers. So it's up to you if you follow this advice or not. But I would still recommend it. I would agree it shouldn't be required. But in my opinion it is.
]]>