| Unique Visitors |
If you have an application that needs very high service levels for availability (99.999%), 24/7/365, including maintenance and patching, then OS Clustering with shared storage is a proven solution. However it can be very complex to set up and maintain Fibre-channel, iSCSI or direct attached shared SCSI solutions. In some cases complex configurations are required not just in the hardware, but also the Operating System of the Guests. If you add in virtualization to the mix, the complexity level can increase more, with the need for physical mode raw device maps. Ironically, increased complexity can decrease overall availability, especially with an increase in the probability of human error. So how do we increase availability, decrease complexity, and provide a simple solution for OS Clustering with shared storage?
Nutanix AOS 5.17 with AHV has the answer. Nutanix AHV allows shared storage for Guest OS Clustering without any complex back end storage support or configuration, unlike other hypervisors from leading vendors, which still require Fibre-channel storage if you wish to use virtual disks. From Nutanix AOS 5.17 onwards you are able to configure a shared volume group and directly attach it to 2 or more VM’s and set up Guest Clustering without any complex in guest OS storage configuration at all. There is no complex storage back end, as that is all provided automatically by Nutanix AOS, and no in guest storage configuration, that might ordinarily be required if using iSCSI. This makes the use of Guest Clustering incredibly simple, as well as being very easy to automate, and significantly less difficult to support and troubleshoot.
The process for creating a Guest OS Cluster has 3 main steps:
Here is an example of how the Volume Group might look in the storage section of Prism for AHV:

While it’s possible to have up to 256 vDisks or Volumes within a Volume Group it is recommended to have 32 or less. If you need more Volumes you can create more Volume Groups.
When you attach a Volume Group to VM’s they will be listed in the Volume Group page within the Storage section of Nutanix Prism Element as follows:

If you wish to have a mixed virtual + physical cluster you can choose to enable external client access to the Volume Group. Any physical / external clients can then use iSCSI Initiator to connect to the clusters Target Data Services IP (DSIP) and mount the volumes.
In the example above I created a 4 node Windows 2016 Cluster, which will host SQL Server 2016 as the primary application. The VM’s are listed below, along with an AD Domain Controller:

After the Failover Cluster Manager components and tools are installed you can configure the Failover Cluster. Note: as part of the cluster creation a verification wizard is executed to ensure compliance with the strict rules needed to form a cluster, including shared storage tests for SCSI fencing and persistent reservations. The nodes in this case were displayed as follows within Failover Cluster Manager:

The next step is to install SQL Server on the cluster nodes and assign all the necessary dependent resources, which would look like the following:

I installed a second SQL Server instance in the same Failover Cluster so I could do comparisons between different configurations. You can see that in the image below:

After I created the cluster I did a series of tests including using tools such as HammerDB and Benchmark Factory for Databases. During the tests I performed live migrations to ensure the cluster didn’t blink in spite of the load, and it worked flawlessly.
Final Word
Nutanix AOS 5.17 and AHV makes creating guest clusters simple and quick and supports both Linux and Windows Guest OS types. You can now configure your fav clustering solutions without the traditional complexity and that means it’s way easier to automate. A Nutanix AHV cluster can now support any number of cluster nodes supported by the OS vendors. The next step in the evolution of this will be when AHV supports Metro Cluster across sites, along with volumes, which will allow for geo distributed guest clusters with greatly reduced complexity compared to the traditional implementations. The cluster example in this article with Windows 2016 and SQL Server 2016 was created just by following the standard Microsoft Documentation and directly attaching a Nutanix Volume Group on AHV directly to the 4 Windows VM’s that would form the cluster. That’s it, no special tuning or complexity needed.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster. Copyright © 2012 – 2020 – IT Solutions 2000 Ltd and Michael Webster. All rights reserved. Not to be reproduced for commercial purposes without written permission.
During VMworld USA 2013 where vSphere 5.5 was launched we heard all about the new enhancements. Some of them were less publicised than others. This article will fill you in on another great reason to consider moving to vSphere 5.5 when it is released. vSphere 5.5 brings with it huge enhancements to the support of Windows Failover Clustering (WFC) previously known as Microsoft Cluster Services (MSCS). This by itself could be a major reason customers choose vSphere 5.5 over previous releases. You may recall that clustering support in vSphere 5.1 was quite a complex matrix to consider, and I tried to explain the various options in my article The Status of Microsoft Failover Clustering Support on VMware vSphere 5.1, which was followed shortly thereafter by Windows Server 2012 Failover Clustering Now Supported By VMware With Some Caveats after the VMware KB (KB 1037959 Microsoft Clustering on VMware vSphere: Guidelines for Supported Configurations) was updated. The release of vSphere 5.5 once again rewrites the rulebook for Microsoft Failover Clustering. So lets dive into it a bit and see what’s changed.
vSphere 5.5 introduces full support for the following:
Clustering support still remains with up to 5 nodes per Windows Failover Cluster, but this isn’t really much of a limitation when you can run as many Windows Failover Clusters as you like on top of a VMware vSphere Cluster, or on top of multiple VMware vSphere Clusters. Provided of course you don’t exceed the 255 SCSI devices per host limit. You may also need to set the perennially reserved flag for the RDM’s once you a reasonable number of RDM’s to ensure your hosts boot up as fast as possible. This is covered in VMware KB 1016106 – ESXi/ESX hosts with visibility to RDM LUNs being used by MSCS nodes with RDMs may take a long time to boot or during LUN rescan.
There are a few things that would be nice to have in VMware vSphere for Windows Failover Clusters and I hope these things are included in future releases:
I haven’t mentioned VMware HA above because that is already supported now and has been for a long time. There were always many good reasons to virtualize Microsoft Failover Cluster systems, including increasing availability and management of the systems, now these benefits have been enhanced to include more options. I’m sure they’ll continue to be enhanced into the future. Just a quick note regarding the new Virtual Hardware v10 in vSphere 5.5. This includes some performance enhancements for Windows VM’s and would definitely be the best option when choosing to virtualize Windows Failover Clusters. I’ll cover Virtual Hardware v10 in more detail in a future article.
On a different note, I often get asked about running Windows Failover Clusters on top of a stretched storage solution such as EMC VPLEX, IBM SVC, HP Peer Persistence, or NetApp Metro Clusters, or just running Windows Failover Geo Clusters. In the case of these configurations VMware relies on Microsoft and the storage vendors support statements, and also on having a stretched network environment underpinning the overall solution. You would need to carefully consider how many nodes would be supported at each site, and if the overall complexity of the solution was justified. I would also strongly recommend an in depth testing and validation process that covers all conceivable failure scenarios. You would also need to have a test instance of the software and infrastructure in order to achieve the high availability you are seeking. On top of this you’d need to consider DR. Geo Clusters are a high availability solution, they are not a DR solution. You have to plan for things like corruption of the cluster and how that would be recovered. However there is no reason you couldn’t operate a stretched Windows Failover Geo Cluster virtualized, and you would achieve additional availability and manageability benefits for doing so.
One of the features introduced in vSphere 5.0 Update 1 was support for a new type of storage behaviour called a Permanent Device Loss or PDL. This is a state common in Stretched Metro Cluster environments (vMSC) where a device becomes unavailable at one site, or where a device is administratively removed and will not be coming back. PDL handles the SCSI sense codes that are sent back from the storage arrays and vSphere then stops sending IO’s to the failed devices. Use of PDL is described in Duncan Epping’s article vSphere Metro Storage Cluster solutions and PDL’s. In vSphere 5.5 the PDL behaviour has been further enhanced.
vSphere 5.5 introduces PDL AutoRemove, which automatically removes a device in a PDL state from a host. A PDL state on a device implies it cannot accept more IOs, but needlessly uses up one of the 256 device per host limit. Now the PDL devices will be automatically removed, which frees up the number of devices per host, given that the devices in PDL state are not coming back. In the case of a vMSC environment the devices may eventually come back at some point in the future and you can simply rescan to pick them up.
Finally before I end this article I would like to mention that vCenter 5.5 now includes support for the SQL Server DB being hosted on a Windows Failover Cluster, however SQL Server 2012 AlwaysOn Availability Groups aren’t supported for the vCenter Server Database. vCenter 5.5 also supports Oracle RAC.
Final Word
In addition to the new Jumbo VMDK support in vSphere 5.5 the enhancements to Windows Failover Clustering support provide even more reason to virtualize your business critical applications with confidence. You can be assured that performance has been greatly enhanced and there will be more articles specifically on performance to come. If you’d like to review more vSphere 5.5 features I recommend checking out the What’s New in vSphere 5.5. Platform-Quick Reference that Alan Renouf has been kind enough to put together.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
It was good to read recently on Duncan Epping’s blog Yellow Bricks that database clustering support for vCloud Director is added in version 5.1. vCloud Director now supports both Oracle RAC and also Microsoft Cluster Services or Microsoft Failover Clustering for MS SQL Server for it’s database. This was previously not supported in vCloud Director 1.0 and 1.5. But what is the support situation for all of the other components when it comes to database clustering, including vCenter Server?
The announcement regarding clustering support for vCloud Director 5.1 databases can be found in kb 2037802. It is really great that VMware has made it clear what the support situation is with no ambiguity.
Here is my understanding of the current state of play with regard to clustered database support (RAC and MSCS/MSFC) based on publicly available information (or lack thereof), this is for the main components of the VMware Cloud Suite that would generally be deployed in most environments:
Single Sign-On (SSO) – Not Supported, vCenter Server Heartbeat is the supported and recommended solution for database high availability
vCenter Server – Not Supported, vCenter Server Heartbeat is the supported and recommended solution for database high availability.
vCenter Server Virtual Appliance – Not Supported, No Database HA currently Supported
vCenter Update Manager – Not Supported, vCenter Server Heartbeat is the supported and recommended solution for database high availability
vCloud Director – Supported from version 5.1 both Oracle RAC and MSCS/MSFC
vCenter Chargeback – Oracle RAC Only from v1.6.2 (refer release notes)
vCenter Orchestrator – Oracle RAC 11g (refer kb 1022828), or vCenter Server Heartbeat
So what do I mean when I say Not Supported? I mean there is no explicit support statement from VMware that they have tested and verified the configuration with clustered databases at the time of writing this article. VMware will provide best efforts support and still help customers troubleshoot their environments. But like most vendors if they suspect a problem caused by the clustered database or clustering technology will refer the customer back to the vendor of that technology. I know many customers that are running vCenter with Oracle RAC databases at the back end successfully and have never had any problems, provided they set it up correctly and tested it, including failover scenarios (should be active/passive configuration). Likewise with a clustered MS SQL database customers have run that successfully as well with vCenter. In many cases the clustering of the database is completely transparent to the application. But that is not always the case depending on how the components connect to and communicate with their database.
When it comes to components that don’t use traditional ODBC connectivity to the database, such as SSO and the vCenter Virtual Appliance, support for clustered MS SQL databases is a bit more difficult. This seem to be primarily why in my opinion SSO is not currently supported with a MS SQL clustered database. For vCenter Server Virtual Appliance it doesn’t support MS SQL at all, but does support Oracle. Oracle RAC Support for the vCenter Server Virtual Appliance is currently going through validation from what I’ve heard and will in the future be supported.
Remember when it comes to support and VMware if something is not included in public documentation, including the product documentation and KB articles, and is not in the product interoperability matrix, then it is not supported. If the support statement isn’t explicit then you can assume that VMware will do everything they can to help (based on my experience and reading relevant kb’s), but it’s on a best efforts basis only.
I’ve mentioned vCenter Server Heartbeat a couple of times above. Some of you may also remember the article I wrote titled Using vCenter Heartbeat to Protect Non-vCenter SQL DB? Think Again!. To ensure you’re aware of what components are supported for protection with vCenter Server Heartbeat I have included them below. This is directly from the vCenter Server Heartbeat Product Documentation. The case remains that if the component isn’t listed below it is not supported and only the databases relevant to the components below (if any) are supported for protection by vCenter Server Heartbeat.
vCenter Server Versions 5.1
■ VMware vCenter Inventory Service
■ VMware ADAM
■ VMware USB Arbitration Service
■ VMware vCenter Server
■ VMware vSphere Client
■ VMware vSphere Web Client
■ VMware vCenter Update Manager
■ VMware vSphere Update Manager Download Service
■ VMware vCenter Orchestrator Configuration
■ VMware vCenter Orchestrator Server
■ VMware vSphere ESXi Dump Collector
■ VMware Syslog Collector
■ VMware vSphere Auto Deploy
■ VMware vSphere Authentication Proxy
■ VMware vCenter Host Agent Pre-Upgrade Checker
■ VMware vCenter Single Sign On
■ VMware vSphere Profile-Driven Storage Service
■ RSA SSPI Service
■ View Composer 1.1, 2.0, 2.7, and 3.0
Note Remote deployment of View Composer is supported starting with View Composer 3.0
■ VMware View Composer
■ VMware Universal File Access
■ vCenter Converter Enterprise
When considering vCenter Server Heartbeat for your environment I would recommend you get VMware involved in the design and deployment, or at least someone (Partner / Consultant) that is experienced with it’s implementation (blatant plug: like my company). It provides great features and excellent availability but does introduce additional complexity. VMware or your preferred partner can help you navigate the complexity and ensure a robust handover to operations to ensure you get what you expect.
Final Word
Depending on the component there are different solutions to provide database high availability. There is no standard across the board between the different VMware components that make up the new vSphere 5.1 Cloud Suite. I expect as the suite develops we will start to see a lot more standardisation in this respect and to see more availability options. Careful consideration needs to be given to providing high availability when deploying the suite within current constraints. As and when the availability options change I will update this post.
Given that this has been based on the easily accessible and publicly available information it’s possible I’ve got something wrong. If so I’d be happy to correct it, so please let me know.
Note: Anyone who is wondering what MSFC is it’s the new term Microsoft use to name their clustering technology. In the past (Windows 2003 and previous) it was called Microsoft Cluster Services, and now it’s called Microsoft Failover Clustering. Many people use the term MSCS to describe both, but MSFC is vastly improved over the previous MSCS.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2012 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.