(cas:72) Google Analyticator was unable to authenticate you with Google using the Auth Token you pasted into the input box on the previous step.

This could mean either you pasted the token wrong, or the time/date on your server is wrong, or an SSL issue preventing Google from Authenticating.

Try Deauthorizing & Resetting Google Analyticator.

Tech Info 400:Error fetching OAuth2 access token, message: 'invalid_grant'
Unique
Visitors
Powered By Google Analytics
Clustering – Long White Virtual Cloudsu by http://longwhiteclouds.com all things Nutanix, VMware, cloud and virtualizing business critical applications Mon, 09 Nov 2020 06:49:28 +0000 en-US hourly 1 https://wordpress.org/?v=6.7.6 45024036 Simple Guest OS Clustering Without Complex Config http://longwhiteclouds.com/2020/11/09/simple-guest-os-clustering-without-complex-config/ http://longwhiteclouds.com/2020/11/09/simple-guest-os-clustering-without-complex-config/#comments Mon, 09 Nov 2020 06:38:06 +0000 http://longwhiteclouds.com/?p=12273


If you have an application that needs very high service levels for availability (99.999%), 24/7/365, including maintenance and patching, then OS Clustering with shared storage is a proven solution. However it can be very complex to set up and maintain Fibre-channel, iSCSI or direct attached shared SCSI solutions. In some cases complex configurations are required […]

]]>


If you have an application that needs very high service levels for availability (99.999%), 24/7/365, including maintenance and patching, then OS Clustering with shared storage is a proven solution. However it can be very complex to set up and maintain Fibre-channel, iSCSI or direct attached shared SCSI solutions. In some cases complex configurations are required not just in the hardware, but also the Operating System of the Guests. If you add in virtualization to the mix, the complexity level can increase more, with the need for physical mode raw device maps. Ironically, increased complexity can decrease overall availability, especially with an increase in the probability of human error. So how do we increase availability, decrease complexity, and provide a simple solution for OS Clustering with shared storage?

Nutanix AOS 5.17 with AHV has the answer. Nutanix AHV allows shared storage for Guest OS Clustering without any complex back end storage support or configuration, unlike other hypervisors from leading vendors, which still require Fibre-channel storage if you wish to use virtual disks. From Nutanix AOS 5.17 onwards you are able to configure a shared volume group and directly attach it to 2 or more VM’s and set up Guest Clustering without any complex in guest OS storage configuration at all. There is no complex storage back end, as that is all provided automatically by Nutanix AOS, and no in guest storage configuration, that might ordinarily be required if using iSCSI. This makes the use of Guest Clustering incredibly simple, as well as being very easy to automate, and significantly less difficult to support and troubleshoot.

The process for creating a Guest OS Cluster has 3 main steps:

  1. Create 2 or more VM’s with the OS of your choice
  2. Create a Volume Group with the number and size of virtual disks that you want for your clustered applications and attach it to the VM’s
  3. Configure the clustering software inside of your chosen OS and deploy the applications

Here is an example of how the Volume Group might look in the storage section of Prism for AHV:

While it’s possible to have up to 256 vDisks or Volumes within a Volume Group it is recommended to have 32 or less. If you need more Volumes you can create more Volume Groups.

When you attach a Volume Group to VM’s they will be listed in the Volume Group page within the Storage section of Nutanix Prism Element as follows:

If you wish to have a mixed virtual + physical cluster you can choose to enable external client access to the Volume Group. Any physical / external clients can then use iSCSI Initiator to connect to the clusters Target Data Services IP (DSIP) and mount the volumes.

In the example above I created a 4 node Windows 2016 Cluster, which will host SQL Server 2016 as the primary application. The VM’s are listed below, along with an AD Domain Controller:

After the Failover Cluster Manager components and tools are installed you can configure the Failover Cluster. Note: as part of the cluster creation a verification wizard is executed to ensure compliance with the strict rules needed to form a cluster, including shared storage tests for SCSI fencing and persistent reservations. The nodes in this case were displayed as follows within Failover Cluster Manager:

The next step is to install SQL Server on the cluster nodes and assign all the necessary dependent resources, which would look like the following:

I installed a second SQL Server instance in the same Failover Cluster so I could do comparisons between different configurations. You can see that in the image below:

After I created the cluster I did a series of tests including using tools such as HammerDB and Benchmark Factory for Databases. During the tests I performed live migrations to ensure the cluster didn’t blink in spite of the load, and it worked flawlessly.

Final Word

Nutanix AOS 5.17 and AHV makes creating guest clusters simple and quick and supports both Linux and Windows Guest OS types. You can now configure your fav clustering solutions without the traditional complexity and that means it’s way easier to automate. A Nutanix AHV cluster can now support any number of cluster nodes supported by the OS vendors. The next step in the evolution of this will be when AHV supports Metro Cluster across sites, along with volumes, which will allow for geo distributed guest clusters with greatly reduced complexity compared to the traditional implementations. The cluster example in this article with Windows 2016 and SQL Server 2016 was created just by following the standard Microsoft Documentation and directly attaching a Nutanix Volume Group on AHV directly to the 4 Windows VM’s that would form the cluster. That’s it, no special tuning or complexity needed.


This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster. Copyright © 2012 – 2020 – IT Solutions 2000 Ltd and Michael Webster. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2020/11/09/simple-guest-os-clustering-without-complex-config/feed/ 7 12273
vSphere 5.5 Windows Failover Clustering Support http://longwhiteclouds.com/2013/09/16/vsphere-5-5-windows-failover-clustering-support/ http://longwhiteclouds.com/2013/09/16/vsphere-5-5-windows-failover-clustering-support/#comments Sun, 15 Sep 2013 20:01:27 +0000 http://longwhiteclouds.com/?p=2355


During VMworld USA 2013 where vSphere 5.5 was launched we heard all about the new enhancements. Some of them were less publicised than others. This article will fill you in on another great reason to consider moving to vSphere 5.5 when it is released. vSphere 5.5 brings with it huge enhancements to the support of […]

]]>


During VMworld USA 2013 where vSphere 5.5 was launched we heard all about the new enhancements. Some of them were less publicised than others. This article will fill you in on another great reason to consider moving to vSphere 5.5 when it is released. vSphere 5.5 brings with it huge enhancements to the support of Windows Failover Clustering (WFC) previously known as Microsoft Cluster Services (MSCS). This by itself could be a major reason customers choose vSphere 5.5 over previous releases. You may recall that clustering support in vSphere 5.1 was quite a complex matrix to consider, and I tried to explain the various options in my article The Status of Microsoft Failover Clustering Support on VMware vSphere 5.1, which was followed shortly thereafter by Windows Server 2012 Failover Clustering Now Supported By VMware With Some Caveats after the VMware KB (KB 1037959 Microsoft Clustering on VMware vSphere: Guidelines for Supported Configurations) was updated. The release of vSphere 5.5 once again rewrites the rulebook for Microsoft Failover Clustering. So lets dive into it a bit and see what’s changed.

vSphere 5.5 introduces full support for the following:

  • Windows 2012 Failover Clustering without having to use in-guest storage access!
  • FC protocol for Windows 2012 Failover Clustering Shared Disks from the ESXi Host for the Cluster pRDM’s
  • FCoE protocol from the ESXi Host for the Cluster pRDM’s
  • iSCSI protocol from the ESXi Host for the Cluster pRDM’s
  • Round Robin Storage Native Multi-pathing Policy

Clustering support still remains with up to 5 nodes per Windows Failover Cluster, but this isn’t really much of a limitation when you can run as many Windows Failover Clusters as you like on top of a VMware vSphere Cluster, or on top of multiple VMware vSphere Clusters. Provided of course you don’t exceed the 255 SCSI devices per host limit. You may also need to set the perennially reserved flag for the RDM’s once you a reasonable number of RDM’s to ensure your hosts boot up as fast as possible. This is covered in VMware KB 1016106 – ESXi/ESX hosts with visibility to RDM LUNs being used by MSCS nodes with RDMs may take a long time to boot or during LUN rescan.

There are a few things that would be nice to have in VMware vSphere for Windows Failover Clusters and I hope these things are included in future releases:

  • Cluster awareness and location awareness within a VMware vSphere Cluster, so that operations make sense from a Clustered VM perspective
  • Support for Shared VMDK’s, rather than having to use pass-through or physical mode RDM’s
  • Support for vMotion and VMware DRS of Windows Failover Cluster Nodes
  • Support for vADP style backups
  • Perhaps support for higher number of Windows Failover Cluster Nodes instead of limiting it to 5, although as I describe above this really isn’t much of a limitation

I haven’t mentioned VMware HA above because that is already supported now and has been for a long time. There were always many good reasons to virtualize Microsoft Failover Cluster systems, including increasing availability and management of the systems, now these benefits have been enhanced to include more options. I’m sure they’ll continue to be enhanced into the future. Just a quick note regarding the new Virtual Hardware v10 in vSphere 5.5. This includes some performance enhancements for Windows VM’s and would definitely be the best option when choosing to virtualize Windows Failover Clusters. I’ll cover Virtual Hardware v10 in more detail in a future article.

On a different note, I often get asked about running Windows Failover Clusters on top of a stretched storage solution such as EMC VPLEX, IBM SVC, HP Peer Persistence, or NetApp Metro Clusters, or just running Windows Failover Geo Clusters. In the case of these configurations VMware relies on Microsoft and the storage vendors support statements, and also on having a stretched network environment underpinning the overall solution. You would need to carefully consider how many nodes would be supported at each site, and if the overall complexity of the solution was justified. I would also strongly recommend an in depth testing and validation process that covers all conceivable failure scenarios. You would also need to have a test instance of the software and infrastructure in order to achieve the high availability you are seeking. On top of this you’d need to consider DR. Geo Clusters are a high availability solution, they are not a DR solution. You have to plan for things like corruption of the cluster and how that would be recovered. However there is no reason you couldn’t operate a stretched Windows Failover Geo Cluster virtualized, and you would achieve additional availability and manageability benefits for doing so.

One of the features introduced in vSphere 5.0 Update 1 was support for a new type of storage behaviour called a Permanent Device Loss or PDL. This is a state common in Stretched Metro Cluster environments (vMSC) where a device becomes unavailable at one site, or where a device is administratively removed and will not be coming back. PDL handles the SCSI sense codes that are sent back from the storage arrays and vSphere then stops sending IO’s to the failed devices. Use of PDL is described in Duncan Epping’s article vSphere Metro Storage Cluster solutions and PDL’s. In vSphere 5.5 the PDL behaviour has been further enhanced.

vSphere 5.5 introduces PDL AutoRemove, which automatically removes a device in a PDL state from a host. A PDL state on a device implies it cannot accept more IOs, but needlessly uses up one of the 256 device per host limit. Now the PDL devices will be automatically removed, which frees up the number of devices per host, given that the devices in PDL state are not coming back. In the case of a vMSC environment the devices may eventually come back at some point in the future and you can simply rescan to pick them up.

Finally before I end this article I would like to mention that vCenter 5.5 now includes support for the SQL Server DB being hosted on a Windows Failover Cluster, however SQL Server 2012 AlwaysOn Availability Groups aren’t supported for the vCenter Server Database. vCenter 5.5 also supports Oracle RAC.

 

Final Word

In addition to the new Jumbo VMDK support in vSphere 5.5 the enhancements to Windows Failover Clustering support provide even more reason to virtualize your business critical applications with confidence. You can be assured that performance has been greatly enhanced and there will be more articles specifically on performance to come. If you’d like to review more vSphere 5.5 features I recommend checking out the What’s New in vSphere 5.5. Platform-Quick Reference that Alan Renouf has been kind enough to put together.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.comby Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.

 


]]>
http://longwhiteclouds.com/2013/09/16/vsphere-5-5-windows-failover-clustering-support/feed/ 4 2355
VMware Products Not Supported with SQL Server AlwaysOn Availability Groups http://longwhiteclouds.com/2013/07/14/vmware-products-not-supported-with-sql-server-alwayson-availability-groups/ http://longwhiteclouds.com/2013/07/14/vmware-products-not-supported-with-sql-server-alwayson-availability-groups/#comments Sun, 14 Jul 2013 04:32:56 +0000 http://longwhiteclouds.com/?p=2164


Recently I posted some updates to the supportability of Windows Server 2012 Clustering and SQL Server in particular in my articles titled Windows Server 2012 Failover Clustering Now Supported By VMware With Some Caveats and The Status of Microsoft Failover Clustering Support on VMware vSphere 5.1. VMware KB 1037959 Microsoft Clustering on VMware vSphere: Guidelines for Supported Configurations explains the configurations […]

]]>


Recently I posted some updates to the supportability of Windows Server 2012 Clustering and SQL Server in particular in my articles titled Windows Server 2012 Failover Clustering Now Supported By VMware With Some Caveats and The Status of Microsoft Failover Clustering Support on VMware vSphere 5.1. VMware KB 1037959 Microsoft Clustering on VMware vSphere: Guidelines for Supported Configurations explains the configurations that are supported and provides guidelines. The reason for this article is because some people might think because these configurations are now supported on top of vSphere that somehow the configurations are also supported with vCenter and other VMware Products. Unfortunately this is not the case. In this article I’ll cover some background and where things are today.

In October 2012 I published an article titled Clustering Support on vCloud Director and vCenter Databases, which explains which products supported clustering of their databases. I didn’t cover every product, just a selection of the popular products, it also didn’t cover SQL Server Replication, Mirroring or AlwaysOn. For things like vCenter, VMware supports vCenter Server Heartbeat as the means of protecting both the vCenter Server itself and the vCenter Server Database (as well as a selection of related components). But you can’t use vCenter Server Heartbeat to protect the vCloud Director database for example, see Using vCenter Heartbeat to Protect Non-vCenter SQL DB? Think Again!.

Clustering of the Database has not been tested or validated by VMware to date prior to vCenter 5.5. VMware will attempt to help customers with their configurations even if it includes a clustered database, unless the clustering technology is believed to be at fault. If the clustering technology is believed to the cause of the fault you would need to contact the vendor of said clustering technology for further assistance.

So where does SQL Server 2012 AlwaysOn Availability Groups come into this picture? SQL Server 2012 AlwaysOn Availability Groups is basically a database availability technique using database replication with automated failover between the members of the availability group (depending on how you configure it). This doesn’t require shared disk clustering, which is why it got added as supported recently on top of Windows Server 2012 and on vSphere 5.1 platform. But in terms of using this for the databases that support your VMware products you’ll be in the same boat as mentioned above with clustering of the databases. VMware has not tested and validated the use of SQL Server 2012 AlwaysOn Availability Groups with their products. It may well work fine, like clustering does for many databases, but it could well break things as well. VMware will likely provide best efforts support up to the point that they believe that AlwaysOn Availability Groups might be causing an issue, at which point it’ll be time to log a call with Microsoft.

[Updated 24/12/2013] As of vSphere 5.5 and vCenter 5.5 VMware has tested and now supported the use of SQL Server Database for vCenter on a traditional Microsoft Failover Cluster solution. This does not however include the use of Always on Availability Groups. Support for Always on Availability Groups may be included in a future release. Two KB articles cover this topic area. KB 1024051 Supported vCenter Server high availability options (Always On Availability Groups would only be covered by third party support as per other cluster configurations), and KB 2059560 Enabling Microsoft SQL Clustering Service in VMware vCenter Server 5.5.

Final Word

So we are all clear. SQL Server 2012 AlwaysOn Availability Groups is not currently tested, validated or supported by VMware for use with VMware’s products. This situation may or may not change in the future. If you wish to use this type of availability technique to protect your VMware products, such as vCenter DB, I would strongly recommend extremely thorough testing and that you weigh up the support risks. It would likely be a lot cheaper and easier to use vCenter Server Heartbeat, which would be my recommended solution for availability of the vCenter Server Database at least, and the other DB’s that vCenter Server Heartbeat supports. SQL Server 2012 AlwaysOn Availability Groups is supported to run on top of the vSphere 5.1 platform to provide database services to any application that supports this method of availability. As always feedback is appreciated.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.comby Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2013/07/14/vmware-products-not-supported-with-sql-server-alwayson-availability-groups/feed/ 6 2164
The Status of Microsoft Failover Clustering Support on VMware vSphere 5.1 http://longwhiteclouds.com/2013/03/22/the-status-of-microsoft-failover-clustering-support-on-vmware-vsphere-5-1/ http://longwhiteclouds.com/2013/03/22/the-status-of-microsoft-failover-clustering-support-on-vmware-vsphere-5-1/#comments Fri, 22 Mar 2013 10:31:23 +0000 http://longwhiteclouds.com/?p=1906


The number of enquiries I’ve been receiving regarding Microsoft Failover Clustering, especially for Microsoft SQL Server Databases has skyrocketed in the past few weeks. I have been receiving a number of enquiries from customers and also from partners including cloud service providers. As a result I thought I’d write this article to help you understand […]

]]>


The number of enquiries I’ve been receiving regarding Microsoft Failover Clustering, especially for Microsoft SQL Server Databases has skyrocketed in the past few weeks. I have been receiving a number of enquiries from customers and also from partners including cloud service providers. As a result I thought I’d write this article to help you understand what the current status is of support for Microsoft Failover Clustering on VMware vSphere 5.1 (GA) and with regard to some VMware products.

Background Reading

Firstly there are two main VMware knowledge base article that outline the support statements of Microsoft Failover Clustering and Microsoft Cluster Services on VMware vSphere. They are as follows:

Microsoft Cluster Service (MSCS) support on ESXi/ESX (1004617)

Microsoft Clustering on VMware vSphere: Guidelines for Supported Configurations (1037959)

This article only applies to vSphere 5.1. The rule book has been rewritten with vSphere 5.5, check out my article on vSphere 5.5 Windows Failover Clustering Support.

Clustering and VMware Solutions

In addition to the above there are specific mention of clustering configurations for the VMware technologies that support it, such as for the vCloud Director SQL Database, which was introduced in vCD 5.1 and covered in my article Clustering Support on vCloud Director and vCenter Databases. The golden rule is this. If VMware does not specifically document a clustering solution as being supported then it is NOT supported. vCenter Server from version 4.0 to current 5.1 GA does not support a clustered Database, be it Oracle RAC or SQL Server. It has not been tested by VMware and is therefore not supported. This may well change in the future as VMware recognises the need to provide alternative high availability solutions for the vCenter Database and I will update this article accordingly. However currently the supported high availability solutions for the vCenter and its database are VMware HA, and vCenter Server Heartbeat. Clustering of the vCenter Server itself is also not supported by VMware but is covered by KB article 1024051 – Supported vCenter Server high availability options.

Customers with production support who wish to run Oracle RAC for the DB for vCenter (not SSO, as that doesn’t work) can get support from the VMware Oracle Support Team under VMware’s Expanded Oracle Support Policy. But they will be limited by the capabilities of vCenter itself, if any. I do know a number of customers running vCenter DB (Not SSO) on Oracle RAC in an active/passive service configuration and it has been fine for years. Also I expect the official support statement to change in the future as the testing for vCenter and RAC is completed.

Not supported does not always mean something doesn’t work. But it does mean it hasn’t been tested by VMware and therefore VMware can’t stand behind the configuration as a supported solution. If it’s not documented as supported, then it’s not supported.

The Status of Microsoft Failover Clustering Support on VMware vSphere 5.1

VMware has done a lot of work to enhance support for Microsoft Failover Clustering and its predecessor Microsoft Cluster Services on VMware vSphere 5.1 to support larger cluster sizes. You can now support up to 5 nodes in a virtual Microsoft Failover Cluster on vSphere 5.1. This is great news for environments where two nodes was not enough, even when combined with the additional availability of VMware HA. I’ve implemented a number of solutions where Microsoft Failover Clustering was used successfully in the cases where it was justified and within the limits that were supported. Strong justification and support constraints are two things I’d like you to think about as you read further.

You can still do hybrid Physical and Virtual clusters, and you can also still do cluster-in-a-box with VMDK’s (dev / test of cluster functionality itself not for high availability). VMware Site Recovery Manager is also supported to protect Microsoft Failover Clusters from a DR perspective and there are a number of different configurations you can use such as multi-node to single node, or multi-node to multi-node. This really does make DR for the cluster easy, less error prone, and of the recovery plan itself once it is initiated is automated and provides audit reporting. VMware HA is fully supported, however VMware recommends you implement anti-affinity rules to ensure cluster nodes are prevented from start up on the same physical host.

So what are the gotcha’s or caveats I hear you ask? Well there are a few gaps in support that you should be aware of when developing your solution architecture. I’ll also cover some of the other valid options you have for high availability later as well and some of the impacts of using Microsoft Failover Clustering. This list is in no particular order.

  1. Clustering Across Boxes (i.e. traditional clustering for high availability purposes) is not supported with the use of VMDK’s or Virtual Mode RDM (vRDM). You must use Physical Mode RDM’s (pRDM) due to the requirement of persistent SCSI reservations.
  2. Due to the requirement to use pRDM’s there is no support for doing backups with vSphere API’s for Data Protection (vADP). So you must use in guest agents for backup.
  3. The is no support for vMotion or DRS with Microsoft Failover Clusters as they use shared disks and a shared SCSI bus. Any attempt to migrate a cluster node will be met with an error message.  This doesn’t mean you can’t deploy a Microsoft Failover Cluster inside a VMware DRS cluster, because you can and it’s fine, it just means that DRS can’t automatically migrate the Microsoft Failover Cluster nodes automatically because vMotion isn’t supported.
  4. Windows Server 2012 Failover Clustering is not supported currently. Period. Not even with in-guest iSCSI. [Updated 21/06/2013] Except with non-shared disk access, in-guest iSCSI, or in-guest SMB storage access. MS SQL Server 2012 on top of Windows Server 2012 with AlwaysOn Availability Groups is supported as it does not require shared disk. See the VMware KB Microsoft Clustering on VMware vSphere: Guidelines for Supported Configurations (1037959) and my article Windows Server 2012 Failover Clustering Now Supported By VMware With Some Caveats.
  5. There is no support for Native iSCSI (where an RDM is presented via the host iSCSI initiator or iSCSI HBA to a guest)
  6. There is no support for Fibre Channel over Ethernet (FCoE). Even if the FCoE Converged Network Adapter (CNA) presents itself as a normal HBA to the host, the use of this configuration with Microsoft Failover Clustering is not supported. With one exception – Two node cluster configuration with Cisco CNA cards (VIC 1240/1280) and driver version 1.5.0.8 is supported on Windows 2008 R2 SP1 64-bit Guest OS in vSphere 5.1 Update 1.
  7. The use of Round Robin Multipathing for your Path Selection Policy (PSP) is not supported.
  8. If you are deploying a hybrid physical node – virtual node Microsoft Failover Cluster the Physical Node can’t use Multipathing software.
  9. No support for VM snapshots, which is one of the reasons that vADP backups don’t work.
  10. No support for Storage vMotion due to the use of pRDM’s.

Some of the above restrictions, especially lack of vMotion and DRS support make it very difficult for cloud service providers that are using vSphere to offer Microsoft Failover Clusters as a service. The reason is obvious. One of the main benefits of having an Infrastructure as a Service is completely non-disruptive hardware upgrades and maintenance. This is not possible with Microsoft Failover Clusters with the current constraints. If cloud service providers wanted to offer a Failover Clustering option they would need to notify customers to shut down their cluster nodes each time the firmware, drivers, or hypervisor version needs to be updated on their hosts. This of course also applies in your private cloud. Downtime would be required on the nodes each time the hosts need to be updated due to the lack of vMotion and DRS capabilities.

Even with the limitation though the advantages of virtualizing your clusters still outweigh the drawbacks. You still benefit from VMware HA, and the performance and reliability you’ve come to expect. You also get the benefit of being able to use VMware SRM for Disaster Recovery.

Options and Alternatives for High Availability

Failover Clustering is inherently complex. It doesn’t always provide high availability either. There are scenarios where downtime is still required and that downtime might be as much as would be expected just using VMware HA. Because the underlying disks are shared any storage loss or corruption will affect the entire cluster. A cluster if very static and hard to move about between hosts or from your private cloud to a cloud provider if you wished. These are some of the reasons it’s not always the best option.

When considering clustering and you think you’re protecting against Guest OS or Host failure think about the last time you saw a Blue Screen of Death (BSOD) from a VM, or you had a host fail. Hardware reliability is greatly improved and most BSOD’s are caused by drivers. With the standard drivers used when you virtualize your servers you are very unlikely to get a BSOD, at least based on the VMware drivers. There will always be exceptions but this is the case based on my experience for the vast majority of workloads, including those with high availability requirements. Microsoft Failover Clustering is not a DR mechanism (generally), so you need additional measures to provide DR, however some of the alternatives can provide HA and DR in the one solution.

Microsoft Failover Clustering can provide flexibility at times around OS patching. But even this use case has alternatives that provide the same level of availability. I give you an option around rolling patch upgrades below.

If you want to provide high availability to vCenter and the vCenter components then the options are VMware HA and vCenter Server Heartbeat. If you are looking to provide high availability to SQL Databases (ones that are not being used for VMware products that don’t support database clustering) then you have a number of options and alternatives, again these are in no particular order.

  1. In guest iSCSI initiation. This is fully supported and will still allow vMotion and DRS migration to occur. Please Refer to the VMware Guide Titled – Setup for Failover Clustering and Microsoft Cluster Service – Update 1, ESXi 5.1, vCenter 5.1. The guide reads as follows on page 9 “Use of software iSCSI initiators within guest operating systems configured with MSCS, in any configuration supported by Microsoft, is transparent to ESXi hosts and there is no need for explicit support statements from VMware.”   Although VMware hasn’t specifically tested this with Windows Server 2012 Failover Clustering there aren’t the same restrictions as this is relying on standard in guest support for the clustering, which as per the guide is transparent to ESXi and does not require any specific support statements from VMware. Provided Microsoft Supports it (Direct Guest Initiated iSCSI for Windows Failover Clustering), which they do, then it’s fine. I would still recommend in guest agents for backups in this case. Incidentally Cisco has a great guide on how to deploy this configuration – Microsoft SQL Server 2012 Failover Cluster on Cisco UCS with iSCSI-Based Storage Access Deployment Guide. This option allows more than the 5 cluster nodes supported by vSphere normally and in fact you could configure a cluster up to the maximum number of nodes supported by Microsoft.  This option is not supported with the use of VM snapshots. You can read more about Snapshot Limitations Here.
  2. If you’re wanting high availability above 99.9% for an application such as Exchange or SQL Server you can use the built-in replication technologies such as DAG’s, Database Mirroring, Log Shipping, or Always On Availability Groups depending on the version. These are fully supported by VMware, have full support for vMotion, DRS and HA, and can also provide a DR mechanism. They also are supported with use of VMware vSphere API’s for Data Protection (vADP) for backups. DAG’s and Mirroring or Always On Availability Groups can be used for high availability as well as disaster recovery. The failover can also be completely automated. They also provide additional protection against disk based corruption where clustering would completely fail. You should check if your software vendor supports the Microsoft SQL Client (if using SQL Server) and these automated failover options. Unfortunately VMware doesn’t support Database Mirroring or Always On Availability Groups at this time for SSO or vCenter databases.
  3. You could choose to keep things simple and just rely on VMware HA. This is a very viable solution for up to 99.9% availability. This is a great solution for the vast majority of cases. I know of a single VM with 528GB RAM and 32 vCPU’s being protected by VMware HA and it runs the entire SAP system and Oracle DB for a very large organization and has done so reliably and performed exceptionally well meeting their SLA’s. When at all possible I recommend keeping things simple. Unjustified and unnecessary complexity adds the risk of downtime and higher probability of human error.
  4. If you need more availability than VMware HA alone can provide  you could add VM and Application Monitoring and Application HA. This will cover the cases of individual application services failing within the guest.
  5. To cover the use case of failover during in guest patching you can use vCenter Orchestrator in combination with hot add and hot remove and clone operations of virtual disks to patch the OS disks or application disks of a single VM while it’s still running and then fail over to the patched version. This requires some advanced understandings about how the OS and hypervisor work together and would best be done along side VMware PSO, but it is possible. This would achieve very similar availability profile during the rolling patch process as Failover Clustering would.
  6. If you wanted to build a Microsoft Failover Cluster inside of a VMware vCloud Director vApp you could also achieve this. You would need to create the cluster nodes connect them to the shared storage by using in guest iSCSI initiators. They could either connect through an external network out to the iSCSI storage or you could make an iSCSI Target VM as part of the vApp. This would give you Microsoft Failover Clusters as a Service inside an Infrastructure as a Service environment running on top of vCloud Director, with self service, and on demand. There would definitely be some scripting involved but this could be a viable solution, and with orchestration you could also set appropriate anti-affinity rules each time one of these clusters was deployed. No restrictions on HA, vMotion or DRS, it would just work. This option allows more than the 5 cluster nodes supported by vSphere normally and in fact you could configure a cluster up to the maximum number of nodes supported by Microsoft. This option is fully supported by VMware. The idea of using an iSCSI VM inside a vApp came from Andrew Mitchell (@amitchell01 also a VCDX) a colleague from the VMware APJ CoE. This option would not support VM Snapshots. Supportability of vADP for backups is unclear as vADP does get around some of the same limitations for snapshots. But this is not recommended as an option for Production vApps, only development and testing. 

 

Final Word

This article covered the current status as of vSphere 5.1 GA. It is very likely that improvements will be made in future releases to address some of the limitations highlighted above. VMware understands what it needs to do in order to deliver the Software Defined Datacenter and to support Business Critical Applications. I’m sure they’re already working hard to improve platform support for Microsoft Failover Clusters, even if it is only needed in a very small minority of cases. In the meantime my recommendations are to use VMware HA and VM Monitoring or App HA unless there is a very strong justification for something in addition to this. If you have that strong justification then leverage the built in application high availability and protection options. Failover Clustering due to it’s complexities and risks is a last resort and inferior in most cases to application level HA.

[Updated 10/09/2013] As of vSphere 5.5 Microsoft Windows 2012 Failover Clustering is fully supported using Fibre Channel, FCoE, iSCSI or any of the in-guest storage IO access methods. Failover Clustering is also supported for the vCenter Database as of vCenter 5.5.  I cover the enhancements to clustering in more detail in a separate article – vSphere 5.5 Windows Failover Clustering Support.

I’d be very interested to get your feedback on this article and hear some of your experiences running Failover Clusters in VMware vSphere.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.comby Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2013/03/22/the-status-of-microsoft-failover-clustering-support-on-vmware-vsphere-5-1/feed/ 29 1906