| Unique Visitors |
Enterprise Strategy Group – ESG, recently completed a evaluation of Nutanix Cloud Platform for Enterprise Workloads and Databases, covering MS SQL Server, Oracle, SAS Analytics, Splunk, and VDI. The results compare performance from 2017 to 2021, and as you expect show some impressive improvements. One thing that shines through is the linear scalability and stability of the Nutanix platform. The performance you get with one node, you get each time you add a node, making performance and capacity scale easily without complications. You can access the full report here, this post will cover some of the brief highlights.
The report (download here) highlights that Nutanix Cloud Platform can simplify infrastructure supporting any enterprise workload or database, and provide consistent, repeatable, and linearly scalable performance with high availability. The elastic scalability and fractional consumption of the Nutanix Cloud Platform can benefit any type of workload, and create a cloud experience wherever a customer chooses to deploy it.
The following image displays the availability graph across Nutanix systems globally and demonstrates high levels of availability can be expected (not included in the report).

Two areas I would like to highlight from the report is the MS SQL Server and VDI Performance, both demonstrate the consistent performance and scalability of the Nutanix Cloud Platform, and low latency that can be expected.


As workloads are added, and nodes are scaled, overall performance increases and response times stay consistent. The report covers many more workloads and has special features to note applicable to big data workloads, such as SAS Analytics, Splunk or Elastic Search. I would recommend you download the full report (download here) and read through it, as it’s only 20 pages.
Final Word
Enterprise workloads and databases need a solid underlying infrastructure to ensure optimal business results. The Nutanix Cloud Platform allows a customer to deploy the same consistent infrastructure in public cloud, co-lo, or on-prem and experience the same reliability and linear scalability. Choose your location, build your cloud, all on your terms.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster. Copyright © 2012 – 2021 – IT Solutions 2000 Ltd and Michael Webster. All rights reserved. Not to be reproduced for commercial purposes without written permission.
Managing large databases is hard. Having changes go through a dynamic environment to production with thorough testing is also hard. When you have a 100TB+ production system, keeping non-production copies up to date is even harder still. Most companies can’t afford to keep exact copies of production due to the size and scale, so compromise the non-production environments. This is no longer necessary. Nutanix Era solves this acute database management problem, and many others!
Nutanix Era is a Database-as-a-Service Solution that solves many of the acute problems facing businesses and DBA’s today when managing large number of databases, or large databases, which impact productivity, reliability, availability and performance.
One of the biggest benefits for customers that have large databases is the ability to take a complete clone of the production system and make multiple copies for non-production uses. Any number of clones can be created without impacting storage capacity requirements, this is due to using native Nutanix API’s in the cloning process. The following 9 minute video shows an example of a 100TB Production Oracle RAC DB being refreshed to 10 clones, for use in testing and development. The clones don’t take up any additional space until the moment data is changed. I hope you like the Jazz music :).
If you like what you see and you’d like to get some hands on experience with Nutanix Era, please take a look at the Nutanix Era Test Drive – Available Online from the comfort of wherever you are.
Recently version 2.0 of Era was released, including the following enhancements.
Hybrid Cloud Flexibility and Multi-cluster Management
With Era 2.0 we have extended our database platform capabilities across clusters and hybrid cloud. This provides our customers the flexibility to build and manage databases on their own terms, giving them the freedom to develop and deploy databases of their choice in the environment they want.
Expanding Database Engine Portfolio with SAP HANA
SAP/HANA has joined Oracle, MS SQL Server, PostgreSQL, MySQL, & MariaDB in the portfolio of database engines supported by Nutanix Era. SAP/HANA customers can now leverage Era’s 1-click capabilities to create an end-to-end sandbox environment on the Nutanix HCI platform.
Extended Choice and Support with PostgreSQL
With Era 2.0, PostgreSQL admins can now take full advantage of all Era capabilities, including patching and in-place restore. In addition, Nutanix will now provide 24×7 support for PostgreSQL databases provisioned by Era, allowing customers to call our support number with any issues.
Enhancing Service Model Delivery with HCL Partnership
For customers seeking fully-managed database services, Nutanix has partnered with global technology company HCL, to offer a joint solution called SKALE DB powered by HCL and Nutanix. The partnership provides a secure, scalable, cloud-ready managed DBaaS offering integrated with Nutanix Era and Prism for automation and management for database & hyperconverged infrastructure.
There is much more to come in future releases!!
Final Word
Good DBA’s are far too valuable to allocate to mundane tasks that can be automated. The DBA’s can create and control the policies, and delegate to competent team members the rights to perform tasks for themselves. DBA’s can then focus on solving real business problems, improving data efficiencies and delivering far greater insights. Utilising Nutanix Era can greatly reduce the storage required to manage database environments and maintaining large numbers of non-production systems. The capital saved from not needing excess storage infrastructure can be redeployed to more business focused investments. There is no need to migrate defects to production because you couldn’t test changes on a full scale copy in your non-production systems! Don’t forget to get some hands on experience with Nutanix Era, please take a look at the Nutanix Era Test Drive – Available Online from the comfort of wherever you are.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster. Copyright © 2012 – 2021 – IT Solutions 2000 Ltd and Michael Webster. All rights reserved. Not to be reproduced for commercial purposes without written permission.
If you have an application that needs very high service levels for availability (99.999%), 24/7/365, including maintenance and patching, then OS Clustering with shared storage is a proven solution. However it can be very complex to set up and maintain Fibre-channel, iSCSI or direct attached shared SCSI solutions. In some cases complex configurations are required not just in the hardware, but also the Operating System of the Guests. If you add in virtualization to the mix, the complexity level can increase more, with the need for physical mode raw device maps. Ironically, increased complexity can decrease overall availability, especially with an increase in the probability of human error. So how do we increase availability, decrease complexity, and provide a simple solution for OS Clustering with shared storage?
Nutanix AOS 5.17 with AHV has the answer. Nutanix AHV allows shared storage for Guest OS Clustering without any complex back end storage support or configuration, unlike other hypervisors from leading vendors, which still require Fibre-channel storage if you wish to use virtual disks. From Nutanix AOS 5.17 onwards you are able to configure a shared volume group and directly attach it to 2 or more VM’s and set up Guest Clustering without any complex in guest OS storage configuration at all. There is no complex storage back end, as that is all provided automatically by Nutanix AOS, and no in guest storage configuration, that might ordinarily be required if using iSCSI. This makes the use of Guest Clustering incredibly simple, as well as being very easy to automate, and significantly less difficult to support and troubleshoot.
The process for creating a Guest OS Cluster has 3 main steps:
Here is an example of how the Volume Group might look in the storage section of Prism for AHV:

While it’s possible to have up to 256 vDisks or Volumes within a Volume Group it is recommended to have 32 or less. If you need more Volumes you can create more Volume Groups.
When you attach a Volume Group to VM’s they will be listed in the Volume Group page within the Storage section of Nutanix Prism Element as follows:

If you wish to have a mixed virtual + physical cluster you can choose to enable external client access to the Volume Group. Any physical / external clients can then use iSCSI Initiator to connect to the clusters Target Data Services IP (DSIP) and mount the volumes.
In the example above I created a 4 node Windows 2016 Cluster, which will host SQL Server 2016 as the primary application. The VM’s are listed below, along with an AD Domain Controller:

After the Failover Cluster Manager components and tools are installed you can configure the Failover Cluster. Note: as part of the cluster creation a verification wizard is executed to ensure compliance with the strict rules needed to form a cluster, including shared storage tests for SCSI fencing and persistent reservations. The nodes in this case were displayed as follows within Failover Cluster Manager:

The next step is to install SQL Server on the cluster nodes and assign all the necessary dependent resources, which would look like the following:

I installed a second SQL Server instance in the same Failover Cluster so I could do comparisons between different configurations. You can see that in the image below:

After I created the cluster I did a series of tests including using tools such as HammerDB and Benchmark Factory for Databases. During the tests I performed live migrations to ensure the cluster didn’t blink in spite of the load, and it worked flawlessly.
Final Word
Nutanix AOS 5.17 and AHV makes creating guest clusters simple and quick and supports both Linux and Windows Guest OS types. You can now configure your fav clustering solutions without the traditional complexity and that means it’s way easier to automate. A Nutanix AHV cluster can now support any number of cluster nodes supported by the OS vendors. The next step in the evolution of this will be when AHV supports Metro Cluster across sites, along with volumes, which will allow for geo distributed guest clusters with greatly reduced complexity compared to the traditional implementations. The cluster example in this article with Windows 2016 and SQL Server 2016 was created just by following the standard Microsoft Documentation and directly attaching a Nutanix Volume Group on AHV directly to the 4 Windows VM’s that would form the cluster. That’s it, no special tuning or complexity needed.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster. Copyright © 2012 – 2020 – IT Solutions 2000 Ltd and Michael Webster. All rights reserved. Not to be reproduced for commercial purposes without written permission.
HammerDB is a very popular benchmark tool for testing multiple different database engines, including Oracle, SQL Server and PostgreSQL. This is a brief article to bring your attention to some ways you may improve your results and get more valid benchmark data.
We do a lot of testing with HammerDB at Nutanix. We run multiple database engines through HammerDB for each release of software we release. One of the great things about running databases on Nutanix is that once you find out how much performance you get from a given configuration of a node or database, you can scale it linearly and get the same experience. This fact was demonstrated by a test that Gary Little (Performance Engineering at Nutanix) did some time ago and published in his article SuperScalin’: How I learned to stop worrying and love SQL Server on Nutanix – Recommended Reading. Over the years we’ve found some things when using HammerDB that can improve consistency and reliability of test results.
Sometimes SQL Server might only use a single NUMA node due to the way connections and transactions are assigned in the engine. Gary describes the problem and solution in his article SQL Server uses only one NUMA Node with HammerDB. The solution is a slight modification to how the HammerDB scripts execute.
Another issue we’ve frequently run into when testing large SQL Server databases that needs lots of cores and memory is that not all the cores get used. Often you might find only 20 CPU’s are being used by the SQL Engine, even though you are using Enterprise edition. This is due to the ISO and installer being used not being the Core Edition of SQL Server Enterprise. You need Core Edition to allow more than 20 CPU’s to be used by the database engine.
Another common issue that we have discovered and Gary has documented in his article HammerDB: Avoiding Bottlenecks In Client, is settings in the actual HammerDB client. Specifically around how the client logs data. Check out Gary’s article for the settings you should uncheck before running a test.
Other basic issues we’ve found is the system under test not being sized properly for SQL Server, or the VM’s running SQL Server not being properly aligned to the NUMA nodes and configurations of the hardware being tested. an 8 vCPU VM isn’t optimal if running on a 12 CPU socket, whereas a 6 vCPU VM would be better. We have found that right sizing the VM’s produces far better results. In this case, 8 vCPU’s might produce less transactions per minute and less new orders per minute and higher response times than if 6 vCPU’s was used. On an 18 core socket, 9 vCPU’s or 6 vCPU’s are also good options. Some people find it difficult to use an odd number of vCPU’s, but the scheduler works fine.
Final Word
Benchmarking SQL Server with HammerDB can help determine what performance you might reasonably expect from a platform running different types of databases and show how a platform scales, but it isn’t exactly the same as your real world workloads. The best way of determining how your workloads will perform is by testing them on a the platform with test cases you have valid comparisons for and under similar conditions. Always account for maintenance tasks, such as backups, stats update, reindex etc and allow headroom for growth. I hope the above help you and wish you all the best benchmarking with HammerDB.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster. Copyright © 2012 – 2020 – IT Solutions 2000 Ltd and Michael Webster. All rights reserved. Not to be reproduced for commercial purposes without written permission.
You are setting up a new template or VM to be used with SQL Server and to get the best performance you are adding multiple disks to the VM. When you first boot the VM the disks are all offline. First thing you do is change the default SAN policy in diskpart to set disks online as follows:
diskpart san policy=onlineall exit
This is great, by default all the connected disks will come online when the VM boots. But what about brining all of the attached disks online all at once so you can do things with them? Well, for that, you need a small script.
I stumbled across an article while Googling, on the Microsoft Technet Forum here. Unfortunately the script didn’t work with a lot of disks, as it had a bug in it. I corrected the bug, so I thought I’d post my modified script here. You can use this to mark all disks on a particular VM or any Windows server for that matter online. I’ve tested this with Windows 2012 R2 and Windows 2016.
#Check for offline disks on server.
$offlinedisk = "list disk" | diskpart | where {$_ -match "offline"}
#If offline disk(s) exist
if($offlinedisk)
{
#for all offline disk(s) found on the server
foreach($offdisk in $offlinedisk)
{
$offdiskS = $offdisk.Substring(2,7)
#Creating command parameters for selecting disk, making disk online and setting off the read-only flag.
$OnlineDisk = @"
select $offdiskS
attributes disk clear readonly
online disk
attributes disk clear readonly
"@
#Sending parameters to diskpart
$OnlineDisk | diskpart
}
}
An alternative suggestion proposed by Barrie in the comments is to use the native PowerShell commands as follows:
$offlinedisks = get-disk | where OperationalStatus -EQ offline foreach ($disk in $offlinedisks) {Set-Disk -Number $disk.Number -IsOffline $false Set-Disk -Number $disk.Number -IsReadOnly $false}
Final Word
The above script is really useful if you are going to add a lot of disks to a VM. You could extend it further to create partitions on all of the disks as well. I used the above to help set up a few SQL Server VM’s for some testing, but it could be used for any type of Windows VM. Don’t forget to turn off write cache by changing the Windows Disk Policy to Quick Remove as I mentioned in this article, before starting any performance testing.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2017 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
At Nutanix Customers’ and Prospects’ often ask about performance data. There is a lot of performance data published, such as in best practice guides, on third party sites such as the SAP SD Benchmark site (where Nutanix has a published certified SAP 2-Tier benchmark), and on blogs such as this. We are able to readily provide insightful data across a wide range of applications, systems and situations. Yet there are still people that say for some reason that we don’t publish performance data (maybe their GoogleFoo is weak?). Nutanix performance is more than just latency and IOPS, we analyze real application characteristics. It’s about time this was resolved with a third party vetted report.
Nutanix has worked with ESG and had them independently review and provide a report on the performance of the Nutanix architecture and platform using real applications as the basis of the analysis. Andy Daniel (at PernixData before joining Nutanix in the acquisition) has done a great job of pulling together the various parties, including my team, to contribute to the work. Andy explains the basis of the testing in his article here, and thanks to Mike Leone from ESG for his expert analysis.
We use real applications because customers run real applications and a micro benchmark isn’t a production application. Martijn Bosschaart explains very well the why of this testing method in his article To Finish First, You First Have To Finish. The applications included cover Microsoft SQL Server, Oracle, Exchange and XenDesktop for VDI. The report shows the typical performance that can be expected from the platforms under testing, some of which are not the latest generation, so performance is likely to have improved on the latest models. The report also shows the predictability and consistency of performance across the test scenarios. If you are interested in the performance of the Nutanix platforms I would highly encourage you to review the ESG Lab Review – Performance Analysis: Nutanix. Hopefully this report can clarify any performance concerns that customers, partners and prospects have.
Final Word
Performance isn’t just one element, it is the result of a whole set of elements across a business solution and should be considered when things are going well, in addition to when things are not going well. It should include upgrades, failed components, recovery operations, and provisioning times and other elements. All of these are part of measuring the success and performance of a business solution. In the absence of specific requirements and a specific solution to a business problem benchmarks and other performance data can give an indication of how a platform might perform or behave under certain specific scenarios. But it’s important to be able to relate those scenarios back to your individual requirements. Dheeraj Pandey, Nutanix CEO, explains how Nutanix is different in many ways to other HCI players in his Quora article. It can be quite nuanced, and architectural decisions of the core platform really make a difference, when things are going well, and when things are not going so well.
The Nutanix team is available to help answer any customer/partner/prospect questions on this report or any other topic relating to the Nutanix platform.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2017 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
It was about this time in 2013 that Michael Corey, Jeff Szastak and I started writing Virtualizing SQL Server with VMware: Doing IT Right (VMware Press) 2014. Microsoft SQL Server was the single most virtualized business critical app in the world then, and it is still the case today. Our book is still as relevant today as it was when we published it and the recommendations we documented still hold true. In spite of it being the most popular critical app to virtualize there are still a lot of cases where some simple best practices are not followed. Best practices that could greatly improve performance. Not all of the best practices apply to all database types at all times, so some care is required. One thing we’ve learned with experience is that the only similarity in customers’ database environments is that they are all different. So this article will focus on the top 5 things you can do to improve performance for many different types of databases and give some examples from performance testing that my team and I have done.
SQL Server Memory Management – Max Server Memory, Reservations, Lock Pages, PageFile
SQL Server, like any RDBMS is really a big cache for data at the end of your storage, regardless of what storage is under the covers. Allocating the right amount of memory to the SQL Server Instance, and ensuring the Operating System has enough memory so it doesn’t cause swapping, are our primary goals. Then we need to look at protecting the SQL Buffer Cache from paging and protecting the VM memory. Paging of the database buffer cache can cause sever performance problems for a busy database server.
Max memory should be changed from the default of 2PB to a value that allows the OS some breathing room. On small SQL Server VM’s with only 16GB or 32GB RAM setting Max Memory to Allocated Memory – 4GB is a good place to start. For larger VM’s the OS may need 16GB or 32GB, and this can be impacted by the number of agents and other tools you have running in the OS. I have seen recommendations that you should leave 10% available for the OS, however this becomes problematic if your VM has a lot of RAM, say 512GB to 2TB.
Reserve the memory assigned to the SQL VM. This guarantees the service levels to the SQL Database, ensures that at least from a memory standpoint it will get the best performance possible, and it will mean the hypervisor page file will not take up any valuable storage. For SQL server VM’s > 16GB memory, and where lock pages are used this is even more critical as hypervisor ballooning is not effective.
Enable the local security policy “lock pages in memory” for the SQL Database service account and set trace flag -T834. This will ensure that the SQL Server instance uses huge pages on the operating system and it protects the database from any OS swapping, as huge pages can’t be swapped. This also reduces the work the OS needs to do with regard to memory management. Huge pages on x86 are 2MB in size, vs the standard 4KB page size that is used by default. Using this setting also prevents ballooning from impacting the SQL Server Instance and is another reason to reserve the memory.
I recommend that you allocate a pagefile big enough for a small kernel dump at minimum, and up to 16GB to 32GB as maximum. If you design your SQL Server VM properly you will not have any OS paging, therefore you shouldn’t need a large pagefile. If you are designing a template that will be used for multiple different sized SQL VM’s then you could consider setting the pagefile to a standard size of 16GB and leaving it at that. The less variation required the better. By reserving the VM memory, locking pages in memory and using -T834 the buffer cache of the DB can’t be paged anyway. These settings will ensure the best possible service level and performance for your database at least in memory.
Split User DB Datafiles and TempDB Across Drives and Controllers
This applies especially to high performance databases. SQL Server can issue a lot of outstanding IO operations and this can cause the queue of a particular drive to become full and prevent the IO’s from being released to the underlying storage for processing. To help alleviate this for large high performance databases you should create the databases with more than one data file (recommended 1 per vCPU allocated to the DB), and allocate each datafile on a different drive. If you have many smaller and less high performance databases you can simply split the database data files across more drives or mount points if you determine a single drive does not provide sufficient performance.
For TempDB the recommendation is 1 datafile per vCPU up to 8 initially and then grow 4 at a time from there as needed. The process of allocating datafiles to TempDB has been automated in SQL Server 2016 so it will choose the correct initial number based on the number of vCPU’s allocated. The reason for doing this is to prevent GAM and SGAM contention.
Each virtual drive and each virtual controller has a queue depth limit, so splitting the datafiles across controllers also helps to eliminate bottlenecks. In a VMware environment you can use up to 4 virtual SCSI controllers, such as PVSCSI and it would be recommended to split the data files across them. You can also tune each controller queue depth by changing registry settings, but be aware of the potential impact on your back end storage. Having really large individual drives / virtual disks might give you extra capacity but it gives you no more performance as the queue depth per device is always limited. This is also the case in cloud environments such as Azure and aligns with Microsoft SQL CAT recommendations.
The image below shows one such design that may be appropriate for splitting data files. This example uses mount points, but you could also use drive letters.
Thanks to Kasim Hansia and Nutanix for the above image.
CPU Sizing, NUMA and MaxDOP
When it comes to CPU sizing for your database VM’s the best size is one that fits within a NUMA boundary. So for a two socket platform this would be a size that fits within a single socket or is easily divisible by the number of cores in a single socket. If you have very large physical servers currently with many databases on them, chopping them up into smaller size VM’s that fit within a NUMA node will help improve performance. The best size of a VM on a system with 8 cores per socket would be 2, 4, or 8 vCPU as an example. In terms of CPU overcommitment a 1:1 vCPU to pCPU core ratio is recommended to start with unless you have a good knowledge of actual system performance, at least for critical production systems, for dev/test a higher ration can be used to start. You can modify increase the ratio as you monitor actual system performance. Production systems general run between 2:1 and 4:1 realistically assuming not all database instances and VM’s running on the host or cluster need the same resources at exactly the same time. You need to design for peak workloads demands and then the averages will take care of themselves.
For very large databases this may not be possible and in that case it is ok to have a VM that spans NUMA nodes as Windows and SQL Server are NUMA aware and will use the processors available to them assuming the correct license, however the scaling of processors across NUMA boundaries in a single VM doesn’t provide linear performance, whereas splitting multiple smaller databases across multiple smaller VM’s that do fit within NUMA boundaries can provide better performance than would otherwise be available on a single physical OS or VM. When it comes to memory and NUMA, more is better for databases and as memory is so much faster than disk or SSD or even NVMe, the penalty for having memory in different NUMA nodes when virtual NUMA is available is not a concern.
With regards to the SQL Server setting Maximum Degree of Parallelism that controls the number of threads or processors that a single query can consume, you need to be careful with modifying it. For OLTP transactional type databases you can get significant performance gains overall when large numbers of users are concurrently accessing the system if MaxDOP is set to a small number or 1, however it is a global setting on the SQL Server instance in versions before 2016 and therefore will impact all databases on an instance. A good rule of thumb may be to set it to an the size of a NUMA node, or some number of processors you are happy to be consumed by a single query. In SQL Server 2016 you can set it per database, so it can be more finely tuned to the individual database workloads. Leaving it at the default of 0 i.e. unlimited can also have negative performance impacts especially when many users access a database as a single query could consume all resources and negatively impact other users.
Networking, Live Migration, Jumbo Frames
When it comes to networking you need to consider more than just the user access to the database, you need to consider management workloads including live migration for maintenance and load balancing, backup, monitoring and out of band management. With very large SQL Server VM’s with 512GB and above the live migration network may have some hefty requirements. Especially with very active SQL Server VM’s. I have seen the live migration networks struggle with evacuating a host for maintenance if they were not designed and implemented correctly. If you have hosts with multiple TB of RAM and enough VM’s to occupy that RAM you should consider multiple 10G networks for live migration traffic. Using LACP network configurations can indeed help, as can using Jumbo Frames. As you start adopting 40GbE, 50GbE and above NIC’s the use of Jumbo Frames to increase performance and lower CPU utilization becomes ever more important. You can achieve up to 10% to 15% additional performance by using Jumbo Frames for live migration traffic depending on CPU type and bandwidth of you NIC. But take care as it does need to be implemented properly. It is fortunate that many enterprise class switches now come with Jumbo Frames enabled by default, but you will still need to enable it in your hypervisor and on the life migration virtual NIC. If you are using Jumbo Frames why not enable SQL Server to use a packet size of 8192 bytes instead of the standard 4096 (same size as a database page although there is no direct relationship), 8192 bytes and it fits nicely into the 8972 byte TCP packet (9000 bytes with overhead included) on the wire. Take into consideration the network impacts of any software defined storage solution especially as adopting modern all flash systems because your network may be too slow for flash.
Maintaining an Accurate and Objective Performance Baseline
Maintaining an accurate and objective performance baseline of your databases is the only real way to measure when things are going wrong or when performance is not acceptable. Before, during and after virtualization you should be updating your baslines whenever major configuration changes are made. This prevents the ‘feeling’ that it’s slow, without defining what slow it, or without being able to quantify the feeling. If you can accurately test acceptable performance and repeat that test then you can be sure your system is behaving as expected throughout its useful life. This is not as easy as it sounds, and due to the many hundreds or thousands of databases that most customers have, a risk based approach is recommended. For the most important or highest risk systems it’s worthwhile making the investment into proper baselines and monitoring, for the great unwashed this might not be practical. There are many tools that can help, and they start from industry standard benchmarks and system monitoring tools to more elaborate enterprise test suites. We cover a number of different options in our book, but a simple option might be to use HammerDB, Record and Replay, and/or SCOM/PerfMon. Without a baseline you have no objective way to measure success.
Performance Results With and Without Best Practices
When you’ve been successfully virtualizing SQL Server for years and you know the best practices it is pretty hard to go back and create a database VM with next, next finish and ignore all that you have learned. But that is just what we had to do in order to measure the difference between the default configuration and applying the best practices. In this case we used a VM with 8 vCPU, 32GB RAM and HammerDB. The only difference between the two tests was the configuration of SQL Server and the operating system. The same number of users in HammerDB are used for each test.
Default configuration without best practices applied:
Configuration after best practices have been applied:
Thanks to Bas Raayman and Bruno Sousa for the two images above.
In this example the difference in performance is 12x between the default configuration and the optimized configuration. The benefits grow as you start to scale out the number of VM’s and number of servers, which is what the next image shows.
Here we have a number of database VM’s being scaled our across a number of servers, in this case using Nutanix systems. The performance growth is linear, as you add more VM’s and more Nutanix nodes you get the same performance per node, and linearly scalability of the overall performance in terms of transactions per minute.
#TBT A blast from the past (2014) Scaling 1M MS SQL transactions per #Nutanix node.https://t.co/CDwiVQKMbQ pic.twitter.com/eViiA3b10m
— Gary Little (@garyjlittle) January 26, 2017
Final Word
We have covered a few best practices that can help improve performance and ensure success of any virtualization project. There is significantly more covered in Virtualizing SQL Server With VMware: Doing IT Right (VMware Press 2014). I also had a hand in crafting the Nutanix best practices for SQL Server, which is freely available. Nutanix has published many best practice guides for many applications and a lot of them are applicable regardless of what system you are running. Hopefully applying some of these simple best practices helps improve the performance of your virtualized SQL Server environments. I’d love to hear any feedback or comments you might have.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2016 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
I’ve been working with a few customers recently that have been using the Microsoft DiskSPD Tool for doing some initial basic tests of their VM storage subsystem. DiskSPD is the tool that replaced SQLIO. Now although it isn’t a completely valid way to test a system that will be used for SQL Server, it can give an indication of how things will perform under certain conditions. But it’s basically just a synthetic IO generator like many others, so it is limited in terms of actual real world applicability to applications. Storage performance is but one factor to consider when designing a system to support a database, and max performance isn’t necessarily indicative of what is required for sustained workloads. But if you want to use it, and to avoid some traps that will invalidate your testing, here are some tips.
Firstly, just like SQLIO, DiskSPD by default will generate zeros if you don’t tell it otherwise. On some new storage systems this will overstate the peak performance numbers as they just acknowledge 0’s and it’s a NOOP at the back end. So be aware of this. In SQLIO we had to use Make-A-File.exe to generate a test file to remove the problem around zero filled data files. This is no longer required with DiskSPD.
To ensure you are always generating real data use the -Z <seed size> parameter. Such as -Z 1G for example. If you use -Z by itself then it will generate just zeros. If you pass it a value such as 1G, it’ll generate a repeating pattern of 1G of data.
To get the most out of your testing I would recommend that you test different IO patterns, different IO sizes, different amount of outstanding IO, and across a single and multiple drive letters.
Be aware that specifying too many outstanding IO operations will just overload the operating system queues and does not allow you to measure the underlying storage subsystem at all. Usually the SCSI driver will have a limited queue depth per drive of 32, potentially higher depending on if you are in a virtual or physical environment and what driver you are using. The Windows Storport driver has a queue depth limit per drive of 255. If you exceed any of these limits your performance suffers and your latency spikes.
SQL Server can drive lots of outstanding IO and can drive lots of queue depth. This is one of the reasons it’s recommended to split very large databases over multiple files and multiple drives. This tries to reduce the queue depth contention that would otherwise result.
For SQL Server typically I would test 8K IO (single page IO) and 64KB IO, and potentially large sequential IO of 512K or 1MB. This will allow you to cover your bases in terms of some of the common IO sizes that you will see in your database. It’s very common to see random 64KB IO sizes with SQL Server.
Here are some example commands with DiskSPD:
\Diskspd-v2.0.15\amd64fre\diskspd.exe -c5000G -d120 -r -w40 -t1 -o16 -b64k -h -L -Z1G f:\TestFile1.dat
\Diskspd-v2.0.15\amd64fre\diskspd.exe -c5000G -d120 -r -w50 -t1 -o16 -b64k -h -L -Z1G f:\TestFile1.dat
\Diskspd-v2.0.15\amd64fre\diskspd.exe -c5000G -d120 -r -w0 -t1 -o16 -b64k -h -L -Z1G f:\TestFile1.dat
The above generate a 5TB file (-c5000G), run for 120 seconds (-d120) use random IO (-r), and have different ratios of write IO (-w). We use a single thread per file (-t1) to generate the IO, and use 16 outstanding IO’s per thread (-o16). The IO size is set to 64kb in this case (-b64k). If you have a striped volume made up of a number of physical disk devices you can use more threads per file. But I’ve found that with a single file on a single drive using more than 1 thread will cause thread contention on the file and impact the validity of testing. Also be aware that the number of outstanding IO’s is per thread, so if you specify that to be too many again you will overload the queues in the OS and invalidate your test. We disable hardware and software caching with -h, and we measure latency statistics with -L.
When testing on large systems with KGroups > 1 and > 64 processors you will get a warning that DiskSPD can’t gather full CPU metrics. You will need to use another tool to gather CPU performance data, such as Perfmon.
When doing sequential tests with the -s option instead of -r you should not use multiple outstanding IO’s, as by definition that makes the test random and not sequential.
If you want to make your testing more realistic you might want to grab the details from using a tool such as procmon in Windows, or if in a VMware environment vscsiStats.
Final Word
As I wrote in my book “Virtualizing SQL Server with VMware: Doing IT Right (VMware Press 2014)“, it’s always a good idea to have a baseline, and to have valid application level tests of performance that take different factors into account, including CPU / RAM usage and NUMA topology impact etc. But to get an indication of storage performance, tools such as DiskSPD are useful, even though they do not give an accurate or full picture of performance that will impact applications. Once you have a baseline you can use it as a point of comparison against any changes or upgrades for your system, and when changing platforms. But to ensure validity of testing you should limit the number of variables between tests and data sets. In addition to using tools such as DiskSPD you should also consider using DB level benchmarks, such as HammerDB, DVD Store (Note: Latest DVD Store benchmark has moved to Github) and others that not only use the database engine, but simulate more of a real world application.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
I’ve been doing some work recently with SQL Server and Nutanix Acropolis Hypervisor on a older Nutanix NX3400 platform in my lab, as my other clusters are busy on other things. The testing involves cloning a number of SQL Server databases, running performance and functionality tests, and then destroying the VM’s and starting over again. The idea is to make one change each iteration and then be able to do comparisons between iterations (keeping a consistent control / baseline). Always being able to return to a consistent state for the Databases is important to have valid test data. To do this repeatedly and quickly using multiple VM’s would not be practical without automation. Similar sorts of scenarios could be used where multiple test environments need to be created rapidly and returned to a consistent state after each test run, such as with a large application project. With all of this being able to do as many tests as possible in a limited period of time is important. The more tests, the more results, the more conclusions, and in a project environment the more defects found earlier that get fixed and prevented from going into production and ending up as support tickets. So in this article I will provide the scripts I used to automate this process, and demonstrate just how rapid it is to clone 12 x 300GB+ databases for testing.
This NX3400 platform I’m using has 4 nodes, each node has 2 x SandyBridge 2.2GHz 8 Core Processors, 256GB RAM, a single 300GB SSD and 5 x 1TB HDD’s. It’s a couple of generations old now, and has long since been replaced by newer and more powerful platforms. But it’s good enough for the tests I’m doing. Here is a screenshot of the cluster:
As you can see the cluster is running Nutanix Acropolis Hypervisor (KVM), and at the time of the test High Availability was not enabled, hence it is in yellow. The cluster is running Nutanix OS 4.1.4.
The Nutanix Acropolis Hypervisor is based on KVM, but has been customized to meet the unique requirements of a distributed linearly scaleable hyper-converged infrastructure platform, and to be as simple as possible to use. One of the major benefits for everyone from Nutanix Acropolis Hypervisor is that you don’t need to be a rocket scientists or computer science PhD to use it, whereas with standard KVM, that’s how it feels. Standard KVM has way to many options, knobs and ways to get things wrong. In my opinion that is why KVM has never been as successful as other hypervisors. Nutanix Acropolis Hypervisor changes all that, itis so easy my 8yr old Son can use it and create his own VM’s. Nutanix Acropolis Hypervisor is available in both the commercial Nutanix platforms from Starter Edition Licenses, as well as the free Community Supported Nutanix Community Edition, which can run on just about any hardware, even nested virtualization. While it doesn’t yet support all of the features customers may have in other hypervisors, it is good for many use cases, and especially for applications where they have built in availability and are natively written to scale out and be highly available (including the likes of SQL Server, Exchange, Oracle and Oracle RAC).
Here is the script I used to rapidly provision 12 SQL Server Databases, which is executed from one of the Nutanix Controller VM’s:
sql_clone.sh:
#!/bin/bash
template="SQLW2K8R2-Template"
# To use batch provisioning uncomment startnum and endnum arrays and modify the values
#startnum=("11" "101" "201" "301")
#endnum=("100" "200" "300" "410")
startnum=1
endnum=12
vm_name="SQLW2K8R2AHV"
batches=${#startnum[@]}
startt=`date +%Y%m%d%H%M%S`
if [ -f vmclone.log ] ; then
rm vmclone.log
fi
for i in `seq 0 $((batches - 1))`
do
echo "Starting batch $((i + 1)) from ${startnum[$i]} to ${endnum[$i]}..."
DATETIME[0]=`date +%Y%m%d%H%M%S`
echo "Creating $vm_name from clone..."
echo "Cloning $template batch $((i + 1)) of $batches started at ${DATETIME[0]}..."
echo "Cloning $template batch $((i + 1)) of $batches started at ${DATETIME[0]}..." >> vmclone.log
completedtasks=`acli vm.clone $(seq -w ${startnum[$i]} ${endnum[$i]} | sed "s/^/$vm_name/" | xargs echo | tr " " ,) clone_from_vm=$template | grep complete | wc -l`
echo "Number of provisioning tasks completed = $completedtasks" >> vmclone.log
echo "Number of provisioning tasks completed = $completedtasks"
DATETIME[1]=`date +%Y%m%d%H%M%S`
echo "Cloning $template batch $((i + 1)) of $batches finished at ${DATETIME[1]}..."
echo "Cloning $template batch $((i + 1)) of $batches finished at ${DATETIME[1]}..." >> vmclone.log
echo "Cloning $template batch $((i + 1)) of $batches took $(( ${DATETIME[1]} - ${DATETIME[0]} )) seconds..." >> vmclone.log
echo "Cloning $template batch $((i + 1)) of $batches took $(( ${DATETIME[1]} - ${DATETIME[0]} )) seconds..."
echo "Powering on $vm_name batch $((i + 1)) of $batches..."
completedtasks=`acli vm.on $(seq -w ${startnum[$i]} ${endnum[$i]} | sed "s/^/$vm_name/" | xargs echo | tr " " ,) | grep complete | wc -l`
echo "Number of Power on tasks completed = $completedtasks" >> vmclone.log
echo "Number of Power on tasks completed = $completedtasks"
DATETIME[2]=`date +%Y%m%d%H%M%S`
echo "Powering on $vm_name batch $((i + 1)) of $batches finished at ${DATETIME[2]}..."
echo "Powering on $vm_name batch $((i + 1)) of $batches finished at ${DATETIME[2]}..." >> vmclone.log
echo "Powering on $vm_name batch $((i + 1)) of $batches took $(( ${DATETIME[2]} - ${DATETIME[1]} )) seconds..."
echo "Powering on $vm_name batch $((i + 1)) of $batches took $(( ${DATETIME[2]} - ${DATETIME[1]} )) seconds..." >> vmclone.log
done
endt=`date +%Y%m%d%H%M%S`
echo "Provisioning operations complete!"
echo "Total time to provisioning and power on all batches took $(( $endt - $startt )) seconds..."
echo "Total time to provisioning and power on all batches took $(( $endt - $startt )) seconds..." >> vmclone.log
To remove the test databases after the test is complete I used this script and executed it from one of the Nutanix Controller VM’s:
remove_sqlvms.sh:
#!/bin/bash vm_name=SQLW2K8R2AHV DATETIME=`date +%Y%m%d%H%M%S` echo "Removing $vm_name started at $DATETIME..." acli vm.off $vm_name\* DATETIME=`date +%Y%m%d%H%M%S` echo "Deleting $vm_name..." echo yes | acli vm.delete $vm_name\* DATETIME=`date +%Y%m%d%H%M%S` echo "Remove $vm_name finished at $DATETIME..."
Here is a video demonstrating what it looks like when I run the script. You can see the timing in the terminal window where I am connected into one of the Nutanix Controller VM’s running the sql_clone.sh script:
The exact same result can be achieved using the Nutanix REST API’s. In this example I have made use of the Acropolis CLI. The functions are supported equally across CLI, REST API, and within the Nutanix PRISM user interface (which leverages the Nutanix REST API’s). The entire process of provisioning the SQL Databases from template, powering them on, having them ready to run tests, from end to end, is completed in less than 90 seconds. As you can see from the video the provisioning and power on process is completed in less than 20 seconds.
It’s important to note that creating these 12 databases from template did not consume any additional storage. This is because the process leveraged the Nutanix smart clone and data avoidance features (no point duplicating data unnecessarily). Storage would only be consumed when new data is written. Storage consumption is further reduced by using Inline Compression (which also improves performance).
Each of the cloned databases is an exact replica of the master template. In these examples I didn’t customize the Windows OS for each clone, other than apply some tuning using a first boot script. Using PowerShell as part of the first boot script you could quite easily customize the OS and register with AD. For the tests I was performing I had set up the template to use a workgroup and not be part of an AD domain. Given I am rapidly creating and destroying these environments I didn’t want to have to register the VM’s and then delete the Computer Objects from AD afterwards. I wanted to treat these more as isolated test instances. This entire process can also be automated as part of a service management or cloud management platform using the Nutanix REST API.
If you want to run Microsoft apps for development and testing, or for production, no problem. The Nutanix Acropolis Hypervisor is supported by Microsoft under the SVVP program. Other application vendors have their own support statements and requirements, so please check with your application vendor or get in touch with anyone at Nutanix to check. Nutanix has many different customers running many different applications on Acropolis Hypervisor, including Oracle RAC, Oracle EBS, OBIEE, Weblogic, SQL Server, Custom Java Apps, Hadoop, Splunk etc. Nutanix has a single customer with a total of > 1500 nodes running Acropolis Hypervisor.
https://twitter.com/andreleibovici/status/614512636936503298
Final Word
By using similar methods you could create rapid test and development environments for multiple projects to use in parallel. Complete end to end application stacks can be created in seconds to minutes, all without consuming any additional space. You can destroy environments and get back to a consistent state every single time. This allows projects to rapidly iterate and provide higher quality business outcomes in much less time. Every developer or tester could have their own environment of a completely integrated end to end application stack, that they can create, test, and then destroy and recreate on demand. On large projects this could literally save months of effort and millions of dollars. All of the infrastructure management is completely integrated into the solution and provides a distributed management framework without any single points of failure. The Nutanix PRISM user interface is the Nutanix Acropolis Hypervisor management interface, and REST API endpoint. The Nutanix solution is so easy to deploy, manage, and use, that you wouldn’t even know it’s based on KVM.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.