(cas:72) Google Analyticator was unable to authenticate you with Google using the Auth Token you pasted into the input box on the previous step.

This could mean either you pasted the token wrong, or the time/date on your server is wrong, or an SSL issue preventing Google from Authenticating.

Try Deauthorizing & Resetting Google Analyticator.

Tech Info 400:Error fetching OAuth2 access token, message: 'invalid_grant'
Unique
Visitors
Powered By Google Analytics
vSphere – Long White Virtual Cloudsu by http://longwhiteclouds.com all things Nutanix, VMware, cloud and virtualizing business critical applications Fri, 12 Feb 2016 18:37:34 +0000 en-US hourly 1 https://wordpress.org/?v=6.7.6 45024036 Heads Up! Do Not Upgrade VMware Tools on Hosts with ESXi 6.0 U1b http://longwhiteclouds.com/2016/01/10/heads-up-do-not-upgrade-vmware-tools-on-hosts-with-esxi-6-0-u1b/ http://longwhiteclouds.com/2016/01/10/heads-up-do-not-upgrade-vmware-tools-on-hosts-with-esxi-6-0-u1b/#comments Sat, 09 Jan 2016 14:00:42 +0000 http://longwhiteclouds.com/?p=11517


Heads Up! If you’ve updated to ESXi 6.0 U1b, build 3380124 and you have lots of templates, you may run into some problems if you update VMware Tools to the latest version. I just upgraded my environments to the latest VMware patches ESXi 6.0 U1b (build 3380124), that has just come out. As you do […]

]]>


Heads Up! If you’ve updated to ESXi 6.0 U1b, build 3380124 and you have lots of templates, you may run into some problems if you update VMware Tools to the latest version. I just upgraded my environments to the latest VMware patches ESXi 6.0 U1b (build 3380124), that has just come out. As you do usually when there is a new hypervisor build you upgrade VMware Tools. Well that proved to be a big problem for my VM templates that I use to provision new systems. But I’ve got a workaround.

As soon as VMware Tools is updated on any templates you will no longer be able to clone those templates. If you’ve updated any templates with the version of VMware Tools that comes with ESXi 6.0 U1b then you need to uninstall it and reinstall the prior version that came with ESXi 6.0 build 3247720. After the couple of reboots that you have to go through with an uninstall and reinstall of VMware Tools you will find that you can now clone VM’s and have them automatically customized. I ran into this problem on Windows 2008 R2 Server. So I know it will impact this guest OS. I haven’t tested other OS’s yet, but others could be impacted. I’ve logged a support call with VMware to address this problem. In the meantime, the workaround is fine. The Official VMware KB Article 2142982 explains the situation.

 

[Updated 14/01/2016] After further testing I have narrowed down the problem area to new installs where the complete option is selected, and any upgrades where the complete options was previously selected, or where the VMCI / NSX Guest Introspection Driver is included. I have been able to successfully clone from a new VM Image that has had a fresh install of Windows 2008 R2 and VMware Tools without the VMCI / NSX Guest Introspection Driver, or where VMware Tools was installed twice / installed and repaired on the same VM, when the complete option was previously selected. This seems to be similar to what other of you have also reported. I have completed the Upgrade Scenario testing as well and confirmed that after an upgrade, if the complete install option was previously selected the VM will not clone due to the same VMCI driver problem. If VMCI driver is removed by running VMware Tools Install again and selecting Modify and unselecting VMCI, then you will be able to close the VM.

This update just in from VMware Support “VMware Engineering have confirmed that the issue is dependent on the install/upgrade sequence. Specifically, the issue is aligned to the version of deploypkg.dll in the vmtools package. GSS and Engineering are mapping the ESX and vmtools update versions to the deploypkg.dll  versions to confirm which upgrade sequences are problematic. A KB article will be published once this information is finalised.

Thanks to VMware Support for getting to this stage very quickly. The VMware KB Article 2142982 has now been published. 

Final Word

I guess someone has to take the risk and patch their systems to the latest versions first, especially as these were security patches with a critical severity. Fortunately like all good IT environments I only did my test systems first. This is the whole point of having infrastructure test systems. You can test infrastructure hardware and infrastructure software changes first before putting them into production. The old saying goes that software eventually works and hardware eventually fails, but these days a lot of your hardware is also software, especially in a virtualized software defined datacenter. It pays to have appropriate test systems and test plans to mitigate the risks associated with software updates and changes of all types, including infrastructure software. Thanks to all of you in the community that contributed to this effort and commented on this blog post. I have been relaying your feedback during my discussions with VMware Support.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2016/01/10/heads-up-do-not-upgrade-vmware-tools-on-hosts-with-esxi-6-0-u1b/feed/ 37 11517
Nutanix: First Hyperconverged Vendor with SAP Certified Platform http://longwhiteclouds.com/2015/10/28/nutanix-first-hyperconverged-vendor-with-sap-certified-platform/ http://longwhiteclouds.com/2015/10/28/nutanix-first-hyperconverged-vendor-with-sap-certified-platform/#comments Tue, 27 Oct 2015 17:37:02 +0000 http://longwhiteclouds.com/?p=11061


Nutanix is the only hyperconverged company in the Leaders of the Gartner Magic Quadrant for Integrated Systems, and recently became the first and only hyperconverged platform to be SAP certified. The official list of hardware platforms certified by SAP are listed here. Customers can run their critical SAP systems, such as ERP, BW, etc. on Nutanix Web-scale […]

]]>


Nutanix is the only hyperconverged company in the Leaders of the Gartner Magic Quadrant for Integrated Systems, and recently became the first and only hyperconverged platform to be SAP certified. The official list of hardware platforms certified by SAP are listed here. Customers can run their critical SAP systems, such as ERP, BW, etc. on Nutanix Web-scale infrastructure to significantly simplify infrastructure, massively reduce deployment times and greatly improve project life-cycle and operations. Customers can have the same confidence with Nutanix as when deploying on any traditional platform, with full end to end SAP and Nutanix Support.  SAP Certification is a big deal because SAP is vary particular about which platforms are allowed to be certified to run their software stack. The certification process is very thorough and very stringent criteria that have to be met. Approximately 74% of the global GDP touches a SAP system, and with over 250K customers SAP is the clear leader in enterprise software, which is an approx $US24B market per annum. SAP is also a significant investor in Nutanix. So why would you want to run SAP virtualized and on Nutanix?

The latest data from VMware shows that virtualizing critical applications is common place. With high rates of virtualization across the globe and for the most critical business systems. The following slide from VMworld 2015 shows the latest results from the VMware survey, specifically focused on critical apps.

vBCA Virtualization Adoption

Note: SAP itself is over 85% virtualized with VMware vSphere and runs many customer facing critical systems virtualized with VMware vSphere. The SAP Case Study Video is a great watch for anyone who is interested in virtualizing SAP systems, or any other applications. There are additional SAP case studies available here. For additional resources from VMware with regards to SAP you can visit this page.

The reason that so many customers are virtualizing their critical SAP systems is quite simple. Improved business service levels, which can’t be achieved as easily on a traditional physical infrastructure, and greater ROI and reduced TCO. When virtualizing any business critical applications it’s important that there are no compromises made to SLA’s and the environment is designed and built to meet the business critical apps requirements. Availability, Performance, Disaster Recover, Management and Project Lifecycle should all improve. Some of the main benefits of Virtualizing SAP are as follows:

  • High Availability, automated fault tolerance, built in by default – over an above what is included with the application stack. This greatly reduces downtime in the case of hardware failures.
  • Automated run-book driven disaster recovery – delivers predictable and consistent business continuity and recovery from disaster every time, with the ability to test DR plans without disruption to production. Audit reports after tests allow the capability to be proven as part of meeting regulatory compliance requirements.
  • Rapid, Automated Provisioning – On demand deployment of completely application environments in minutes from templates, ensure consistency and scalability across SAP environments. This capability can greatly improve quality of releases to production in less time, while at the same time allowing dynamic increase in capacity for production environments to handle cyclical peak load demands.
  • Zero downtime live migration for proactive hardware maintenance and non-disruptive hardware migration without a requirement to modify or upgrade the applications at all.
  • Simplified driver stack for the operating system reduces complexity and risk for the applications and removes the need to manage storage multi-pathing and network teaming compared to a traditional physical infrastructure.
  • Greatly improved return on infrastructure investment – Optimizing infrastructure investments means more CAPEX/OPEX to invest in applications, end user experience and business / customer value added activities.

 

Nutanix Plus SAP

This brings us to why Nutanix and SAP? We’ve already covered that Nutanix is the first and only hyperconverged leader in the Gartner Magic Quadrant for Integrated Systems and is also the most visionary vendor of them all, SAP is a substantial investor in Nutanix, and that Nutanix is the only hyperconverged vendor delivering a SAP certified platform. But why should you choose Nutanix for SAP environments, especially production?

  • Nutanix is a SAP Global COIL Member: Bringing customers joint co-innovations and demonstrations from of joint SAP and Nutanix technology from SAP Global Co-Innovation Labs (COIL) facilities. The picture below is of the Nutanix systems at SAP COIL Singapore and we are going to expand in COIL globally soon. There are automation scripts being co-developed right now between SAP COIL and Nutanix for operational tasks.
    Nutanix COIL Lab Setup
  • Invisible Infrastructure, Reduced Complexity, Reduced Risk: Deploy in minutes and expand on demand, upgrade non-disruptively / transparently at lunchtime, self-healing – no more nights and weekend callouts, focus on more business value adding activities, and spend more time with your friends and family. Would you like some Spaghetti Carbonara with your upgrade?
    NutanixOneClickUpgrade
  • Consistent, Predictable Performance, Reduced Risk: Because of the way the Nutanix architecture is designed it provides much more predictable and consistent high performance initially (at deployment), but also as the environment grows. No longer does one system monopolies resources and performance to the detriment of other systems. Every workload gets its fair share of resources, performance and resiliency improves as the environment grows. Nutanix also provide much more predictability when components fail, and eliminate some of the negative performance impacts of component failure as every node, controller and disk participates to balance recovery operations. This is due to every node / controller and disk participating in recovery operations, so recover is balanced, consistent, and no single point for a bottleneck exists. Data Locality ensures that as the environment grows all workloads have access to the shortest IO path to their data, without adding unnecessary network congestion. The graphs below show consistent, predictable performance, both in terms of database transactions per second, but also database transaction latency, as the environment scales. SQL server is used in this case to demonstrate the capability and the test was performed on a Dell XC630 platform powered by the Nutanix software.
    3NodeXCSQL-ScaleOut6NodeXCSQL-ScaleOut
  • Reduced Project Deliver and Production Defect Risk: One of the greatest benefits of virtualizing is greatly improving the efficiency of application release cycles and thereby revolutionizing the ability of the IT team to deliver business outcomes (while greatly reducing project timelines and project labour costs). With Nutanix, operations that used to take weeks to a month in a traditional environment (even a virtualized one) can now be done in minutes (without incurring heavy storage consumption due to Nutanix data efficiency techniques). Multiple independent SAP Landscapes can be deployed in such a way that they consume almost no more storage than a single environment. This makes Nutanix All Flash Platforms economics and performance attractive. This can also provide much more realistic SAP Landscapes, which can allow early detection of defects and avoid defects being found later in production. It provides the ability to always get back to a known good state quickly. Again using SQL Server as a demonstration, the below video shows how quickly multiple test environments can be deployed, applications customized, and be ready for testing to a known good state, all in just a few minutes. The automation can be extended to SAP and through tools such as SAP Landscape Virtualization Manager.

 

  • Pay as you Grow, Never Do A Forklift Upgrade Again:Investments can be aligned with projects and operation needs and avoid excess capacity, while providing just in time, infrastructure on demand with predictable scalability and cost (cloud on your terms, under your control). Start small, buy just what you need today and very simply and consistently expand as demands dictate for applications, database and infrastructure. No longer do you need a crystal ball to anticipate what will happen in 3 to 5 years. This also means you can take immediate advantage of technological advances in hardware, without having to do a forklift upgrade. Speaking of which, there is never again a forklift upgrade and long project planning cycle for system replacement in a Nutanix environment. You simply non-disruptively deploy new nodes, and non-disruptively removed old nodes when the time comes. Data movement is completely transparent and automatic.NutanixLinearScalability
  • Improved Enterprise Data Protection, Availability and Business Continuity, Secure by Default: Keep SAP application servers and databases protected and running with frequent, VM-centric, easy-to-restore backups and affordable, simple disaster recovery with Nutanix proprietary out of the box replication/metro-availability solutions.
    NutanixCapabilities
  • Single Point of Global Management, Powerful API’s and Analytics: Global management of all Nutanix platforms from a single screen, accessible from anywhere via a standard HTML5 browser, including powerful API’s and built in analytics to pinpoint any anomalies. Both REST and PowerShell API’s are available for complete infrastructure and application automation.
    NutanixPRISMCentralNutanixPRISMNutanixAnalytics
  • Robust Platform for Any Applications: Runs the most robust and highly available database platforms with ease. Including Oracle RAC. The following demonstration shows an extreme load test on an Oracle RAC platform with a mid range Nutanix system while all nodes of the Oracle RAC cluster are simultaneously live migrated using VMware vMotion.

 

Final Word

Robust enterprise infrastructure for critical applications doesn’t have to be complex. You no longer have to suffer through months of upgrade and deployment planning, or worry about doing project changes in the middle of the night. Nutanix offers a robust platform that reduces risk for critical applications, provides high performance with predictable and consistent scalability, all without hard limits. A platform that self heals and is non-disruptively upgraded and expanded. Your applications deserve the benefits that invisible web-scale infrastructure from Nutanix can deliver. We are extending the same level of simplicity, resilience and worry free life to our upcoming SAP HANA appliances as well. Stay tuned for details in the coming weeks and months. You can visit the SAP page on the Nutanix site for further SAP and Nutanix specific details.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2015/10/28/nutanix-first-hyperconverged-vendor-with-sap-certified-platform/feed/ 3 11061
Heads up! Backing up a virtual machine with Changed Block Tracking (CBT) enabled fails in ESXi 6.0 http://longwhiteclouds.com/2015/10/11/heads-up-backing-up-a-virtual-machine-with-changed-block-tracking-cbt-enabled-fails-in-esxi-6-0/ http://longwhiteclouds.com/2015/10/11/heads-up-backing-up-a-virtual-machine-with-changed-block-tracking-cbt-enabled-fails-in-esxi-6-0/#respond Sun, 11 Oct 2015 07:41:28 +0000 http://longwhiteclouds.com/?p=11053


This problem is now over 4 months old but seems to have started being reported more. Over the past week I’ve heard a number of reports of databases being corrupted after backups where the VMware API’s for Data Protection (VADP) were used with change block tracking on ESXi 6.0 if you are not on the […]

]]>


This problem is now over 4 months old but seems to have started being reported more. Over the past week I’ve heard a number of reports of databases being corrupted after backups where the VMware API’s for Data Protection (VADP) were used with change block tracking on ESXi 6.0 if you are not on the patched or latest build. I think one of the reasons this hasn’t been noticed much sooner is that customers are now starting to upgrade to ESXi 6.0 in greater numbers. This may also be because some customers have an N-1 policy, so install 6.0 GA when U1 is out for example. Upgrading to ESXi 6.0 has many benefits for almost all environments, and it pays to upgrade to the latest patched version (to try and avoid these types of problems) and thoroughly test before putting into production. This problem demonstrates a need to check KB articles for known issues, in addition to release notes, when performing an upgrade. VMware published a KB article on this problem some time ago and the patch to address this problem has been available since 25th June 2015.  If you are in the process of upgrading to ESXi 6.0 or have done already, it is my recommendation to be running on the latest build if possible. I hope the majority of you have no encountered this problem.


]]>
http://longwhiteclouds.com/2015/10/11/heads-up-backing-up-a-virtual-machine-with-changed-block-tracking-cbt-enabled-fails-in-esxi-6-0/feed/ 0 11053
The Need To Reboot ESXi Hosts After vCenter Upgrade http://longwhiteclouds.com/2015/01/21/the-need-to-reboot-esxi-hosts-after-vcenter-upgrade/ http://longwhiteclouds.com/2015/01/21/the-need-to-reboot-esxi-hosts-after-vcenter-upgrade/#comments Wed, 21 Jan 2015 01:42:57 +0000 http://longwhiteclouds.com/?p=10559


Most of you will be intimately familiar with vCenter Server. It provides the main management capabilities in a VMware Virtualized Datacenter, including provisioning, monitoring, patching etc for all your hosts and virtual machines. You’ll spend a lot of your time using it if you’re a VMware Admin. But you may not know that it is advisable to […]

]]>


Most of you will be intimately familiar with vCenter Server. It provides the main management capabilities in a VMware Virtualized Datacenter, including provisioning, monitoring, patching etc for all your hosts and virtual machines. You’ll spend a lot of your time using it if you’re a VMware Admin. But you may not know that it is advisable to reboot your hosts after you upgrade the version of vCenter Server. Otherwise you may get a little surprise the next time you go to patch or upgrade your hosts. I ran into this myself just recently after upgrading to the latest version of vCenter and then immediately attempting to upgrade my ESXi hosts in my lab. Fortunately this isn’t really a problem and is very easy to fix. This article will explain why.

During a vCenter Upgrade the VMware HA Agent on all of the ESXi hosts will be updated. After this happens the “require reboot” flag is set on the hosts, as explained in VMware KB 2034945. There is no visual cue that this is the case, so it is very easy to miss. The most likely time you’ll pick this up is if you are using Update Manager to scan your hosts for updates, or you’re attempting to upgrade ESXi. If you upgrade ESXi manually and haven’t already rebooted, then it’s very likely that the HA agent will fail to initialize the first time when the host restarts. This is very easily resolved by right clicking on the host and selecting the “Reconfigure for vSphere HA” option. This takes about a minute and you’re hosts HA Agent will be reconfigured and working again. Until the HA Agent is fully working and initialized successfully you will not be able to run any VM’s on the host.

If you want to run a PowerCLI Command to see if your hosts need a reboot or not prior to an upgrade check out this article titled How to find VMware ESX(i) servers that need a reboot using PowerCLI.

Final Word

This is really only a minor inconvenience and is easily fixed by reconfiguring HA. But it is better to catch this and know about it first, rather than finding out only after HA fails to initialize, especially in a highly automated environment where the problem might not be caught for a couple of days and you could have problems starting VM’s.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2015/01/21/the-need-to-reboot-esxi-hosts-after-vcenter-upgrade/feed/ 4 10559
Performance Testing MySQL and PostgreSQL with HammerDB and PGBench http://longwhiteclouds.com/2014/11/27/performance-testing-mysql-and-postgresql-with-hammerdb-and-pgbench/ http://longwhiteclouds.com/2014/11/27/performance-testing-mysql-and-postgresql-with-hammerdb-and-pgbench/#comments Thu, 27 Nov 2014 00:45:02 +0000 http://longwhiteclouds.com/?p=9937


 There are a lot of environments that are running MySQL and PostgreSQL to support their systems. My team at Nutanix and I have been getting a lot of enquiries about how to set up these databases for best performance, and customers have also been using them to benchmark and baseline different systems. One of the […]

]]>


postgresql-hdr_left HammerDB LogoThere are a lot of environments that are running MySQL and PostgreSQL to support their systems. My team at Nutanix and I have been getting a lot of enquiries about how to set up these databases for best performance, and customers have also been using them to benchmark and baseline different systems. One of the challenges with these databases is that they give only limited control over where data files and transaction logs can be placed, and this makes increasing parallelism of IO a little bit of a challenge. Your database is just an extension of your storage and all storage devices, even virtual ones, have a limited queue depth that you can work with. Unlike Oracle, SQL Server, Sybase, DB2, etc you can’t just create a whole bunch of mount points and spread your data files over them (which increases available queue depth and potential IO parallelism). But the solution to this problem is made quite simple with Linux LVM (Logical Volume Manager). I’ll take you through some of the steps I took to set up a test VM for MySQL testing with HammerDB and PostgreSQL for testing with PGBench.

Firstly what is HammerDB and PGBench?

HammerDB is an open source database testing tool that simulates a test that is styled along the lines of TPC-C. But it doesn’t implement a full TPC-C test and can’t be compared to published audited results. There is documentation on the HammerDB Test here, and the binaries can be downloaded from SourceForge here. Although these tests don’t compare to published audited results they can be used as a comparison between two systems that have run the same test. The output is two values, a TPM (Transactions Per Minute), which is database specific and can’t be compared between two different database types, and a NOPM value (New Orders Per Minute), which can be compared between different HammerDB tests and database types. NOPM is the more relevant metric to keep an eye on.

PGBench is a built in performance test benchmark that comes with PostgreSQL and it simulates a TPC-B type test, which is a database stress test. TPC-B as a test has been obsoleted by the TPC org, and the TPC-B test within PostgreSQL couldn’t be used to compare against any audited results, but it can be used to compare against like tests using the same tool. This should be used only to compare between different PostgreSQL database performance results.

Setting up a Linux LV for use with PostgreSQL or MySQL (Example based on PostgreSQL):

For the tests I was doing I used a striped logical volume (8k stripe size) across 8 disks of 10G each and copied the data directory of the postgresql to the new location. All eight disks were equally distributed across 4 vSCSI controllers (PVSCSI recommended with VMware vSphere). The total volume size was about 80GB.

1. vgcreate vol0 /dev/sdb /dev/sdc /dev/sdd /dev/sde /dev/sdf /dev/sdg /dev/sdh /dev/sdi

2. lvcreate -i8 -I8 -l 100%FREE -n lvol0 vol0

3. mkfs.ext4 -m 0 /dev/vol0/lvol0

4. mkdir /home/postgres

5. Add “/dev/vol0/lvol0         /home/postgres          ext4    defaults,noatime,nodiratime        0 0” in /etc/fstab

6. mount -a 

7. By default, postgresql installs the data directory at /var/lib/pgsql/9.3/data. I copied the data directory to 

/home/postgres/data

8. Now, when we start the postgres server, we should use the -D option to point the new location of the data directory. 

/usr/pgsql-9.3/bin/postgres -D /home/postgres/data

This should now use the database at /home/postgres and write it across all 8 disks. vgdisplay and lvdisplay should look like this.

[root@RT17-CentOS-1 9.3]# vgdisplay

  — Volume group —

VG Name               vol0
System ID
Format                lvm2
Metadata Areas        8
Metadata Sequence No  4
VG Access             read/write
VG Status             resizable
MAX LV                0
Cur LV                1
Open LV               1
Max PV                0
Cur PV                8
Act PV                8
VG Size               79.97 GiB
PE Size               4.00 MiB
Total PE              20472
Alloc PE / Size       20224 / 79.00 GiB
Free  PE / Size       248 / 992.00 MiB
VG UUID               hmmbRZ-xoHF-69K6-7DkM-EIqo-jqoY-3X8qz8

[root@RT17-CentOS-1 9.3]# lvdisplay

  — Logical volume —

LV Path                /dev/vol0/lvol0
LV Name                lvol0
VG Name                vol0
LV UUID                ZHija9-OXek-59Ud-Q4uJ-gThy-55Zf-CJRIaD
LV Write Access        read/write
LV Creation host, time CentOS-1.local, 2014-10-14 10:34:17 -0400
LV Status              available
# open                 1
LV Size                79.00 GiB
Current LE             20224
Segments               1
Allocation             inherit
Read ahead sectors     auto
– currently set to     256
Block device           253:2

One additional step you could take it to create another logical volume and move the pg_xlog directory. This can also increase performance for PostgreSQL. For MySQL the steps are similar, except you move the MySQL Home Directory to the LVM, which you might mount at /home/mysql (check /etc/my.cnf for location details of the default install).

An example of a performance results based on PostgreSQL 9.3 from a single VM of 4 vCPU’s, 32GB RAM, running on a Nutanix 3050 node is below. The templates used to do the MySQL HammerDB and PostgreSQL PGBench testing are available from Nutanix via your SE or Nutanix Support.

PostgreSQL Test Results with PGBench:

Running tests using: psql -d pgbench
Script tpc-b.sql executing for 120 concurrent users
transaction type: Custom query
scaling factor: 1500
query mode: simple
number of clients: 120
number of threads: 6
duration: 300 s
number of transactions actually processed: 1021500
tps = 3402.171378 (including connections establishing)
tps = 3403.517277 (excluding connections establishing)

An example of the TPM for a single MySQL Database VM with 4 vCPU’s, 32GB RAM, running on a Nutanix 3050 node is below. NOPM is around 31K. Note that TPM from one database type can’t be compared to another database type. So the MySQL TPM below can’t be compared to the TPM of an Oracle database. But the NOPM can be compared. If you’re running a test with MySQL with a similar configuration though then you can compare the TPM.

MySQL HammerDB 100 Warehouses 4 vCPU Tuned 2014-11-19_22-38-18

Final Word

Before migrating any system to a new platform you should record or baseline the exiting performance and then validate the performance after the migration. You should make sure to read and follow any relevant vendor best practices to get the best out of your platforms. Any improvement in performance means less hardware overall needed to achieve acceptable performance, or the longer you can run a system before an upgrade is required. Although the HammerDB test doesn’t fully stress out the storage, the performance of your storage will make a difference to your test results. As will the set up of your OS and the CPU/Memory configuration of your VM’s and platform overall. It’s a good test of the overall platform and all it’s sub-components.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/11/27/performance-testing-mysql-and-postgresql-with-hammerdb-and-pgbench/feed/ 1 9937
Installing Cumulus Linux from a MacBook Pro http://longwhiteclouds.com/2014/11/13/installing-cumulus-linux-from-a-macbook-pro/ http://longwhiteclouds.com/2014/11/13/installing-cumulus-linux-from-a-macbook-pro/#comments Thu, 13 Nov 2014 11:35:22 +0000 http://longwhiteclouds.com/?p=9283


Cumulus Linux is the answer to companies that want to run software defined networking on a range of open networks industry standard switches, without necessarily being locked into one physical switch hardware vendor. But unlike network virtualization solutions such as NSX, Cumulus Linux is the Network OS (NOS) for the physical switches, rather than a […]

]]>


CumulusTurtleLogoCumulus Linux is the answer to companies that want to run software defined networking on a range of open networks industry standard switches, without necessarily being locked into one physical switch hardware vendor. But unlike network virtualization solutions such as NSX, Cumulus Linux is the Network OS (NOS) for the physical switches, rather than a virtualization layer on top. Cumulus is part of the NSX ecosystem and integrated into NSX, so essentially you can use Cumulus to run on the physical switches and integrate it to NSX to provide the network virtualization (termination and VXLAN switching/routing in hardware also supported on some switches). Cumulus is Linux for network switches, so it’s easy to manage, and very easy to automate. I happen to be working on a project now to build the best practices for Cumulus Linux with Nutanix and VMware vSphere. So I needed an easy way to get Cumulus installed on my lab switches, from my MacBook Pro, which is what the remainder of this article is about.

You can choose open network (ON) switches from a variety of vendors that are on the Cumulus HCL. In my case I chose Dell Force10 S4810-ON‘s. The -ON is an important part, as that is the Open Network variety. The Dell Force10 Switches are enterprise class low latency switches. The -ON switches come with the Open Network Install Environment included, so that you can install Cumulus Linux. This is the default boot environment for the switches and starts automatically.

There are 6 ways you can install Cumulus Linux NOS on your Open Network switches as follows:

  1. Passed from the boot loader.
  2. Search locally attached storage devices for one of the ONIE default installer filenames (for example,
    USB).
  3. Exact the URLs from DHCPv4.
  4. Inexact URLs based on DHCPv4 responses.
  5. Query to IPv6 link-local neighbors using HTTP for an installer.
  6. TFTP waterfall — from DHCPv4 option 66

In my case I had DHCP configured for my management network so I just needed an easy way to start up a web server so the switches could discover the Cumulus Linux firmware and download and install it. As a tip, if you don’t have a DHCP server already you can configure your MacBook Pro for Internet sharing, which then starts a DHCP server. Setting up the web server was a lot easier than I thought it would be. It turns out there is a very easy way to start up a temporary HTTP server on a MacBook Pro from any directory through the terminal. I stumbled across an article titled Start a Simple Web Server from Any Directory on Your Mac. All I had to do was change to the directory containing the Cumulus Linux package, which I had renamed to work with the ONIE process. But there was one key element missing from the article that I needed in order to get it to work.

If you attempt to run the web server exactly as mentioned in the Life Hacker article, such as python -m SimpleHTTPServer 80, you will get the following:

michael2012mbp:Cumulus michaelwebster$ python -m SimpleHTTPServer 80

Traceback (most recent call last):

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py”, line 162, in _run_module_as_main “__main__”, fname, loader, pkg_name)

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/runpy.py”, line 72, in _run_code exec code in run_globals

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/SimpleHTTPServer.py”, line 224, in <module> test()

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/SimpleHTTPServer.py”, line 220, in test BaseHTTPServer.test(HandlerClass, ServerClass)

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/BaseHTTPServer.py”, line 595, in test httpd = ServerClass(server_address, HandlerClass)

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/SocketServer.py”, line 419, in __init__ self.server_bind()

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/BaseHTTPServer.py”, line 108, in server_bind SocketServer.TCPServer.server_bind(self)

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/SocketServer.py”, line 430, in server_bind self.socket.bind(self.server_address)

  File “/System/Library/Frameworks/Python.framework/Versions/2.7/lib/python2.7/socket.py”, line 224, in meth return getattr(self._sock,name)(*args)

socket.error: [Errno 13] Permission denied

This was easily fixed by using the sudo command, entering the admin password, and running the operation as root. Which resulted in the following:

michael2012mbp:Cumulus michaelwebster$ sudo python -m SimpleHTTPServer 80

WARNING: Improper use of the sudo command could lead to data loss or the deletion of important system files. Please double-check your typing when using sudo. Type “man sudo” for more information.

To proceed, enter your password, or type Ctrl-C to abort.

Password:

Serving HTTP on 0.0.0.0 port 80 …

Then it was just a matter of kicking off the discovery process on my S4810 switches, which I did by rebooting them via the console cable (which happens to be connected via USB to a VDI desktop running on one of my Nutanix hosts, and it even still works with vMotion).

Once the switches restarted they found my MacBook on the network and found the web server and began to search for the firmware. ONIE goes through a standard process to identify and download the correct firmware by using the following naming conventions from the most specific to the lest specific:

  1. onie-installer-<arch>-<vendor>_<machine>-r<machine_revision>
  2. onie-installer-<arch>-<vendor>_<machine>
  3. onie-installer-<vendor>_<machine>
  4. onie-installer-<arch>
  5. onie-installer

In my case it looked like this from the web server on my MacBook Pro:

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] “GET /onie-installer-powerpc-dni_7448-r0 HTTP/1.1” 404 –

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] code 404, message File not found

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] “GET /onie-installer-powerpc-dni_7448 HTTP/1.1” 404 –

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] code 404, message File not found

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] “GET /onie-installer-dni_7448 HTTP/1.1” 404 –

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] code 404, message File not found

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] “GET /onie-installer-powerpc HTTP/1.1” 404 –

xxx.xxx.xxx.11 – – [13/Nov/2014 19:31:34] “GET /onie-installer HTTP/1.1” 200 –

As you can guess from the above I had named the Cumulus Linux package onie-installer. I could have named it onie-installer-powerpc or any one of the other specific naming conventions. If this was a large scale environment it would be best to set up the distribution point using specific names for each model of system in the environment. Although the Dell S4810’s are PowerPC based some switches are x86_64 based, such as the Dell S6000-ON 40GbE switches.

Final Word

The process I used to get Cumulus Linux installed on my lab switches is probably fine for small scale environments and PoC’s. For large scale you will want something more robust and automated. One of the great things about Cumulus being Linux for switches is that it can fit into you existing Linux management and automation frameworks, such as Puppet, Chef and CFEngine. You can completely automate the configuration and management of a large scale switching environment, which when combined with network virtualization by NSX can become incredibly agile and flexible for any type of application workload. I’m working with Cumulus and I will be documenting the official best practices for Cumulus with Nutanix and it will be published on the Nutanix web site once we’re done. Along the way I will bring you any interesting things that I find.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/11/13/installing-cumulus-linux-from-a-macbook-pro/feed/ 1 9283
VMworld 2014 Wrap Up http://longwhiteclouds.com/2014/10/19/vmworld-2014-wrap-up/ http://longwhiteclouds.com/2014/10/19/vmworld-2014-wrap-up/#respond Sat, 18 Oct 2014 20:43:07 +0000 http://longwhiteclouds.com/?p=4963


Another VMworld event is over and it’s hard to believe it’s been a whole 12 months since the last one. Certainly during the keynotes there was a lot of coverage about what VMware has achieved over the last 12 months and it is impressive especially in the end user computing and hybrid cloud spaces. But […]

]]>


Another VMworld event is over and it’s hard to believe it’s been a whole 12 months since the last one. Certainly during the keynotes there was a lot of coverage about what VMware has achieved over the last 12 months and it is impressive especially in the end user computing and hybrid cloud spaces. But overall I felt that VMworld USA 2014 lacked some of the sparkle of last year. But I guess it’s hard to top last year considering it was the 10th anniversary. This year seemed much more about building a solid foundation for a software defined datacenter, a software defined enterprise and a hybrid cloud model integrating applications with infrastructure, providing ability and flexibility, but without compromise. Although attendance was flat or a little down on last year the breakout sessions were packed, right up to the last session on Thursday. Instead of having our heads in the clouds this year it was all about the vCloud Air, and we vRealized the product naming is about to be changing. So lets dive into what I think are some of the highlights.

My VMworld started on Sunday with a Nutanix sponsored VCDX study group. Nutanix is a big supporter of the VCDX program for the entire community. The study group was put on for candidates that wanted to know more about the VCDX process and practice the design and troubleshooting scenarios. It was completely vendor agnostic, and it needs to stay that way. Nutanix understands that the only way sponsoring a VCDX study group can be of value is if the content is vendor agnostic and covers a wide range of topics. There were many VCDX helping in the room and giving advice from across many companies. This really is what the community is all about. Everyone helping each other.

Then I moved on to opening acts at VMunderground that was put on by vBrownBag. I was on the storage panel and it was a good discussion around Virtual Volumes, Hyper Convergence and Flash. I even agreed with a traditional SAN vendor that hyper converged appliances will not help SuperDomes and Mainframes, but then again I can always migrate the workloads and processes off those systems, and the Unix mid range systems as well, to a Hyper Converged world. SuperDomes, Mainframes and Unix systems is where the legacy SAN technology will stay for the foreseeable future and it will be a decline over a number of years, just like we’ve seen with the traditional big iron systems themselves. The move away from traditional SAN for x86 connected environments isn’t going to happen over night, but it’s a trend that is starting to take hold, but honestly it’s not even scratching the surface of the potential opportunity yet. The announcements from VMware and EMC around their hyper-converged offerings are just more validation of that. Flash is definitely the way of the future, and it opens up things that were previously not possible. I have a section on flash technology in the storage chapter of Virtualizing SQL Server with VMware.

There were a number of VMware announcements during the keynotes that are worth mentioning. But before I do I have to get something off my chest. vRealise is the worst name ever thought of for anything. My initial reaction to the new name for VMware’s hybrid cloud, vCloud Air was somewhat similar, but at least Air has a cool ring to it, like iPad Air for example. vRealize, just NO! I feel sorry for the sales team who have to try and sell that now. Ok, rant over. The overall themes about this being a brave new world and requiring bravery from all of the customers and the community participants was interesting. I’ve been doing virtualization for a very long time and even for business critical apps it’s a very safe bet. But SDDC and the Software Defined Enterprise going to further reduce silos and this will require some organisational changes and maturity. This is really where the bravery comes in, most of the challenges are not technical.

Two major overall themes were used during the keynotes. Firstly – “Compatibility, Compliance, Choice”, and secondly the “Power of &”. VMware has done a great job of building a partner ecosystem across a number of technologies, including the vCloud Air Network, which has 3900 partners, and the broader ecosystem around the hypervisor and NSX. This is where compatibility, compliance and choice really comes in. Seamless compatibility, compliance with regulatory and industry requirements, and choice of multiple partners and technologies. This is then extended to the OpenStack, NSX and Containers, which can run extremely well in a VMware environment, and this is the Power of &. Have your containers without compromise. Have your OpenStack on a platform easily fit on top of VMware vSphere.

By far the biggest highlight was meeting a lot of people who regularly read my blog and have benefited from the work that I and others in the community have done over the years. This is why we keep doing it. Because it makes a difference. It was also great to meet a lot of people who had bought Virtualizing SQL Server with VMware: Doing IT Right, and had got a lot out of it too. My co-authors, Michael Corey, Jeff Szastak and I were blown away by the stories that were relayed to us about how the book had helped people, especially when it was being used to explain to DBA’s how virtualization works and that SQL is a great candidate for virtualization. The book was so popular that it actually sold out at VMworld, and we had a lot of people come up to us during the meet the authors session and book signing.

Here is a photo of my co-authors and I with the happy customer who purchased the very last copy of our book at VMworld.

VMworld 2014 - Last SQL Book Sold

 

Shortly after the above photo was Michael Corey and I recorded an interview with VMworld TV’s Eric Sloof regarding virtualizing SQL Server Databases on VMware.

I was lucky again this year to present a session that was included in the top 10 sessions of VMworld for the second consecutive year – VMworld 2014 SDDC1600 Art of IT Infrastructure Design The Way of the VCDX Panel.

 

Final Word

It was another great VMworld and a very successful VMworld.  I’m very much looking forward to next year. Hopefully we’ll see the return of the Monster VM sessions and some other business critical apps sessions from me in next years VMworld.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/10/19/vmworld-2014-wrap-up/feed/ 0 4963
VMware Turns Off TPS Taps in vSphere ESXi and vCloud Air to Avoid Rare VMescape Security Bug http://longwhiteclouds.com/2014/10/19/vmware-turns-off-tps-taps-in-vsphere-esxi-and-vcloud-air-to-avoid-rare-vmescape-security-bug/ http://longwhiteclouds.com/2014/10/19/vmware-turns-off-tps-taps-in-vsphere-esxi-and-vcloud-air-to-avoid-rare-vmescape-security-bug/#comments Sat, 18 Oct 2014 19:24:58 +0000 http://longwhiteclouds.com/?p=7709


VMware has announced that it will turn off TPS in upcoming version of it’s hypervisor ESXi and vCloud Air hybrid cloud service. This is due to a security bug, considered a very rare possibility and only exploitable in very controlled and largely misconfigured environments.  TPS also known as Transparent Page Sharing is a memory management technique that allows multiple […]

]]>


VMware has announced that it will turn off TPS in upcoming version of it’s hypervisor ESXi and vCloud Air hybrid cloud service. This is due to a security bug, considered a very rare possibility and only exploitable in very controlled and largely misconfigured environments.  TPS also known as Transparent Page Sharing is a memory management technique that allows multiple VM’s to share a read only copy of the same memory page. When a VM needs to update or write to a page a new copy is created. The idea is that if there are many VM’s with similar memory pages on the same physical host server it will de-duplicate the pages and only store one copy. The result is that you can run more VM’s per physical server while still achieving very good performance.

TPS has for a long time been used as a competitive advantage by VMware over all of the other hypervisors. But realistically it hasn’t been in wide use by most customers for some time (since ESX 3.5) as the amount of RAM per host has increased, because of the use of large memory pages (2MB instead of 4KB) in Nehalem and above processors, and because most customers don’t want to run their systems at 100% utilization so that they can handle bursts of activity. When using large pages TPS only kicked in when systems were over 96% memory utilization, at which point large pages would be broken down into small pages that could be shared. However this has been a popular technique with service providers and with virtual desktop environments, and in some test and development environments, where over commitment of memory may have been acceptable.

The security problem was found by recent research that leverages Transparent Page Sharing (TPS) to gain unauthorized access to data under certain highly controlled conditions. The research demonstrated that by forcing a flush and reload of cache memory, it is possible to measure memory timings to try and determine an AES encryption key in use on another virtual machine running on the same physical processor of the host server, if Transparent Page Sharing is enabled. This is effectively a VM escape, where code executed within one VM can break the hypervisor isolation and read data from another VM’s memory. Certainly not a good situation if said VM contains credit card data, as we’ve already had enough breaches recently. The conditions under which this could be exploited would be rare in the real world, especially as most environments don’t use TPS actively, even if it is enabled. Even so, I believe in being secure by default, and even though the number of conditions that have to simultaneous by true for this to be exploited would be very rare, if this were exploited the impact could be high. So I believe that VMware is taking the right approach to this research by disabling TPS.

I have been a proponent for leaving TPS enabled in the past, even though a few others have previously recommended it be disabled for performance reasons. My argument was that TPS is a good safety net if all else fails, even if during normal operations it is not used. Also performance was never proven to be a factor. I put this argument in my article Blueprint for Successful Large Scale Oracle Virtualization on vSphere when an EMC paper recommended disabling TPS. To quote that article “Disabling TPS can have disastrous consequences, including causing additional host swapping, which can result in extremely poor performance, much worse than disabling it could ever possibly gain.” So this begs the question, now that it’s being disable by VMware what impact will it have?

Without TPS you will have to have much more conservative memory usage per host. If you business requirements dictate, you will have to be able to sustain maintenance and failure without causing memory overcommitment. If there is a failure or maintenance that causes temporary or prolonged overcommitment of memory you will have a lot more guest OS swapping, due to ballooning, and also host swapping may occur, which would greatly impact performance. Memory swapping is the enemy of performance, and this also adds significantly to poor performance on shared storage if it occurs. But this is possibly better than the alternative security bug.

If you have an existing VMware vSphere environment this will mean you need to evaluate the level of resource usage you have today, your standard operating procedures for maintenance, and the settings of VMware HA Admission Control for failure. If you don’t have sufficient available memory to operate your environment in the case of failure or maintenance, then you may need to upgrade the amount of RAM per host or purchase additional hosts. With any additional hosts you’d need additional licenses. Frank Denneman has a good take on the capacity planning implications in his article here.

TPS will be disabled by default from the following VMware vSphere Releases:

  • ESXi 5.5 Update release – Q1 2015
  • ESXi 5.1 Update release – Q4 2014
  • ESXi 5.0 Update release – Q1 2015
  • The next major version of ESXi

VMware’s official statement on this problem is contained within KB 2080735 Security considerations and disallowing inter-Virtual Machine Transparent Page Sharing. This KB also contains the steps to disable TPS on older versions of VMware vSphere that will not be covered by patches.

If you want to check whether you have TPS enabled or not on your existing versions, and if you want to disable it you can use the following PowerCLI examples (explicitly provided without any warranty, use at your own risk):

 

Check if TPS is Enabled on all hosts connected to a vCenter Server, Mem.ShareScanGHz returns > 0 if enabled.

Connect-VIServer <YourvCenter>
Get-VMHost –State Connected | Get-AdvancedSetting –Name Mem.ShareScanGHz | Format-Table –Property Entity,Name,Value -AutoSize
Disconnect-VIServer

 

Disable TPS on all hosts connected to a vCenter Server by setting Mem.ShareScanGHz = 0, check the setting has been applied correctly

Connect-VIServer <YourvCenter>
Get-VMHost –State Connected | Get-AdvancedSetting –Name Mem.ShareScanGHz | Set-AdvancedSetting –Value 0
Get-VMHost –State Connected | Get-AdvancedSetting –Name Mem.ShareScanGHz | Format-Table –Property Entity,Name,Value -AutoSize
Disconnect-VIServer

 

So if TPS is vulnerable to data leakage and VM escape attacks what about the recently announced Project Fargo, AKA VMFork? VMFork allows a running VM to be quiesced and rapidly cloned by using a similar copy on write technique to share a read only copy of the parent VM memory, and sharing the parent VM’s read only disk, with updates being written to a delta disk. This allows a VM to be cloned and get up and running on the network with it’s own personality in a matter of a few seconds, with the VM memory and disk effectively being deduped at the same time. This doesn’t just have applicability to VDI environments, but web server environments, Dev and Test environments and many other use cases. I’m sure VMware won’t let VMFork out in the wild until issues such as the VM escape bug with TPS are addressed. Kit Colbert, VMware CTO for End User Computing, has said to me that VMFork is much more secure than TPS, so it may not suffer from the same problems.

VMware is not alone with a VM escape vulnerability being discovered. There was also a security bug made public regarding the Xen hypervisor that allowed a VMescape, where code executed within one VM could escape the encapsulation of the hypervisor to a neighbour VM or dom0. This is covered at the VUPEN Vulnerability Research Team’s blog site.

 

Final Word

Nothing is fully secure. You can never guarantee that your system isn’t vulnerable to attack. All you can do is take appropriate measures to reduce the risk of attack, implement technical controls and monitoring and auditing processes. Implement separation of duties, least privilege access, and role based access controls. Implement the guidelines that make sense based on your business requirements from the VMware and other vendors hardening guides. Comply with the security standards for your industry / company that make sense. Stay on top of critical security patches and implement them as soon as practicable, especially for any environments containing public facing or highly secure systems.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/10/19/vmware-turns-off-tps-taps-in-vsphere-esxi-and-vcloud-air-to-avoid-rare-vmescape-security-bug/feed/ 2 7709
Exchange Best Practices and 1.4Million Mailboxes on Nutanix NX-8150 http://longwhiteclouds.com/2014/09/28/exchange-best-practices-and-1-4million-mailboxes-on-nutanix-nx-8150/ http://longwhiteclouds.com/2014/09/28/exchange-best-practices-and-1-4million-mailboxes-on-nutanix-nx-8150/#comments Sun, 28 Sep 2014 10:43:45 +0000 http://longwhiteclouds.com/?p=6383


Nutanix has recently published a Best Practice Guide for Mictosoft Exchange on VMware vSphere and Josh Odgers explains some of it’s contents  and benefits of Exchange on Nutanix in his blog article here. If you are interested in virtualizing Exchange, and/or using Nutanix, you might want to get hold of the guide and have a read […]

]]>


Nutanix has recently published a Best Practice Guide for Mictosoft Exchange on VMware vSphere and Josh Odgers explains some of it’s contents  and benefits of Exchange on Nutanix in his blog article here. If you are interested in virtualizing Exchange, and/or using Nutanix, you might want to get hold of the guide and have a read through it. It explains how to simply set up Exchange on Nutanix, the benefits of it, how it compares to a traditional physical JBOD approach and much more. The paper introduces the capability of running Exchange on the Nutanix NX-8150 nodes, which have been specially designed to run large applications, such as Exchange, SQL Server, Oracle and SAP. This is the node type Josh Odgers and I used as part of a design capable of hosting 1.4 million Exchange 2013 Mailboxes, which demonstrates the building block architecture of Nutanix and the ability to scale to meet requirements for large environments. Let’s take a look at that design at a high level.

This design was created as part of a response to an RFI (Request For Information) for a customer in Europe. They wanted to support 1.4 million Exchange 2013 mailboxes at 500MB each with 50 messages per day. The number of messages per day and the size of mailboxes can be adjusted by adding more nodes to the design.

The Nutanix NX-8150 nodes  (1 node per 2RU) used in the below example are configured with 2 x E5-2690 v2 sockets (10 cores each @ 3GHz) and 256GB RAM. They each have 10TB usable capacity per node before any dedupe or compression benefits, and can be configured with up to 6.4TB of raw flash capacity per node (3.2TB raw flash per node used in this example). Note: Diagrams are not exactly to scale as there are additional RU remaining unused per rack.

NX8150

Assumptions:

1. 500MB Mailbox for 1.4 Million users

2. 50 Messages / Day

3. Following MS recommendations as follows:

a. 10% Hypervisor Overhead

b. 80% CPU utilization per server maximum

4. No Capacity savings from Compression or deduplication taken into consideration

5. Site Resiliency is mandatory

Configuration

1. 16 x 9 node DAGs with mixed server roles (MBX+CAS server roles combined) per VM with 24 vCPUs (PreferHT) and 96GB RAM

2. Total of 144 x NX-8150 w/ 256GB RAM inc 32GB RAM for CVM plus 128GB remaining to support MBX server vMotion for maintenance (per site)

3. Secondary Datacenter also requires 144 x NX-8150 nodes to host 1 DAG and 1 LAG copy

4. 512MB mailboxes @ 100% capacity requires 6939GB per node which leaves >2TB free per node in the worst case scenario

5. Mailbox Servers will run at <80% utilization (at both sites) as per MS recommendations

6. Exchange Transactional IOPS Per Node required is 474 – Note: Jetstress results in N+1 is over 600 IOPS/Node

7. 1 x Active Directory Global Catalog Server per node with 2 vCPUs / 4GB RAM (or equiv number of cores in larger AD VMs) Note: 313 GC Cores required per site

8. A total of 10 racks required per site (20 racks in total). The entire Nutanix environment can be managed across both data centers from a single management console (PRISM Central).

Recommendation

1. Nutanix/vSphere cluster sizes of 9 nodes striped across racks, 16 nodes per rack (each a member of a different Nutanix cluster) – Note: 16 x NX-8150 is 32RU leaving 10RU – 2 x 1RU ToR Switches , 2 x 1RU Cable Management, 6RU remaining (additional 1RU in each of the network spine racks required)

2. Mailbox Server placement being 1 DAG member per Nutanix Cluster i.e. extremely high resiliency as 1 cluster failure does not bring down a single DAG

3. Use 2 x 48 port 10GbE MoR (Middle of Rack) switches per 2 Rack pair and 2 x 48 Port 1GbE switches per rack pair (one switch per rack with servers cross connected) – 32 ports per 10GbE switch for mailbox servers and 16 per 1GbE switch for IPMI (Out of Band Management), optionally could also connect servers with 1GbE in addition to IPMI

4. Start with 8 nodes per DAG and 4 dags (32 nodes per site) for a total of 64 Nodes across both active/active sites (~300K mailboxes @ ~4688 per node), grow by additional 64 nodes, then 16 nodes (~75K mailboxes) at a time until final configuration is met.

Initial Exchange Infrastructure Layout (~300K Mailboxes)

Exchange8150LargeScale-initial

Next Step Exchange Infrastructure Layout (~600K Mailboxes)

Exchange8150LargeScale-Scale2

Third Step Exchange Infrastructure Layout (~675K Mailboxes)

Exchange8150LargeScale-Scale3

Fourth Step Exchange Infrastructure Layout (~750K Mailboxes)

Exchange8150LargeScale-Scale4

Subsequent Steps for Exchange Infrastructure Layout (Adding ~75K mailboxes per increment)

Exchange8150LargeScale-ScaleN

Final Exchange Infrastructure Layout (1.4M Mailboxes)

Exchange8150LargeScale-Final

Exchange Network Infrastructure Layout

Exchange8150LargeScale-Networking

Final Word

It is easy with Nutanix to start much smaller and design exchange environments for much lower number of mailboxes (or much larger mailbox size) and for smaller growth increments (<500 – > 2000 mailboxes). This example was for a very large environment and specific customer requirements, and shows for this large environment a small(ish) starting point and growth increments, until it reaches it’s final size. The diagrams include all compute, storage and networking equipment required to host the specified number of mailboxes. At the time of publishing this article the outcome of the RFI that this design was created for is unknown. If Nutanix is successful and the customer agrees to be a reference you’ll be sure to hear about it. The purpose of this article was to show at a high level how you could use Nutanix NX-8150’s and a building block approach to design a large scale environment suitable for Exchange. Although the target for this design, and the Exchange Best Practices paper is VMware vSphere you could also use Microsoft Hyper-V (Nutanix Exchange Best Practice Guide on Microsoft Hyper-V is in the works).

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/09/28/exchange-best-practices-and-1-4million-mailboxes-on-nutanix-nx-8150/feed/ 3 6383
Is Your Heart and Data Bleeding From the Shell Shock? http://longwhiteclouds.com/2014/09/26/is-your-heart-and-data-bleeding-from-the-shell-shock/ http://longwhiteclouds.com/2014/09/26/is-your-heart-and-data-bleeding-from-the-shell-shock/#respond Fri, 26 Sep 2014 04:42:31 +0000 http://longwhiteclouds.com/?p=6262


If you thought Ebola was deadly to humans wait till you get a load of the latest security issue impacting the world wide web and most everything connected to it including potentially your phone, lights, servers and the list goes on (excluding Windows systems). If Heart Bleed wasn’t bad enough at the start of the […]

]]>


If you thought Ebola was deadly to humans wait till you get a load of the latest security issue impacting the world wide web and most everything connected to it including potentially your phone, lights, servers and the list goes on (excluding Windows systems). If Heart Bleed wasn’t bad enough at the start of the year the new Shell Shock bug certainly is. It is what I would term the Mother of All Bugs (MOAB). It impacts almost all Unix, Linux and Mac systems and allows a remote attacker to execute arbitrary code and potentially steal your data, credit cards and other information. So how serious is this? Well the NIST CVE Alert Rating on this is a 10 for severity, and a low for complexity to exploit (read my 7yr old could exploit this bug). So basically the worst possible kind. Oh, but wait, there’s more…

According to this article, there are already worms exploiting this bug. So the impact could be wide spread for the vulnerable systems. I would expect most major vendors to come out with security advisories very promptly for this, after they have assessed their systems. For those of you running VMware, they have posted a blog here, and an advisory here. As things stand if you’re running VMware tools on top of Windows, such as vCenter for example, then you are not vulnerable. Also ESXi is not vulnerable as it uses ash shell via BusyBox instead of Bash. However any virtual appliances may well be vulnerable, including the vCenter Server Appliance. I would recommend keeping and eye on VMware KB 2090740 for the latest updates. For home users, lock up your networks tight and try to prevent anyone getting in the virtual front door, until such time as there are widely available fixes.

This bug highlights the importance of keeping patches up to date and staying across the alerts from the likes of NIST. Be aware of this bug and get patched and protected as soon as you can. Not everyone has a vaccine for this one yet, but hopefully it’s not far away. This bug could cause a lot of change to the way systems are designed, implemented and secured. What’s to say another bug of this nature isn’t just around the corner? Better to be prepared.

Final Word

This is probably one of the highest impact and most wide spread bugs with the highest severity that I’ve seen in over 20 years in IT (reminds me of the original internet worm). As the Internet of Things (IOT) spreads bugs of a similar nature will have a much wider impact and much more sever consequences. Security of your systems is going to become an ever more serious issue and this is why Micro Segmentation, and using technologies such as VMware NSX and vCloud Networking and Security will become so important. In addition to more intelligent firewalls, such as from Palo Alto Networks. As much as we give Microsoft a hard time over security and patches, neither Heart Bleed and Shell Shock impacted Windows systems.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/09/26/is-your-heart-and-data-bleeding-from-the-shell-shock/feed/ 0 6262