(cas:72) Google Analyticator was unable to authenticate you with Google using the Auth Token you pasted into the input box on the previous step.

This could mean either you pasted the token wrong, or the time/date on your server is wrong, or an SSL issue preventing Google from Authenticating.

Try Deauthorizing & Resetting Google Analyticator.

Tech Info 400:Error fetching OAuth2 access token, message: 'invalid_grant'
Unique
Visitors
Powered By Google Analytics
SDDC – Long White Virtual Cloudsu by http://longwhiteclouds.com all things Nutanix, VMware, cloud and virtualizing business critical applications Fri, 13 Apr 2018 21:56:56 +0000 en-US hourly 1 https://wordpress.org/?v=6.7.6 45024036 Stop Playing Russian Roulette With Your Data http://longwhiteclouds.com/2018/02/17/stop-playing-russian-roulette-with-your-data/ http://longwhiteclouds.com/2018/02/17/stop-playing-russian-roulette-with-your-data/#comments Fri, 16 Feb 2018 23:49:57 +0000 http://longwhiteclouds.com/?p=12095


Some vendors in the storage, hyper-converged, and cloud industries may be playing Russian Roulette with their customers’ data. Solutions are not created equally, some turn off basic data integrity features such as data checksum by default, or when there are performance problems. Some don’t have background consistency checks and scrubbing to protect against silent data […]

]]>


Some vendors in the storage, hyper-converged, and cloud industries may be playing Russian Roulette with their customers’ data. Solutions are not created equally, some turn off basic data integrity features such as data checksum by default, or when there are performance problems. Some don’t have background consistency checks and scrubbing to protect against silent data corruption or latent sector errors. Others might use consumer grade devices that may have a higher risk of error and higher failure rate. In the age of software defined solutions, the customer has become the storage platform architect. There is enough rope to hang yourself (your data and your platform availability) any number of different ways. Which is why having a software foundation and integrated solution that has been properly validated from end to end, and that contains data integrity and enterprise data protection features at it’s core, should be the highest priority. Return of data, in the form it was originally written, at any scale, while protecting against known data and device risks, is of upmost importance. How important is performance (IOPS, Latency and Throughput) if you can’t even read back the data you originally wrote? Here are the top 10 questions you can ask potential vendors to find out if they really have protecting your data as their top priority.

Before we get started with the questions, it’s always good to have some science and evidence to back things up. Here is one paper – An Analysis of Data Corruption in the Storage Stack. Another paper – Characterizing Private Clouds: A Large-Scale Empirical Analysis of Enterprise Clusters. Both papers cover large scale studies. Any study across a small population of devices or a very small sample size is going to be invalid. Any conclusions from something like a 30 drive study isn’t going to be valid when you have tens of thousands, hundreds of thousands, or millions of devices.

Questions to ask your potential solution vendor:

  1. Does the solution include data checksums to ensure that data written is the same as data read / returned, if so, are they on by default or optional?
  2. Do the checksums have a performance impact on random or sequential IO operations, if so, what is the impact?
  3. Does the solution include consistency checks or scrubbing to protect against silent data corruption, silent bit rot, and latent sector errors, if so, are they on by default or optional?
  4. Does the solution include SMART checks and predictive failure analysis, which could include predictive replacement and automated support case generation?
  5. What is the annualized return rate or failure rate of the devices used in the solution and over what number of devices and duration has that been measured?
  6. How does the solution protect data between different components (disks, servers/nodes, clusters) and is this tunable based on different requirements?
  7. How does the solution protect against multiple concurrent component failures and which type of component failures are protected against?
  8. Are user defined failure domains supported to protect against situations such as chassis failure, rack failure, multiple storage device failure?
  9. How does the solution recover from single and multiple component failure, this could be single storage device failure, multiple devices on the same shelf or node, or multiple node failures, and what is the expected recovery time and performance impact? Does this scale linearly as the solution continuously grows over time?
  10. Do recovery options rely on a single device, such as a hot spare, or in the case of an object store, a single device holding the replica of a large component, or do recoveries utilize all devices in the system equally and fairly?

There are plenty more questions that could be asked, but the 10 questions above cover the most common areas of risk in terms of data integrity, data protection and data loss prevention, and that are not always protected against, at least not by default, with some systems.

Final Word

From a Nutanix point of view, as a leader in the Gartner Magic Quadrant for Hyperconverged Infrastructure, we take data integrity seriously and it’s our top priority. We protect against all of the areas highlighted in the questions above and we have a paper that explains the Infrastructure Resiliency of Nutanix Solutions, which compliments the research paper on enterprise clusters. We also have many hundreds of thousands of devices in production that are proactively monitored from which we can draw real world data, and a very thorough device qualification and QA process, which limits risk. We use similar high standards across all hardware platforms that our software supports. Our software is built based on a philosophy that hardware will eventually fail, so we must deal with these failures gracefully. Your data deserves better protection!

 


This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2018 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2018/02/17/stop-playing-russian-roulette-with-your-data/feed/ 2 12095
Heads Up! Do Not Upgrade VMware Tools on Hosts with ESXi 6.0 U1b http://longwhiteclouds.com/2016/01/10/heads-up-do-not-upgrade-vmware-tools-on-hosts-with-esxi-6-0-u1b/ http://longwhiteclouds.com/2016/01/10/heads-up-do-not-upgrade-vmware-tools-on-hosts-with-esxi-6-0-u1b/#comments Sat, 09 Jan 2016 14:00:42 +0000 http://longwhiteclouds.com/?p=11517


Heads Up! If you’ve updated to ESXi 6.0 U1b, build 3380124 and you have lots of templates, you may run into some problems if you update VMware Tools to the latest version. I just upgraded my environments to the latest VMware patches ESXi 6.0 U1b (build 3380124), that has just come out. As you do […]

]]>


Heads Up! If you’ve updated to ESXi 6.0 U1b, build 3380124 and you have lots of templates, you may run into some problems if you update VMware Tools to the latest version. I just upgraded my environments to the latest VMware patches ESXi 6.0 U1b (build 3380124), that has just come out. As you do usually when there is a new hypervisor build you upgrade VMware Tools. Well that proved to be a big problem for my VM templates that I use to provision new systems. But I’ve got a workaround.

As soon as VMware Tools is updated on any templates you will no longer be able to clone those templates. If you’ve updated any templates with the version of VMware Tools that comes with ESXi 6.0 U1b then you need to uninstall it and reinstall the prior version that came with ESXi 6.0 build 3247720. After the couple of reboots that you have to go through with an uninstall and reinstall of VMware Tools you will find that you can now clone VM’s and have them automatically customized. I ran into this problem on Windows 2008 R2 Server. So I know it will impact this guest OS. I haven’t tested other OS’s yet, but others could be impacted. I’ve logged a support call with VMware to address this problem. In the meantime, the workaround is fine. The Official VMware KB Article 2142982 explains the situation.

 

[Updated 14/01/2016] After further testing I have narrowed down the problem area to new installs where the complete option is selected, and any upgrades where the complete options was previously selected, or where the VMCI / NSX Guest Introspection Driver is included. I have been able to successfully clone from a new VM Image that has had a fresh install of Windows 2008 R2 and VMware Tools without the VMCI / NSX Guest Introspection Driver, or where VMware Tools was installed twice / installed and repaired on the same VM, when the complete option was previously selected. This seems to be similar to what other of you have also reported. I have completed the Upgrade Scenario testing as well and confirmed that after an upgrade, if the complete install option was previously selected the VM will not clone due to the same VMCI driver problem. If VMCI driver is removed by running VMware Tools Install again and selecting Modify and unselecting VMCI, then you will be able to close the VM.

This update just in from VMware Support “VMware Engineering have confirmed that the issue is dependent on the install/upgrade sequence. Specifically, the issue is aligned to the version of deploypkg.dll in the vmtools package. GSS and Engineering are mapping the ESX and vmtools update versions to the deploypkg.dll  versions to confirm which upgrade sequences are problematic. A KB article will be published once this information is finalised.

Thanks to VMware Support for getting to this stage very quickly. The VMware KB Article 2142982 has now been published. 

Final Word

I guess someone has to take the risk and patch their systems to the latest versions first, especially as these were security patches with a critical severity. Fortunately like all good IT environments I only did my test systems first. This is the whole point of having infrastructure test systems. You can test infrastructure hardware and infrastructure software changes first before putting them into production. The old saying goes that software eventually works and hardware eventually fails, but these days a lot of your hardware is also software, especially in a virtualized software defined datacenter. It pays to have appropriate test systems and test plans to mitigate the risks associated with software updates and changes of all types, including infrastructure software. Thanks to all of you in the community that contributed to this effort and commented on this blog post. I have been relaying your feedback during my discussions with VMware Support.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2016/01/10/heads-up-do-not-upgrade-vmware-tools-on-hosts-with-esxi-6-0-u1b/feed/ 37 11517
Configuring Scalable Low Latency L2 Leaf-Spine Network Fabrics with Dell Networking Switches http://longwhiteclouds.com/2015/03/26/configuring-scalable-low-latency-l2-leaf-spine-network-fabrics-with-dell-networking-switches/ http://longwhiteclouds.com/2015/03/26/configuring-scalable-low-latency-l2-leaf-spine-network-fabrics-with-dell-networking-switches/#comments Thu, 26 Mar 2015 08:36:19 +0000 http://longwhiteclouds.com/?p=10603


In December 2014 I got an early Christmas present from Dell. They shipped me the latest 40G and 10G S series (Force 10) switches so that I could begin to test, validate and document the integration and reference architectures between Dell Networking and Nutanix. I’m starting with a L2 MLAG (Multi-chassis Link Aggregation Group)  configuration […]

]]>


Logo-JPG-Dell_Force10_Dell-BlueIn December 2014 I got an early Christmas present from Dell. They shipped me the latest 40G and 10G S series (Force 10) switches so that I could begin to test, validate and document the integration and reference architectures between Dell Networking and Nutanix. I’m starting with a L2 MLAG (Multi-chassis Link Aggregation Group)  configuration and I will work my way through to a full ECMP (Equal Cost Multi-path) L3 configuration including VMware NSX. This will be a journey and I’ll include the different options and their key considerations, and configurations in the eventual white papers that Nutanix publishes. I’ve had a few weeks to configure them and do some initial testing (after the Christmas Holiday break), so I thought I’d write about what I’ve found so far. On the NSX front, you’ll be interested to know that Nutanix already has customers running NSX and that the platforms work extremely well together, as they both scale out linearly and predictably. But we’ll leave NSX specific discussion till another day. This article will contain some highlights without steeling the thunder of the white papers I’m working on.

Traditionally datacenter networks were usually designed with 3 layers, Access – where servers connected, Aggregation or Distribution – where the access switches connected and also normally the L2 demarcation, and Core, where everything was bought together at L3. Spanning Tree Protocol (STP) is enabled and redundant links will be blocked to prevent network loops on the L2 segments. However with this design you are not able to utilize all the links, due to STP, and it can be complex to scale and achieve consistent latency between the different points in the network. However these problems can be solved by taking a leaf and spine architecture approach.

As you can see in the diagram above I have 2 Spine Switches – Dell S6000’s with 32 x 40GbE ports, and 4 Leaf Switches – Dell S4810 with 48 x 10GbE Ports and 4 x 40GbE Ports. The Spine and two of the Leaf Switches are running Dell FTOS (Force Ten OS) 9.6, while two of the Leaf switches are the S4810-ON (Open Networks) model and are running Cumulus Linux 2.5. Check out the details of the Dell Networking Force 10 S series switches.

cumulus-networks-dell-sdn-bare-metal-white-box-switchesCumulusTurtleLogo

 

 

The reason I like Cumulus Linux, is well, it’s just Linux, but with hardware accelerated switching and routing. If you know Linux (Debian is the distribution it’s based off), then it’s very easy to use. No need for additional training, fits into the same management frameworks, such as Puppet, Chef, Ansible and others. So it’s great to have the choice of this or FTOS on the Dell Networking switches. I found both Cumulus and FTOS easy to use, and the documentation is very good.

The port to port latency on the S6000’s is ~500ns, while the S4810’s are ~800ns (so about 1.3us Leaf to Spine), even when routing. With the overhead in the IP stack of each of the hypervisors and VM’s I’m seeing latency across the network between VM’s of < 90us end to end, this is using standard Intel 10GbE NIC’s, VMXNET3 vNIC and without using the latency sensitive settings or tuning within VMware vSphere 5.5.

The main benefit of using MLAG is that you don’t have to have a whole lot of links disabled to prevent network loops and have STP interfering with it, it’s also very easy to set up. You still have STP enabled to prevent loops during switch boot, but when things are up and running all of the links are available to pass traffic and you benefit from the combined bandwidth. Each switch is managed and updated independently, so doesn’t become a single point of failure, as it could if you’d chosen to stack the switches. Switch stacking might be ok if you needed multiple stacks anyway and the host links to the stack themselves were redundant (you can’t mix stacking and MLAG together). While this does add some slight management overhead, it gains the ability to update the switch firmware independently without causing any disruption. With the standard automation and management frameworks such as Puppet, Ansible, Chef, CFEngine etc, the management overhead is greatly reduced or eliminated in any case.

Here is a diagram of what my lab network looks like at a high level. My lab hosts are directly connected to the leaf switches.

NZ Performance Lab - Networking

In FTOS an MLAG is called a Virtual Link Trunk (VLT), and in Cumulus it’s called CLAG (Chassis Link Aggregation Group). The process to configure them is fairly similar at a high level.

Configure Out of Band Management Interfaces (after initial switch boot / install)
Create a port channel between the adjacent switches (Spine or Leaf)
Configure Spanning Tree (RSTP)
Create a Peer Link on top of the port channel between the adjacent switches so that inter-switch communications can take place to sync mac addresses etc. Set up a backup address in case the primary fails, this will prevent split brain scenarios.
Configure and enable VLT or CLAGD (examples will follow)
Configure and enable the other port channels, edge ports, VLAN’s, routing etc

FTOS Spine Example:

! Note: Peer Link is recommended to be static port channel not LACP.
!
lacp ungroup member-independent vlt
lacp ungroup member-independent port-channel 100
!
default vlan-id 4000
!
protocol spanning-tree rstp
no disable
bridge-priority 16384
!
vlt domain 1

peer-link port-channel 100
back-up destination xxx.xxx.xxx.xxx <- IP Address of Backup Destination
primary-priority 16384
peer-routing
peer-routing-timeout 1
!
interface Port-channel 10
description Cumulus Leaf-Link
no ip address

mtu 9216
portmode hybrid
switchport
lacp fast-switchover 
vlt-peer-lag port-channel 10
no shutdown
!
interface Port-channel 100
description Peer-Link
no ip address
mtu 9216
channel-member fortyGigE 0/120,124
no shutdown
!
interface fortyGigE 0/112
description Leaf1 – Port Channel 10
no ip address
mtu 9216
flowcontrol rx on tx off
!
port-channel-protocol LACP
port-channel 10 mode active
no shutdown
!
interface fortyGigE 0/116
description Leaf2 – Port Channel 10
no ip address
mtu 9216
flowcontrol rx on tx off
!
port-channel-protocol LACP
port-channel 10 mode active
no shutdown
!
interface fortyGigE 0/120
description Peer-Port 1 – Port Channel 100
no ip address
mtu 9216
flowcontrol rx on tx off
no shutdown
!
interface fortyGigE 0/124
description Peer-Port 2 – Port Channel 100
no ip address
mtu 9216
flowcontrol rx on tx off
no shutdown
!
interface Vlan 500
description Host VLAN
no ip address
mtu 9216
tagged Port-channel 1,10
no shutdown
!
interface Vlan 4000
mtu 9216
!untagged Port-channel 100
no shutdown
!

Cumulus Leaf Example with CLAGD:

# This file describes the network interfaces available on your system
# and how to activate them. For more information, see interfaces(5), ifup(8)
#
# Please see /usr/share/doc/python-ifupdown2/examples/ for examples
#
#
# The loopback network interface
auto lo
iface lo inet loopback
# The primary network interface
auto eth0
iface eth0
address xxx.xxx.xxx.xxx/24
broadcast xxx.xxx.xxx.255
# Spine Link
auto spn1-2
iface spn1-2
bond-slaves swp49 swp50
bond-mode 802.3ad
bond-miimon 100
bond-use-carrier 1
bond-min-links 1
bond-xmit_hash_policy layer3+4
clag-id 1
# clag-id needs to be unique on each clag, like a vlt domain id on FTOS
mstpctl-portnetwork no
mtu 9216
# Peer Link to Other LeafSwitch
auto pl
iface pl
bond-slaves swp51 swp52
bond-mode 802.3ad
bond-miimon 100
bond-use-carrier 1
bond-min-links 1
bond-xmit_hash_policy layer3+4
mstpctl-portnetwork no
mtu 9216
# CLAGD Peer Int Config
auto pl.4000
iface pl.4000
address xxx.xxx.xxx.xxx/30
clagd-enable yes
clagd-priority 8192
clagd-peer-ip 172.16.0.2
clagd-backup-ip 192.168.255.12
clagd-sys-mac 44:38:39:ff:00:01
# Switch Port Interface Configuration
auto swp1
iface swp1
mtu 9216

auto swp2
iface swp2
mtu 9216

auto swp3
iface swp3
mtu 9216

auto swp4
iface swp4
mtu 9216

auto swp49
iface swp49
mtu 9216

auto swp50
iface swp50
mtu 9216

auto swp51
iface swp51
mtu 9216

auto swp52
iface swp52
mtu 9216

# Bridge Configuration
#
auto br0
iface br0
bridge-vlan-aware yes
bridge-ports pl spn1-2 glob swp[1-4]
bridge-stp on
bridge-pvid 1
bridge-vids 500
mstpctl-portadminedge swp1=yes swp2=yes swp3=yes swp4=yes
bridge-mcsnoop 1
mtu 9216
# Bridge VLAN
# Host VLAN
auto br0.500
iface br0.500
address xxx.xxx.xxx.xxx/23
broadcast xxx.xxx.xxx.255
up ip route add 0.0.0.0/0 via 192.168.1.230
mtu 9216

The above are examples and not complete configs and you can’t just copy and paste it all and expect it to work in your environment. But it could be used as a starting point.

Bringing this all together. The Dell S6000 / S4810 combination allows you to create a scalable network design that provides predictable and consistent low latency and high throughput from end to end in the network. The configuration of the MLAG/CLAG/VLT is straight forward, and it provides management flexibility. You get to choose either FTOS or Open Networking such as Cumulus Linux as your switching software. With Cumulus Linux, it’s just like Linux, but wire speed non-blocking networking accelerated in hardware and fits into the normal management frameworks. For environments, such as Hyper-converged or Web-scale infrastructure, the network scales linearly, as do the systems that connect to it. Each time you grow, you get consistent, predictable and linear performance.

Here is an example of a high level diagram that might be appropriate for a small scale deployment. In this case the Dell S4810’s are used as both Leaf and Spine switches in a VLT configuration, with Dell N3048 or N3024 providing 1G connectivity. With this you could easily start with a single rack and scale to 8 racks. Each rack would have full 10G and 1G redundant connectivity. With say 24 Servers or Hyper-converged nodes per rack you would be able to support 192 servers across 8 racks. 40G QSFP+ ports are used for Peer Links, while the Leaf connects to the Spine using 40G to 10G break out cables.

NX3000 Small Scale Networking

 

Here is an example high level diagram of a medium density Nutanix Web-scale Converged Infrstructure deployment with Dell S4810 Leaf switches connected to S6000 Spine, which could be using FTOS VLT or Cumulus CLAG. As you can see this design is capable of scaling to 12 racks (576 nodes) with the S6000 Spine switches, and up to 52 racks and 2496 nodes, with the Dell Z9500 Spine switches. Dell Z9500 supports 128 x 40GbE ports. Both options have spare 40GbE ports still available for Boarder-Leaf connectivity (Cross Datacenter, or Internet routers etc).

NX3060 Medium Density

If you wanted a higher density design you can combine S6000 in middle of rack or top of rack configuration for Leaf switches with S6000 or Z9500 Spine. This diagram provides a high level example of a high density configuration on a standard 48U rack. In this example the design would support 10 racks with 880 nodes with S6000 Spine, and 48 racks with 4224 nodes with a Z9500 Spine. With enough ports to accommodate the Boarder-Leaf nodes as well.

NX3000 High Density

With densities in the datacenter increasing and the power consumption of the servers decreasing I can see an explosion of 40GbE Top of Rack (ToR) or Middle of Rack (MoR) Leaf switches. This would also provide an option to easily allow different bandwidth oversubscription models, 6:1, 4:1 etc as requirements change.

So far we’ve covered the Dell Networking switches by themselves, with FTOS and Cumulus, and then some examples combined with Nutanix. Dell is an OEM partner of Nutanix software and delivers the Dell XC Series Web-scale Converged Appliances. With Dell XC you have a number of hardware options for different use cases, and you can build a complete solution, including networking.

The following example shows Dell S4810 Leaf switches and Dell S6000 Spine switches with Dell XC series appliances. These appliances are 2U each and contain one node. Dell XC also has 1U options.

Dell XC with Dell Networking

 

Final Word

Those of you who spend time in a modern datacenter will see I’ve drawn the diagrams with the switch ports facing forward, which is actually backwards. I did this to make it easy to draw, and because I like flashing lights. In real world environments the air flow of the switches would be reversed and the back of the switches, where the PSU is, would be facing the front of the rack, so that the cabling can be nice and tidy. In summary, where predictable, low latency, high throughput, linearly scalable networking is required, Leaf Spine architectures are becoming increasingly common (and are simpler than three tier IMHO). Web-scale converged or Hyper-converged Infrastructure benefits from the low latency, high throughput and linearly scalability the Dell Networking switches can provide. As does network virtualization such as VMware NSX. Dell Networking switches combined with Nutanix appliances, or Dell XC series appliances powered by Nutanix can deliver a unified and simplified high performance virtual infrastructure with greatly reduced complexity compared to a traditional three tier architecture. A great foundation for a private cloud or software defined datacenter.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2015 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2015/03/26/configuring-scalable-low-latency-l2-leaf-spine-network-fabrics-with-dell-networking-switches/feed/ 16 10603
Nutanix Web Scale IT Now with All Flash and Metro Availability http://longwhiteclouds.com/2014/10/28/nutanix-web-scale-it-now-with-all-flash-and-metro-availability/ http://longwhiteclouds.com/2014/10/28/nutanix-web-scale-it-now-with-all-flash-and-metro-availability/#comments Tue, 28 Oct 2014 07:24:18 +0000 http://longwhiteclouds.com/?p=8265


Nutanix Web Scale NoSAN now meets NoDisk. I didn’t know that the band Queen could predict the future of IT when I first listened to their song Flash Gordon. But the lyrics I’ve quoted above seem to suggest they could somewhat predict the future of the storage industry. Flash will undoubtedly have a big impact […]

]]>


Flash Saviour of the Universe

Nutanix Web Scale NoSAN now meets NoDisk. I didn’t know that the band Queen could predict the future of IT when I first listened to their song Flash Gordon. But the lyrics I’ve quoted above seem to suggest they could somewhat predict the future of the storage industry. Flash will undoubtedly have a big impact on IT, even if it is only just starting to penetrate the datacenter now (only a small percentage of total deployed storage is flash). So it is probably no surprise that eventually Nutanix Web Scale Converged Infrastructure platform would include options for all flash. Then on top of that we add Metro Availability, the metro storage cluster type availability that is only a few clicks to set up, and significantly simpler to operate and test compared to traditional metro solutions. So you can have your all flash and you don’t need to compromise on any data services. Of course Metro Availability is just a software feature so is available in any of the Nutanix platforms, it will just take a software upgrade once the new version of the Nutanix OS is available (Available from 4.1). So why all flash?

NX9240SPECs

Scale-Out Storage Processing Power with Your Flash:

Flash requires storage processing power to drive IOPs and performance.  Why put all your flash behind two controllers? Dual-controller architectures cannot sufficiently drive large amounts of flash.  Each controller has limited performance, plus you need to run at only 50% utilization to ensure performance is available during maintenance and failure. By spreading flash devices across many controllers, you can drive higher aggregate performance.  This performance also increases as you scale out the number of controllers, with no technical upper limit.  All within a single datastore, namespace, and management domain, and without any single point of failure.

Scale-Out trumps Rip and Replace:

Traditional storage vendors live by a three year rip-and-replace lifecycle.  Storage Controllers need to be swapped out in order to take advantage of advancements in Intel x86 processor capabilities.  With Nutanix’s revolutionary file system, new and old storage controllers can co-exist in the same cluster, allowing you to immediately employ advancements in Intel computing technologies.  More importantly, you can increase performance without a destructive and risky Rip and Replace. You don’t have to wait three years, you can just add a single node at a time, when needed, on demand, without any disruption, and get all the benefits straight away. Being software defined means with a simple software upgrade you also get any new software enhancements and continued investment protection on the same hardware. Your same hardware just keeps getting faster. This is true for all flash as it is for hybrid disk and flash systems. 

Put flash next to your VMs with Nutanix’s data-locality:

Data-locality matters.  Getting the performance and data closer to your VM’s greatly improves performance and provides performance isolation from noisy neighbours. The Virtual Machine data-locality built into the Nutanix Distributed File System keeps the majority of write and read storage I/O on local flash.   Read I/O does not need to traverse the network. This results in an improvement of read latency while also reducing Network bandwidth consumption.

Density, Performance, and Scale:

NX9240Block

Up to 32 Nodes of the Nutanix 9040 per rack with 288TB of enterprise grade flash storage, 768 CPU Cores, 16TB RAM. Each node containing up to 9.6TB flash, 2 x Intel Xeon CPU’s (10 Core 3GHz, or 12 Core 2.7GHz), and 512GB RAM. Get all the flash and compute you need to run all of your high performance VM’s. This platform is built for serious workloads that need lots of consistently low latency storage access and high throughput. Especially where software licensing means you want to scale up performance on fewer systems (Like Oracle DB’s and App servers for example). All with < 16KW power consumption per rack!

NX9240Rack

Final Word

Nutanix is constantly evaluating how we can bring uncompromising simplicity and web scale converged infrastructure to more use cases to meet our customers requirements. We are squarely focused on the future of the software defined datacenter and new storage technologies and we can bring these to market very quickly. This is the first all flash platform, but I’m sure it won’t be the last. No need to compromise on data services, such as metro availability, replication, snapshots and DR to go all flash. From NOS 4.1 you’ll be able to have metro availability with a few clicks of a button on any Nutanix platform.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/10/28/nutanix-web-scale-it-now-with-all-flash-and-metro-availability/feed/ 2 8265
Nutanix and vCloud Automation Center http://longwhiteclouds.com/2014/10/19/nutanix-and-vcloud-automation-center/ http://longwhiteclouds.com/2014/10/19/nutanix-and-vcloud-automation-center/#comments Sat, 18 Oct 2014 22:12:04 +0000 http://longwhiteclouds.com/?p=7811


My colleague Magnus Andresson (VCDX-56 and Double VCDX DCV/Cloud) has put together some short videos showing some example solutions with Nutanix and vCloud Automation Center working together. vCloud Automation Center has recently been renamed vRealize Automation also known as vRA (vee Raa! – intentionally not used in the title). I hope you enjoy these videos […]

]]>


My colleague Magnus Andresson (VCDX-56 and Double VCDX DCV/Cloud) has put together some short videos showing some example solutions with Nutanix and vCloud Automation Center working together. vCloud Automation Center has recently been renamed vRealize Automation also known as vRA (vee Raa! – intentionally not used in the title). I hope you enjoy these videos and it gives you some ideas of how you can integrate vCloud Automation Center into your solutions with Nutanix.

 

vRA and vRO (vRealize Orchestrator) Use to Report Nutanix Cluster Status:

Creating Nutanix Containers with vRA:

Deploying VM’s on Nutanix using vRA:

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/10/19/nutanix-and-vcloud-automation-center/feed/ 2 7811
VMware Turns Off TPS Taps in vSphere ESXi and vCloud Air to Avoid Rare VMescape Security Bug http://longwhiteclouds.com/2014/10/19/vmware-turns-off-tps-taps-in-vsphere-esxi-and-vcloud-air-to-avoid-rare-vmescape-security-bug/ http://longwhiteclouds.com/2014/10/19/vmware-turns-off-tps-taps-in-vsphere-esxi-and-vcloud-air-to-avoid-rare-vmescape-security-bug/#comments Sat, 18 Oct 2014 19:24:58 +0000 http://longwhiteclouds.com/?p=7709


VMware has announced that it will turn off TPS in upcoming version of it’s hypervisor ESXi and vCloud Air hybrid cloud service. This is due to a security bug, considered a very rare possibility and only exploitable in very controlled and largely misconfigured environments.  TPS also known as Transparent Page Sharing is a memory management technique that allows multiple […]

]]>


VMware has announced that it will turn off TPS in upcoming version of it’s hypervisor ESXi and vCloud Air hybrid cloud service. This is due to a security bug, considered a very rare possibility and only exploitable in very controlled and largely misconfigured environments.  TPS also known as Transparent Page Sharing is a memory management technique that allows multiple VM’s to share a read only copy of the same memory page. When a VM needs to update or write to a page a new copy is created. The idea is that if there are many VM’s with similar memory pages on the same physical host server it will de-duplicate the pages and only store one copy. The result is that you can run more VM’s per physical server while still achieving very good performance.

TPS has for a long time been used as a competitive advantage by VMware over all of the other hypervisors. But realistically it hasn’t been in wide use by most customers for some time (since ESX 3.5) as the amount of RAM per host has increased, because of the use of large memory pages (2MB instead of 4KB) in Nehalem and above processors, and because most customers don’t want to run their systems at 100% utilization so that they can handle bursts of activity. When using large pages TPS only kicked in when systems were over 96% memory utilization, at which point large pages would be broken down into small pages that could be shared. However this has been a popular technique with service providers and with virtual desktop environments, and in some test and development environments, where over commitment of memory may have been acceptable.

The security problem was found by recent research that leverages Transparent Page Sharing (TPS) to gain unauthorized access to data under certain highly controlled conditions. The research demonstrated that by forcing a flush and reload of cache memory, it is possible to measure memory timings to try and determine an AES encryption key in use on another virtual machine running on the same physical processor of the host server, if Transparent Page Sharing is enabled. This is effectively a VM escape, where code executed within one VM can break the hypervisor isolation and read data from another VM’s memory. Certainly not a good situation if said VM contains credit card data, as we’ve already had enough breaches recently. The conditions under which this could be exploited would be rare in the real world, especially as most environments don’t use TPS actively, even if it is enabled. Even so, I believe in being secure by default, and even though the number of conditions that have to simultaneous by true for this to be exploited would be very rare, if this were exploited the impact could be high. So I believe that VMware is taking the right approach to this research by disabling TPS.

I have been a proponent for leaving TPS enabled in the past, even though a few others have previously recommended it be disabled for performance reasons. My argument was that TPS is a good safety net if all else fails, even if during normal operations it is not used. Also performance was never proven to be a factor. I put this argument in my article Blueprint for Successful Large Scale Oracle Virtualization on vSphere when an EMC paper recommended disabling TPS. To quote that article “Disabling TPS can have disastrous consequences, including causing additional host swapping, which can result in extremely poor performance, much worse than disabling it could ever possibly gain.” So this begs the question, now that it’s being disable by VMware what impact will it have?

Without TPS you will have to have much more conservative memory usage per host. If you business requirements dictate, you will have to be able to sustain maintenance and failure without causing memory overcommitment. If there is a failure or maintenance that causes temporary or prolonged overcommitment of memory you will have a lot more guest OS swapping, due to ballooning, and also host swapping may occur, which would greatly impact performance. Memory swapping is the enemy of performance, and this also adds significantly to poor performance on shared storage if it occurs. But this is possibly better than the alternative security bug.

If you have an existing VMware vSphere environment this will mean you need to evaluate the level of resource usage you have today, your standard operating procedures for maintenance, and the settings of VMware HA Admission Control for failure. If you don’t have sufficient available memory to operate your environment in the case of failure or maintenance, then you may need to upgrade the amount of RAM per host or purchase additional hosts. With any additional hosts you’d need additional licenses. Frank Denneman has a good take on the capacity planning implications in his article here.

TPS will be disabled by default from the following VMware vSphere Releases:

  • ESXi 5.5 Update release – Q1 2015
  • ESXi 5.1 Update release – Q4 2014
  • ESXi 5.0 Update release – Q1 2015
  • The next major version of ESXi

VMware’s official statement on this problem is contained within KB 2080735 Security considerations and disallowing inter-Virtual Machine Transparent Page Sharing. This KB also contains the steps to disable TPS on older versions of VMware vSphere that will not be covered by patches.

If you want to check whether you have TPS enabled or not on your existing versions, and if you want to disable it you can use the following PowerCLI examples (explicitly provided without any warranty, use at your own risk):

 

Check if TPS is Enabled on all hosts connected to a vCenter Server, Mem.ShareScanGHz returns > 0 if enabled.

Connect-VIServer <YourvCenter>
Get-VMHost –State Connected | Get-AdvancedSetting –Name Mem.ShareScanGHz | Format-Table –Property Entity,Name,Value -AutoSize
Disconnect-VIServer

 

Disable TPS on all hosts connected to a vCenter Server by setting Mem.ShareScanGHz = 0, check the setting has been applied correctly

Connect-VIServer <YourvCenter>
Get-VMHost –State Connected | Get-AdvancedSetting –Name Mem.ShareScanGHz | Set-AdvancedSetting –Value 0
Get-VMHost –State Connected | Get-AdvancedSetting –Name Mem.ShareScanGHz | Format-Table –Property Entity,Name,Value -AutoSize
Disconnect-VIServer

 

So if TPS is vulnerable to data leakage and VM escape attacks what about the recently announced Project Fargo, AKA VMFork? VMFork allows a running VM to be quiesced and rapidly cloned by using a similar copy on write technique to share a read only copy of the parent VM memory, and sharing the parent VM’s read only disk, with updates being written to a delta disk. This allows a VM to be cloned and get up and running on the network with it’s own personality in a matter of a few seconds, with the VM memory and disk effectively being deduped at the same time. This doesn’t just have applicability to VDI environments, but web server environments, Dev and Test environments and many other use cases. I’m sure VMware won’t let VMFork out in the wild until issues such as the VM escape bug with TPS are addressed. Kit Colbert, VMware CTO for End User Computing, has said to me that VMFork is much more secure than TPS, so it may not suffer from the same problems.

VMware is not alone with a VM escape vulnerability being discovered. There was also a security bug made public regarding the Xen hypervisor that allowed a VMescape, where code executed within one VM could escape the encapsulation of the hypervisor to a neighbour VM or dom0. This is covered at the VUPEN Vulnerability Research Team’s blog site.

 

Final Word

Nothing is fully secure. You can never guarantee that your system isn’t vulnerable to attack. All you can do is take appropriate measures to reduce the risk of attack, implement technical controls and monitoring and auditing processes. Implement separation of duties, least privilege access, and role based access controls. Implement the guidelines that make sense based on your business requirements from the VMware and other vendors hardening guides. Comply with the security standards for your industry / company that make sense. Stay on top of critical security patches and implement them as soon as practicable, especially for any environments containing public facing or highly secure systems.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/10/19/vmware-turns-off-tps-taps-in-vsphere-esxi-and-vcloud-air-to-avoid-rare-vmescape-security-bug/feed/ 2 7709
Evaluating A HyperConverged Virtualization Platform – Questions To Ask http://longwhiteclouds.com/2014/08/01/evaluating-a-hyperconverged-virtualization-platform-questions-to-ask/ http://longwhiteclouds.com/2014/08/01/evaluating-a-hyperconverged-virtualization-platform-questions-to-ask/#respond Thu, 31 Jul 2014 12:51:31 +0000 http://longwhiteclouds.com/?p=4250


As more and more companies start to look at hyperconverged, web scale (such as Nutanix – where I work) or just converged solutions for your vitualization platform there is a need to make sure you know what your looking at and go in eyes wide open. There are a number of different aspects to evaluate, none of the […]

]]>


As more and more companies start to look at hyperconverged, web scale (such as Nutanix – where I work) or just converged solutions for your vitualization platform there is a need to make sure you know what your looking at and go in eyes wide open. There are a number of different aspects to evaluate, none of the different options are the same. There are no right or wrong answers, as many different solutions may meet your requirements, or be best suited to your requirements. The idea behind this article is to give you a list of questions to consider and to ask any potential vendor. This list is designed to be vendor neutral, and there are really no right or wrong answers, it’s just so you understand what you’re getting for some important aspects, based on my experience.

This list is going to be a starting point and isn’t going to be completely exhaustive. You will have your own priorities and additional questions that need to be answered. Let’s first look at some key aspects of any hyperconverged virtualization platofrm.

  1. Data Protection. Any system in my opinion that doesn’t have data protection as the highest priority has no place in the enterprise. Every system should take great steps to protect data and protect against data loss in common and uncommon failure scenarios. This is not backup, but primary data protection and data resiliency of online in flight data. Data protection is important through all aspects of operations, including upgrades. Make sure that all upgrade scenarios are non-disruptive and non-data destructive.
  2. Availability. Making sure the system reduces single points of failure, can continue operations in spite of component failure (including management components), reduce single points of contention, hot spots, denial of service conditions and includes components that are reliable.
  3. Manageability. Being able to operate, monitor, troubleshoot, provision new components, scale (up or down), provide security and auditability.
  4. Performance. Having enough performance to meet your requirements, being able to adapt performance to meet future requirements, having predictability of performance and reducing and containing impacts of runaway workloads or noisy neighbours.

I have put the above key characteristics of any enterprise hyperconverged solution, in what I believe is the priority order that is most appropriate.

Some additional aspects to consider:

  1. Architecture Effort. How much time do you need to invest in architecting the solution, sizing the initial solution, and determining how the solution would best fit your requirements? If the solution is engineered from the factory this will reduce your effort. Are you expected to be the platform architect and build all the individual components and do all the integration?
  2. Purchasing and Acquisition. Local partner, ease of acquisition, initial units or quantity required, cost alignment to consumption and utilization, purchase on demand, pay as you grow, flexibility to change components or mix and match components after initial purchase based on business or technical requirements.
  3. Delivery and Deployment. Expected delivery timeframes, expected deployment timeframes, expected timeframes and effort required to scale up, or scale out a solution, expected time to productivity.
  4. Support. Integrated support for full solution stack from single or multiple points, spare parts, availability of parts, location of parts, service and response times, location of support, frequency of software updates, consequences for missed SLA’s. Do you really need 24/7/365 with 4hr response or can you settle for a lower service level due to the resiliency, redundancy and reliability of the solution?

Some questions to consider:

  • How is data resiliency and protection achieved?
  • Does it support hypervisor storage acceleration or offload and space saving features (VAAI for vSphere, ODX for Hyper-V)?
  • Does it Support API’s for Data Protection and Backup (Such as VMware VADP)?
  • Are more than one hypervisor supported and if so which ones?
  • Is it possible to change hypervisors for an existing running environment and how much effort is involved?
  • Are multiple hardware platforms supported and if so which ones?
  • How many other environments is the solution deployed into that have similar requirements to yours?
  • What happens when hypervisor HA features have to be disabled for troubleshooting or if HA has to be turned off for some other reason?
  • What happens if the whole cluster needs to be shut down and it runs your key management components (vCenter or SCCM etc)?
  • Is there any ability to encrypt disks or to encrypt communications between components of the architecture?
  • Does it support Fault Tolerance (VMware vSphere)?
  • Does it support Jumbo Virtual Disks (2TB+ virtual disks, VHDX etc)?
  • Is there any way a failure of a single node or component could cause degradation of productive workloads to the point that they become non-responsive?
  • Is there any built in support for DR Replication, backup, snap shots?
  • What happens in the case of an SSD failure?
  • What happens in the case of a hard disk failure?
  • Is the upgrade process non-disruptive and can it be completed without a reboot of the physical host and migration of the virtual machines?
  • Does the upgrade process in any way impact data availability, data protection or data integrity?
  • What happens if a node is unavailable for more than a set period of time minutes (for example half an hour or an hour)?
  • What is the impact on productive end user workloads in the case of a data rebuild / re-protection scenario?
  • How is a failure of a hard disk alerted?
  • How is health of the environment checked, monitored and alerted?
  • What happens if a physical host is rebooted and it has a failed disk?
  • How does it integrate with enterprise backup products?
  • Can Change Block Tracking techniques for incremental backups be used?
  • How are virtual machine recoveries from backups be achieved and can individual files be restored?
  • Can the management infrastructure itself be protected easily for DR, snap shotted and backed up without any third party components or add ons?
  • How easy is it to restore the management infrastructure in the case of a DR event?
  • Does the solution require multicast support on the network for it to work?
  • Does the solution require Jumbo Frames (>1500 byte MTU)?
  • What happens if management components (vCenter or SCCM) is down or completely destroyed and needs to be rebuilt?
  • Where are Hypervisor Core Dump and other troubleshooting data stored?
  • Are any other storage locations required for troubleshooting data (placement of core dumps, logs, scratch locations etc) outside of the hyperconverged storage provided?
  • What happens if different types of management traffic share the same IP subnet?
  • If a decision is made to change hardware platforms or hypervisors in the future can it be done and how easy is it to achieve?
  • How does the solution provide the balance of resource requirements to fit your needs (CPU, RAM, Storage, Network)?
  • What are the units of scale and how is performance impacted when the solution is scaled up or down?
  • Can the solution be scaled down as well as being scaled up and if so can it be done non-disruptively?
  • If your solution is tightly coupled to the hypervisor is there any chance that environments that don’t leverage this feature or solution could be impacted by the code that is coupled to the hypervisor?
  • How many patches have been released for core hypervisor and management components that are a direct result for tight coupling to the hypervisor but not related to core hypervisor functionality?
  • What built in monitoring, management, and alerting capabilities exist out of the box?
  • How does the solution integrate with existing monitoring and management platforms?
  • Can the solution be extended or information be made available to other systems through plug-ins or API’s?
  • If you make a miscalculation or can’t accurately predict the required performance what is the consequences of changing the solution at a later stage and how costly would it be?
  • Does the solution allow you to reduce the dependence on specialised skills and resources and reduce training requirements for your environment?
  • How long does it take to recover / rebuild from various component failure scenarios, such as node, hard disk, SSD etc?
  • Are there any single points of failure?
  • What data services, such as compression, data deduplication are included in the solution and what is the expected performance and use cases for them?
  • How long does it take to clone or provision new virtual machines and how much space do the clones take up?
  • How does the solution prevent a single workload from monopolising all resources?
  • How are points of congestion handled and how is data and network congestion dealt with by the solution?
  • How does the solution limit the impact of component failure or performance problems?

Final Word

Although I put this list together for hyperconverged virtualization platforms and solutions it would apply to many other converged solutions and even traditional architectures. This list is by no means exhaustive and there are many more things to consider. This is just some of the things I could come up with off the top of my head . It would be great if you contribute to this list by providing feedback and comments below on other aspects that are important to consider and questions to be asked.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/08/01/evaluating-a-hyperconverged-virtualization-platform-questions-to-ask/feed/ 0 4250
Virtualizing SQL Server with VMware: Doing IT Right! http://longwhiteclouds.com/2014/07/08/virtualizing-sql-server-with-vmware-doing-it-right/ http://longwhiteclouds.com/2014/07/08/virtualizing-sql-server-with-vmware-doing-it-right/#comments Mon, 07 Jul 2014 14:31:05 +0000 http://longwhiteclouds.com/?p=4121


  The printing presses at VMware Press have been churning out copies of Virtualizing SQL Server with VMware: Doing IT Right, which I’ve co-authored with Michael Corey and Jeff Szastak. We had a great technical reviewer for this book in Mark Achtemichuk (Mark A – VCDX-050). There will be enough copies for everyone, and we’ll be signing […]

]]>


 

9780321927750

The printing presses at VMware Press have been churning out copies of Virtualizing SQL Server with VMware: Doing IT Right, which I’ve co-authored with Michael Corey and Jeff Szastak. We had a great technical reviewer for this book in Mark Achtemichuk (Mark A – VCDX-050). There will be enough copies for everyone, and we’ll be signing them at VMworld(s), and vForum Sydney if you’d like your copy autographed (and hopefully SQL PASS events also). This is 512 pages that combines decades of experience into a book, in a single place, which is the definitive guide for virtualizing SQL Server on VMware. This project took us over 12 months to complete, and many late nights and full weekends. If you ask my family they almost didn’t see me for the entire year. This is the third book project I’ve been involved with, after being technical reviewer for VCDX Bootcamp and Virtualizing and Tuning Large-Scale Java Platforms, both also by VMware Press. Duncan Epping has done us a great honour and written a fantastic foreward for the book.

Although I have dedicated this book to my wife, Susanne, and my four sons, Sebastian, Bradley, Benjamin, and Alexander , for their ongoing support (and putting up with my absence during this project). I’ve also dedicated this book to the VMware Community.

I was also lucky to have some great sounding boards during this project in addition to my fantastic co-authors. Kasim Hansia, VMware Strategic Architect and SAP expert, Cameron Gardiner, Microsoft Senior Program Manager Azure and SQL, and Josh Odgers (VCDX-090), Nutanix Senior Solutions and Performance Architect.

Below I have included the full Table of Contents so you can take a look at what we cover and see the value that we’ve packed into this book. I hope you enjoy the book, get a signed copy, and succeed when virtualizing SQL Server.

 

Foreword xvii

Preface xix

About the Authors xxiii

About the Technical Reviewer xxv

Acknowledgments xxvii

Reader Services xxix

1 Virtualization: The New World Order? 1

Virtualization: The New World Order 1

Virtualization Turns Servers into Pools of Resources 3

Living in the New World Order as a SQL Server DBA 3

A Typical Power Company 6

Summary 7

2 The Business Case for Virtualizing a Database 9

Challenge to Reduce Expenses 9

The Database Administrator (DBA) and Saving Money 10

Service Level Agreements (SLA) and the DBA 11

Avoiding the Good Intention BIOS Setting 12

DBAs’ Top Reasons to Virtualize a Production Database 13

High Availability and Database Virtualization 14

Performance and Database Virtualization 16

Provisioning/DBaaS and Database Virtualization 17

Hardware Refresh and Database Virtualization 20

Is Your Database Too Big to Virtualize? 22

Summary 23

3 Architecting for Performance: The Right Hypervisor 25

What Is a Hypervisor? 25

Hypervisor Is Like an Operating System 26

What Is a Virtual Machine? 28

Paravirtualization 29

The Different Hypervisor Types 29

Type-1 Hypervisor 30

Type-2 Hypervisor 31

Paravirtual SCSI Driver (PVSCSI) and VMXNET3 31

Installation Guidelines for a Virtualized Database 32

It’s About Me, No One Else But Me 33

Virtualized Database: It’s About Us, All of Us 34

DBA Behavior in the Virtual World 34

Shared Environment Means Access to More If You Need It 35

Check It Before You Wreck It 36

Why Full Virtualization Matters 36

Living a DBA’s Worst Nightmare 37

Physical World Is a One-to-One Relationship 38

One-to-One Relationship and Unused Capacity 38

One to Many: The Virtualized World 40

The Right Hypervisor 40

Summary 41

4 Virtualizing SQL Server: Doing IT Right 43

Doing IT Right 43

The Implementation Plan 44

Service-Level Agreements (SLAs), RPOs, and RTOs 45

Baselining the Existing vSphere Infrastructure 46

Baselining the Current Database Workload 48

Bird’s-Eye View: Virtualization Implementation 50

How a Database Virtualization Implementation Is Different 51

Summary 55

5 Architecting for Performance: Design 57

Communication 58

Mutual Understanding 59

The Responsibility Domain 60

Center of Excellence 61

Deployment Design 63

SQL Workload Characterization 64

Putting It Together (or Not) 65

Reorganization 68

Tiered Database Offering 70

Physical Hardware 73

CPU 74

Memory 76

Virtualization Overhead 76

Swapping, Paging? What’s the Difference? 78

Large Pages 79

NUMA 79

Hyper-Threading Technology 85

Memory Overcommitment 87

Reservations 87

SQL Server: Min/Max 90

SQL Server: Lock Pages in Memory 92

Storage 93

Obtain Storage-Specifi c Metrics 94

LSI Logic SAS or PVSCSI 94

Determine Adapter Count and Disk Layout 95

VMDK versus RDM 96

VMDK Provisioning Type 96

Thin Provisioning: vSphere, Array, or Both? 98

Data Stores and VMDKs 99

VMDK File Size 100

Networking 100

Virtual Network Adapter 100

Managing Traffi c Types 101

Back Up the Network 103

Summary 104

6 Architecting for Performance: Storage 105

The Five Key Principles of Database Storage Design 106

Principle 1: Your database is just an extension of your storage 106

Principle 2: Performance is more than underlying storage devices 107

Principle 3: Size for performance before capacity 107

Principle 4: Virtualize, but without compromise 108

Principle 5: Keep it standardized and simple (KISS) 109

SQL Server Database and Guest OS Storage Design 109

SQL Server Database File Layout 110

Number of Database Files 110

Size of Database Files 114

Instant File Initialization 120

SQL Server File System Layout 122

SQL Server Buffer Pool Impact on Storage Performance 129

Updating Database Statistics 130

Data Compression and Column Storage 132

Database Availability Design Impacts on Storage Performance 135

Volume Managers and Storage Spaces 136

SQL Server Virtual Machine Storage Design 136

Virtual Machine Hardware Version 137

Choosing the Right Virtual Storage Controller 138

Choosing the Right Virtual Disk Device 143

SQL Virtual Machine Storage Layout 152

Expanding SQL Virtual Machine Storage 158

Jumbo VMDK Implications for SQL Server 159

vSphere Storage Design for Maximum SQL Performance 164

Number of Data Stores and Data Store Queues 165

Number of Virtual Disks per Data Store 170

Storage IO Control—Eliminating the Noisy Neighbor 173

vSphere Storage Policies and Storage DRS 177

vSphere Storage Multipathing 184

vSphere 5.5 Failover Clustering Enhancements 185

RAID Penalties and Economics 187

SQL Performance with Server-Side Flash Acceleration 198

VMware vSphere Flash Read Cache (vFRC) 199

Fusion-io ioTurbine 201

PernixData FVP 204

SQL Server on Hyperconverged Infrastructure 207

Summary 213

7 Architecting for Performance: Memory 217

Memory 218

Memory Trends and the Stack 218

Database Buffer Pool and Database Pages 219

Database Indexes 222

Host Memory and VM Memory 225

Mixed Workload Environment with Memory Reservations 226

Transparent Page Sharing 228

Internet Myth: Disable Memory TPS 229

Memory Ballooning 230

Why the Balloon Driver Must Run on Each Individual VM 232

Memory Reservation 232

Memory Reservation: VMware HA Strict Admission Control 233

Memory Reservations and the vswap File 233

SQL Server Max Server Memory 234

SQL Server Max Server Memory: Common Misperception 235

Formula for Confi guring Max Server Memory 236

Large Pages 237

What Is a Large Page? 237

Large Pages Being Broken Down 238

Lock Pages in Memory 239

How to Lock Pages in Memory 241

Non-Uniform Memory Access (NUMA) 241

vNUMA 243

Sizing the Individual VMs 244

More VMs, More Database Instances 244

Thinking Differently in the Shared-Resource World 246

SQL Server 2014 In-Memory Built In 246

Summary 247

8 Architecting for Performance: Network 249

SQL Server and Guest OS Network Design 250

Choosing the Best Virtual Network Adapter 250

Virtual Network Adapter Tuning 252

Windows Failover Cluster Network Settings 254

Jumbo Frames 256

Configuring Jumbo Frames 259

Testing Jumbo Frames 262

VMware vSphere Network Design 264

Virtual Switches 265

Number of Physical Network Adapters 267

Network Teaming and Failover 270

Network I/O Control 274

Multi-NIC vMotion 276

Storage Network and Storage Protocol 279

Network Virtualization and Network Security 281

Summary 286

9 Architecting for Availability: Choosing the Right Solution 287

Determining Availability Requirements 287

Providing a Menu 288

SLAs, RPOs, and RTOs 290

Business Continuity vs. Disaster Recovery 291

Business Continuity 291

Disaster Recovery 291

Disaster Recovery as a Service 292

vSphere High Availability 294

Hypervisor Availability Features 294

vMotion 296

Distributed Resource Scheduler (DRS) 297

Storage vMotion 297

Storage DRS 297

Enhanced vMotion X-vMotion 298

vSphere HA 298

vSphere App HA 299

vSphere Data Protection 300

vSphere Replication 300

vCenter Site Recovery Manager 301

VMware vCloud Hybrid Service 302

Microsoft Windows and SQL Server High Availability 302

ACID 302

SQL Server AlwaysOn Failover Cluster Instance 304

SQL Server AlwaysOn Availability Groups 306

Putting Together Your High Availability Solution 308

Summary 310

10 How to Baseline Your Physical SQL Server System 311

What Is a Performance Baseline? 312

Difference Between Performance Baseline and Benchmarks 315

Using Your Baseline and Your Benchmark to Validate Performance 318

Why Should You Take a Performance Baseline? 319

When Should You Baseline Performance? 320

What System Components to Baseline 320

Existing Physical Database Infrastructure 321

Database Application Performance 323

Existing or Proposed vSphere Infrastructure 325

Comparing Baselines of Different Processor Types and Generations 328

Comparing Different System Processor Types 328

Comparing Similar System Processor Types Across Generations 330

Non-Production Workload Influences on Performance 331

Producing a Baseline Performance Report 332

Performance Traps to Watch Out For 333

Shared Core Infrastructure Between Production and Non-Production 333

Invalid Assumptions Leading to Invalid Conclusions 334

Lack of Background Noise 334

Failure to Considering Single Compute Unit Performance 335

Blended Peaks of Multiple Systems 335

vMotion Slot Sizes of Monster Database Virtual Machines 336

Summary 337

Contents

11 Configuring a Performance Test—From Beginning to End 339

Introduction 339

What We Used—Software 341

What You Will Need—Computer Names and IP Addresses 341

Additional Items for Consideration 342

Getting the Lab Up and Running 342

VMDK File Configuration 345

VMDK File Configuration Inside Guest Operating System 352

Memory Reservations 355

Enabling Hot Add Memory and Hot Add CPU 356

Affinity and Anti-Affinity Rules 358

Validate the Network Connections 359

Configuring Windows Failover Clustering 359

Setting Up the Clusters 362

Validate Cluster Network Configuration 368

Changing Windows Failover Cluster Quorum Mode 369

Installing SQL Server 2012 374

Configuration of SQL Server 2012 AlwaysOn Availability Groups 387

Configuring the Min/Max Setting for SQL Server 392

Enabling Jumbo Frames 393

Creating Multiple tempdb Files 394

Creating a Test Database 396

Creating the AlwaysOn Availability Group 399

Installing and Configuring Dell DVD Store 406

Running the Dell DVD Store Load Test 430

Summary 436

Appendix A Additional Resources 437

 

Note: Guidance specific to number of data files has been extensively reviewed and is based on extensive implementation experience as the lowest common denominator of many versions of SQL Server that the book covers, from SQL 2000 onwards, which are in common use across the globe in enterprises today. This could be considered overly conservative for the most recent versions of SQL (such as SQL 2014) and you should determine the best approach based on your specific circumstances and requirements. When we refer in the book to keeping a backup of the transaction log on SAN or alternative storage to local flash this can easily be achieved with a transaction log backup – see msdn.microsoft.com/en-us/library/ms191429.aspx. There is an error on page 309 where the failover time of a cluster is listed as 3 seconds. A 0 (zero) was omitted and the failover time should read 30 seconds. We apologise for this error and it will be corrected.

 

 

Final Word

We put a lot of time and effort into this project in the hope that you get a lot of value out of it. Your comments and feedback are as always welcomed. Our book is now generally available in both paperback and eBook form. If you like what you see in the table of contents then please consider purchasing your copy today.

This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/07/08/virtualizing-sql-server-with-vmware-doing-it-right/feed/ 9 4121
VMware VCAP-CIA Exam Experience and Recommendations http://longwhiteclouds.com/2014/05/24/vmware-vcap-cia-exam-experience-and-recommendations/ http://longwhiteclouds.com/2014/05/24/vmware-vcap-cia-exam-experience-and-recommendations/#comments Sat, 24 May 2014 04:52:30 +0000 http://longwhiteclouds.com/?p=3636


On Wednesday 14th May (NZST) I sat the VMware Advanced Professional – Cloud Infrastructure Administration (VCAP-CIA) exam. After receiving my results on Saturday  17th, only 3 days after sitting the exam, I was very relieved to have passed. I would like to thank VMware and the hard work by people including Joshua Andrews and the certification […]

]]>


On Wednesday 14th May (NZST) I sat the VMware Advanced Professional – Cloud Infrastructure Administration (VCAP-CIA) exam. After receiving my results on Saturday  17th, only 3 days after sitting the exam, I was very relieved to have passed. I would like to thank VMware and the hard work by people including Joshua Andrews and the certification team for getting the results through so fast. This is a massive improvement over previous advanced live lab exams. I had already passed VCAP-CID during the exam beta process, so for me this puts me one step closer to VCDX-Cloud, which is my goal. This article will cover my exam experience and tips and recommendations for others that wish to attempt the exam.

To give you some background, I have architected about half a dozen cloud environments prior to sitting VCAP-CID and VCAP-CIA. Half were enterprise clouds and half were public clouds based on VMware vCloud Director or VMware vSphere. So I had quite a bit of experience with the solution stack that makes up a VMware Cloud environment, in addition to my VCDX Datacenter Virtualization. But I still try to approach any exam with the same methodology.

Exam Prep Method

My usual exam prep method is to review the exam blueprint and do an analysis of where I’m strong and where I need to study further to increase my knowledge. This exam was no different. There are quite a lot of areas to cover as there are a lot of technologies that are included in VMware’s cloud stack and with VMware vCloud Director. I found some additional areas to study and I made sure I reviewed the documentation, reviewed the vCloud Architecture Toolkit (vCAT), and set up a lab environment (using virtual ESXi hosts) to test things and get hands on. It’s important to be familiar with the command line inside your vCD Cells just as it is knowing how to navigate the GUI’s of the various tools.  So make sure you’re prepared. Know the VCD cells intimately, the logs, the config files, the config tools, troubleshooting. Know vApps, Networking, configuration, resources and all the underlying infrastructure like the back of your hand. You will have to fix some broken things and you don’t have much time. vCD lends itself to this type of exam, and they cover the blueprint areas well. When it comes to your lab environment make sure you have all of the vCD stack components, including Chargeback, vCNS with VXLAN, vCD etc.

My First Attempt

This is a tale of how not to do it. I arrive back in Auckland, New Zealand after a business trip at around midnight with the VCAP-CIA exam the following morning. I got about two hours sleep as my youngest son wasn’t sleeping well. I didn’t read the first question correctly and instead of having to fix something, I really fixed it by corrupting the cell. I tried for about 20 minutes to troubleshoot it and get it working. I eventually gave up and just went as fast as I could through the questions I did know. Unfortunately time was very much against me and I wasn’t thinking straight due to having little sleep. Tip: Get plenty of sleep the night before the exam and read the questions carefully.

My Second Attempt

Fortunately when I failed the exam the first time there was a free retry voucher available. So I booked the exam, but due to work commitments had to keep putting it off as I would never get time to study it correctly. 14th May 2014 I sat the exam again. This time I had gotten a good rest the night before and I was much more prepared. But unfortunately the connection from the testing centre wasn’t as prepared as I was. Latency was my enemy, as was the speed of the vCD environment responding to my inputs. I performed the required tasks as fast as I could but I found I could only think about two or three questions at a time so I couldn’t get too far ahead. I had been moving forward and backward to do multiple questions at the same time to try and compensate for the performance of the vCD environment and also the latency challenges. But this strategy can only achieve so much. You have to keep track of all these questions in your head and where each part of the environment is at. In my opinion this makes the exam very challenging, especially with such limited time. By the time I had run out of time I had not even attempted about 5 or 6 questions, so I was just hoping I’d done enough in the rest of the exam to get a pass.

The Results

4 days was all I had to wait for the official results and the news that I’d passed the VCAP-CIA exam. One of the guys behind the test scoring automation reached out to me to let me know that there had been a lot of work going on behind the scenes on automating the exam scoring. It still wasn’t as fast as they’d like, so they were still looking to improve it. Given this quick turnaround the torture that is the advanced exams is made quite a bit better. At least you know where you’re at a lot sooner. In the future they plan to make the results much faster. So I hope you get an even faster experience when you sit the VCAP exams. I can also tell you there are plans to improve the interface to largely eliminate the lags and delays that we experience when doing these live lab exams from all over the world. This is great news, as the exams are hard enough without the latency, and VMware need more people sitting advanced exams to build a pipeline for VCDX and to ensure customers have access to skilled professionals to provide integrated solutions across all of VMware and ecosystem partner solutions.

Update: Joshua Andrews has just posted an article regarding the great improvements in exam marking time that VMware has implemented. Now you might receive your score for DCA and DTA in the same day. CIA will also see improvements. Check out the article titled Improved VCAP DCA DTA Score Reporting.

Final Word

Overall I really like the live lab exam format and the way they deeply test your abilities to administer and troubleshoot an advanced environment. The latency does make it hard, but it is possible to pass with the right approach, which hopefully this article helps you with. vCD is still my preferred Cloud tool while it’s supported as it solves a lot of problems that vCAC doesn’t yet, and it’s also more relevant for service providers, which is a market I understand after working at an ISP for so many years and architecting a few public clouds. I would encourage people who are in an organisation with vCD or working for a Cloud Service Provider to give VCAP-CIA a go. My plan now is to brush up a vCloud design I did a while ago and submit it for VCDX-Cloud. If I’m successful hopefully I’ll be in the first five double VCDX’s.

This post appeared on the Long White Virtual Clouds blog at longwhiteclouds.comby Michael Webster +. Copyright © 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/05/24/vmware-vcap-cia-exam-experience-and-recommendations/feed/ 2 3636
Master the Monster VM Performance in Your Software-Defined Datacenter at VMworld 2014 http://longwhiteclouds.com/2014/05/06/master-the-monster-vm-performance-in-your-software-defined-datacenter-at-vmworld-2014/ http://longwhiteclouds.com/2014/05/06/master-the-monster-vm-performance-in-your-software-defined-datacenter-at-vmworld-2014/#comments Tue, 06 May 2014 09:55:18 +0000 http://longwhiteclouds.com/?p=3494


At VMworld in 2013, which had 25,000 attendees, it was fantastic that one of my sessions got into the top 10 sessions of the entire event. With big thanks to my co-speakers Andrew Mitchell, Frank Denneman, Mark Achtemichuk, Emad Benjamin and Mostafa Khalil on the VAPP4679 – Software Defined Datacenter Design Panel for Monster VM’s. The audience agreed […]

]]>


At VMworld in 2013, which had 25,000 attendees, it was fantastic that one of my sessions got into the top 10 sessions of the entire event. With big thanks to my co-speakers Andrew Mitchell, Frank Denneman, Mark Achtemichuk, Emad Benjamin and Mostafa Khalil on the VAPP4679 – Software Defined Datacenter Design Panel for Monster VM’s. The audience agreed we took the Technology to the Limits for High Performance, High Utilization Workloads. We had some excellent questions and feedback during this event. This year I am upsizing the Monsters and brining you even more of what you asked for. Performance data, design and architecture guidance, Monster VM sink hole avoidance strategies. The Monster VM panel this year is super sized and I’m lucky to be joined in this session with Frank Denneman, Mostafa Kahlil, Josh Odgers. If you liked what you saw last year, which you can watch the recording of below, then you’ll love what we have for you this year. We’re certainly aiming to be in the top 10 sessions again of VMworld 2014 (assuming we are selected of course). Your support really helps.

VMworld 2013 VAPP4679 – Software-Defined Datacenter Design Panel for Monster VM’s

You can see all of the Top 10 Sessions of VMworld 2013 here.

 

VMworld 2014 Monster VM Sessions:  

Here are three Monster VM focused sessions that I hope you’ll like and vote for. You can vote for your favourite sessions on the VMworld Session Voting web site.

Session 1302 – Mastering Monster VM Performance in Your Software-Defined Datacenter

Top 10 Session at VMworld 2013, now even more of a Super-Sized Monster for 2014. Master your Monster VM Performance. Master your Software-Defined Datacenter. Designing and architecting a Software-Defined Datacenter that will host Monster VM’s requires a different approach and different considerations to smaller VM’s to ensure you get the best benefits from the platform. Mostafa Khalil (VCDX-002), Frank Denneman (VCDX-029), Michael Webster (VCDX-066), and Josh Odgers (VCDX-090) will lead a panel discussion the key best practices that will allow you to master design and performance of many high utilization, high performance Monster VM’s. The panel will share real world performance data and answer your questions on topics such as host and cluster design and sizing, VM sizing, performance tuning, CPU scheduling, memory management, resource management, storage design and much more. Come prepared with your toughest Monster VM questions and a willingness to share you experiences.

  • Michael Webster, Sr Solutions and Performance Engineer, VCDX-066, Nutanix
  • Mostafa Khalil, Director, VMware Technical Support, VCDX-002, VMware
  • Frank Denneman, Technology Evangalist, VCDX-029, PernixData
  • Josh Odgers, Sr Solutions and Performance Engineer, VCDX-090, Nutanix

Session 1729 – 47 Things  to Know About Database Virtualization

This advanced technical, no fluff session will cover 47 unique items that individuals looking to virtualize their databases should know. The session will provide items that apply for SQL and Oracle databases running on vSphere. Presenting the session will be VMworld veterans, co-authors of Virtualizing SQL Server on VMware: Doing IT Right (VMware Press), and database virtualization experts Michael Corey, Michael Webster (VCDX 066), and Jeff Szastak. 
Topics covered will be, but not limited to, physical CPU virtual CPU, NUMA, physical RAM, virtual RAM, PVSCI tuning, storage (lots of storage), physical server configuration, ESXi configurations, vSphere Cluster design, OS tuning, database specific tuning, VMXNET tuning, jumbo frames, and more!

  • Jeff Szastak, Staff Systems Engineer, VMware
  • Michael Webster, Sr. Solutions & Performance Engineer, VCDX-066, Nutanix
  • Michael Corey, President, NTirety
  • Don Sullivan, Staff Systems Engineer, Database Specialist, VMware

1376 Temenos T24 Migrated from AIX to vSphere 5.1 and SQL Server 2012

This sessions would be one I recommend. Termenos is a core banking system used throughout the world. See how Rene helped Bank AlBilad virtualize their core banking platform on VMware vSphere with a Mosnter VM SQL Database. I was fortunate to be able to work with Rene on this project specifically with SQL Server and vSphere tuning.

  • Rene Van Den Bedem, VCDX-133, CIO Advisor, Bank AlBilad

 

Non-Monster VM sessions I’m involved with:

Session 1600 – Art of IT Infrastructure Design, The Way of the VCDX – PanelNon-Monster VM Sessions:

This session covers IT Infrastructure Design Methodology as practiced by both VCAP Design and VCDX certified architects. The session covers the what, how, and why of the methodology. Examples are provided and the panel discusses their experience as panelists reviewing and validating hundreds of designs. Case studies are presented for designs with customers. One example demonstrates design success. The other example demonstrates design failure. The panel discusses and contrasts the approaches of both.

  • Mark Gabryjelski, VCDX-023, Worldcom Exchange, Inc.
  • John Arrasjid, Principal Architect, VCDX-001, VMware
  • Chris McCain, Director, VCDX-079, VMware
  • Michael Webster, Sr. Solutions and Performance Engineer, VCDX-066, Nutanix
  • Mostafa Khalil, Director, VCDX-002, VMware

Session 2631 – The Way of the VCDX, Testing Architects on IT Infrastructure Design

  • John Arrasjid, Principal Architect, VCDX-001, VMware
  • Fabio Raposelli, Sr. Consultant, VCDX-058, VMware
  • Matt Vandenbeld, Solutions Architect, VCDX-, VMware
  • Michael Webster, Sr Solutions and Performance Engineer, VCDX-066, Nutanix
  • Mark Brunstad, VCDX Program Manager, VMware

 

Final Word

Please vote for all your favourite session. It helps the selection committee make sure they have enough of the content that you want.  Vote Here, Vote Now!

This post appeared on the Long White Virtual Clouds blog at longwhiteclouds.comby Michael Webster +. Copyright © 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.


]]>
http://longwhiteclouds.com/2014/05/06/master-the-monster-vm-performance-in-your-software-defined-datacenter-at-vmworld-2014/feed/ 1 3494