| Unique Visitors |
Customers love choice, and with Nutanix Enterprise Cloud platform, Customers’ can have choice and also simplicity. This is something that has been very difficult before now. This week Nutanix announced support of XenServer on its Enterprise Cloud Platform. This adds a 4th Hypervisor for on premises use and will benefit those customers wishing to have an option for a full Citrix integrated stack. Without compromising the simplicity, manageability, serviceability, low TCO and overall customer delight that Nutanix has become known.
Here is a tweet from Sudheesh Nair, President of Nutanix, about the announcement. There will definitely be more to come as well as we head towards Nutanix .Next User Conference in Vienna from 8th to 10th November.
https://twitter.com/sudheenair/status/785822620679671808
Nutanix and Citrix have built a close partnership together for VDI and this is one of the results of that partnership. The Citrix announcement of the availability of the XenServer Tech Preview on Nutanix Enterprise Cloud can be found here.
Now customers can have the full Citrix VDI experience they love, which will support GPU, PVS, NetScaler and all the goodness they want, all with Citrix XenServer on the Nutanix Enterprise Cloud.
http://www.nutanix.com/2016/10/11/putting-focus-end-users-citrix-xenserver-nutanix-enterprise-cloud-platform/
Final Word
Nutanix aims to bring more choice with greater simplicity and lower TCO to your datacenter so you can run your applications and focus on your applications while your infrastructure is invisible. The result is a better quality of life for your IT people and a better business outcome for your organization. Nutanix is really only just beginning to scratch the surface of the possibilities form this platform, which melds an AWS or Azure like consumption model with the security and controls of an on premises infrastructure. Allowing a real hybrid choice and freeing you to make better business decisions and not be tied to any one place.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2016 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
Usually when you go up against Oracle it is a herculean effort, a David vs Goliath battle. By being prepared, even if you are the David, you can prevail. When it comes to Oracle licensing there is a lot of FUD as I have written about many times before. You need to read your contract very carefully and understand it and in most cases it is a good idea to get independent legal advice. The consequences of a configuration error or mistake can be very high. What is in the contract isn’t as interesting as what is not in the contract. There is no mention of virtualization, clustering (other than RAC), partitioning etc. There is actually nothing in the contract that prevents you from configuring a VM with just 2 vCPU’s and provided you can prove the software never used more than 2 vCPU’s, only licensing those 2 vCPU’s. But nobody has had the courage to configure just 2 vCPU’s for an Oracle system and pay only the license based on those 2 vCPU’s in a larger environment. They would have had the partitioning guide thrown at them (not contractural and not referred to in the contract, for education purposes only). That is until now!
Daniel Hesselink of License Consulting, a global licensing consulting business with HQ based in the Netherland, is putting his own money where his mouth is. He is taking on Oracle, or hassling Oracle as I put in the title. He recently wrote about an article where he has purchased a 2 vCPU VM in a VPS service and has installed a licensed version of Oracle on the system. As there is nothing against this in the contract, and he can prove he has only ever paid for 2 vCPU’s, there is no way Oracle can say that their software has run anywhere else. This could open up the flood gates of Oracle software in cloud environments. It would also allow for pay as you grow licensing across any hypervisor and any cloud.
So the question is, can you do the same thing? The answer depends on your contract and the wording you have signed up to. Each contract may be slightly different and contain custom wording. You should seek independent advice based on your individual circumstances. Maybe talk to Daniel as this is his business and he’s been doing it successfully for years. If you do decide to go down this road, and do it properly with proper documentation, operational process, audit controls etc, I’d love to hear about your experience, I’m sure Daniel would too.
Final Word
It’s time to take the power and control back over your business and your use of Oracle software. By allowing a pay as you grow model across any cloud environment Oracle software usage would likely skyrocket. Few people have a problem with Oracle technology, most people have a problem with their licensing practices. Their use of their dominant market position to try and get customers on their platform by offering special privileges that are not allowed on other platforms. Only you as their customer can start to turn the tide. But you have to have the courage, the right advice, and the right processes in place to do it safely. You also have to be willing to stand up to Oracle during the inevitable audit when it arrives.
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2016 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
Another VMworld event is over and it’s hard to believe it’s been a whole 12 months since the last one. Certainly during the keynotes there was a lot of coverage about what VMware has achieved over the last 12 months and it is impressive especially in the end user computing and hybrid cloud spaces. But overall I felt that VMworld USA 2014 lacked some of the sparkle of last year. But I guess it’s hard to top last year considering it was the 10th anniversary. This year seemed much more about building a solid foundation for a software defined datacenter, a software defined enterprise and a hybrid cloud model integrating applications with infrastructure, providing ability and flexibility, but without compromise. Although attendance was flat or a little down on last year the breakout sessions were packed, right up to the last session on Thursday. Instead of having our heads in the clouds this year it was all about the vCloud Air, and we vRealized the product naming is about to be changing. So lets dive into what I think are some of the highlights.
My VMworld started on Sunday with a Nutanix sponsored VCDX study group. Nutanix is a big supporter of the VCDX program for the entire community. The study group was put on for candidates that wanted to know more about the VCDX process and practice the design and troubleshooting scenarios. It was completely vendor agnostic, and it needs to stay that way. Nutanix understands that the only way sponsoring a VCDX study group can be of value is if the content is vendor agnostic and covers a wide range of topics. There were many VCDX helping in the room and giving advice from across many companies. This really is what the community is all about. Everyone helping each other.
Then I moved on to opening acts at VMunderground that was put on by vBrownBag. I was on the storage panel and it was a good discussion around Virtual Volumes, Hyper Convergence and Flash. I even agreed with a traditional SAN vendor that hyper converged appliances will not help SuperDomes and Mainframes, but then again I can always migrate the workloads and processes off those systems, and the Unix mid range systems as well, to a Hyper Converged world. SuperDomes, Mainframes and Unix systems is where the legacy SAN technology will stay for the foreseeable future and it will be a decline over a number of years, just like we’ve seen with the traditional big iron systems themselves. The move away from traditional SAN for x86 connected environments isn’t going to happen over night, but it’s a trend that is starting to take hold, but honestly it’s not even scratching the surface of the potential opportunity yet. The announcements from VMware and EMC around their hyper-converged offerings are just more validation of that. Flash is definitely the way of the future, and it opens up things that were previously not possible. I have a section on flash technology in the storage chapter of Virtualizing SQL Server with VMware.
There were a number of VMware announcements during the keynotes that are worth mentioning. But before I do I have to get something off my chest. vRealise is the worst name ever thought of for anything. My initial reaction to the new name for VMware’s hybrid cloud, vCloud Air was somewhat similar, but at least Air has a cool ring to it, like iPad Air for example. vRealize, just NO! I feel sorry for the sales team who have to try and sell that now. Ok, rant over. The overall themes about this being a brave new world and requiring bravery from all of the customers and the community participants was interesting. I’ve been doing virtualization for a very long time and even for business critical apps it’s a very safe bet. But SDDC and the Software Defined Enterprise going to further reduce silos and this will require some organisational changes and maturity. This is really where the bravery comes in, most of the challenges are not technical.
Two major overall themes were used during the keynotes. Firstly – “Compatibility, Compliance, Choice”, and secondly the “Power of &”. VMware has done a great job of building a partner ecosystem across a number of technologies, including the vCloud Air Network, which has 3900 partners, and the broader ecosystem around the hypervisor and NSX. This is where compatibility, compliance and choice really comes in. Seamless compatibility, compliance with regulatory and industry requirements, and choice of multiple partners and technologies. This is then extended to the OpenStack, NSX and Containers, which can run extremely well in a VMware environment, and this is the Power of &. Have your containers without compromise. Have your OpenStack on a platform easily fit on top of VMware vSphere.
By far the biggest highlight was meeting a lot of people who regularly read my blog and have benefited from the work that I and others in the community have done over the years. This is why we keep doing it. Because it makes a difference. It was also great to meet a lot of people who had bought Virtualizing SQL Server with VMware: Doing IT Right, and had got a lot out of it too. My co-authors, Michael Corey, Jeff Szastak and I were blown away by the stories that were relayed to us about how the book had helped people, especially when it was being used to explain to DBA’s how virtualization works and that SQL is a great candidate for virtualization. The book was so popular that it actually sold out at VMworld, and we had a lot of people come up to us during the meet the authors session and book signing.
Here is a photo of my co-authors and I with the happy customer who purchased the very last copy of our book at VMworld.
Shortly after the above photo was Michael Corey and I recorded an interview with VMworld TV’s Eric Sloof regarding virtualizing SQL Server Databases on VMware.
I was lucky again this year to present a session that was included in the top 10 sessions of VMworld for the second consecutive year – VMworld 2014 SDDC1600 Art of IT Infrastructure Design The Way of the VCDX Panel.
Final Word
It was another great VMworld and a very successful VMworld. I’m very much looking forward to next year. Hopefully we’ll see the return of the Monster VM sessions and some other business critical apps sessions from me in next years VMworld.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com. By Michael Webster +. Copyright © 2012 – 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
Today we experienced the second outage to this blog in the last 12 months. The outage today was 15 hours in total, unacceptably long in my view for a professional hosting company. Very inconvenient to me as I’m sure it was to all of you as well. But this does serve as a good reminder about the maxims of cloud computing. Mind you the maxims don’t just apply to cloud computing, they are general system design principles that need to be considered and dealt with. But particularly so in Cloud Computing, Outsourcing or Hosting when you are delegating some form of control over your systems to a third party. Maxims are generally established propositions or principles, at least if you’re talking about legal maxims, which I find fascinating to read. But in the context of Cloud Computing the title of this article gives them away. Hardware Fails, Software Has Bugs and People Make Mistakes. It’s up to us to deal with these, as they are guaranteed to happen. This article covers some lessons that can be learned form this recent experience in the hope that we can all have higher availability.
High Availability Doesn’t Seem So High When It Happens At The Wrong Time
Even though this was a 15 hour outage, if this was the only outage this year (and I hope it is), the hosting provider availability would still be 99.8%. That’s quite high isn’t it? Well it depends. It depends what the impact on your system is. The impact will likely depend on the type of business you operate and also the type of function the system performs for your business. It will also depend on the time of month that the outage occurs.
I had a customer example once where a 3 day outage happened at exactly the wrong time. The impact of this outage resulted in a loss in the hundreds of millions of dollars. It worked out to something like $10million per hour of downtime. Fortunately although this impact was incredibly significant for this organisation it didn’t cripple it to the point of extinction. But even with this three day outage, that happened at the wrong time, the system availability was still measured at 99%. The point is you can’t just take an annualized availability metric on it’s own. You must understand the context, impact, and other variables.
Annualized system availability is one thing, maximum tolerable downtime might be something completely different. For example, if you have an availability SLA of 99.5%, on an annual basis you are accepting 1.83days of downtime. But what happens if this is during the busy Christmas retail period and you’re running an online store expecting a lot of customers? Almost 2 days downtime would be catastrophic to sales, and could potentially put you out of business. So it’s unacceptable for all that downtime to happen in one incident, perhaps your maximum tolerable downtime is 2 hours or 4 hours. Even this would be a lot, but it might be tolerable. This means that your system must be able to recover from any failure and be operational again within 4 hours of any individual incident, no matter what has failed or for what reason, even if over the course of a year you accept it might be unavailable for almost 2 days in total. These sort of metrics in your SLA’s are absolutely critical if you are to achieve from your suppliers what you actually expect.
The Maxims of Cloud In Action
In the case of this site I have no such SLA for MTD. However my hosting provider Bluehost.com does have a network availability guarantee. They guarantee 99.9% availability. They have multiple levels of redundancy, skilled staff, their own datacenters etc. So what happened to cause a second outage in 12 months and in this case a 15 hour outage in a single incident?
I suspect there are multiple reasons and hopefully a full root cause analysis is made available to customers, along with a preventative action plan that gives customers confidence that steps are being taken to reduce the risk of another similar incident. This post incident process is standard procedure in many environments and the transparency helps everyone learn from the incident and work together to prevent it happening again.
The information that has been made available suggests that a firmware bug in the networking equipment caused the outage. Because so much of our hardware is now software, or software controlled (software-defined), bugs can have disastrous cascading impacts. However not only is the actual bug a problem (and some of them can occur after a period of time, long after an update), but the complexity of the systems involved and the troubleshooting process that you have to go through to even determine this is a bug rather than something else, is immense. Complex systems in and of themselves can produce a higher probability of downtime and also increase the duration and impact of downtime.
Usually this level of complex troubleshooting and systems outage has a contributing human error element. It’s just natural and happens in almost all environments at some time or another. We’re all human and we all make mistakes some time. This is just something that has to be accepted and we have to put measures in place to reduce the risk and the impact of making mistakes. This is where testing and validation become so important in non-production environments.
This 15 hour outage is a breach of the Bluehost.com SLA’s, even if the first outage wasn’t. Their network availability from this one incident has reduced on an annualized basis to roughly 99.8%.
Lessons and Takeaways
Validation of all system components is important to ensure they function as expected, fail as expected, and any updates are properly vetted prior to going into production. Knowing how a system behaves when everything is going well is not enough, you need to thoroughly understand what happens when things go wrong. But even this will not completely eliminate all incidents. The systems must be designed to be able to tolerate incidents and be restored within the SLA’s set by the systems. With critical systems where impacts must be limited staggered updates may be required to ensure that not all system components can be impacted at one point in time.
Complex systems are hard to troubleshoot, can increase the probability of downtime, and can increase the duration of downtime. Keep everything as simple as possible, and no more complex than absolutely necessary to meet your requirements.
Understand your availability requirements and impacts, not just annualized downtime, and the impact associated with downtime at different times of your business cycle. A thorough availability impact analysis and requirements should be part of every system design.
Hardware fails, software has bugs and people make mistakes. This is a reality, you need to plan for it. It’s up to you to put measures in place to ensure your system meets availability requirements. This means reducing single points of failure, which may include not relying on a single organizations to run or host your systems. This could be as simple as having a primary site and a backup site that can be activated by a simple DNS change. The measures appropriate for your situation need to be financially justifiable, and based on a proper risk analysis. Another good saying is that ‘Hardware eventually fails, and software eventually works’.
Clearly defined and understood standard operating procedures, troubleshooting procedures, and communications plans can help a lot in a crisis. But you also need to have well trained staff that know all of these procedures and have them tested regularly. Having them written down and not part of the organisation culture is no good.
Your hardware is now software, it probably has bugs. Now that your hardware has so much software, is software controlled or software defined, it probably has lots of bugs. This means you not only have to test your applications software, but you have to test your hardware as well.
Use the concept of availability domains to limit risk. An availability domain is basically the same thing as a failure domain. It’s conceptually a way to limit the risk of an outage or incident to a subset of system components. The idea being that not all of the system will fail at the same time if it’s split or separated into more than one availability domain. If you design availability domains into your systems you can increase availability. Even with availability domains, in almost all systems there is always a few single points of failure that could still impact availability, you just have to ensure the probability is as low as possible and the risks are understood and properly managed.
Transparency and owning up to your mistakes is the best policy in a crisis. Transparent, direct and frank communication can buy you a lot of good will in a crisis and allow you the time you need to get to resolution. Transparency after the crisis is over is also important. Handling a disaster or a crisis well can enable you to retain customers you would have otherwise lost. Having transparent root cause analysis and preventative action plans can build confidence in your customers so they know you’ve learned from the experience and it is much less likely to happen again. This is not just a technical or IT problem, this is a business problem. The communications plan and transparency needs to be organisation wide.
Final Word
Although I’m disappointed about the outage I understand the complexity involved in troubleshooting and resolving something like this. I also take responsibility as I know the buck stops with me. The decision to host with this provider was mine. I could host with a different provider or with multiple providers. But unfortunately that’s not economic. Sometimes you just have to accept a risk, which is hopefully reduced by vetting and good SLA’s. But I apologise to you, my readers, for the site being down for 15 hours, this outage was unacceptable. I hope we can all learn some lessons from this experience, my hosting provider included, and that we can provide much better availability to our customers as a result. As always your feedback is appreciated.
—
This post appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2014 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
This week I’m brining you this article from my hotel in Shanghai, I’m in China to present Unix to VMware Migration workshops. Last week I made the treck to Las Vegas like around another 12,000 people for the annual HP Discover US Event (June 11th – 13th). This was my first time to HP Discover and I was very grateful to HP and Ivy Worldwide for making this opportunity possible. HP Discover is not like any other IT event that I’ve been to and HP certainly know how to put on a magnificent show. Almost all of the Sands Expo Center at the Venetian was taken up by the exhibition stands of different HP divisions with the remainder for Sponsors. In terms of number of attendees it’s around half the size of VMworld, but lacks nothing in terms of spectacle. The big themes of the event were aimed around creating a better enterprise and HP’s slogan for the event was “Build a Better Enterprise Together”. So the big themes in the order of my interest were Big Data, Software Defined Networking, Software Defined Storage, Converged Infrastructure and Moonshot. I will briefly cover what I consider the highlights in this article.
Just to give you an appreciation for the size of this event here is a photo that shows a fraction of the show floor at the Sands Expo Center at the Venetian Las Vegas:
Kevin Bacon opened up the event by describing the concept of the 6 degrees of separation and the 6 degrees of Kevin Bacon, which was an internet sensation for some time. He quickly then explained the explosive growth of data (Big Data) and how the world is creating more data every day than was created prior to 2003. My first thought was it must be good to be in the storage business. Here is a photo of Kevin on Stage with some of the data statistics. With the explosive growth of social networking it’s more like 4.7 degrees of separation according to Bacon’s presentation.
Meg Whitman, HP CEO, then took the stage and continued to explain the significance of Big Data and HP’s overall strategies, including sharing a compelling case study from NASCAR. NASCAR are using HP Big Data, Software and Hardware solutions to make NASCAR and even more intense experience and to monitor and adapt their business in real time. Brining the fans closer to the action and taking in feeds from video, voice and social media. Meg also spoke about the transition from the Mainframe world to the client server world and now how the world is changing again in another revolutionary step to a mobile, social and real time business environment. I know from first hand experience how the new paradigm is taking shape. Traditional Mainframe and Unix platforms are being aggressively migrated to industry standard servers running VMware vSphere and applications are being rewritten to adapt to the cloud and mobile era. HP’s argument is that this new era requires a revolution in server infrastructure and this is where their Moonshot Servers come in.
Moonshot servers are ultra compact and need less power than that of a 60w light bulb. HP.com, which receives more than 3million visitors per day is running on 12 Moonshot servers consuming approximately 720W of power. HP believes this will save the Internet due to estimations 10million servers are required in the next few years, data is doubling every six months, and every million servers not only requires a huge investment in datacenters but also requires an entirely new power plant to be built. The Moonshot servers not only consume less power, but they also cost a lot less and consume significantly less space. In the day 2 keynote we heard that currently you can get 450 Moonshot servers per rack, which is set to increase to 1800 servers per rack in the future. The Moonshot servers run one of three different CPU options, ARM, low powered Intel or Nvidia GPU. Future versions will also be suitable for virtualization, which will open up a all new possibilities for hosting and cloud environments to scale dynamically. By my estimations you can already get approx 3072 Virtual Machines with 2 vCPU and 8GB RAM per rack, if not more, if you’re running VMware vSphere on blade infrastructure. So some more number crunching will be required on Moonshot.
So taking a lot of power requirements out of the servers is great, but my overriding question is what about the storage that is needed to power these monster clusters and the huge amount of data growth. Storage has traditionally consumed a low of space and power. The storage side of the equation wasn’t covered during the keynote presentation, and this shouldn’t detract from the overall achievements of the Moonshot platform for the niche applications that it can be used for currently. One thing is for sure the advances in flash based storage will greatly help with power consumption and performance per watt, more on HP Storage later.
HP’s message during the keynote really showed they are trying to be all things to all people (mission impossible?). They seem to be trying to beat the likes of IBM, and to a certain extent Oracle, at their own game. They want to partner with their customers to build a better enterprise together, which is the right approach in my opinion and a good goal. Some of the products and solutions presented come across as little ‘me too’ without real innovation. Only time will tell if they can actually pull off the aim of being the best company to provide a solution to every problem, execution will be key. There is still duplication and lack of integration between solutions and offerings that will need to be sorted out.
Here are a couple of the key images from Meg Whitman’s keynote presentation:
The day 2 keynote was all about the technology. HP showcased Servers, Storage, Networking, Software and Services. What I found very interesting from the day 2 keynote was how little attention was paid to the traditional Unix systems. However all the attention is going towards what HP term Industry Standard Servers and even their Business Critical / Mission Critical x86, such as the DL980’s. I guess this is acknowledgment of a well established trend away from Unix platforms. One of the coolest products shown during this keynote from a server perspective was a very compact ROBO server termed ‘baby’s first datacenter in a box’. I think this could actually be a hit, I know my kids already have their first datacenters, but the real use case is for remote office / branch office where you don’t want to have a computer cupboard anymore. The small server was very compact and also very quiet. Moonshot and the roadmap of the platform and it’s results were interesting to see from a revolutionary server platform perspective as I mentioned also from the day 1 keynote. The Gen 8 Servers got a good showing and there were some impressive industry leading benchmarks published during the event such as the VMMark results where HP Gen 8 servers took the lead powered by Fusion-io.
I missed the morning Storage Press Conference due to a conflicting appointment but there was plenty of storage as part of the day 2 keynote. The key storage platform announcements from my perspective were the all flash 3Par array, the 7450, and what HP is doing in the software defined storage space. Overall the strategy isn’t quite complete and the vision is good but some components aren’t quite there yet technically. But they will be before too long. The good thing about the 3Par 7450 is that it leverages the same 3Par architecture so can leverage the same management tools and software capabilities as the other arrays. It’s performance from the benchmarks discussed were pretty good – 554K IOPS at 4KB IO Size and 0.7ms latency for random read. This was from a fully populated system. This isn’t quite as fast as some competing storage systems, but you need to consider the platform includes all the other capabilities of 3Par including all the replication support etc. 3Par has good scalability in the platform, includes a lot of features as standard, but still lacks a few high end capabilities that some competitors bring to the table (more on this later). As a comparison I can get 187K random read IOPS at 0.5ms latency from a single VM connected into one of my vSphere hosts from a single Fusion-io ioDrive2 card, but of course this doesn’t have the features of a full 3Par array, such as replication etc. 3Par does seem to be gaining ground though with 1500 new customers and over $1billion in revenue in 2012. Plus they have the switch to 3Par guarantee (provided you’re not using thin provisioning or dedupe already) that you’ll get 2 x the VM density.
Next up came networking with HP claiming #2 position to Cisco and bigger than the next 5 network vendors combined in terms of revenue . Really the #1 or #2 spot would go to VMware if it were purely on number of ports (if you include physical and virtual ports), but the comparison was based on revenue. I’ve always liked the HP datacenter switches since I first used them more than 10 years ago. They’ve always been easy to manage, very feature rich, and lightning fast. This year they introduced a number of enhancements to flex fabric where a single 11900 switch could support 92 C-Class Servers. HP was also flexing the SDN muscle touting 40 OpenFlow switches and 20M installed ports. I was particularly interested in HP’s announcements around the 5900V virtual switch for VMware vSphere. I thought this would be a serious play against the Cisco Nexus 1000v. Unfortunately all it does is bypass the hypervisor networking stack and implement the VM networking on a physical switch. This not only adds latency to traffic that would otherwise pass between VM’s on the same host, it doesn’t allow the performance of in-hypervisor security modules, and it means you’re locked into the physical HP network hardware stack, which greatly reduces your flexibility. At least with the VMware Distributed Switch or the Cisco Nexus 1000v you can use any hardware switches you like, including HP’s.
HP covered their new converged and open Cloud offering based on OpenStack briefly. But unfortunately they forgot to mention that it’s not even compatible with HP’s other Cloud service, which is offered by HP Enterprise Services, called the HP ES VPC. Compatibility between the two clouds is coming in the future apparently. In the meantime HP ES VPC Cloud runs on VMware vSphere and the new HP Converged Cloud runs on OpenStack. This seems similar to the IBM OpenStack Cloud that was announced recently and adds additional competition to the other offerings in the cloud market.
Big Data was big news at HP Discover and during the day 2 keynote we got to dig a little deeper into HP’s Big Data solution called HAVEn. HAVEn is an acronym that stands for Hadoop Autonomy Vertica Enterprise Security for n applications. It’s a platform that can integrate many different types and streams of data, both structured and unstructured, and analyse them in near real time. It was especially interesting to review the NASCAR case study again to see how they are using real time big data analytics from all sorts of different data feeds, including social networking, to make business decisions and improve their business. The HP Discover event itself was also using HAVEn to measure sentiment and influence and other factors around the event to gain insight into what was going on, which is a great demonstration of the platforms real world capabilities. I’ll show you a photo of the HAVEn implementation for HP Discover below.
The most interesting aspect for me was the way that Autonomy could take unstructured data and make sense of it, giving it context, while Vertica could analyse structured data at lightning speeds. But what sort of data? Well it’s not just limited to social networking and business data but also other relevant and important data. The HAVEn platform can take huge amount of data from different sources and correlate them all. The types of data HP described was machine data, such as sensors, logs and system metrics, business data such as from CRM systems, and human information such as voice feeds and video. Another real life example HP gave during the keynote is how the Venetian|Palazzo Casino Hotel used the HP HAVEn platform to monitor all the gaming systems, Casino gaming floor and identify possible issues in real time, including integration with all the camera systems etc. They gave a demonstration of how HAVEn could be used to solve a system performance issue, but to be honest this demo really missed the mark as it didn’t demonstrate HP’s unique capabilities. VMware vCenter Operations could have done the same thing and a lot more quickly for the use case HP presented. However if they had integrated the demo system with data from the call center voice recordings and online social networking to determine severity and trends before it became a major issue, and then solve the problem, it would have been much more relevant. Given recent media events it was hard to avoid comparing what HAVEn could do to what had come out about PRISM. No comment either way was forthcoming from HP, which is not surprising. The ability to scale your big data systems up and down and make testing them efficiently lends itself to good virtualization use cases. There is a lot of work going into optimizing VMware vSphere environments to run Hadoop and there is no reason why HAVEn couldn’t run just as well in a virtual environment, provided you design it properly and test and verify it to make sure it’s working to your requirements. It will be interesting to see how this technology develops in the future. But the electronic big brother is already watching you. So you can’t assume any privacy for any information you make available online. Just assume that anything you say or do on the Internet could be on the front page of the news or someone’s blog tomorrow.
Here are some of the key slides from the day 2 keynote:
Here is a photo of the HAVEn system that was monitoring HP Discover. You might notice that I made a bit of an impact on day 2. I was the major influencer for pretty much all of day 2.
One of the best parts of being invited to be a blogger at HP Discover was the coffee talks that were arranged so that we could meet all the key executives and subject matter experts for the different areas across HP. This allowed us to drill down on key points not covered during the keynotes. Below I’ll give you some very brief highlights from each of the coffee talks.
Big Data and HAVEn (Hadoop, Autonomy, Vertica, Enterprise Security for n applications)
Machine info, business info, human info, correlated and available in one place providing context and support for real time decision making. Examples include determining from voice calls caller sentiment (if they are likely to stop doing business with you) and whether they are lying, allows offers and decisions to be made in real time to prevent lost customers and improve customer experience. Take data from camera feeds to determine in real time if someone is cheating (Casino for example) or if there is a health concern or possible security event, such as a door opened at the wrong time of day. This technology could be described to be similar to a number of scientific discoveries. It has the potential to do immense good but also could be used and abused for the wrong reasons. It also gives hackers a high value target. If you hack the big data systems you potentially have access to every piece of important data all in one place and correlated between many sources. The security and privacy of the data is a great concern, and not just access to the raw data but also the search results across the data. We discussed how HAVEn could be used to make systems more secure by providing context based security such as location data, and detecting duress in someones voice. There is massive data growth being driven by machine data, such as seas of sensors and massive telemetry data, and human information, such as voice and video data.
Mobile Apps
The overall conclusion from the Mobile Apps coffee talk was that if you really want to take advantage of mobile and give a good user experience then you need to completely rewrite your apps and you can’t really leverage your existing investments. I don’t agree with this and I think you should be able to take a hybrid approach that gets even greater value out of your existing apps while making your business much easier and more responsive to your customers using a mobile and social online strategy. HP’s argument was that you mobilise transactions, you don’t mobilise applications.
HP gave an example of Mary K cosmetics where 20% of agents are joining the business because of how easy the mobile app is to use and the support across any device. Having the mobile app has massively boosted Mary K’s business. The App supported 240K contractors. This example was impressive and quite compelling. Really shows the benefits of developing purpose built mobile apps can have a positive impact for the right transactions and interactions.
We did have an interesting discussion on security and the death of the password in favor of multi-factor authentication. There are a number of problems with traditional password security and verification (remembering many different passwords is one such challenge) and I think it is high time it was replaced by multi-factor means of verifying identity. Only time will tell how this plays out. But contextual security will be important also, such as determining your location and authenticity of your identify based on usual location, usual usage patterns, and the usual devices you use in addition to traditional factors for authentication.
Storage
We spent a bit of time going through the 3Par storage system architecture and especially the details of the new 3Par 7450 All Flash Array. I like the way the 3Par architecture is highly parallelised. This makes it great at handling lots of random IO’s, which is very normal for virtualized environments. HP covered the new Peer Persistence feature in the 3Par system, which will allow 3Par to attain the VMware vMSC certification for deployment in stretched cluster environments (similar to what HP P4000 system can do now). HP wasn’t sure how long the VMware vMSC certification process would take, but it is underway. There were a couple of gotchas with this new Peer Persistence feature such as you must implement a uniform storage access configuration and each LUN is active/passive not active/active across sites. The uniform storage access is much more complex to implement and maintain and doubles the number of paths zoned to each host. This effectively halves the maximum number of LUNs you can have configured per host as you need redundant paths to each array at each site. The LUNs being only active for read / write at one site is also not ideal as any VM’s executing on the opposite site will have additional latency for all their storage accesses back to the primary site for the LUN. 3Par Peer Persistence is trying to compete against EMC VPLEX and other vMSC solutions (such as NetApp), but it’s not quite on par yet due to no support for SRM on top of the stretched cluster (unless using hypervisor based replication), and the lack of support for active / active read/write LUNs and non-uniform storage access. All in all though if you had an existing 3Par implementation and wanted to have a stretched cluster between two active datacenters (provided they are within the distance limitations <10ms) then you could now do it. Many 3Par customers I spoke to were very interested in this new enhanced feature, which also includes provision for a witness, they were also very happy with the performance and manageability of their existing 3Par systems. Most 3Par customers like the fact that they can start with an entry level 3Par and grow it as their needs increase while keeping consistency of management platforms.
Networking
There was quite a bit of discussion around network management and monitoring and using OpenFlow combined with Blueprinting to roll out applications on top of virtual networks. The blueprinting capabilities, which basically allow you to deploy an application anywhere and have it’s networking automatically configured, are very compelling. This was however limited to OpenFlow based networks.
I was interested in how the networking hardware was going to be made better to support the new Software Defined Networks, Network Overlays and Network Virtualization technologies. But unfortunately there appears to be nothing in the pipeline to build mechanisms into network hardware to reduce the overheads associated with network layer virtualization or overlay technologies, nor increased MTU sizes, which will increasingly become important as Ethernet reaches large scale 40G and 100G adoption. It appears to be HP’s opinion that =>9K MTU is not required and in fact most people are / should still run 1500. This is in spite of a 10% performance improvement when using Jumbo Frames on 10G links and this performance improvement is only going to increase as the Ethernet bandwidths scale up.
Converged Cloud
HP launched what they call their Converged Cloud, this cloud equals OpenStack. It isn’t currently compatible with HP Enterprise Services VPC cloud offering (note the duplication and lack of integration I mentioned earlier), which is based on vSphere. You can’t migrate between the two different HP clouds, but this will apparently be changing in the future. This appears to be a play in competition to RackSpace, Amazon and VMware vHCS. With VMware vHCS customers can easily move their existing workloads from any VMware environment to a hybrid cloud model, and to any of the many Partner’s VMware based clouds easily, without any modifications. This level of operability isn’t yet available with the OpenStack clouds, but no doubt it will be there in the future. OpenStack right now seems very much like the early days of Linux where there are so many different distributions and choices.
Converged Infrastructure
HP’s converged infrastructure should really be called pre-packaged infrastructure rather than converged as it lacks converged management and many operations still require standalone management tools, doesn’t update as an entire platform as one whole, but can potentially scale and change more easily than other converged platforms. Still this makes it every easy for customers to purchase and deploy HP’s technology. This is really HP’s first step to converged infrastructure and they’re working hard to converge the management and lifecycle processes together. It’s still early days yet but vBlock and newer converged offerings such as Nutanix have got the jump on HP and still lead while HP catches up. Some of the people at the conference were liking the HP converged infrastructure to what VMware provides in the vCloud Suite currently. A single SKU to purchase but it’s really a number of products that aren’t really integrated. If you’re an existing HP customer the new converged SKU’s will make purchasing HP technology much easier and much quicker to implement as it comes pre-packaged to your specifications from the factory / distribution. Overall this will decrease your time to market and also decrease your TCO. Once they have the management side of things sorted out it’ll break down the IT silos and also decrease further the operational costs of the infrastructure as you’ll be able to manage the complete platform as one.
Here is a photo of the HP Virtual Convered Infrastructure Rack (right). This is a virtual demo controlled by an HP tablet where they can show you everything about the converged infrastructure platform, but without having to carry around a full implementation. This is rather a great bit of kit. The rack on the left is the actual converged platform based on an entry level solution that can then be easily expanded.
I took some time out to visit some of the sponsor stands at the show and I’m really glad I did. There were some excellent sponsors at the event and I learned quite a lot. Here are some of the highlights from the sponsors at HP Discover.
AVT
AVT is a company that provides a solution that allows customers to migrate OpenVMS, VMS Clusters and Alpha systems to VMware on HP systems without code changes. This means that any existing customers that want to retire their old hardware can do it but without the massive costs associated with redeveloping their application. AVT is the only HP supported solution that allows this type of migration. You can run multiple VAX or Alpha systems per host easily. The main advantages are reduced risk of running on old or retired hardware, prolong the value of your VAX and Alpha software, avoiding costly and complex software migration costs. There is also a good potential for performance improvements when going to the latest generation hardware platforms. The company estimates there are still 500k VAX and Alpha systems in existence and that’s not surprising. So if you do happen to have a VAX or Alpha in your datacenter it would pay to get in touch with AVT.
Fusion-io
It was good to see Fusion-io at HP Discover and especially good to learn about their spectacular performance results (as mentioned earlier). HP along with Fusion-io set some world record VMMark Results, and achieved performance for an Oracle workload against an on ION storage system of 2.2million IOPS and 24GB/s. Fusion-io had the performance demo running live on the show floor and I hope they bring this to VMworld this year as well and run a Virtual Machine benchmark against it on the latest VMware release. To reach this spectacular performance HP and Fusion-io used a Gen 8 DL980 with 6 x dual port 16Gb/s Qlogic HBA’s going via a Brocade 6510 16Gb/s FC switch to 3 x Gen 8 DL380 servers each with 2 x dual port Qlogic 16Gb/s HBA’s and 4 x 2.4TB HP IO Accelerators (Fusion-io Cards), 28.8TB in total. This is a surprisingly small configuration for such a massive performance result. This certainly demonstrates not only the potential of the Fusion-io technology but also of the HP Gen 8 servers also.
TIBCO
I had been getting a lot of requests recently from customers that wanted to virtualize varios types of TIBCO workload including DataSynapse Federator and GridServer. I thought the best way to find out what the support situation was would be to ask TIBCO themselves. I went to the stand and they were extremely helpful. Within a few minutes I was connected to the VP of GridServer and he confirmed that indeed large scale deployments of GridServer were supported on VMware vSphere. Here is what they said:
* Yes, you could run 6000 GridServer engines (or any number), each on a virtual machine, provided the underlying physical infrastructure was sufficient. (This is pretty much what TIBCO DSPG QA lab looks like.)
* Yes, you could have 6000 separate grids, each on its own virtual infrastructure, again provided the underlying physical infrastructure was sufficient.
* You should not attempt to run 600 or 6000 anythings on a single VM.
This last comment is of course common sense. GridServer is Java based so the best practices with regards to Java and Low Latency Workloads on VMware vSphere apply. If you weren’t already aware there is a low latency setting in the VMware vSphere 5.1 Web Client available for those workloads that need low latency.
Mellanox
I went past the Mellanox stand to find out what the latest was from their product lines. I was mainly interested in their Ethernet products. Previously their 10G NIC’s had been quite hard to configure and get working on VMware vSphere and they would show up as phantom NIC’s and sometimes not be available immediately at boot time. This problem has been fixed with the latest firmware and drivers I’ve been told. It was very interesting to see that Mellanox is providing 40GbE and 56GbE card, which includes support for VMware. The 40GbE NIC’s and Switches have a latency of 220ns, compared with 270ns for 10GbE ports.
The Mellanox SX1018 40GbE C Class Blade Switch, which provides 220 ns latency, supports 18 full speed 40GbE ports out of the C Class Chassis. That’s 720Gb/s of low latency bandwidth from a single chasis. I didn’t go into the costs but I’m sure it’s not going to be all that cheap, but certainly impressive from a performance point of view.
Based on what I discussed with Mellanox at their stand I think converged network / storage fabrics are the way of the future with clear leadership in Ethernet over straight FC. When it comes to networking history has shown that Ethernet always wins. With low latency and non-blocking QoS available on high speed Ethernet ports there really is no reason to have traditional FC fabrics and switches if you’re making new investments. 40GbE is here now and 100GbE will probably be available before the end of 2013, at which point 40GbE will start to come down in price.
You can forget about running 40GbE on a Gen 2 PCIe slot. With these sorts of speeds you’re going to have to be running PCIe v3 to get able to get the bandwidth, like you get on the HP Gen 8 servers. With all the advances in networking that are going to be upon us shortly it’s certainly going to make things interesting in the datacenter.
VMware
VMware had a great stand at HP Discover covering Software Defined Datacenter and End User Computing. They were also represented well on various HP stands and other partner stands around the event. Here is a photo of the VMware stand.
Final Word
HP Discover was a great event and I highly recommend it to any current or potential future HP customer. If you can’t make the USA event HP discover is also on in Europe. This year it’s in Barcelona in December. Apart from all the great technology on display at HP Discover one of the highlights was the HP Discover Party on Wednesday night. HP hired out the MGM Grand Arena and had Los Lonely Boys and Santana entertain the conference attendees. This was one awesome show and the entire crowd really enjoyed it. They played all the crowd favourites and HP put on plenty to eat and drink as well. This is one of the best, if not the best, vendor sponsored parties that I’ve ever been to.
I and the other bloggers were also fortunate to get to spend some time with the HP CEO Meg Whitman. Here is a photo of all of us together.
Note: Travel to HP Discover 2013 was paid for by HP; however, no monetary compensation is expected nor received for the content that is written in this blog.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
VMware’s vCloud Director is a very good way for organisations to start to take advantage of cloud computing, including private, public and hybrid models. Cloud computing can offer new efficiencies and cost savings for organisations that optimise it’s use. But where do you start? If you go searching for best practices, design considerations and references for designing and building a vCloud Director environment you will find plenty for large scale deployments. But it might seem difficult to find much in the way of design considerations for starting off with a small vCloud Director environment, such as for a proof of concept, small lab or pilot. I have previously written about Considerations For Designing a Small vCloud Director Environment – Allocation Models, which discusses takes you through which allocation models there are and which ones you might want to use. In this article I hope to offer some advice that will allow you to dip your toes in the water and get up and running quickly without much complication, while allowing you to scale up in the future. This article is a continuation in the series and looks at the different options for storage at a high level and offers some recommendations that you might want to consider. Future article will discuss the rest of the components you need to consider.
Even if designing for a small environment one resource I highly recommend you review if the vCloud Architecture Toolkit (vCAT), which is a VMware Validated Architecture for Cloud Computing and supporting tools. The vCAT is fully supported by VMware, so if you leverage the design considerations and guidance contained within it you know you can get support.
This article focuses on Storage Design in vCloud Director. In future articles I will cover additional design considerations.
Pre-requisites and assumptions
Classes of Storage and Provider Virtual Datacenters
In versions prior to 5.1 of vCloud Director your storage design and storage allocation would have had a major impact on how many Provider Virtual Datacenters (PvDC’s) and Clusters or resource pools you would require. This is because each PvDC could have only a single tier or class of storage. So the storage tier or class became directly linked (tightly coupled) to the overall service definition of the PvDC, i.e. the characteristics that define what a particular class of service is made up of, such as standard, enhanced or premium (Bronze, Silver, Gold) etc. As an example in your Standard PvDC you might have had 2.4GHz Xeon CPU’s, 96GB RAM (1066MHz) per Host and NFS Datastores backed by SATA disks (RAID 6 or RAID DP). In your Enhanced PvDC you might have had 2.93GHz Xeon CPU’s, 256GB RAM (1333MHz), and FC Connected SAN with FC 10K disks in RAID 5. As soon as you wanted to offer another tier of storage you would have needed to offer another PvDC, even if the rest of the service model wasn’t changing.
So what is the impact of adding another PvDC if you want to add another class or tier of storage? Well if you were to follow standard VMware Design Guidance (this is the new term that has replaced best practice), then you would have to set up an entirely new cluster of hosts. Each PvDC should be backed by a Cluster of 2 or more hosts with HA and DRS enabled. In theory it is possible to use a single host in a PvDC, but then you can’t test HA or DRS functionality with vCD. You may be able to see a problem here for a small environment that will be used for a Pilot, PoC or Lab.
The answer to this conundrum in versions of vCD prior to 5.1 is to use resource pools inside of a cluster of 2 or more hosts, instead of using different clusters of hosts, to back your PvDC’s. Each PvDC is then mapped to one of these resource pools. The different tiers or classes of storage are allocated to all hosts in the cluster and the allocated to the correct PvDC. OrgVDC’s are then created in the correct PvDC to consume the different classes of storage. There was and still is no option (as of vCloud Director 5.1) to have a single VM utilise resource from multiple service tiers or classes of storage.
Using resource pools instead of clusters does have drawbacks. Firstly there is the resource pool priority-pie paradox, which may impact the allocation of resource for any given VM or sibling resource pool. Secondly a resource pool is not allowed to consume 100% of the parent resource pools resources. Depending on versions of vSphere backing the resource pool it might only be able to consume 94% of the parents resource pool. This will potentially leave 6% unallocated. You’re also not able to segment service definitions by different type of compute characteristics as resource pools span the cluster, which isn’t a problem if you’re doing this for the sole reason of allowing multiple classes of storage. Multiple PvDC’s may be competing for resources though. It will be more complicated when you come to expand out the environment, for example if you wanted to move a PvDC to another cluster.
The workaround to provide different tiers or classes of storage to a single cluster is only required in versions of vCloud Director prior to 5.1. From vCloud Director 5.1 you can assign multiple storage tiers to a single PvDC and even use Storage Profiles, Storage Clusters and Storage DRS, which was also not possible prior to vCD 5.1. Thus the class or tier of storage is now not required to be coupled or linked to the overall service definition. You can now have multiple storage service definitions per service class, such as Big Data (SATA) and Fast Data (15K FC), in standard, enhanced and premium PvDC’s. This means you can have a single PvDC in a single cluster of 2 hosts, and still have multiple tiers of storage. This is a very good reason for using vCloud Director 5.1 for your vCloud environment.
The above assumes that you are not using Auto Tiering on your arrays to automatically allocate the different tiers of storage. If you are using Auto Tiering then you would only have multiple service tiers and therefore PvDC’s (prior to vCD 5.1) or Storage Clusters and Storage Profiles (vCD 5.1 and later) if you are offering different auto tiering policies. When using array auto tiering with vSphere 5.1 and vCloud Director 5.1 it is recommended that IO load balancing be disabled. This is because the array is handling that for you and SDRS may make false recommendations.
Datastore Sizing
Because of the somewhat unpredictable size of VM’s in a vCD environment it’s generally recommended to size datastores to be fairly large and have fewer of them. In a normal vSphere environment you may have had datastore sizes of 292GB, 512GB or 1TB etc for enterprise workloads, however in vCD you might want to use 2TB or 4TB sizes (or larger), and have fewer. This allows better placement of VM’s when the size of each VM varies a lot. This also helps a lot when using Fast Provisioning (Linked-Clone Technology) explained below. The lager size datastores also help reduce the risk of running out of space during rapid growth. This guidance is generally applicable to test/dev, pilot, PoC or lab environments and is what I’ve found generally works well. Having the larger size datastores allows potentially more efficient use of storage, especially when combined with thin and fast provisioning (explained below) but may trade off ultimate performance. A lot will still come down to the storage technology you’re using.
Thin Provisioning and Fast Provisioning
Thin Provisioning and Fast Provisioning are both techniques that when implemented allow for more efficient use of storage resources and for the potential of storage overcommit. Provided you take care both can be used safely and allow for significantly better storage economics. Like all techniques that allow improved efficiency due to overcommitment the underlying assumption is that not everything needs the same scarce resources all at the same time. I’ll briefly explain how they work and what you need to watch out for.
Thin Provisioning works by only allocating blocks to a VM as they are needed by the guest operating system, unlike a thick provisioning which allocates all storage assigned to a VM up front. This means that in cases where you know 20% or 30% of allocated storage is only there ‘just in case’ or ‘because the OS said we had to’, you can effectively use that elsewhere for other VM’s. Provided not all of the VM’s that are using thin provisioning need all their storage all at once you can get effective overcommitment of the underlying datastore and this boosts your ROI. All without any significant performance penalty (assuming modern hardware environment). Using Thin Provisioning can introduce an additional management overhead as you have to ensure your datastores don’t run out of space. Storage DRS in vCloud and vSphere 5.1 environment can help with this though and reduce the management overhead. Even with the additional management overhead in many cases the improved economics can far outweigh the risks.
Fast Provisioning is the vCD term for Linked-Clones (Similar to vCD’s predecessor Lab Manager and VMware View). With Fast Provisioning you have a parent image (a.k.a. Shadow VM) and multiple children are linked back to the parent. So you effectively have only one main copy of a VM, each other VM is only a set of configuration files and delta files. The parent image is read only and any changes are written to the delta files. You could think of this similarly to single instance storage in an email system, where only one copy of an email is stored on disk and multiple mailboxes link back to the main copy. If you’re linked clones have a parent image that contains the OS and base applications and don’t change much you could literally have hundreds of copies with only a fraction of the assigned storage used. When using Fast Provisioning each parent image can have up to 30 linked clone images. Once you get past this limit another parent image is cloned and the linked clone chain starts again. There are also considerations when your chains need to cross datastores. All of which is explained in the vCD documentation so I won’t go into it here. The potential storage efficiency gains by using Fast Provisioning’s linked-clone technology are stagering. 10x storage overcommitment is possible based on my experience, which is like reducing your storage investment by 90%, or reducing the cost of storage 90%. But like all things there are tradeoffs and risks.
A couple of things to note about Fast Provisioning. Firstly disk alignment is an issue. Images are initially aligned, and the parent image is aligned, but as the images start to grow the IO’s will be unaligned. This will have an impact on performance. Secondly as the parent image is shared by all the linked clones is on VMFS and you’re using a version of vSphere prior to 5.1 your cluster size will be limited to 8 hosts. This is because the maximum number of hosts that can have the same file open is 8 on VMFS prior to vSphere 5.1. This is not a problem with NFS storage, or with VMFS in vSphere 5.1 and later, both of which support cluster sizes up to the maximum of 32. In saying this though, it’s not going to be an issue for a small environment of a couple of hosts, which is what we’re discussing here.
Both Thin Provisioning and Fast Provisioning can be used together and complement each other. By using both you can massively improve the efficiency of your storage and boost your return on investment. But because of the potential for up to 10x storage efficiency or overcommitment when both thin provisioning and fast provisioning are combined (based on my experience in a number of vCD environments) there are risks that need to be managed. You should seriously consider having burst capacity available to cater for any peak demands and also pay careful attention to how much storage is allocated to each OrgVDC.
Recommendations
If you are using vCD prior to 5.1 then I recommend you implement PvDC’s based on resource pools in your small cluster of 2 hosts so that you can demonstrate the use of multiple classes or tiers of storage. This allows you to keep the environment simple and small for Pilot, PoC, or Lab use. If you are using vCloud Director 5.1 then you don’t have to create multiple PvDC’s and assign them to resource pools. You can simple create a single PvDC and assign the entire cluster to it. You can then use Storage Profiles, Storage Clusters, and Storage DRS to present the different classes or tiers of storage.
For detailed explanation and considerations of using clusters or resource pools to back PvDC’s I recommend you read Frank Denneman’s article Provider VDC – Cluster or Resource Pool.
Size your datastores so you have fewer of them but they are generally large (2TB to 4TB is fairly common). This allows the best storage utilization and efficiency, especially when using Fast Provisioning and Thin Provisioning. Start out with a small number of datastores and then grow as required.
Use thin provisioning to make the most efficient use of your storage, there is no significant performance impact for doing so, but beware of the risk of sudden growth and plan accordingly. Use Fast Provisioning for dev/test workloads and for situations where storage space efficiency is more important than performance. There is a performance overhead to using Fast Provisioning’s linked-clone technology, but the economics are compelling in a lot of cases.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
VMware’s vCloud Director is a very good way for organisations to start to take advantage of cloud computing, including private, public and hybrid models. Cloud computing can offer new efficiencies and cost savings for organisations that optimise it’s use. But where do you start? If you go searching for best practices, design considerations and references for designing and building a vCloud Director environment you will find plenty for large scale deployments. But it might seem difficult to find much in the way of design considerations for starting off with a small vCloud Director environment, such as for a proof of concept, small lab or pilot. In this article I hope to offer some advice that will allow you to dip your toes in the water and get up and running quickly without much complication, while allowing you to scale up in the future. This article looks at the different allocation models. Future article will discuss the rest of the components you need to consider.
Even if designing for a small environment one resource I highly recommend you review if the vCloud Architecture Toolkit (vCAT), which is a VMware Validated Architecture for Cloud Computing and supporting tools. The vCAT is fully supported by VMware, so if you leverage the design considerations and guidance contained within it you know you can get support.
This article focuses on Resource Allocation Models in vCloud Director. In future articles I will cover additional design considerations.
Pre-requisites and assumptions
Allocation Models
vCloud Director allows resources to be allocated from a Provider Virtual Datacenter (PvDC) to Organisation Virtual Datacenters (OrgvDC) for consumers to use. The allocation is done in accordance with three different models. Pay as You Go (PAYG), Allocation Pool, and Reservation Pool.
The advantage with PAYG is that you get a defined reserved amount of resource per VM when the VM’s are powered on only, so they don’t consume any resources from the OrgVDC or PvDC while they are powered off. The system admin defines the overcommitment and service levels for compute resources. New in vCloud Director 5.1 the PAYG Virtual Datacenter can have a limit set to prevent one OrgVDC consuming all cloud resources.
With Allocation Pool the OrgVDC consumes the guaranteed portion from the PvDC regardless if VM’s are powered on, but this at least gives them a defined amount of resources. Each VM then is set a reservation equal to the percent that is guaranteed and the VM’s deployed can consume the resources up to the defined limit of the allocation model. Powered Off VM’s don’t consume resources from the OrgVDC. The system admin defines the overcommitment and service levels for compute resources. New in vCloud Director 5.1 all vCPU’s have a defined limit set which will impact how many VM’s can be powered on within an OrgVDC. This vCPU limit is defined by the System Administrators.
With the reservation pool the you set a defined limit and then it’s up to the tenant of the OrgVDC to define the level of overcommitment and how much resources are reserved for each VM/vApp. You guarantee the resources from the PvDC to the OrgVDC and the OrgVDC consumes those resources regardless if a vApp is powered on. VM’s or vApps can have different reservations or no reservation. It’s up to the tenant to choose. But the catch is the tenant or consumer can cause performance problems within their OrgVDC if they overcommit too aggressively. The tenant needs to take some care with capacity planning around their usage.
In some ways an Allocation Pool with 100% guaranteed is superior to a Reservation Pool, as the system admin / provider is defining the SLA of the pool and VM’s to be 100% guaranteed, I.e. Allocation = Reservation. The tenant then can’t get themselves into trouble as easily. It’s up to the system admin to manage capacity.
The Allocation Pool and Reservation Pool offer more predictability of billing to the tenant, but are less efficient from a service provider perspective in terms of resource utilization. PAYG offers less predictable billing to the tenant, but more efficient resource utilization potentially, and more flexibility in a small environment.
Recommendation
When designing a small environment I generally recommend you start with PAYG resource models, especially where usage is unknown. This allows the best sharing of resources between OrgVDC’s and Organisations. PAYG provides a dynamic on demand environment, but can still have an upper limit set per OrgVDC. The PAYG model supports an Elastic VDC, which means it can grow across clusters. This makes it easy to expand the environment as and when needed. The Allocation Pool Model also supports elastic vDC from v5.1 with some restrictions, but it will depend which update version is actually being used on what functionality of the Allocation Pool model you receive.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.
I was talking to my friends at Mako Networks, provider of PCI DSS certified and secure network equipment to telco’s, payment card processing and retail industry (and others) in late 2012 on how they were using virtualization to solve their business requirements. One of the interesting use cases I thought I’d share is how they’re using what they term Lightweight Virtualization to produce a Multi-Tenanted Device (MTD), which is a virtual multi-tenanted network concentrator. This fits in nicely with their cloud based central network management system (CMS) and allows their partners to serve thousands of customers without thousands of individual network concentrator devices.
I’ve been using the Mako Networks systems in my primary and DR locations for over 10 years. They’ve always been rock solid and fast. The thing that really got me interested though was how simple the routers were to get set up and to manage via the cloud based central management system, they just work, and the quality of reporting that is built in is great. Delivering complex things like secure networking in a very simple way is of great value in my opinion.
The idea of a multi-tenanted device (MTD) for Mako’s Network VPN Concentrator appliances is quite similar in concept to how VMware vCloud Director provides multi-tenanted edge gateway services in public cloud environments based on VMware vCloud Director (i.e. multiple tenants being serviced by virtualized edge gateway devices). The difference here is that the MTD is a component of a physical network concentrator appliance being used to deliver PCI DSS compliant network services for Mako Networks’ Partners to potentially thousands of customers and tens of thousands of individual merchant devices and sites.
I asked Kevin Ptak, Mako’s Communications Manager and Murray Knox, Mako’s R&D Director to write this article. I have added some emphasis and detail in bold italics below.
Disclosure: Mako Networks are not paying to have this article published on my blog, I’m publishing this because I think it will be on interest to my readers, however one of my companies does have a very minor shareholding in their parent company (<1%).
—
Introduction
Mako Networks provides a managed network service, comprised of (primarily physical) network appliances installed at customer locations and a cloud-based Central Management System. This includes both small-office appliances and central office concentrators to allow a number of branch offices to securely communicate with a central office.
Some time ago, one of our partners approached us with a problem. They wanted to deploy our VPN concentrator service to multiple customers to allow access to our partner’s SaaS offerings. However, their customers weren’t big enough to warrant an individual concentrator each. Additionally, the target market was between 500 and 2,000 customers and it was not feasible to deploy this many concentrators in their data center. Therefore, our partner requested that we develop a version of the concentrator that can be used by multiple customers, a Multi-Tenanted Device (MTD).
One of the reasons they requested a MTD was to be able to provision and enable a new customer in an eight-hour timeframe. They could not achieve this Service Level Agreement (SLA) using alternative technologies but their experience with our platform showed that the Mako System makes it very quick and simple to provision a new customer.
Multi-Tenanted Device
The Mako System is designed around the concept of one customer per device. It would be a significant change to our management system and appliances to allow multiple customers per appliance. Therefore, we decided a better approach would be to support multiple virtual appliances within a single physical appliance.
Our initial idea was to use standard virtualization technology and deploy our Mako appliance firmware as a guest on the virtualization platform. This would have provided the required capabilities and could be developed in a relatively short time. However, we quickly found that this would not meet partner requirements to provision new customers within eight hours. Also, initial capacity planning showed the memory requirements for this platform were significant.
We had experience with lightweight virtualization in other areas and decided to investigate the feasibility of using it for the MTD. Lightweight virtualization differs from standard virtualization in that there is a single instance of the operating system that executes and provides services to the individual guests. Each guest is really no more than a set of processes to which the operating system provides a virtual filesystem and network namespace.
Lightweight virtualization is useful when you want to deploy a large number of very similar (or identical) applications that you need to keep segregated. It uses less memory and disk than full virtualization, as you don’t need a full copy of the guest operating system for each instance. The host operating system can provide core services (e.g. scheduling, clock synchronization, memory management, logging) allowing each guest to have a smaller memory and CPU footprint.
Lightweight virtualization is now a mainstream technology on the Linux platform in the form of LXC and network namespaces. We created a number of prototypes using this technology to prove that it could support our Mako appliance requirements and that the network namespaces provide the required levels of separation between appliances. After proving the viability of LXC, we then proceeded to design and implement the MTD around this technology.
We had to design a deployment process for the MTD to allow provisioning of a new customer in a simple, error-free manner without requiring complex network changes. We settled on a process whereby a customer is provisioned and configured in the Mako CMS in the standard way, then we copy the customer’s configuration data to the MTD and supply it to a deployment script which creates a lightweight container and configures the network interfaces to be used.
Another issue we encountered was the time required to create up to 500 network interfaces when the system is booted from cold. Early versions of the LXC system required significant amounts of time to create the required number of network interfaces. Eventually, with careful tuning and patching we managed to get the time required down to an acceptable level, aided by the rapid development of LXC and namespace systems.
Our standard software update process for appliances was not going to be suitable for the MTD. Normally, when we deploy a software update we supply every component including the operating system to ensure the integrity of the update. However, with the MTD, the host operating system is pre-deployed and would be supplying many of the capabilities and we only needed to deploy a limited number of subsystems. Further, each virtual Mako appliance uses identical copies of the firmware. Therefore, we decided that all virtual appliances share the same instance of the firmware. This signififcantly reduced storage requirements and made it so that only one copy of the firmware needed to be maintained.
Another system requirement was high availability. We had to provide a system that was immune to both hardware and software failures. Our nework appliances (or ‘Makos,’ as we call them) firmware already has mechanisms for managing software failures. However, in a virtualized environment it could not manage failures of the host operating system or hardware. Therefore, we adopted a two-pronged approach to resiliency. Our CPE firmware would handle its own internal failures. For hardware and host problems we decided that clustering was an appropriate solution. The MTD would be deployed on multiple physical nodes within a cluster and each node capable of handling the full load of 500 customers when required. (This removed the host OS being a single point of failure, which would otherwise have been the case)
Scaling and Capacity Planning
Capacity planning is an essential stage of deploying a system in a lightweight environment, as with any other virtualization system.
Overcommitting resources is a standard technique with virtualization. However, it is essential that overcommitting not jeopardize either SLAs or the ability of the guests to function. We have very good performance profile information for our standard Makos, which was helpful in sizing the MTD. We know how much memory, CPU, disk I/O and network I/O each Makos typically uses (even during peaks).
Network response time was an important SLA for us. if a virtual Mako was paged out of memory then performance would suffer dramatically. Therefore we decided to ensure that there was enough RAM so that an MTD appliance could run with 500 virtual CPEs all in memory at once.
We also had good knowledge of the disk I/O requirements of the Mako in both normal and peak conditions. Multiplying the I/O requirements of a single Mako by 500 showed that standard hard disk technology would not provide the throughput we required, therefore we adopted SSDs.
Lastly, we developed a series of load a stress tests and tested the MTD thoroughly up to and beyond its design limits. We tested the MTD multiple times in a wide variaty of scenarios until we and our partner were satisfied that it would operate reliably and meet the required SLAs.
Resource capping is used within each lightweight virtualization container to ensure that one container can’t spin out of control consuming resources and adversely affecting others. In addition to this Mako Networks has advanced monitoring across all of the containers so that they can detect any adverse behaviour and take action before there is a system impact. The practices applied here are the same as would be applied to other business critical workloads in other virtualization environments, especially where it comes to testing and ensuring SLA’s are met.
Summary
Like standard virtualization, lightweight virtualization is a rapidly evolving and maturing technology. However it allowed us to develop a new product in a relatively short timeframe and offer another level of scalability while retaining the simple management of our core offering. It allows our partners to offer a new service to their customers with a significant reduction on hardware requirements compared to a physical deployment.
—
Final Word
I hope you found this of interest. There are quite a few lessons that can be taken from this in terms of developing virtualiation architecture strategies for all sorts of business critical workloads. I also found this interesting myself due to Mako Networks cloud based Central Management System (CMS) that controls the thousands of devices deployed globally, and the multi-tenancy device use case for the VPN concentrator. If you found this interesting and want to know more about Mako Networks I would encourage you to visit their web site – www.makonetworks.com. Their system is incredibly cost effective and can provide secure PCI DSS certified merchant IP based payment processing over public networks. Even if you’re not interested in PCI DSS networking for payment processing, their system can be used to architect secure and affordable corporate networks with branch offices, with the bonus of being PCI DSS certified in the future. As always your comments and feedback are welcomed.
—
This post first appeared on the Long White Virtual Clouds blog at longwhiteclouds.com, by Michael Webster +. Copyright © 2013 – IT Solutions 2000 Ltd and Michael Webster +. All rights reserved. Not to be reproduced for commercial purposes without written permission.