(cas:72) Google Analyticator was unable to authenticate you with Google using the Auth Token you pasted into the input box on the previous step.

This could mean either you pasted the token wrong, or the time/date on your server is wrong, or an SSL issue preventing Google from Authenticating.

Try Deauthorizing & Resetting Google Analyticator.

Tech Info 400:Error fetching OAuth2 access token, message: 'invalid_grant'
Unique
Visitors
Powered By Google Analytics
Comments on: Hardware Fails, Software Has Bugs, People Make Mistakes – Usually You Get All At Once! http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/ all things Nutanix, VMware, cloud and virtualizing business critical applications Sat, 19 Apr 2014 02:41:03 +0000 hourly 1 https://wordpress.org/?v=6.7.6 By: vcdxnz001 http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32731 Sat, 19 Apr 2014 02:41:03 +0000 http://longwhiteclouds.com/?p=3334#comment-32731 In reply to Luca Dell'Oca.

Hi Luca, I agree. I think the RPO and RTO as well as the maximum tolerable downtime should form part of the SLA. The system then needs to be designed to meet these SLA's and the right procedures and processes put in place. Often times the problem with the SLA's and the services is that the customer doesn't actually know the risk they're taking and therefore can't adequately plan. They have to base everything on the SLA contract with know real way to know if the service provider can meet it. That's why i like Jeramiah's way of determining the risk by looking over the old incident root cause and corrective action reports. At least then you can measure how they respond to incidents. Otherwise you may find that the SLA isn't worth the paper it's written on and is at best a target or aspirational goal, rather than something that can be relied upon.

]]>
By: Luca Dell'Oca http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32711 Fri, 18 Apr 2014 23:36:33 +0000 http://longwhiteclouds.com/?p=3334#comment-32711 From a Data Protection perspective, I’ve felt SLAs are often only numbers good for the marketing of the service provider. It’s an agreement, and the customer has no way to check if the SP is effectively able to guarantee that value. We agree on an SLA, as a customer I hope the real SLA would always be 100%, and if the SP violates it, at least I want some money back.

Also, in your examples there is the other problem with SLA: it’s an average value usually on a year base, while an outage impacts a business in every single event. I’d prefer to discuss (and write down on the contract) RPO and RTO with a Service Provider rather than SLA.

]]>
By: Brett Weaver http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32699 Fri, 18 Apr 2014 21:35:59 +0000 http://longwhiteclouds.com/?p=3334#comment-32699 " hopefully a full root cause analysis is made available to customers, along with a preventative action plan that gives customers confidence that steps are being taken to reduce the risk of another similar incident."
In my experience infrastructure is the best area in IT for adopting this approach. I am quite sure you will find things improve.
<RANT>Unfortunately Software Projects crash and burn -or- succeed and no one works out why! That's how you have project after project failing and costing a fortune.
Trying to get clients to run Post Implementation Reviews seems like the hardest thing in the world. What are the Universities teaching people in Business courses? </RANT>

]]>
By: » Heads Up Alert: vSphere 5.5 U1 NFS Random Disconnection Bug! Long White Virtual Clouds http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32563 Fri, 18 Apr 2014 01:26:46 +0000 http://longwhiteclouds.com/?p=3334#comment-32563 […] The timing of this is a coincidence that it comes right on the coat tails of my previous article Hardware Fails, Software Has Bugs and People Make Mistakes – Usually You Get All At Once! During the disconnects VM’s will appear frozen and the NFS datastores may be greyed out. […]

]]>
By: Jeramiah Dooley http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32558 Fri, 18 Apr 2014 01:06:59 +0000 http://longwhiteclouds.com/?p=3334#comment-32558 In reply to vcdxnz001.

There's a *great* Google+ (I know, I know) community for nothing but postmortem reports of all kinds. It's fascinating reading:
https://plus.google.com/u/0/communities/115136140

]]>
By: vcdxnz001 http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32557 Fri, 18 Apr 2014 01:03:16 +0000 http://longwhiteclouds.com/?p=3334#comment-32557 In reply to Jeramiah Dooley.

Absolutely spot on. The SLA doesn't help you when the service has already been down way beyond the limits you've signed up for. But you've got to know what the limits and are manage the risks. Even with a highly professional and competent third party things can still go wrong. The only thing that can help you then is the transparency and learning from it and hopefully the compensation clauses in the contract. But as we know they will not usually cover consequential losses. I like the idea of reviewing the previous outage communications and how they're handled. That is a good way of measuring risk. The $10m an hour example wasn't the result of a third party but a combination of hardware failures, software bugs and human error. The organisation concerned did a thorough root cause analysis, created an in depth preventative action plan, performed a risk analysis and prioritised each risk area and mitigation steps and worked methodically over a period of time to reduce and eliminate the causes of the service interruption and implemented mitigations of additional risks. An additional outcome of the incident was additional business continuity processes outside of technology to reduce the impact of system outages. Not all solutions have to be technical or be part of IT. The Business Continuity plan, processes and risk assessments need to drive the requirements for IT solutions.

]]>
By: Jeramiah Dooley http://longwhiteclouds.com/2014/04/18/hardware-fails-software-has-bugs-people-make-mistakes-usually-you-get-all-at-once/comment-page-1/#comment-32485 Thu, 17 Apr 2014 15:44:33 +0000 http://longwhiteclouds.com/?p=3334#comment-32485 There are two interesting bits here:

1) You aren't sure what actually happened and don't know whether a full postmortem will be made available.
2) Your SLA was violated, but will that make you whole?

It's interesting to see how various service and services providers handle these things differently. So many times I've heard an out-sourced IT org say "We are paying for the SLA" while at the same time the provider is saying "Make the SLA 100% for marketing purposes, we're never on the hook for actual damages." In your example of the company losing $10M an hour, I promise if there was a 3rd party data center or service provider involved, the SLA didn't compensate them for that loss!

Postmortems are the best evidence of transparency between a provider and a customer. Before I sign up, I want to see where they have published all of the communication related to previous outages. I want to see how they communicate with their customers. I want to *know* that if there's an outage I'll know the good, bad and ugly of the event, because that's what I have to gauge risk and whether I want to stay.

Having been on the service provider side, the idea that customers buy SLAs is the single best piece of misdirection ever marketed.

]]>