| Unique Visitors |
Hi Luca, I agree. I think the RPO and RTO as well as the maximum tolerable downtime should form part of the SLA. The system then needs to be designed to meet these SLA's and the right procedures and processes put in place. Often times the problem with the SLA's and the services is that the customer doesn't actually know the risk they're taking and therefore can't adequately plan. They have to base everything on the SLA contract with know real way to know if the service provider can meet it. That's why i like Jeramiah's way of determining the risk by looking over the old incident root cause and corrective action reports. At least then you can measure how they respond to incidents. Otherwise you may find that the SLA isn't worth the paper it's written on and is at best a target or aspirational goal, rather than something that can be relied upon.
]]>Also, in your examples there is the other problem with SLA: it’s an average value usually on a year base, while an outage impacts a business in every single event. I’d prefer to discuss (and write down on the contract) RPO and RTO with a Service Provider rather than SLA.
]]>There's a *great* Google+ (I know, I know) community for nothing but postmortem reports of all kinds. It's fascinating reading:
https://plus.google.com/u/0/communities/115136140…
Absolutely spot on. The SLA doesn't help you when the service has already been down way beyond the limits you've signed up for. But you've got to know what the limits and are manage the risks. Even with a highly professional and competent third party things can still go wrong. The only thing that can help you then is the transparency and learning from it and hopefully the compensation clauses in the contract. But as we know they will not usually cover consequential losses. I like the idea of reviewing the previous outage communications and how they're handled. That is a good way of measuring risk. The $10m an hour example wasn't the result of a third party but a combination of hardware failures, software bugs and human error. The organisation concerned did a thorough root cause analysis, created an in depth preventative action plan, performed a risk analysis and prioritised each risk area and mitigation steps and worked methodically over a period of time to reduce and eliminate the causes of the service interruption and implemented mitigations of additional risks. An additional outcome of the incident was additional business continuity processes outside of technology to reduce the impact of system outages. Not all solutions have to be technical or be part of IT. The Business Continuity plan, processes and risk assessments need to drive the requirements for IT solutions.
]]>1) You aren't sure what actually happened and don't know whether a full postmortem will be made available.
2) Your SLA was violated, but will that make you whole?
It's interesting to see how various service and services providers handle these things differently. So many times I've heard an out-sourced IT org say "We are paying for the SLA" while at the same time the provider is saying "Make the SLA 100% for marketing purposes, we're never on the hook for actual damages." In your example of the company losing $10M an hour, I promise if there was a 3rd party data center or service provider involved, the SLA didn't compensate them for that loss!
Postmortems are the best evidence of transparency between a provider and a customer. Before I sign up, I want to see where they have published all of the communication related to previous outages. I want to see how they communicate with their customers. I want to *know* that if there's an outage I'll know the good, bad and ugly of the event, because that's what I have to gauge risk and whether I want to stay.
Having been on the service provider side, the idea that customers buy SLAs is the single best piece of misdirection ever marketed.
]]>