Kudos? For nothing but the words "heroku takes 100% of the responsibility ..."?
Sorry, but that's not cutting it for me right now. I pay Heroku $250 a month and I was down for 60 hours (not 16). Our app isn't even out of private beta so I fully expected to be paying Heroku $2-3K/month by the end of the year. Now, I'm not sure I'll stay.
If you're really taking 100% responsibility, then consider pro-rating the bills of affected paying customers (based on the downtime).
I've run both cloud and non-cloud applications. In my experience, you won't get 99.95% annual uptime over 5 years without a full-time sysadmin, the ability to provision a complete offsite infrastructure and fail over to it within a few hours, and a backup/restore process that you rigorously test every month or so.
You generally won't get all that for $2K to $3K a month. Sure, you can drop $15K on an expensive database server, and co-locate it somewhere. But that only works until somebody takes a backhoe to your fiber, your RAID controller fails catastrophically, somebody pwns your production server, your sysadmin flakes out, or you discover that your backup scripts have been broken for months.
Realistically, if you're only spending $2-3K per month on hosting and administration, you'll eventually experience one or more of the above, and your site may be down for a day or more.
This isn't to say that I'm happy about Heroku's long downtime. One of my clients was offline almost as long as you were. But I'm pleased that Heroku recognizes just how badly they screwed up, and that they're taking the two most important steps they can to prevent a recurrence: multi-region support, and continuous backups for everyone. Multi-region support may not be sufficient to protect against cascading Amazon outages, but it's a good start.
The company where we host most of our servers has an SLA thats starts paying 10% monthly refund per 10 minutes of downtime that is their fault. I can't believe that you guys will take anything at this magnitude of downtime and still stick around.
Does Heroku even have an SLA? I can't find it. If they did maybe they would have been more proactive to prevent this kind of problem.
"Downtime that is their fault" is kind of a giant caveat, no? Is it their fault if they lose transit or power, for instance? With that level of refunds I suspect "their fault" basically only covers one of them accidentally running over a server with their car. The problem is, that guaranty isn't getting anyone anything of value.
I suspect Heroku has SLAs for their bigger customers, but don't really know for sure. I do think you're overestimating what kind of incentive an SLA is for a provider, though. SLAs are basically an on paper way of showing your commitment to keeping things running and responding to problems. If you don't have that commitment already, the paper isn't going to change anything.
Pointy haired bosses and lawyers love SLAs, but smart people who shop for this stuff don't care all that much about them. An SLA isn't going to convince me to go with one provider over another, nor is lack of an SLA going to make me avoid a provider I already like and respect.
I dont know about other peoples SLA's but seeing as you're hinging on my simplified description my SLA provides 100% uninterrupted transit to the Internet and 100% uninterrupted electricity so if the power goes out it is still 'their fault' but if I rm -rf / it is my fault.
I am not a lawyer or a PHB but I run a small business that has customers that pay for a service so if that service goes down I look bad and they are upset.
Oh, well 100% uptime for power and bandwidth is pretty standard then, I figured you were comparing an SLA for similar type services as you'd get from Heroku and/or EC2.
Comments
Kudos? For nothing but the words "heroku takes 100% of the responsibility ..."?
Sorry, but that's not cutting it for me right now. I pay Heroku $250 a month and I was down for 60 hours (not 16). Our app isn't even out of private beta so I fully expected to be paying Heroku $2-3K/month by the end of the year. Now, I'm not sure I'll stay.
If you're really taking 100% responsibility, then consider pro-rating the bills of affected paying customers (based on the downtime).
I've run both cloud and non-cloud applications. In my experience, you won't get 99.95% annual uptime over 5 years without a full-time sysadmin, the ability to provision a complete offsite infrastructure and fail over to it within a few hours, and a backup/restore process that you rigorously test every month or so.
You generally won't get all that for $2K to $3K a month. Sure, you can drop $15K on an expensive database server, and co-locate it somewhere. But that only works until somebody takes a backhoe to your fiber, your RAID controller fails catastrophically, somebody pwns your production server, your sysadmin flakes out, or you discover that your backup scripts have been broken for months.
Realistically, if you're only spending $2-3K per month on hosting and administration, you'll eventually experience one or more of the above, and your site may be down for a day or more.
This isn't to say that I'm happy about Heroku's long downtime. One of my clients was offline almost as long as you were. But I'm pleased that Heroku recognizes just how badly they screwed up, and that they're taking the two most important steps they can to prevent a recurrence: multi-region support, and continuous backups for everyone. Multi-region support may not be sufficient to protect against cascading Amazon outages, but it's a good start.
"If you're really taking 100% responsibility, then consider pro-rating the bills of affected paying customers (based on the downtime)."
That's implied, at the very least, by the phrase "heroku takes 100% of the responsibility. "
I would be very surprised if they didn't offer more than that.
Your biggest concern is you want a $20 refund?
They can keep their $20 in my view, as long as they ensure it never happens again.
The company where we host most of our servers has an SLA thats starts paying 10% monthly refund per 10 minutes of downtime that is their fault. I can't believe that you guys will take anything at this magnitude of downtime and still stick around.
Does Heroku even have an SLA? I can't find it. If they did maybe they would have been more proactive to prevent this kind of problem.
"Downtime that is their fault" is kind of a giant caveat, no? Is it their fault if they lose transit or power, for instance? With that level of refunds I suspect "their fault" basically only covers one of them accidentally running over a server with their car. The problem is, that guaranty isn't getting anyone anything of value.
I suspect Heroku has SLAs for their bigger customers, but don't really know for sure. I do think you're overestimating what kind of incentive an SLA is for a provider, though. SLAs are basically an on paper way of showing your commitment to keeping things running and responding to problems. If you don't have that commitment already, the paper isn't going to change anything.
Pointy haired bosses and lawyers love SLAs, but smart people who shop for this stuff don't care all that much about them. An SLA isn't going to convince me to go with one provider over another, nor is lack of an SLA going to make me avoid a provider I already like and respect.
I dont know about other peoples SLA's but seeing as you're hinging on my simplified description my SLA provides 100% uninterrupted transit to the Internet and 100% uninterrupted electricity so if the power goes out it is still 'their fault' but if I rm -rf / it is my fault.
I am not a lawyer or a PHB but I run a small business that has customers that pay for a service so if that service goes down I look bad and they are upset.
Oh, well 100% uptime for power and bandwidth is pretty standard then, I figured you were comparing an SLA for similar type services as you'd get from Heroku and/or EC2.
Do they have a SLA? I presume they can pass on some of the refunds from Amazon.
You're going to complain about a refund of what, ~ $22, from a service with no SLA for a site in 'private beta'?
You must think pretty highly of your app. Why don't you take it out of 'private beta' and let the rest of us look at it?