Kudos to Heroku for taking full responsibility, and for planning to engineer around these kinds of Amazon problems in the future.
In particular, I'm delighted to hear that they plan to perform continuous backups on their shared databases:
3) CONTINUOUS DATABASE BACKUPS FOR ALL. One reason why we were able to fix the dedicated databases quicker has to do with the way that we do backups on them. In the new Heroku PostgreSQL service, we have a continuous backup mechanism that allows for automated recovery of databases... We are in the process of rolling out this updated backup system to all of our shared database servers; it’s already running on some of them and we are aiming to have it deployed to the remainder of our fleet in the next two weeks.
Combined with multi-region support, this should make Heroku far more resilient in the future.
Kudos? For nothing but the words "heroku takes 100% of the responsibility ..."?
Sorry, but that's not cutting it for me right now. I pay Heroku $250 a month and I was down for 60 hours (not 16). Our app isn't even out of private beta so I fully expected to be paying Heroku $2-3K/month by the end of the year. Now, I'm not sure I'll stay.
If you're really taking 100% responsibility, then consider pro-rating the bills of affected paying customers (based on the downtime).
I've run both cloud and non-cloud applications. In my experience, you won't get 99.95% annual uptime over 5 years without a full-time sysadmin, the ability to provision a complete offsite infrastructure and fail over to it within a few hours, and a backup/restore process that you rigorously test every month or so.
You generally won't get all that for $2K to $3K a month. Sure, you can drop $15K on an expensive database server, and co-locate it somewhere. But that only works until somebody takes a backhoe to your fiber, your RAID controller fails catastrophically, somebody pwns your production server, your sysadmin flakes out, or you discover that your backup scripts have been broken for months.
Realistically, if you're only spending $2-3K per month on hosting and administration, you'll eventually experience one or more of the above, and your site may be down for a day or more.
This isn't to say that I'm happy about Heroku's long downtime. One of my clients was offline almost as long as you were. But I'm pleased that Heroku recognizes just how badly they screwed up, and that they're taking the two most important steps they can to prevent a recurrence: multi-region support, and continuous backups for everyone. Multi-region support may not be sufficient to protect against cascading Amazon outages, but it's a good start.
The company where we host most of our servers has an SLA thats starts paying 10% monthly refund per 10 minutes of downtime that is their fault. I can't believe that you guys will take anything at this magnitude of downtime and still stick around.
Does Heroku even have an SLA? I can't find it. If they did maybe they would have been more proactive to prevent this kind of problem.
"Downtime that is their fault" is kind of a giant caveat, no? Is it their fault if they lose transit or power, for instance? With that level of refunds I suspect "their fault" basically only covers one of them accidentally running over a server with their car. The problem is, that guaranty isn't getting anyone anything of value.
I suspect Heroku has SLAs for their bigger customers, but don't really know for sure. I do think you're overestimating what kind of incentive an SLA is for a provider, though. SLAs are basically an on paper way of showing your commitment to keeping things running and responding to problems. If you don't have that commitment already, the paper isn't going to change anything.
Pointy haired bosses and lawyers love SLAs, but smart people who shop for this stuff don't care all that much about them. An SLA isn't going to convince me to go with one provider over another, nor is lack of an SLA going to make me avoid a provider I already like and respect.
I dont know about other peoples SLA's but seeing as you're hinging on my simplified description my SLA provides 100% uninterrupted transit to the Internet and 100% uninterrupted electricity so if the power goes out it is still 'their fault' but if I rm -rf / it is my fault.
I am not a lawyer or a PHB but I run a small business that has customers that pay for a service so if that service goes down I look bad and they are upset.
Oh, well 100% uptime for power and bandwidth is pretty standard then, I figured you were comparing an SLA for similar type services as you'd get from Heroku and/or EC2.
Is this reasonable? I'm sure a lot of amazon hosted companies are thinking similar thoughts.
But being on multiple Availability Zones was supposed to be bulletproof (according to Amazon). Now that we know that wasn't the case, is being hosted on multiple regions going to provide the necessary level of protection?
Is it an over-reaction to say that relying completely on Amazon could now be seen as irresponsible to your users, given the magnitude of this event?
The one problem I have with such concerns is this: what other viable options are there? Google AppSpot? Windows Azure? Perhaps. But AWS is flexible, very few stack limitations. The only other alternative I think is to go back to the pre-cloud era, when hosting was much more expensive, and outages were still possible, especially when you couldn't keep up with big traffic spikes.
Honestly, I would prefer this kind of mass outage than the alternative. It's cheaper, easier, and I bet you there's still better uptime overall.
However obviously it's good PR, and we all appreciate the Mea Cupla from Heroku, the fact is, they are proposing to migrate to a situation where they are still completely reliant on AWS for their hosting.
I'm just not sure you can really say "We don’t want to ever put our customers through something like this again and we’re working as hard as we can on making sure that we won’t ever have to.", when at the end of the day, you are again relying on a company that has failed you in the past.
Not trying to attack Amazon or Heroku, I'm honestly intrigued by this issue; not to mention the fact that we are facing the exact same decision at work.
Regarding Heroku's plan to continue relying on a company, Amazon, that failed them before:
If Heroku evolves to an architecture in which they utilize multiple AWS regions (as they mention in lesson #1 of their post-mortem) and if each region has a distinctly partitioned API "control plane," this should result in a materially improved availability situation for Heroku. EC2 Availability Zones guard against machine, power, and building failures. EC2 Regions should theoretically guard against API infrastructure and AWS software code failures.
Heroku need not necessarily ditch their current single-IaaS-provider architecture in order to achieve significantly better control over their service's uptime.
On the other hand, when downtime does occur, the ability for Heroku to prioritize their incident response manpower to first handle paying customers has its limits based on their downstream dependencies. If all the broken bits are within Amazon's black box, Heroku doesn't have much control over prioritization (Amazon fixes your stuff whenever it gets around to fixing your stuff). If Heroku operated over multiple cloud providers, even with the added complexity of such an approach, at least Heroku would have control over choosing which of their most important customers to migrate first to a working cloud, away from a broken and black box cloud.
In the end, I certainly don't see these considerations as simple. It's easy to cry when things go wrong, but I think the level of scalability and availability that has been achieved up to the present is quite noteworthy.
Interesting that this is essentially all stemming from yet again a communication failure from AWS. Once they have a post-mortem and can explain the multi-AZ issue, we may have a better idea of whether multi-region spread is sufficient redundancy. Or they could completely fail to communicate enough information, and adequately wary customers will be left with no choice but to assume that regions are not sufficiently independent.
I'm currently trying pretty hard to get one support answer out of Google to help prevent 6 hours of downtime next week. Go with the smaller companies who will at least respond when you need it.
Agreed. I'm trying to, but need them to link my old app to the new one for SSL API purposes, but that means a mystical form-filling-waiting-might-happen-might-not. I'll focus on the things that I can change while I wait!
Fair enough, point well taken. But I don't think acknowledging that difference changes the crux of the Q&A here. What IaaS does it better than AWS? Seriously. It's not a surprise that Heroku, Engineyard, and various other PaaS options are built on AWS.
We could just as easily say that relying completely on Heroku is irresponsible to our own users.
Just as Heroku took responsibility for the unexpected weaknesses their reliance on a single region created, I believe their customers should take responsibility for the unexpected weaknesses our reliance on a single hosting provider has created.
Heroku still has the value of added resiliency, even if it's not 110% bulletproof. Ultimately, we're responsible for the architecture design of our own sites.
In particular, I'm delighted to hear that they plan to perform continuous backups on their shared databases:
3) CONTINUOUS DATABASE BACKUPS FOR ALL. One reason why we were able to fix the dedicated databases quicker has to do with the way that we do backups on them. In the new Heroku PostgreSQL service, we have a continuous backup mechanism that allows for automated recovery of databases... We are in the process of rolling out this updated backup system to all of our shared database servers; it’s already running on some of them and we are aiming to have it deployed to the remainder of our fleet in the next two weeks.
Combined with multi-region support, this should make Heroku far more resilient in the future.