Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

"When things start to go wrong, your automatic recovery code will increase the load on your system, and commonly lead you into a spiral of death."

Here is a Battlebots story; When I was competing we would see new teams in the pit area at the beginning of the competition (there were 3 days of preliminaries) that had really nice looking bots. I'd ask them, "So, have you run it full speed into a concrete wall?" And they would either say "Yeah, wow it was amazing ... " and tell some story of mayhem, or "No." (sometimes with a prognostication of confidence in their design skills or their simulations).

Teams that said "No" never made it out of the preliminaries. Not once in my experience.

That story underpins a fundamental truth in systems analysis, "Beat it until it fails before you depend on it."

This is something that Google does really really well by the way, I've watched them turn of 25 core routers simultaneously carrying hundreds of gigabits worth of data, just to verify that what they think will happen, does happen.

You learn not only what breaks, but if you go through the fire drill of bringing it back online, and you take copious notes when people say "Dammit! I need to do 'x' and I can't." your ability to respond will improve.

For very large systems, this can sometimes be the only way to develop this information.

Amazon has clearly had an event of extraordinary magnitude in their data centers. They've no doubt discovered all sorts of tools that they could use to recover more quickly. I would love it if someone from there would post a complete post mortem, but my expectations are low (there is a lot of proprietary benefit in knowing some of this stuff).



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: