Had a bunch of timeout alerts. Some machines of an application server array have the packet drop issue. ELB says that everything is peachy. Folks, we're experiencing yet another "EC2 flavored SNAFU™".
Initial though was: f#*k, I ran out of network I/O, since EC2 simply states "low, mid, high" as performance specs, therefore some proper planning is out. Turned out that with load avg under 0.02 on all machines, the I/O wait wasn't to blame. Average response time, per Pingdom, went up from 170ms to 450ms. New Relic isn't happy either. I guess we should all thank Amazon. Again.
I see the same thing. Pingdom reports higher response-time, but no downtime (meaning no alert). Also no alerts from AWS Cloudwatch. I first became aware of the issue when internal api-tests started failing at 9:56am CET. I see users accessing the site, but I don't know how many it's failing for.
No issues in the internal EC2 network. At least, none that I could find. I guess that's the reason why ELB doesn't shift any traffic. The whole issue seems to be on the Internet facing network. Failing routers, maybe.
Pingdom still claims 100% uptime, but New Relic (which includes an equivalent pinging service) reports downtime from time to time. Around 25 timeout alerts into the last couple of hours.
Comments
Had a bunch of timeout alerts. Some machines of an application server array have the packet drop issue. ELB says that everything is peachy. Folks, we're experiencing yet another "EC2 flavored SNAFU™".
Initial though was: f#*k, I ran out of network I/O, since EC2 simply states "low, mid, high" as performance specs, therefore some proper planning is out. Turned out that with load avg under 0.02 on all machines, the I/O wait wasn't to blame. Average response time, per Pingdom, went up from 170ms to 450ms. New Relic isn't happy either. I guess we should all thank Amazon. Again.
I see the same thing. Pingdom reports higher response-time, but no downtime (meaning no alert). Also no alerts from AWS Cloudwatch. I first became aware of the issue when internal api-tests started failing at 9:56am CET. I see users accessing the site, but I don't know how many it's failing for.
No issues in the internal EC2 network. At least, none that I could find. I guess that's the reason why ELB doesn't shift any traffic. The whole issue seems to be on the Internet facing network. Failing routers, maybe.
Pingdom still claims 100% uptime, but New Relic (which includes an equivalent pinging service) reports downtime from time to time. Around 25 timeout alerts into the last couple of hours.
Down vote? Really? I guess I should be thankful that I didn't have to explain this to a client: http://i.imgur.com/URf0H.png