Skip to content

Comment on Cloud Platform at Google I/O – new Big Data, Mobile and Monitoring products

Comments

This is one space where Google really excels.

We're in the AWS ecosystem, and the database offerings are really subpar. DynamoDB, which I originally expected to be somewhat comparable to MongoDB, is an incredibly frustrating (and expensive) product to use. AWS Data Pipeline is extremely confusing and very expensive as well.

AWS's offerings really lag behind Google's offerings (like BigQuery) in this space. Hopefully AWS can catch up because I'd rather not have requests bouncing between data centers.

If you are in AWS there is also the four RDS (Oracle, MySQL, PostresSQL, SQL Server) options as well as RedShift. Also the best thing about AWS is that there are so many third party choices e.g. MongoLab, MongoHQ, Instaclustr, Cloudant.

Databases is not the area I would be choosing Google for.

I think App Engine's Datastore is generally an under-appreciated gem. Possibly because you have to use App Engine to use it without sacrificing performance, maybe because it's not easy enough to use if all you have is some JavaScript + JSON and don't want or know how to write Python/Java/Go.

But it's actually the only generally available product I know of that solves all the hard problems (availability, partition tolerance, some - but well defined - consistency with cross entity transactions) with zero hassle for you.

If you read through http://aphyr.com/tags/Jepsen, you get some appreciation for how hard this is to pull off without running into operational nightmares (massive data loss, split brains, etc).

Disclaimer: I work for Google, though not on Datastore.

We've had good luck with DynamoDB, but it could be that it just fits our use case very well. What sort of frustration were you running into? (Honestly interested to avoid trouble down the line)

Most recently: hot hash key. DynamoDB uses the object to be persisted's hash key to route it to the right data cluster.

We're a SaaS company with lots of tiny customers and a few very large customers. We need to keep an index to show a specific customer only their data. That means the index for our largest customers gets hit a lot. The problem with this structure is we have to pay as if all of our customers were as popular as our biggest customers, or we get throttled. And even though the DynamoDB interface shows that you have provisioned 10x above your current usage, you still get throttled, because you're being throttled only in a single cluster.

So, let's say you solve that problem, but now you need to drop the troublesome index on a billion+ row table. With DynamoDB you can't change a table's indexes, so you have to migrate your table to a new table. Doing that without downtime is an incredible challenge.

Which reminds me of when they announced indexes. We were so excited only to find out we couldn't add indexes to our tables, but instead had to recreate them all.

The whole point of SaaS is to make our lives easier, but with DynamoDB our lives were much more difficult than just using Mongo.

Anyway, I need to do a blog post on this -- it's a bit too complicated for a HN comment. :-)

Hot keys are going to be the same with Mongo. The issue sounds more like that you're using a single key per customer than anything else.

Yeah, I think all nosql db's will have that issue if you have extremely unbalanced sharding. This is an application level fault and should be solved there.

But, the thing that is extra bad about dynamo db is how they are paying for 10x higher provisioning as a stop gap, and still getting throttled. That sucks.

FWIW we're using dynamo db and we love it. Pro tip: setup dynamic-dynamodb and let it autoscale for you in realtime. http://aws.amazon.com/blogs/aws/auto-scale-dynamodb-with-dyn...

Please do. It sounds like a very interesting adventure.

What do you think is the analogous product to Cloud Dataflow in the AWS ecosystem? SWF? http://aws.amazon.com/swf/

I believe it's Kinesis: http://aws.amazon.com/kinesis/

+1. It's Kinesis.

DataFlow is not Kinesis. It's more like Kinesis plus Esper plus BigQuery and you still wouldn't have one set of queries to run against streaming and batch data like you do with DataFlow.

I presume AWS Data Pipeline. But they have their differences, so perhaps there is no true analog.

http://aws.amazon.com/datapipeline/

Simple Workflow is new to me, so thanks for putting it on my radar!

I've been using Simple Workflow (in particular, the Flow framework: http://aws.amazon.com/swf/details/flow/) a lot recently to manage complex asynchronous distributed workflows, and it's a revelation. Some people laugh about the "Simple" part in the name, but it's kind of true - once you get over the initial learning curve. It's a bit like Git that way (First a lot of banging your head, then a lot of bang for buck)

Have you looked at AWS Redshift? It is somewhat analogous to BigQuery.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.