7 ms·
Designing Schemaless, Uber Engineering’s Scalable Datastore Using MySQL (2016)
- verst 9y agoThis was discussed here 2 years ago. https://news.ycombinator.com/item?id=10894047 https://news.ycombinator.com/item?id=10894047
- alpb 9y agoIndeed discussed 4 times before https://hn.algolia.com/?query=schemaless%20datastore%20mysql&sort=byPopularity&prefix&page=0&dateRange=all&type=story https://hn.algolia.com/?query=schemaless%20datastore%20mysql... This also should be titled "(2016)"
- ronnier 9y agoSummary: > We ended up building a key-value store which allows you to save any JSON data without strict schema validation, in a schemaless fashion (hence the name). It has append-only sharded MySQL with buffered writes to support failing MySQL masters and a publish-subscribe feature for data change notification which we call triggers. Lastly, Schemaless supports global indexes over the data.
- samstave 9y agoLet’s also not be smug, let’s explain for everyone who comes here: why is this good or why is it bad. So describe why this is good or describe why it’s bad That comment was edited with more info...
- deleted 9y ago[deleted]
- rubyn00bie 9y agoCan this have a 2016 added to it, please?
- dang 9y agoYes.
- stmw 9y agoGreat catch!
- mangatmodi 9y agoOP here, I had no idea that it had been discussed here. I am not able to edit the post now.
- trans_exUsual 9y agoAnd it still needs a different name.
- crtqt3 9y agopronounced "she-males"
- whalesalad 9y agoStill blows my mind that it took Uber so long to migrate away from a single db solution. The bit about wanting an event system to handle downstream trip processing w/o having one failure block the whole job was shocking. I’m all for avoiding premature optimization but this was taken to the extreme. PostgreSQL is capable of all of this out of the box. Wonder why a custom tool was built instead?
- sidlls 9y agoPerhaps they had an existing system they thought might cost more to rewrite with a different DB.
- subway 9y agoStill blows my mind that it took Uber so long to migrate away from a single db solution. The bit about wanting an event system to handle downstream trip processing w/o having one failure block the whole job was shocking. I'm pretty sure they've seen a few iterations on their data stores. I remember attending a meetup at Urban Airship in 2012 or so where Uber engineers gave a presentation about a data store migration (I think from Mongo to MySQL)
- mangatmodi 9y agoAbout postgres: Uber has bashed it really bad here https://eng.uber.com/mysql-migration/ https://eng.uber.com/mysql-migration/
- gaius 9y agoPostgreSQL is capable of all of this out of the box. Wonder why a custom tool was built instead? They have over-hired engineers is the obvious answer.
- crondog 9y agoso, has anyone ever pointed out that 'Schemaless' looks like how a spammer would spell 'shemales'?
- dcposch 9y agoFYI, MySQL has a fresh new JSON data type now. It has some great properties. It lets you mix data with a strict schema and data without a strict schema, getting some of the benefits of both worlds. The JSON datatype avoids many of the annoying legacy considerations that other SQL column types have. You don't have to specify a length--so you won't make a VARCHAR(255), then get burned when one day a value has more than 255 characters. You don't have specify a character encoding--JSON is always utf8mb4, the right one. (MySQL's 'utf8' encoding, perversely, supports only a subset of utf8 and will break if you try to write an emoji.) Here's a table that illustrates some of the power: create table unitType ( id bigint not null auto_increment, buildingId bigint not null, info json, name varchar(255) as (info->>'$.name') not null, primary key(id), foreign key (buildingId) references building(id) on delete cascade, unique key(buildingId, name) ); We're modeling unit types in a building. For example, one building might contain 1-bedrooms, some nicer 1-bedrooms, and some 2-bedroom units. - It's very easy to add new fields. If, tomorrow, we decide that each unit type needs a `minSqft` and `maxSqft`, I can add them with no database migration. - We still get most of the benefits of a schema. The database makes it impossible for a unitType to exist that does not belong to a building. The database also makes it impossible for a single building to have two unitTypes with the same name. (With a truly schemaless DB like Mongo, the complexity of preventing or dealing with those kinds of invalid data end up in the application code.) - It makes it easy to use SQL directly, with no ORM. SQL is a powerful language; ORMs are often a leaky abstraction and a source of unessential complexity. With JSON columns for extensibility, you end up with way fewer migrations and way less need for auto-generated SQL. - Computed columns (like name above) are really powerful. Most of the above is possible in Postgres as well. Postgres does not have computed columns, as far as I can tell. -- This is just to say: 99% of people on Hacker News are closer to where we are (rapid prototype phase) than where Uber is (Web ScaleTM). If that's you, consider just using JSON columns to maximize your development velocity! You can always do something fancier (like Schemaless) later on.
- pritambaral 9y ago> You don't have to specify a length--so you won't make a VARCHAR(255), then get burned when one day a value has more than 255 characters. Does MySQL not have a TEXT data type, or is it not well-supported or otherwise disadvantaged? > It's very easy to add new fields. If, tomorrow, we decide that each unit type needs a `minSqft` and `maxSqft`, I can add them with no database migration. TBF, column-adding being a pain is really a MySQL-specific problem. > It makes it easy to use SQL directly, with no ORM. SQL is a powerful language; ORMs are often a leaky abstraction and a source of unessential complexity. With JSON columns for extensibility, you end up with way fewer migrations and way less need for auto-generated SQL. ORMs being a "leaky abstraction" is a good thing; good ORMs are not too far removed from SQL precisely because they're leaky. Pushing schema management to the app layer (as opposed to migrations) is also a source for "unnecessary complexity".
- mewse 9y agoI wish I had video of the faces I was undoubtedly pulling, during the few seconds I spent puzzling out the pronunciation and meaning of the word "Schemaless". /Shema-leez/ ? /Szhee-males/ ? Naming products is demonstrably a hard problem.
- deleted 9y ago[deleted]
- singularity2001 9y agoGod forgive me for reading 'she-males'
- mangatmodi 9y agoAny critique on Uber's use of Triggers for triggering billing service? I have been reading that Triggers shouldn't be used to esp, call external services as, the external service might not be ACID compliant(no rollback?) and if expensive, they can hold the DB lock on the row for really long time.