8 ms·
Announcing MoSQL
- jabagonuts 14y agoAt what point do you abandon mongodb and just use postgresql?
- danielpal 14y agoAt no-point. Stripe is doing it right. They are using the right tool for each job. Mongo for storage speed etc and then postgres to analyze query etc. This kind of comment shows how little knowledge you have about NoSQL and SQL. Is not a SQL vs NoSQL, it's about using the right technology for the job.
- seanwoods 14y ago> This kind of comment shows how little knowledge you have about NoSQL and SQL. The question is perfectly valid. In many scenarios (not necessarily Stripe's), PostgreSQL is fast enough to do the job. Stop putting people down for legitimate engineering questions.
- dennis82 14y agoare you kidding me? There is absolutely NO reason whatsoever to use a NoSQL database for a financial services company. Postgres is more than capable of sustaining the necessary speeds of a startup. Relational databases were created in the first place to solve these very problems around transactionality and analytics for finance. This library is a beautiful example of reinventing the wheel, and otherwise creating a patchwork of unnecessary - and ultimately brittle - infrastructure.
- rdouble 14y agoEverywhere I've worked that did high volume transaction processing had an architecture that required a piece like this. Even if you use a relational database for intake, you still need to move the data to another database for analytics. Moving the data automatically via replication sounds a lot better than the typical batch process running at 4am.
- gdb 14y ago(I work at Stripe.) Where we use MongoDB, it's not because of speed. PostgreSQL is certainly capable of fast performance. MongoDB is useful for its ability to log freeform data as well as for its replication model. (We use sharded MongoDB in a few places, but mostly use straight replica sets.) We use MySQL, MongoDB, PostgreSQL, and Impala. They're all useful in different places.
- spicyj 14y agoWhy do you use MySQL over Postgres and vice versa?
- dennis82 14y ago> "We use MySQL, MongoDB, PostgreSQL, and Impala." Thanks for the clarification, but this makes it even more obvious your engineering team is introducing needless complexity into your organization. Postgres can store unstructured data just fine, so you have a 'solution' that uses 3 OLTP stores instead of one.
- taligent 14y agoPostgreSQL is awful for storing unstructured data. It is the most cumbersome, clunky syntax I've seen for a while and it lacks ORM support meaning you are forced to manually write it. Making developers productive is an important aspect for choosing a database.
- gfodor 14y agoChoosing a data store based upon syntax and slightly limited ORM support isn't exactly a great idea. Both of these things can be improved rapidly with a little code. More important questions are how is the data stored, how is it accessible, how can you scale the system, what operational constraints are there, how fast is it, what types of data modeling can be done, what consistency/transaction guarantees does it provide, etc. These are the things that will make developers productive because they will not be putting out fires all the time.
- Ingaz 14y agoTell this FIS Global. There is absolutely no reason to make banking system on GT.M but they did. Although: GT.M is the only(?) NoSQL that is ACID-compliant.
- taligent 14y ago> There is absolutely NO reason whatsoever to use a NoSQL database for a financial services company Yes there is. PostgreSQL doesn't support multi master replication which makes it a terrible choice if you really want to make sure every transaction gets written. I really wonder at what point people that keep recommending PostgreSQL are going to wake up and realise what is happening in the industry. People are scaling OUT not UP. Especially startups.
- nirvdrum 14y agoIn all fairness, you could use something other than Postgres that's also ACID.
- shawn-butler 14y agoI'm sorry, postgres-xc doesn't work for you needs? [0] It has worked for me in the past. [0] http://postgres-xc.sourceforge.net/ http://postgres-xc.sourceforge.net/
- knightni 14y agoI would imagine that for your average startup, using solutions that don't even support transactionality will cause greater complexity issues. Especially given the enormous window before db scale out/up becomes an issue on well-designed applications.
- taligent 14y agoEnormous window ? Many startups would be using AWS and it is not inconceivable that you would have Multi-AZ/Multi-Region VPSs. Scaling out != Expensive.
- j-kidd 14y ago> People are scaling OUT not UP. Especially startups. Startups need to scale out because many of them like to deploy on mediocre EC2 instances with the slowest SAN storage ever. People that keep recommending PostgreSQL are rightfully ignoring this industry.
- lucian1900 14y agoThe only advantage MongoDB has over Postgres is built-in sharding, and even that is of dubious value.
- nelhage 14y agoTo pick one, we like the fact that MongoDB lets you change your schema and add new fields to your documents without having to worry about migrations or keeping track of schema versions, or any of that. You could build something like that on top of SQL, but it's nice to have a tool where you don't have.
- djb_hackernews 14y agoHow does MoSQL handle schema changes and new fields in mongo? I'm imagining with this tool you start to need to be a bit more careful with the flexibility which initially drew you to mongo.
- nelhage 14y agoMoSQL will just throw any fields it doesn't recognize into a JSON "extra_props" field (if you ask it to). So everything will work fine, and existing SQL code (which doesn't know about those fields) will continue to be fine. If you need the data in SQL, you can either parse the JSON somehow, or rebuild the SQL table with a MoSQL schema that knows about the new fields.
- psaintla 14y agoSerious question to you or anyone else who uses schemaless databases. Why is the ability to change schemas on the fly a good thing? Having worked at two companies that did, it was nothing but a recipe for disaster in large groups. Code that was dependent on expecting an integer or a string and not a collection would constantly break because a developer in some other group decided to store a collection instead of a the original data type that was expected. Schemaless databases required more documentation to track changes made between groups and led to more bugs because we could never be guaranteed of what kind of data we would be receiving. I've always thought of a database schema as a contract that makes guarantees to all applications. Why would you want to be able to break that contract?
- eksith 14y ago>This kind of comment shows how little knowledge you have about NoSQL and SQL. Try not to be condescending and your point will be better received. "Right technology" as I'm sure you're aware, has as much to do with subjectivity as appropriateness. Familiarity, workflow, ease of use (and did I mention familiarity?) cannot be overstated even when the perceived benefits are considered. Read: religion. Some of the people who rally against NoSQL may be deriding it from a knee jerk reaction, however others are simply frustrated with developers who, as Ted Dziuba would say, "value technological purity over gettin' shit done".
- deleted 14y ago[deleted]
- nodesocket 14y ago10gen also has a nice python app which syncs by tailing the MongoDB oplog to an external source. Most common is Solr. https://github.com/10gen-labs/mongo-connector/tree/master/mongo-connector https://github.com/10gen-labs/mongo-connector/tree/master/mo... Seems to be high quality, and supports replica sets.
- Ensorceled 14y agoNice. Real businesses need a data warehouse and SQL is the right tool for that job. I thank them for releasing this.
- thesis 14y agoMaybe I'm misunderstanding your comment but... Real businesses need real solutions for their use cases. SQL is not necessarily the right tool for "that" job.
- PommeDeTerre 14y agoIf such a "solution" involves safely querying, analyzing, storing and manipulating data in any way, SQL and relational databases are usually the best option in practice. It's much more effective and efficient to use a SQL query than it is to throw together a huge amount of imperative JavaScript code (that's usually very specific to a single NoSQL database, as well) merely to perform the equivalent query. It's much safer to use a database that offers true support for transactions and constraints, rather than trying to hack together that functionality in some Ruby or PHP data layer code, or relying on some vague promise of "eventual consistency", for instance. It's much more maintainable, and leads to higher-quality data, to spend some time thinking about a schema, rather than just arbitrarily throwing data into a schema-less system, and then having to deal with the lack of a schema throughout any application code that's ever written. Aside from an extremely small and limited handful of situations (Google and Facebook, for instance), relational databases are the best tool for the job.
- taligent 14y ago> Real businesses need a data warehouse and SQL is the right tool for that job. Honestly. I don't think you could be more misinformed if you tried. Hint: Google "Big Data".
- knightni 14y ago...data warehouses in general mostly use SQL, and lots of businesses use data warehouses successfully. Teradata, Netezza, Oracle, DB2, etc. I'm not sure why his statement was controversial - SQL's a great language for reporting and analytics.
- physcab 14y agoThis is pretty cool but I'm struggling to see what the use cases are, atleast for analysis. There might be quite a bit of benefits for running application code that I'm not aware of. With regards to analysis though, their own example question is "what happened last night?" but then they go on to say that it is a near real-time data store. Does it matter that it is a real-time mirror then? I've always liked the paradigm of doing analysis on "slower" data stores, such as Hadoop+Hive or Vertica if you have the money. Decoupling analysis tools from application tools is both convenient and necessary as your organization and data scales.
- deleted 14y ago[deleted]
- bradleyland 14y agoThat's your preference, but I've not often found an occasion were someone (a client/stakeholder) said, "Yeah, go ahead and give that to me slower rather than faster." It just never seems to happen. I'm thinking of this as something like polyglot memoization. Pretty cool when you think about it. Frequently need something that is slow in NoSQL, but fast in SQL? Memoize it to your SQL datastore. The alternative has always been to write it to two places. I kind of dig moving this out to the datastore to figure out. I'm thinking that plenty of people will find this useful.
- physcab 14y agoThat's why I'm curious what sort of questions they are answering with this tool. If the bulk of their questions are a variant of SELECT COUNT(DISTINCT user_id) FROM table, then yes, this would be convenient to have. But if their questions start to revolve around transaction cohorting or path analysis where there are potentially hundreds of millions to billions of transaction_ids with some gnarly JOINs thrown in for good measure, I would be surprised to see this scale.
- nelhage 14y ago(I wrote MoSQL) PostgreSQL scales surprisingly well for this purpose, and is much nicer for interactive queries than Hadoop/Hive. We use Impala[1] for some larger datasets, but Impala is comparatively new, and it's nice to have something as battle-tested as postgres here. As for the "why do we need realtime?": In my mind the benefit of a near-realtime replica is not that you actually often need it, but that it means you never have to ask the question of "Was this snapshot refreshed recently enough?", and never end up having to wait several hours for an enormous dump/load operation, when you realize you did need newer data. [1] http://blog.cloudera.com/blog/2012/10/cloudera-impala-real-time-queries-in-apache-hadoop-for-real/ http://blog.cloudera.com/blog/2012/10/cloudera-impala-real-t...
- e1ven 14y agoVery neat project. I can see several use-cases for this where I work- It'd be nice to have alternatives means of searching through data. I'd also like to mention a project I've been contributing to, Mongolike [My fork is at https://github.com/e1ven/mongolike https://github.com/e1ven/mongolike , once it's merged upstream, that version will be the preferred one ;) ] It implements mongo-like structures on TOP of Postgres. This has allowed me to support both Mongo and Postgres for a project I'm working on.
- danso 14y agoOut of curiousity, but what is the rest of Stripe's stack like? Ruby, apparently, but I'm assuming they don't use any kind of Mongo ORM at all.
- hgimenez 14y agoAuthor of MoSQL, did you consider just using the MongoDB FTW instead? https://github.com/citusdata/mongo_fdw https://github.com/citusdata/mongo_fdw
- nelhage 14y ago(I wrote MoSQL) I actually played with mongo_fdw. At this point, it's a really cute hack, and useful for some things, but it doesn't give Postgres enough information and knobs to really let the query planner work effectively, so it ends up being really slow for complex things. I do love the concept, though.
- BlackJack 14y agoWhat were your thoughts on MongoConnector? (https://github.com/10gen-labs/mongo-connector/tree/master/mongo-connector https://github.com/10gen-labs/mongo-connector/tree/master/mo...)
- andrewjshults 14y agoIs there currently support for "unrolling" arrays or hashes into tables of their own? If not, would definitely be interested in helping to add that on (we use arrays on documents quite a bit, but have run into a number of situations where a simple SQL query for analysis could have quickly replaced a bunch of mongo scripts).
- nelhage 14y agoThere isn't support. It's definitely something I've pondered. If you're interested in adding support, I'd be happy to hear from you at (my username) AT stripe.com.
- andrewjshults 14y agoEmail sent!
- dcraw 14y agoI've added that capability to mongo_fdw, which I use for getmetrica.com. I'll be contributing it back soon (after that 9.2 API conversion). Would be happy to talk to you about the wrapper or Metrica. Email's in my profile.
- andrewjshults 14y agoFYI, the email field from the profile doesn't actually get displayed publicly. Mine is (username) @gmail.com
- dcraw 14y agoWhoops. Okay, emailing.
- dugmartin 14y agoReading the headline I thought they were introducing a SQL like interface to their API, sort of like FQL for Facebook and I got a little excited. Something like this to get the email addresses of all your active trial subscribers: SELECT c.email FROM customers c, subscriptions s WHERE c.subscription_id = s.id AND s.status = "active" and s.trial_start IS NOT NULL; (where of course the customer and subscription tables would be a virtual view on your customers and subscriptions)
- gfodor 14y agoFYI you can store unstructured data in PostgreSQL (and query it) with the introduction of hstore. So knock one more reason to use MongoDB instead of PostgreSQL off your list. (Disclaimer: the length of my list to use MongoDB has always been a constant that is less than one.) http://www.postgresql.org/docs/9.1/static/hstore.html http://www.postgresql.org/docs/9.1/static/hstore.html
- mrkurt 14y agoWow, hstore really isn't a great alternative to an actual document DB. The "better" Postgres option would be a JSON type and functional indexes.
- shawn-butler 14y agoThere is a JSON type but it just validates content. HSTORE can be fully indexed (gIST and GIN). Just have to roll your own object graphs for nesting if that's what you need to do. I swear I have typed this exact same comment previously. Deja vu, maybe
- mrkurt 14y agoJSON type gives you some typed values within the doc, multi-level nesting, etc. You can add functional indexes (http://www.postgresql.org/docs/9.1/static/indexes-expressional.html http://www.postgresql.org/docs/9.1/static/indexes-expression...) to index specific attributes within the JSON, do legit sorts over values, reasonable array queries, etc. It seems much, much closer to what Mongo does than anything you can do with hstore.
- shawn-butler 14y agoI think you just restated my comment. Do you believe expression indexes do not apply to HSTORE? I consider both HSTORE (key/value) and the current JSON type and record functions are just intermediate steps to a fuller API [0]. [0]: http://www.postgresql.org/message-id/50EC971C.3040003@dunslane.net http://www.postgresql.org/message-id/50EC971C.3040003@dunsla...
- Ingaz 14y agoI thought that "young" NoSQLs sometime in will got SQL interface. Look at old NoSQLs: Intersystems Cache got SQL interface, GT.M (in PIP-framework) also got SQL. My impression that MongoDB looks a lot like MUMPS storage with globals in JSON.
- scragg 14y agoSomeone should write a client library so you can do ad hoc data aggregation queries without using SQL. You can call it NoMoSQL :)
- bryanjos 14y agoI love this idea. I can see myself using MoSQL pretty soon. Does it handle geospatial data? Can it replicate geospatial data from Mongo to a Geometry data type in Postgres?
- meaty 14y agoAlso useful when MongoDB blows chunks because it was a crap architectural decision and you quickly port your app to raw SQL...
- govindkabra31 14y agohow do you deal with sharded mongo clusters?
- umur 14y ago(disclosure: I'm one of the founders at Citus Data) hey, one way to do that is to use the MongoDB foreign data wrapper - also mentioned in some of the earlier threads. mongo_fdw (https://github.com/citusdata/mongo_fdw https://github.com/citusdata/mongo_fdw) allows you to run SQL on MongoDB on a single node. Citus Data allows you to parallelize your SQL queries across multiple nodes (in this case, multiple MongoDB instances) by just syncing shard metadata. So you would effectively run SQL on a sharded mongo cluster without moving the data anywhere else. another idea could be to use MoSQL to neatly replicate each mongo instance to a separate PostgreSQL instance, and then use Citus Data to run distributed SQL queries across the resulting PostgreSQL cluster.
- ElGatoVolador 14y agoIf you need to make a tool(and use twice the amount of storage) to be able to "query your data" in a SQL manner while using noSQL, it probably means you are using the wrong tool for the job.
- j-kidd 14y agoActually, it is pretty common to replicate the transactional data into another data store for analytical purpose. However, using PostgreSQL as the OLAP data store may not be the wisest move.
- dschiptsov 14y agoMongoDB is great for a lot of reasons - record-level locking? multiple concurrent writes? append-only journals? I have read than in version 2.x they announce some features, so, it is greatness?
- arthulia 14y agoCan't wait for NoMoSQL
- Uchikoma 14y agoWaiting for BroSQL.