10 ms·
How to make MongoDB not suck for analytics
- jrochkind1 8y agoWhat is the benefit of having it in mongo in the first place, in this scenario?
- twblalock 8y agoThe people who write the business logic and the people who do the analytics have different concerns. It's sometimes better to make different database choices for those two systems and just copy the data into the analytics system, rather than make a substandard choice of database to try to accommodate both. If the devs want to use Mongo, it's their problem -- it shouldn't matter much to the analytics people, because they can just copy the data into a different database that fits their needs.
- jrochkind1 8y agoFair enough. Perhaps there are other uses of mongo going on in addition to exporting it to a different database/format? Otherwise I'd be curious what justifies the mongo choice. Certainly sometimes you are in a position where you just gotta take what other units in the org give you and deal with it. But _someone_ in the org is hopefully in the position to be able to articulate why they are using mongo in the first place...
- AznHisoka 8y ago"they can just copy the data into a different database that fits their needs." That's easier said than done when your database is over 10 TB big.
- twblalock 8y agoYou only need to copy the stuff that changed since the last time you copied stuff. If you are really generating 10TB of data more than a few times a day, you can look into putting it in something like Kafka for real-time consumption by the analytics team instead of batch copying.
- ayw 8y agoYou just need to read the oplog, so it only needs to track your saves. In general, you probably should have at least something in your stack which reads all changes from your DB, at the very least for backup reasons.
- dizzystar 8y agoExcept that devs have to do ETL every day so analysts can do their query work.
- twblalock 8y agoThat can be automated.
- dizzystar 8y agoIn theory, yes; in practice, not really.
- sdoering 8y agoI tend to disagree. Having multiple automated ETL processes running for different projects/clients/colleagues I see that the code does not change as often, as I had anticipated. Automation here (in my case) is a net win on time.
- marcinzm 8y agoWhy? Technically speaking, simple ETL is easy to automate and not too much maintenance headache.
- marcinzm 8y agoWhich you'll likely want with any DB since analytics workloads are very different than the usual production DB workloads.
- mitchell_h 8y agoThis is such a common issue there's an entire architecture pattern developed to solve it. https://en.wikipedia.org/wiki/Lambda_architecture https://en.wikipedia.org/wiki/Lambda_architecture . No classic ETL, and any number of folks/systems can plug into the messaging system and get all the data.
- sbr464 8y agoI agree with this perspective, and have been researching it more lately.
- ayw 8y agoFor better or for worse, MongoDB tends to be easier for developers move quickly, so it ends up getting adopted quite a bit. This is more about how to deal with it after it's already in your stack.
- omeid2 8y agoRethinkDB blows MongoDB on easy to use factor out by a large margin, with the upside of being a project focused on actual quality rather than pure marketing.
- chiaolun 8y agoRethinkDB doesn't get enough love. It's rare to see anything pass the Jepsen tests to the degree that Rethink did: https://aphyr.com/posts/329-jepsen-rethinkdb-2-1-5 https://aphyr.com/posts/329-jepsen-rethinkdb-2-1-5 It's sad that, for a backend DB, correctness can be trumped by marketing.
- scalablenotions 8y agoWhat about Managed solutions, like DynamoDB? What could be easier than that - with cloud scale analytics opportunities to boot.
- riboflavin 8y agoDremio helps with a lot of this, particularly the speed aspect – uses Parquet as well as Apache Arrow. (I work at Dremio.) Speeding things up: https://docs.dremio.com/acceleration/reflections.html https://docs.dremio.com/acceleration/reflections.html
- alextheparrot 8y agoI’m only familiar with speeding up Parquet - it looks like you’re mainly sorting or partitioning the data into different views so that you can choose the best view format at runtime based on your desired query or aggregation? We’ve seen this increase speeds by many orders of magnitude, so I wouldn’t be surprised that this creates meaningful speed-ups when done automatically (Which is cool!). Random aside, how do you handle the consistency problems that can occur when you have multiple views when doing deletes?
- nevi-me 8y agoDremio quickly becomes useless with MongoDB given that for a while it's not been possible to join data from two MongoDB collections by their object IDs. Last time I checked, Dremio mangled the id into some string that can't even be matched to the same id on a separate collection. I had data in PG and Mongo, but couldn't join it together. I asked about this on the forum, was told it's a known issue; and it seemed to end there. I resorted to doing my analytics by hand in the end, MongoDB's aggregation framework is good enough. Create views from aggregation queries, and it becomes easier The downside is that one needs a business license to use the BI connector.
- oneweekwonder 8y ago> The downside is that one needs a business license to use the BI connector. Have you looked the postgres mongo fdw[0] before? [0]: https://github.com/EnterpriseDB/mongo_fdw https://github.com/EnterpriseDB/mongo_fdw
- larrydag 8y agoThis is a huge concern for me at my current organization. Dev has decided to put all data into mongoDB. Yet all decisions are based on that data and the tools we have do not allow for seamless flow (ETL) from mongoDB. That data is important for deriving decisions that affect revenue and costs. Where are solutions for the data analysts and scientists? Frankly I'm pretty sick of hearing it can just be automated. In my mind there has to be a decent "business intelligence stack". I'm not sure I'm coining that because I didn't get good search results from that phrase. Believe me I've been trying to find solutions. I believe there is big opportunity in building out this sort of stack that bridges data management and data analysis. Sure you can call IBM, Microsoft, Dell, HP but be prepared for big costs and huge software bloat. I would like simplified solutions and options that can fit with most industry standard tools. I'm also willing to work with anyone on this as well.
- ty64 8y agoThis exists: https://www.mongodb.com/download-center#bi-connector https://www.mongodb.com/download-center#bi-connector
- oneweekwonder 8y ago> The MongoDB Connector for BI is available as part of the MongoDB Enterprise Advanced subscription, So may one only use it with a subscription? I got mongodb postgres foreign data wrapper[0] working in a previous life. [0]: https://github.com/EnterpriseDB/mongo_fdw https://github.com/EnterpriseDB/mongo_fdw
- threeseed 8y agoYou can connect Hadoop/Spark directly to MongoDB so in some cases you may not need to do an ETL at all. You can also use something like NiFi which supports MongoDB and will allow you to shift data out to Avro/Parquet on HDFS/S3 for your data scientists to use. As for BI stack. Not sure what you mean. There are hundreds of tools which blend data management and data analysis. You can do this with Hortonworks (Atlas + Spark) or Alteryx for example.
- mindingdata 8y ago
- squirrelicus 8y agoOkay so... To make MongoDB not suck for analytics, ETL it in a different format. For engineers trained in backed systems, this is pretty obvious. After reading this, I also don't know why I'd choose Pequot things over any other thing. Baby's first ETL -- just scan the db with a cursor and analyze the data in a script -- tends to cover 90% of the use cases for BI db analytics with almost zero resource consumption anyway. Point being don't write a query to do analytics if your db can't answer your questions performantly, and don't build [latent, stale, slow] Enterprise ETL unless you really need it.
- mac01021 8y agoAs someone who grew up around the home if the Pequot tribe, I'm amused by the choice made here by your input device's autocorrect feature.
- minitoar 8y agoWe use a similar technique at Interana. Our DB is a column store, but we break things up over the time dimension to keep file sizes of individual columns reasonable. One of these time buckets is essentially analogous to a single parquet file. In addition we split/sort these buckets into smaller buckets as more events are added.
- georgewfraser 8y agoThere are several companies, including mine (Fivetran) that will replicate MongoDB into a columnar data warehouse for analytics. For most people, a commercial replication tool + a commercial columnar data warehouse is the best trade off of cost/ease of use. Commercial DWHs deal with all the details of patching columnar formats under-the-hood, and commercial replication tools like us will deal with all the complexity of things like the mongo oplog. For not that much $ you can have a working system in like a day.
- stickfigure 8y agoI tried using MongoDB for the customer-facing analytics of a large e-commerce marketplace. It didn't work very well. The problem is that at some point you end up wanting joins. MongoDB was actually the third try. My first two attempts were BigQuery and Keen, neither of which worked out because they support only one index - time. Users want to slice and dice by various axes! And there's an obvious additional index you need - "merchant" - which column stores usually say propose setting up isolated partitions for. If you do that, you can't ask questions across the whole system! We ended up with Postgres. It was actually faster than MongoDB for simple aggregations, and joins made it much better/faster for complicated queries. Of course it only works quickly if your dataset fits in RAM, but terabyte-size instances are pretty affordable and give you a lot of headroom. That was a couple years ago. I don't know what they're using now, probably the same. It was a frantic few weeks figuring out what was going to work - each of those systems made it to production and quickly discovered to be inadequate in vivo. If you're in a startup, even if you're using exotic NoSQL systems like Google Cloud Datastore or DynamoDB - just use Postgres or MySQL for analytics. It will work long enough for you to figure out something else when you need it.
- deleted 8y ago[deleted]
- threeseed 8y agoYou are completely contradicting yourself. On one hand you complain about using technologies before you have done a prototype and evaluated the product. Then you blindly tell startups to just use MySQL/PostgreSQL without having any idea of their use case or whether it matches their query patterns. If you are a startup the right way to go is to document your use case, understand what queries those use cases demand and then find the right database that satisfies it e.g. don't pick MongoDB if you are doing lots of joins and don't pick PostgreSQL if you are doing wide-table feature engineering type analytics. Right tool for the right job.
- nostalgeek 8y ago> Right tool for the right job. I would argue that since both Mysql and PostgreSQL are JSON document stores with mostly the same capabilities when it comes to querying and aggregation I don't see the advantage of using MongoDB at all. I wouldn't even use MongoDB for caching when redis does a better job at it. Logs? I don't see why logs cannot be shoved into a RDBMS. Prototyping? create a table with a JSON field and a primary key. Distributed file system? I don't know any business which uses gridFS as a CDN, full text search? PostgreSQL does it better. So what is the job your are talking about? PostgreSQL is so much powerful for analytics because of the power of SQL.
- drej 8y agoFor those seeking tl;dr: The answer is not to use MongoDB.
- drej 8y agoA more complete answer is to dump your data into a columnar format into S3 and then use one of plethora analytics tools that can work with this format (AWS Athena and Drill are mentioned, other tools like Presto, Spark, Redshift Spectrum or BigQuery can help).
- drej 8y agoBut I don't want to be just snarky. We faced the very same dilemma and solved it in a similar way - we use Apache Spark, which can connect to MongoDB directly. It loads fairly quickly and we can save it to Parquet on S3 directly, the whole thing is about 5 lines of code. If you have a Spark platform in place, it's a decent solution for this.
- twblalock 8y agoThat doesn't get you out of having to face the problem. This is not a challenge unique to MongoDB or other NoSQL databases. Oracle or Postgres might be ideal for your transactional data store, and a columnar database might be ideal for your analytics. I suppose you could choose one of those options and sacrifice either your customer experience or your analytics, but it's probably better to use the best database for each use case.
- icedchai 8y agoAmazing! I didn't even have to read the article to know that.
- endymi0n 8y agoProtip: MongoDB works absolutely best for analytics when it is replaced with a sane and scaleable column-oriented database like Redshift or BigQuery right before serving that report.
- dmitriid 8y ago> How to make MongoDB not suck for analytics Easy: you don't use Mongo
- sztanko 8y agoJust try this out: https://github.com/EXASOL/docker-db https://github.com/EXASOL/docker-db and you will be impressed. This is an embryo of a real analytical database. Pros: - an 8 CPU installation with 64gb memory will probably be hundred times faster then postgres. -it supports full sql - It is super stable, even as docker container Cons: - it does not support nested data - once you reach volumes of around 2Tb, you will probably have to switch to a paid version (I mean, you still can continue running on a 200gb ram box, but it will be suboptimal) P.s. I am not affiliated with Exasol.
- arghwhat 8y ago> an 8 CPU installation with 64gb memory will probably be hundred times faster then postgres. "Probably" not. The way this usually goes down is that there may be a few synthetic benchmarks show a large performance benefit over existing established databases (x2, not x100), with any non-synthetic benchmark showing very poor performance (1/10th, 1/100th, sometimes even worse), and also often very unstable performance. The product is then also usually beta quality, as it is hard to compete with the 36 years Postgres has been in development since its inception in 1982 (and that's not counting the 9 years of Ingres development, which Postgres—"Post-Ingres"—spawned from). Important features are usually also quite lacking. If someone claims x10 or x100 performance improvement over established databases, they better have published a few papers about all the computer science research they must necessarily have done to get there.
- DrummerDaveS 8y agoFull disclosure - I currently work for Exasol.. but I thought I'd just clarify that Exasol has been around for over 15 years and is far from 'beta' (currently on version 6 with hundreds of production installations worldwide). I've also been in the industry for > 40 years and worked with many database products (including Ingres and Postgres) - and all I can say is download the free community edition from the Exasol website or the Docker image as described above and try it for yourself - you will be up and running very quickly and I think you will be pleasantly surprised regarding both functionality and performance.
- kockic 8y agoI see that most of the `don't use mongodb for analytics` are being down-voted, however I tend to agree with them. For all the people out there looking for the database for analytics please check Clickhouse from Yandex, it's easy to get started, amazingly fast and open source. Disclaimer: I am not affiliated with Yandex in anyway, just a happy customer
- eddd 8y agoI kind of a hoped it'll end up a joke saying "Don't use mongo". Last time I used it was 2.4 and it was the worst db experience ever. Back then It was more sane to craft a solution with PG and HSTORE. Now, I think RedShift does the job, why would anyone use mongo on production for anything today?
- nickserv 8y agoIt's not too far from that joke. It's like if you ask "how do I drive my car downtown" and I answer, "Easy, just park at the station and take the train". To answer your other question, their marketing goes a long way. I recently started at a new company, and the lead was proudly telling me how the project was developed using Mongo... So I start explaining how it's basically shit after using it professionally for a few years. His answer? But SQL doesn't scale well enough!
- 198394549 8y agoWhy is it basically shit? It appears to store and retrieve the data as per my instructions.
- nickserv 8y agoExcept when it doesn't. We've had data corruption issues related to oplog, out of sync secondaries and excessive resource usage on the primary. As far as major problems. There were also a bunch of smaller problems but in fairness those were on the nodejs/mongoose side of things. Would not recommend.
- codingdave 8y agoI'm not aware of any analytics platform that runs directly from the source data. There is just about always some kind of ETL process, or at the very least, a data transformation process to shape the data as needed, to provide data that works well for the reporting. So while information on making MongoDB performant for such things is mildly interesting... it just isn't how analytics are generally architected.
- jrs95 8y agoOkay, we get it, Mongo sucks. Or at least that seems to be the consensus. From what I can tell it seems they've improved their tech a lot though, and I have to wonder if a lot of the "mongo sucks" sentiment comes from either 1. Using early versions of Mongo that really did suck or 2. people having used Mongo at companies where nobody really knew how to use Mongo that well.
- notoriousp 8y agoLittle bit offtopic but what product did you use to create those visualizations?
- manigandham 8y agoThis is called ETL, to a data warehouse. Regardless of the choice of primary database, this is nothing new and just shows how a lot of startup technical talent seems to be discovering the same things all the time, usually with needlessly convoluted approaches, and writing blog posts about it.