7 ms·
Startups should use a relational database
- dclara 13y agoI'm totally with you. We've experienced to have Object database or XML database and NoSQL database. Now we understand that relational database is just the right way to go for web applications, because it deals with structured data so well, keeps querying and sorting, filtering seamlessly and effortlessly. It's a must. It is the same thing for choosing Linux distribution and JDK mode. See the references here: http://bingobo.info/blog/table-of-contents.jsp http://bingobo.info/blog/table-of-contents.jsp BTW, your title should have "relational" instead of "relation".
- luzero 13y agoStartups should use the right tool. A relational database might be just as wrong as a nosql one if all you need is redis.
- RyanZAG 13y agoDepends what your startup is doing. If you are only using your database to store some basic transactions, then a relational database is a very good fit. This is really the case for most startups tackling common problems. However, if your startup is tackling a problem with unique technical challenges, then you can't just ignore the issue. For example, a geo-location startup tracking the location in real time of users with a free app is simply not going to be able to use a relational database.
- charliesome 13y ago> For example, a geo-location startup tracking the location in real time of users with a free app is simply not going to be able to use a relational database. Why not?
- deleted 13y ago[deleted]
- nl 13y agoRapidly growing, infrequently queried data is not the ideal scenario for most relational databases. 1) Relational databases typically aren't optimised for write-throughput. It's quite possible to do it, but you'll need fast and large disks (eg, FusionIO in a SAN or something). 2) Location-tracking applications typically don't require interactive queries - generally it is more a batch-based system that can be run offline. Saying you are not going to be able to use a relational database is overstating it a bit in my view. Clearly you can make it work, but something like Cassandra will give you better write thoughput, won't force you to rely on a SAN/NAS for data storage and will let you use Map/Reduce to batch process the data.
- robbles 13y agoI don't see why not - use a relational database for storing long-term data, such as users, friendships, preferences, etc. and store ephemeral data such as location in a key-value store like Redis. Store summary statistics of the location data in your primary data store in scheduled background tasks. Just because you're storing a huge amount of one specific type of data, that doesn't prevent you from taking advantage of the features of a relational database.
- Sanddancer 13y agoPrecluding yourself from using a relational database means that you're not going to be able to use some of the best tools available for geographical data. Things like PostGIS are built around, and heavily dependent on, the fact that it is Postgres. Now, you may not want to use an RDBMS for everything, but at the same time, you don't want to pull it completely out of the system either.
- VLM 13y agoTotal disagree. The most important job of a startup at startup time is probing the market. If you have a billion customers each with a million records you are the Google Maps location thingy and are not a startup anymore and a relational solution may, or may not, work. If you have ten users the ideal database is probably ten interns and some whiteboards. I'm not kidding. I observe there is a common claim brought up every week on HN about fake it till you make it. You don't automate (insert menial task here) until you have ten customers. I observe there is another common claim brought up every week on HN about how you need a massively scalable design which can never change from day one, because your schema, either formal or informal, is perfect and unchanging, LOL. Those two stylistic outlooks are not compatible. Start with something that has too many features, too many abilities, too much room for expansion, like a relational DB, and then later on if you need to, put some stuff on another platform. IF you need to. IF your company survives. Lots of IF.
- RyanZAG 13y agoFake it till you make it is why you'd use a quick and dirty mongodb with json dumped straight from your clients, though. Maybe we have different ideas on the speed of development on mysql vs mongodb though.
- VLM 13y agoAlso probably different outlooks on which letters to emphasize in CRUD, and probably different assumptions about what "all" startups are doing with a database, anyway.
- adamnemecek 13y agoA blog post titled "Startups should use NoSQL databases" in 3, 2, 1, ...
- guard-of-terra 13y agoDon't they? For some tasks relational databases are good, for some they are worse. Call me captain. However relational databases will have hard time with big data because your dataset is bigger than your database and you have no relational integrity.
- pkolaczk 13y agoHe forgot one of the very important reasons to use (some) NoSQL databases: high availability. Relational database systems are very poor at providing that. Most often the availability options are limited to resistance to node failures. RDBMSes have several SPOFs and must use failover which is not dependable, hard to test, and in many times needs manual intervention. Forget resistance to network partitions.
- coolsunglasses 13y agoRarely matters for a startup.
- pkolaczk 13y agoFor the startup I once worked for, it mattered much more than we had thought at the beginning. The investors were smart enough to notice we had some considerable periods of downtime. Additionally, once we got first million of users (not really that much and nowhere near the scale of Google or FB) we ran into performance problems which couldn't be easily solved just by indexing, optimizing queries or adding more hardware, and "buying" a beefy Oracle superserver was not an option as we didn't have enough revenue yet. So we had to dump joins, relax transactions, denormalize a lot and ended with a half-baked, bug-ridden NoSQL store on top of PostgreSQL, that couldn't even do horizontal partitioning well. I wished we had a proper solution like Cassandra right from the start. It would save us lots of pain.
- coolsunglasses 13y agoFirst million users. Come on. 99% of your audience is never going to have that problem. Especially if they spend their early days fucking with a Cassandra cluster instead of talking to customers. And it should be noted, you made it anyway. When you make it by the skin of your teeth, that means you probably timed it right. Preempting a problem far-ahead of time in startups means time and effort was wasted, especially if it was done before the existence of the problem was established. It is unreal to me that people still can't figure out how to apply Maslow's hierarchy to startups.
- robconery 13y agoOne thing that could likely get you fired rather quickly is running analytics on your live transactional system. Yes, your business needs to make decisions based on data, this is not terribly new. To think that you only have one data store is a bit short-sighted. Many businesses (including startups) have moved to using document stores for high read environments and scraping nightly drops to their backend analytics systems. This is smart - you don't want to run summing/aggregation on a live transactional system for (hopefully) obvious reasons. EDIT: it's also worth noting that map/reduce is typically much more powerful when aggregating large datasets. When trying to run analytics on top of a transactional system, developers like Ray here would end up with multiple joins and groupings - all of which slow everything down. Map/reduce certainly isn't perfect, but the author dismisses it as difficult witchcraft when, in practice, parallel execution of MR queries can greatly decrease resources and time to information. I sort of think we've moved beyond this discussion.
- nostromo 13y agoThe disconnect between your comment and the article is the term "startup" now means giant companies like Airbnb and tiny two person companies that haven't yet created an MVP. I think this article is targeted at the latter: pre-MVP and just post-MVP. For those startups, having two databases with one dedicated to a backend analytics system reeks of premature optimization.
- karterk 13y agoHaving taken this path from a 2-person running a "nobody cares about this" app to an app with some decent traction, there are always other things more worthwhile to do in the early stages than trying to get insights from the invariably small amounts of data in the DB. A 2-person company should be talking to users rather than trying to analyze patterns from database tables. Sample size is just too small, and you will be evolving so fast that the trends will almost be meaningless. If you grow to be more than 2-people, then taking a mirror dump of the prod db to run queries against is pretty trivial effort.
- VLM 13y ago"having two databases with one dedicated to a backend analytics system reeks of premature optimization." Its free software, you don't have to pay for two instances of Oracle. One thing that will quickly kill a biz is combining the functions of PROD and DEV/TEST. Making the DEV/TEST box the DEV/TEST/REPORTS box is not a big deal, and you can't run a (real) biz without a DEV/TEST box.
- HorizonXP 13y agoSerious question: what are NoSQL databases really good for? I'm only really used to relational DBs, and I'm unclear about which problems a NoSQL database is useful for.
- BadassFractal 13y agoTake Redis for example: simple KV database with in-memory perf and a very comfortable API with option to flush data to disk periodically. Addresses a lot of interesting scenarios such as session storage, request throttling by a certain key etc. Not used as a replacement for a RDBMS most of the time, but rather as a specialized tool for a certain use case. At some point your RDBMS is already hammered hard enough, you don't need to dump everything into it, but ultimately you could, especially at first. Comes with clustering for free as long as you're aware of the risks and use it appropriately. This is not to say you couldn't turn Postgres into something similar. Put the data on a ramfs, relax various write guarantees (lots of knobs in PG) etc.
- danieldk 13y agoFor example, XML databases are handy when you have large XML datasets that you like to query. The database indexes the XML, allowing you to execute most XPath queries and XQuery programs quickly.
- taspeotis 13y agoI had a quick Google and MSSQL [1], DB2 [2] and Oracle [3] all support indexing XML. [1] http://technet.microsoft.com/en-us/library/ms191497.aspx http://technet.microsoft.com/en-us/library/ms191497.aspx [2] http://publib.boulder.ibm.com/infocenter/db2luw/v9r5/index.jsp?topic=%2Fcom.ibm.swg.vs.addins.doc%2Fhtml%2Fibmdevori-DB2XMLIndexes.htm http://publib.boulder.ibm.com/infocenter/db2luw/v9r5/index.j... [3] http://docs.oracle.com/cd/B28359_01/appdev.111/b28369/xdb_indexing.htm#CHDJECDA http://docs.oracle.com/cd/B28359_01/appdev.111/b28369/xdb_in...
- danieldk 13y agoIt's not about indexing paths you indicate, XML databases are about indexing the whole structure (remember, XML is not relational, but are graphs). Besides that they provide XPath/XQuery processing and optimization. You can query large XML documents or sets of documents, like you'd query them with e.g. XQilla. Of course, it is possible to implement all of this on top of existing database technology. E.g. Oracle's Berkeley DB XML is implemented on top of Berkeley DB. But, a relational database with some indexing of XML does not provide the same functionality as an XML database.
- yeukhon 13y agoI don't know. It's a hard question. MongoDB is pretty much used in any hackathons simply because it's easy to setup, driver support is good, and schemaless. The last one is really why people use MongoDB over SQL DBMS. For startup, there might be a concern that schema migration is tough. But one can argue that not careful with schema design can break api and make codebase messy. I guess I will stick with the hard work now... I guess not careful with schema will definitely bite me.
- danieldk 13y agoI don't know. It's a hard question. MongoDB is pretty much used in any hackathons simply because it's easy to setup, driver support is good, and schemaless. The last one is really why people use MongoDB over SQL DBMS. I find that a poor argument. One can use an ORM that automatically creates a schema based on classes. E.g. I like Ebean with DDL generation. You just write classes and add @Entity annotations. Ebean automatically creates the schema. Combine this with an embedded database, such as h2, and there is virtually nothing to set up. Once you are out of the rapid iteration phase, you can take the latest Ebean generated schema and use to proper migrations for later changes.
- est 13y agoAnother reason to choose MongoDB is built-in array and nested dict support, with good enough indexing. So you don't have to create bullshit m2m tables with tedious joins for a fucking tagging system
- sedlich 13y agoStrongly disagree with the article as simplification always looks shiny. Start-Ups should sit back for a few hour and days and invest the work to answer some serious questions as these http://nosql-database.org/select-the-right-database.html http://nosql-database.org/select-the-right-database.html (there are other cataloges like this one). Then you get a little closer to the truth.
- jwilliams 13y agoThis comes up every now and then on HN. There are plenty of NoSQL horror stories. Thing is. Most SQL database at scale is a bit of a horror too. Have you seen real-life production relational databases? Gawd. Hacks on hacks. Then you add another database. And another analytics database. And a bunch of point to point data feeds. Argh. But hey. That's data. If you think choosing SQL will solve your analytics woes down the line -- it's just not true. You're in for some pain no matter what you do. ... That's unless you get a porcelain schema first time. Which, if you're in a startup, probably means you're working on the wrong problem. That's not an argument for using NoSQL (I used MongoDB daily, but I've got plenty of love for PostgreSQL). It's a rebuttal that SQL magically solves a different problem.
- lampe3 13y agoI often read that argument that NoSQL Databases are Schemaless and yes the Database is but your Data is or it isn't. YOU must know your Data. "All the while moving work onto the developers to standardize how they handle different migration cases." I know a startup is fast and bla bla... BUT your team should know the tools that you are using... For me SQL DB's force me to add a new field and some kind of value and i don't like to be forced to a solution. "In document stores, you have two choices: store related data as sub-documents, or store related data as separate documents with references. It is up to the developers to understand the trade-offs of both approaches. Selecting one over the other can lead to performance gains or issues, scalability issues and above all, make asking certain questions of the data a lot harder." Again know the tools you are using. And for example MongoDB has good ORM's too. "But that takes much more forethought and is dependent on a particular problem." If your startup is doing something new and shiny you don't have the knowledge and forethought and you often dont know what particular problem will come at you. Most of the point's look like: You learned at your University SQL now you know it(but in really life you don't) and now use it because you know how to normalize a Database. This argumentation is often used to say why java is so great or why javascript is bad. I personally started with php then moved to rails and now to meteor(uses MongoDB) and we never before meteor could make so fast a good prototype which for a startup is very important. So yeah if you are comfy with SQL use it if your comfy with NoSQL use it.
- lmm 13y agoThe unhappy truth is that for many startups, relational integrity and transaction safety are simply not very valuable. Customers of an early-stage startup are by definition willing to take a risk on whatever they're getting from that startup. So simply not thinking about these problems - accepting that occasionally a partial write will happen, or two writes will collide, or a migration will not quite work correctly and your pages will crash until it's fixed - is a worthwhile sacrifice to increase development speed.
- raverbashing 13y agoExactly Relational DBs may even be the wrong choice for some specific problems. And about this quote from the article: "At some point, you will need to ask your primary database questions. If you chose the wrong database, this is where things get tricky. " Yes, this is correct. However, I know how to read the manual of whatever db I'm using and maybe code something simple to process the output to the format of my liking.
- stiff 13y agoThis is just a very bad excuse for doing ridiculously shitty systems that stay shitty way after the startup phase. If you at all bother to understand your problem domain, and are not just a monkey at the typewriter, writing down a good domain model and adding the constraints is not going to decrease development speed, quite to the contrary, it's going to increase it, web developers often spend whole workdays just tracking down in the logs the "story" of some now angry customer that happened to violate some unformalized assumption of the system and got mishandled later in the process, especially if this customer happened to pay already. Not to mention that with a good domain model, the code for the individual functionality flows out naturally, while with a shitty one, you might end up with three times as much code for the same thing. Also, those mistakes in modelling the domain and in enforcing the constraints are often there to stay and slowly become impossible to fix, once you have 10000 records that do not fall into a few well specified states, it's hard to go through all of them, find some common denominators, and migrate the database. Not to mention that with the mess people can do in the code, and with the messy stack in use today, it's easy to introduce bugs that might be hard for anyone to notice but seriously harm your business. The amount of fashionable nonsense in software engineering seems to be higher than ever, unfortunately.
- joshguthrie 13y ago"Blog titles should stop using should."
- VLM 13y agoThere can only be one DB much like the LOTR can only have one ring. Why? Thats the only area the linked article falls down on. Its a pretty good article other than that. So you properly normalized your entire system, customer billing transaction records all the way up to article tags. Then article tags gets too huge. So next version looks at RDBMS and Redis, and the next version after that only looks at Redis. Customer billing transactions remains on a "real" DB and the tag cloud lives on redis. And the problem with that is... what exactly? Its obsolete thinking. I can't have two databases because we're a poor startup and the only databases that exist are DB2 and Oracle and everyone knows they're super expensive so super expensive times two is unaffordable. Dude, its almost 2014 not 1980, Postgres/mysql/redis its all free.
- raycmorgan 13y agoThank you for your comment. I agree that there is nothing wrong with using a plurality of systems when needed. Your example of moving from one to another is great! Start with a simple system, and once you find bottlenecks, optimize with specialized stores.
- Gulthor 13y agoPeople too often forget about graph databases when talking about NoSQL solutions. Graph databases offer an interesting and elegant alternative to relational databases and I could definitely see a startup decide to use this kind of technology. As far as I know, most graph databases support transactions and offer great scalability. Such databases are also schema-less and can be queried with Gremlin, a powerful graph traversal language (see www.tinkerpop.com). With respect to scalability and transactions, Titan (http://thinkaurelius.com/ http://thinkaurelius.com/) looks very promising: it supports various backends for storage (Cassandra, HBase, etc.) and indexing (currently Elastic Search and Lucene). Graph analytics can be done via Faunus (http://thinkaurelius.github.io/faunus/ http://thinkaurelius.github.io/faunus/), backed by Hadoop. There are other vendors out there (Neo4J, OrientDB, etc.) which offer interesting solutions worth looking at - I'm just a bit less familiar with them. The major downside I see with graph databases is that most of them are fairly recent and their ecosystem is tiny (though growing). Should a startup venture on such young technologies, or stick to mature and battle-tested solutions (ie. relational databases)? Could startups use this kind of graph "NoSQL" databases? I don't see why not. If your startup is some kind of social network, graph databases are certainly an option worth considering. If I were to create a startup, I'd hardly use a document database like MongoDB but I will really consider using a graph database. In the end, it's all about having the right tool in hand, and knowing how to assert what is "right" for you.