12 ms·
ORM is an anti-pattern
- pspeter3 15y agoI think that there are definitely some valid points about efficiency while using ORMS when the queries get more complicated. However on a simpler scale, like a blog, ActiveRecord or other ORMS aren't horribly inefficient and are faster and easier for the programmer which is why people ultimately use them.
- deleted 15y ago[deleted]
- sawsaw 15y agoThe point Seldo was making, however, is that for simple applications like this a relational database doesn't make sense, and is in fact overkill.
- ianterrell 15y agoExcept RDBMSs are ubiquitous. Whether or not they're overkill, they're cheap, easy, and well understood. What web host in the world doesn't run MySQL? What platforms don't support Sqlite3? What PaaS doesn't give you easy access to Postgres or another database. Compare that to support and sysadmin knowledge of Mongo, Redis, Couch, whatever: there's no comparison. It's only "overkill" in the very limited sense of, what, CPU cycles? Unused relational algebra potential?
- Joeri 15y agoI think the argument is that SQL+ORM is not that different from NoSQL. You have a point that it's probably easier to use an ORM than a NoSQL store in the real world.
- MatthewPhillips 15y agoI don't know... I've spent a couple of days writing code for Redis to store a list of 15 objects whose date fields are the 15 lowest. Granted it's my first time and I didn't/don't know what I'm doing, but it's quite a lot code for something that would be trivially simple in SQL.
- johnzabroski 15y agoIn high quality RDBMSes like Oracle and SQL Server, there is not a whole lot you can do to tune your queries and the majority of your bottleneck is likely to be in how you are using your data. If you have inefficiencies, you either have a poorly designed schema or out of date statistics/indices. The story is slightly different for open source RDBMSes which have inferior cost-based optimizers and multi-version concurrency control implementations that fail on a wider range of corner cases like: select count(*) from Billing_And_Accounts_Receivable.Transaction which is slow in PostgreSQL for reasons I'd rather not take the time to teach about here.
- foamdino 15y ago"there is not a whole lot you can do to tune your queries" What does this even mean? For any query which relies on data from different tables, yes there are opportunities to tune the query, and there are opportunities to build different kinds of indexes, and there are opportunities to de-normalize the data and remove the join completely. This is independent of your RDBMS vendor. The statement that Oracle/SQL Server are 'high quality' and the implication that open source RDBMSes are not, is flame-bait. All RDBMSes (that I have experience with - including Oracle and SQL server - so caveat emptor) will hit a performance wall when an index becomes invalid and the engine has to table-scan - and indexes can become invalid automatically as the amount of data in the tables changes - typically this will happen at 3am in the morning when you are on call. There is also the hardware that the files are stored on, and the balancing act of price/performance of splitting tables onto separate spindles, using SSDs for semi-hot tables, cramming in more ram, upgrading the network connection between db and app servers etc etc. Every real world application needs constant tweaking from query to hardware as bottlenecks appear in different layers. Again these issues are almost entirely independent of the RDBMS (apart from with SQL Server that only runs on Windows, so the deployment platform is limited by what Windows server supports). Nonrelational datastores can also hit performance walls too - the fact that some of the nonrelational stores are newer also has the added 'fun' that the performance characteristics are not completely known.
- johnzabroski 15y agoCommercial DBMSes have way better customer support than open source DBMSes. Not only do they better understand the performance characteristics of their clients, but they have dedicated teams whose job it is to isolate corner cases where the DBMS is not optimized for a particular environment. My statement -- there is not a whole lot you can do to tune your queries -- is about rewriting the queries themselves, such as join hints, etc. Changing RAID configurations, organizing the data physically on disk differently, etc. have no bearing on the quality of SQL an ORM generates. So what does? Well, 20 years ago many DBMSes required you to compile stored procedures in order for the cost-based optimizer to cache anything. Today, not only does SQL Server 2008 R2 support fine grained plan caching and allow you to adjust the size of that cache, but you can remove from the cache any plan you dislike. You can also force a poorly performing query to use a specific plan cache. This is one of the ideas behind the LINQ Re-Motion project for .NET. I am really arguing about where to put work effort into, and really NOT disagreeing with you. We are talking past each other. I personally believe, as I wrote in the author's blog comments, that an ORM based on an algebraic model would likely be better than the big and irregular ORM APIs we have today.
- ianterrell 15y ago"In the long term has more bad consequences than good ones." I would hypothesize that one of the long term good consequences of healthy ORM options is the existence of the vast majority of database backed applications we all know and love. Sure, when/if they got popular someone had to tune some SQL, but how many of those projects would have even been started without ActiveRecord or Hibernate or EJB3 or CoreData? The opinion that "ORM is an anti-pattern" is ridiculous nonsense.
- stcredzero 15y agoDeath by a thousand queries Wait a moment here. In my experience, most problems like this can be solved by noting something like: When we get X, we also get all of the associated Y's and their Z's. Declarative association of a Batch Query with certain retrievals isn't a newbie project, but it's something a lone programmer can put together in a week in a good ORM. (I've done it.) I would expect this library feature to be common in the Ruby/Python world.
- jeswin 15y agoIt has been a supported feature in just about every other ORM too (Hibernate, Entity Framework, Linq-to-sql..), for ages. The author doesn't seem to have understood ORMs well enough.
- johnzabroski 15y agoGood observation. Anyone writing a blog article on ORM should probably read the c2 wiki entry with the huge feature comparison chart of ORMs first.
- Swannie 15y agoI think the author's point is that by the time you have gone to learn all of the complexities of an ORM, and all the steps you must take to sidestep when you can't bend it to shape... you could have much more easily written it by hand.
- stcredzero 15y agoIf you have a Declarative Batch Query, then you you get a bunch of SQL and plumbing code for free. The join and the iterative code you'd need to do that writing it "by hand" is 10x more.
- gnaritas 15y agoThe author couldn't be more wrong.
- keithnoizu 15y ago
- InclinedPlane 15y agoI've seen a lot of the problems that ORM creates with big projects. The most egregious is lack of control. You'll run into some problem caused by some quirk of your ORM system and you'll dig down into the SQL and learn precisely what's causing it, but you still won't be able to fix it because you don't know the magic voodoo incantations to change your config or the ORM client code in the right way to fix it. When ORM starts to get in the way like that it really makes you wonder whether it's worthwhile.
- sunchild 15y agoI would love to hear about an example of this kind of case. My experience is that ORM queries, used properly, deliver predictable results. If they don't because of a bug in the ORM, can't you just drop down to the native query and move on?
- chrisjsmith 15y agoIt's mainly due to the fact that MORE abstraction means MORE complexity. I currently have to deal with an NHibernate mess of over 1000 domain objects (!) at the moment and it's an absolute nightmare on the performance and maintainability front.
- smhinsey 15y agoTo be fair, 1,000 (or however many you need) stored procedures probably wouldn't be all that much better. I've seen systems taking both approaches and I've seen both go right and wrong. It has more to do with the team than the technology, I think.
- chrisjsmith 15y agoWe replaced 15,000-odd stored procedures with NHibernate. It's the same turd in a different coat.
- Brad_Smith 15y agoif More abstraction is leading to more complexity, then there is something very wrong going on.
- sunchild 15y agoWhat is it about coders and blogs that brings out the "cranky old man" vibe? An ORM is an insanely convenient way for newbies to use various data stores while avoid learning umpteen different query languages. The teaching value of ActiveRecord for newbies is hard to overstate. It's also a damn nice way to move your application closer to platform independence – a valuable thing in today's PaaS integrated stacks. If you seek efficiency and performance, don't use an ORM. Lick the freezing cold metal, if you want. Nothing is stopping you from doing what you like! (Also, I haven't dropped down to SQL since Rails 3.x and meta_where. Yes, I realize that my applications "won't scale". They are appropriately scaled for their intended purposes.)
- enjo 15y agoIt's not just for newbies. I'm competent with SQL, but that doesn't mean I enjoy writing query after query. We use Django and utilize the ORM... we jump down into SQL in the (rare) case that we really need too. I'm sure at higher scale (we're mid-level at this point) that we might be hand tuning more and more, but that hasn't happened yet.
- bluekeybox 15y ago> An ORM is an insanely convenient way for newbies to use various data stores while avoid learning umpteen different query languages Let's count these umpteen different query languages: (1) SQL, (2) ... ? > If you seek efficiency and performance, don't use an ORM. If you don't need efficiency and performance, why use a DBMS at all? Isn't the whole point of a DBMS to make data access not only secure but also efficient (B-tree indexes, etc.)?
- true_religion 15y agoThe variants of SQL are different between different databases. Even the data-types supported are different. It's not uncommon for consumer and enterprise products to have to work with different databases. In the case of a consumer product, it'd be a client having only one database type installed and requiring that your software use it. You benefit by writing in an ORM because you can have the same codebase for multiple db installs. With enterprise customers it'll be having multiple databases installed at once, and having your code interface between all of them. You benefit with the ORM by not having to remember and handcode all of the query differences.
- gerardo 15y agoThe relational-object problem is called Object-Relational impedance mismatch, duh!(http://en.wikipedia.org/wiki/Object-relational_impedance_mismatch http://en.wikipedia.org/wiki/Object-relational_impedance_mis...) For me, fast Web Application development is worth the tradeoff. I usually begin to hate sql on the second month of a project.
- wccrawford 15y agoI love SQL. What I hate is the tedium of turning the results into something useable. ORMs eliminate that pretty handily. As they say, "If it didn't exist, I'd have to invent it."
- gerardo 15y agoBTW, how are you working around these days the most obvious problems of an ORM?
- bni 15y agoBy not using them
- div 15y agoLabeling an ORM as an anti-pattern is throwing the baby away with the bathwater. Sure, you will encounter some cases in which your ORM will be a pain in the ass or even actively work against you, but most good ORM's will allow you to talk to the database directly. For example, both Hibernate and ActiveRecord allow you to just throw straight sql to your database, returning a bunch of key value data. Which is exactly what a good solution does: provide large gains for the common cases, and get out of the way for the edge case.
- adelevie 15y agoI'd be interested in the author's take on ActiveRecord's implementation of Relational Algebra with ARel[1]: > To manipulate a SQL query we must manipulate a string. There is a string algebra, but its operations are things like substring, concatenation, substitution, and so forth–not so useful. In the Relational Algebra, there are no queries per se; everything is either a relation or an operation on a relation. Connect the dots and with the algebra we get something like “everything is named_scope” for free. Also, if I couldn't use something like ActiveRecord in my Rails apps, I'd end up re-writing most of its functionality in my model code somewhere. If I don't get Model#find_by_some_attribute() for free, then I have to spend time writing it. [1] http://magicscalingsprinkles.wordpress.com/2010/01/28/why-i-wrote-arel/ http://magicscalingsprinkles.wordpress.com/2010/01/28/why-i-...
- mixonic 15y agoEmbedding strings of one language in a second language is an anti-pattern. I've been at a bunch of NYC dev events recently, and people at both Goruco and Percona Live were hating on ORMs. ORMs have gotten really good in the last few years, I think the haters just haven't been using them. Show developers a good alternative and they will go there. Some of the basic points made in this article ring true, but the suggested alternatives are weak. ARel is a great start to a non-orm database wrapper in Ruby! Somebody just needs to go there.
- deleted 15y ago[deleted]
- maresca 15y agoIs this guy talking about object relational mapping or object role modeling?
- andybak 15y agoI'm guessing the former as I've never heard of the latter. Is it time for a Central Registry for TLA's? (CRT! Damn. It's taken...)
- eftpotrm 15y agoPersonally, in developing quite a lot of different data-backed apps, I've never really found the problem ORMs are solving to be a hugely significant one; it seems like a 'quick fix' for coders who don't really understand SQL anyway, which always felt to me to be attacking the problem in the wrong place. SQL isn't that hard.... In any case, while I don't dispute that it might offer speed of startup advantages for some developers, it seems no-one is so far disputing that it simply doesn't scale and, if your project really takes off, it will be creating problems. Call me a fogey if you will but I don't like the idea of launching a project that I know will need very substantial rearchitecting too early in its life.
- MartinCron 15y agoI'm definitely an old-school code-sql-by-hand guy and I used an ORM (Entity Framework 4.0 Code-First) for the first time for my startup project, a data-intensive online strategy game. I've found that for many things, it's so much faster in terms of dev time, especially with the super-cool code-first approach, to get things out there and working using the ORM. I can create a new fully-functioning and reliably-working repository class, along with its test double, in around a minute. Really. Of course, I have found that I've had to replace bits of it with hand-coded SQL for performance reasons. But I've decided to stick this general approach for now because I don't need to substantially rearchitect everything, I can just replace the very few bits that have proven to be an actual performance bottleneck, and keep the development speed for the many places where runtime speed just doesn't matter as much.
- smhinsey 15y agoI don't mean to start an argument over which ORM you want to go with, but you might want to give this a look, if you're interested in being able to quickly iterate on a data model and still have test coverage: http://wiki.fluentnhibernate.org/Persistence_specification_testing http://wiki.fluentnhibernate.org/Persistence_specification_t....
- MartinCron 15y agoCool. Thanks for sharing.
- vertice 15y agogod. thank you. I have never met an ORM that didnt eventually rub me the wrong way.
- geekfactor 15y agoAs they say, familiarity breeds contempt. I can name a hundred things my wife does that annoy me, but that doesn't mean I'm not worlds better with her than without.
- true_religion 15y agoTrue, but some people might be better off looking for new partners---and new alternatives to ORMs.
- deleted 15y ago[deleted]
- clistctrl 15y agoYesterday morning I would've called this guy a cranky old man... but doing some coding last night made me want to kill something. Coming from an Active Record background I tried using Linq to SQL. My application has a WPF front end and a Windows service on the backend. Passing the same object between the 2 is driving me insane. The problems are so much more cryptic, and the code I had to write to go around it completely negates any reason for using it in the first place. I'll be ripping it out tonight.
- kprobst 15y agoToList() usually solves most marshalling problems (not that that is your problem, but it's quite common). Nothing wrong with Linq2SQL, it's just another technology that has a learning curve. There's more than enough information out there to solve most problems I've ever run into with Linq or EF. Whatever you're dealing with, someone out there probably already dealt with and blogged about it or asked a question and got an answer somewhere.
- pnathan 15y agoWell, I can't speak for others, but in my small-scale use of databases, hand-coded SQL has never been an issue.
- ry0ohki 15y agoSame here, ORM just adds another layer of frustration for me.
- SeoxyS 15y agoThe main problems with ORMs is that they're trying to work around non-object-oriented data stores. Layers of abstractions and ORMs in particular are generally good things—but they can't do magic when it comes to dealing with SQL. If you're going to be using an ORM, I'd strongly recommend rethinking your data store. Object databases such as MongoDB is a perfect fit, but even a key-value store like Cassandra would be a much better option than SQL. I think it's interesting to note that Core Data, Cocoa's ORM, is one of the fastest data store out there. It uses SQLite, but defines its own schemas. I believe it'll also let you store pure binary data.
- chc 15y agoSmall nitpick: If you're using a non-relational database, the library you use to connect to it is not an object-relational mapping layer.
- SeoxyS 15y agoYou have a point! I guess I'm bastardizing the definition of an ORM. What I mean by it was a library that automatically maps the data layer to objects in your application code, and (often) gives you tools to work with these objects. I wrote an "ORM" for MongoDB which adds functionality such as transparent relationships.[1] Basically, even though it's not a relational db, it lets you do things like this: foreach ($author->books as $book) // ^ ^ this is a Book object // ^ this is an iterator, it loads a // Book object lazily every iteration. echo $book->author->name; // ^ this is an Author object, auto- // matically & lazily loaded & cached. [1]: https://github.com/kballenegger/MongoModel https://github.com/kballenegger/MongoModel Among other many cool features. The point of highlighting that though, was to illustrate that giving up traditional RDBMS doesn't mean giving up on awesome relationships. The only thing missing is subqueries—but honestly, I don't think that's a very big loss.
- dgallagher 15y agoYes, Core Data supports storing raw binary data. IIRC is also uses caching (for SQLite stores), and lazily-loads related objects as needed. You can customize this behavior to optimize its memory usage for your code.
- Lagged2Death 15y agoI'm a noob when it comes to designing and implementing programs that interface with a relational database. Based on my small experience so far with an ORM, I'd say this post is spot-on, clearly articulating the frustrations I've felt on my project. A friend of mine even wrote a blog post about the problems I've had: http://es-cue-el.blogspot.com/2011/03/entity-framework-and-inheritance.html http://es-cue-el.blogspot.com/2011/03/entity-framework-and-i... That said, though, I do wish there were more detail on this point: The programming world is currently awash with key-value stores that will allow you to hold elegant, self-contained data structures in huge quantities and access them at lightning speed. I'd love to know more about such libraries, frameworks, or tools, but this isn't a lot to go on.
- ebiester 15y agosearch for noSQL. He's right that it's a misnomer, but it's hard to give a recommendation because they are optimized for different use cases. http://en.wikipedia.org/wiki/NoSQL http://en.wikipedia.org/wiki/NoSQL gives a good set for the various use cases.
- rahoulb 15y agoI suspect a major problem with ORMs is that a number of people are using them without a full knowledge of SQL beforehand. Relational databases are complex and you need to understand what's going on if you want to make best use of them. I've never had any issues with ActiveRecord (although I agree that it's mixing of data access and business logic can be a problem) - but I also know what needs joins, which columns to include, which indexes to use, when to drop to raw SQL; all from years of writing complex SQL and stored procedures by hand (and I never want to go back to that). And I don't ActiveRecord to magically guess that stuff for me.
- Spyro7 15y agoI think that designating ORMs as "anti-patterns" is a bit strong. Perhaps my understanding of what an ORM is supposed to accomplish is different from the author's, but I think that some of the criticisms that the author levels against ORMs are a bit off. Inadequate abstraction - I would make the argument that it doesn't make sense to expect an ORM to be able to completely abstract away from the underlying database. The reason that the documentation of the various ORMs is sprinkled with SQL concepts is that the ORM is providing a window into an SQL-based environment. I would never Incorrect abstraction - I actually agree with this point, but this does not really seem to reflect on ORMs. This point has much more to do with the ongoing debate between the NoSQL movement and relational databases. Death by a thousand queries - I hardly think that this is a knock against all ORMs. Different ORMs have different solutions (or a lack thereof) to this problem. I use Django a lot, and Django's built-in ORM offers a lot of "frills" that can help to protect against this (lazy loading, selective loading of columns, selected loading of related models). I know that, in the Ruby world, Datamapper seems to have some ways of dealing with this problem as well. It really isn't as simple as saying all ORMs do this therefore all ORMs are bad. The reality is more nuanced. Ultimately, my principle problem with this piece is that it seems to conflate its argument for NoSQL and its argument against ORMs. NoSQL is wonderful, but it seems to be somewhat orthogonal to the value of ORMs. ORMs are not perfect, and there is plenty of room for improvement; however, writing everything in SQL solely due to performance fears will usually turn out to be a case of premature optimization.
- riffraff 15y agoFWIW even rail's ActiveRecord has the same frills (lazy/eager loading, selective loading of columns, loading related models via join or via grouped selects). It even has modules that do this fixes automatically :)
- mistermann 15y agoOne aspect that always comes up is the "inefficiency" of doing a select * from a table with 30 columns when you only need 4 columns. 99% of the time the millisecond performance difference doesn't matter, and if it does, there is a standard non-default way to handle it in most ORM's. However, one aspect that is usually conspicuously absent in anti-orm blog posts is that of development time and cost. ORM usage practically guarantees known coded efficiencies, but it lets you implement and pivot really quickly, the time and money saved is easily more than enough to pay for a bump in hardware to overcome the 10% slower code. But to do so is heresy for these people....selecting columns from the database that you do not use is just not done, full stop. Which is cheaper, in dollars, is irrelevant.
- epscylonb 15y agoThis surprised me as well, I was told that in Postgres at least, there is no difference in query speed between selecting a subset of a record and the entire thing.
- chrisb 15y agoThe query speed may be the same (I don't know if this is true or not), but fetching the data from disk and returning it over a network will certainly be slower, especially if the unnecessary columns contain large strings or blobs.
- seldo 15y agoIt depends. In MySQL, if the only columns you select are indexed columns, the entire query will be pulled from the index, which is usually entirely in-memory, so enormously faster: http://dev.mysql.com/doc/refman/5.5/en/where-optimizations.html http://dev.mysql.com/doc/refman/5.5/en/where-optimizations.h... (The columns currently all have to be numeric for this to be true, but that's surprisingly often the case)
- nerfhammer 15y agoShort answer: Any relational database is going to read in the entire row when you select it. You will save network traffic and probably space in some internal buffers only. Longer answer: Any relational database will read the entire page the row resides on when you select a row. This means that tables with more columns will have less efficient pages which will need to be read/buffered more frequently per row. This also means there is an exception: columns that are stored in overflow pages away from the rest of the row like blob or text columns or (in some db's) very long varchars may not necessarily be read if you don't select them. There is an additional, very important exception: if you use a covering index then the db will not need to read the data page the row resides on. For example, if you have an index on (username, user_id) and you select "select user_id from table where username=xxxxx" then it will be able to read the user_id from the leaf node of the index and no bookmark lookup to the data pages will be needed. In some db's the primary key is always "covered" and you never need a bookmark lookup to get it.
- gte910h 15y agoThis is a person who doen't write many large scale systems: You will have considerably more (sometimes serious) bugs if you write all your SQL by hand all the time in a app that uses a lot of DB queries. Yes, you still need to understand what the ORM does when you do certain things, you still need to understand what nasty joins you're writing and all that. But you can let all the minutiae of what you DO write work out well in a rote, well tested manner. The article smells a bit of a guy who didn't know SQL or had a team member who didn't, and they though just using and ORM would work. If your app is successful, you will usually need to optimize things. But this is true for SQL or any time saving abstraction as well, not just ORMs
- hahainternet 15y agoWhat I found most telling was the talk about doing complex joins in your application. You do them in views and stored procedures. Your application code should be distinct from the data it sources, so people don't have to wade through 400 lines of terrible PHP just to change 'fullname' to 'concat(etc)'. Don't torture your DBAs.
- rimantas 15y agoStored procedures are the sure way torture both, DBA and devs.
- dazzer 15y agoI agree with rimantas. Views and Stored Procedures should not be used to perform business logic stuff i.e. I should not have to create a view with massive joins just because a logic need requires it. And I should not have to dive into my database when a business rule changes!!! Views and Stored Procedures are useful when you lack a layer of abstraction (e.g. An MS Access FrontEnd) where you may want to put security restrictions on the data that is exposed to a particular group of users, OR for performance reasons where a reasonably complex query can be run faster as a stored procedure. Of course this is my personal view, and is definitely a point of contention for many people. The whole point of an ORM is to abstract the data from application code. Business Logic can be built on top of it with minimal knowledge of the underlying data storage system except in exceptional cases. ORM frameworks aim to simplify the process of writing these boilerplate code and continue to fulfil most common use cases.
- DavidMcLaughlin 15y agoThis seems incredibly naive. ORMs reduce code duplication. They speed up development, especially when you're treating the underlying data storage as a "dumb" datastore that could just as easily be sqlite or H2 as MySQL or Postgres. As for ORMs having some sort of negative impact on the queries sent to the underlying database - it really depends on what ORM you use but any ORM I've used had support for pre-loading relationships in advance when required, removing that N+1 problem. I also want to add that I wrote an ORM for the first company I worked for and when it was finished it was a drop-in replacement for 90% of the queries in our application - and I mean that literally the SQL generated by the ORM was exactly the same as the SQL being replaced. The queries that it couldn't replace (mainly reporting queries) already had an aggressively tuned caching layer in front of them anyway because they were so hairy. But the real point is this: the performance of the ORM didn't really matter because we were a database driven website that needed to scale - so we had layers upon layers of caching to deal with that issue. And that is an extremely important point - the way ORMs generalise a lot of queries (every query for an object is always the same no matter what columns you really need) lends itself to extremely good cache performance. Take the query cache of MySQL for example - it stores result sets in a sort of LRU. If you make n queries for the same row in a DB but select different columns each time - you store the same "entity" n times in the query cache. Depending on how big n is, that can cause much worse cache hit performance than simply storing one representation of that entity and letting all n use cases use the attributes they need. Now, relying on MySQL's query cache for anything would not be smart, but replace it with memcached or reddis or whatever memory-is-a-premium cache and the same point stands. Another example to drive the point home is a result set where you join the result entities to the user query so that you can get all the results back in a single query. In theory this is a great way to reduce the number of queries sent to your DB but if you have caching then there are many times where you could have very low cache hit ratios for user queries since they tend to be unique (for example they use user id) but where you could still get great cache hit performance if certain entities appear often across all those result sets by leaving out the join and doing N+1 fetches instead. ORMs prevent you from scaling as much as using Python or Ruby over C does. So I guess that leaves the point about leaky or broken abstractions. Well I would never claim that you can abstract across a whole bunch of databases anyway, I think that's a ridiculous claim that most ORMs make. These types of abstractions when people try to hide the underlying technology are really just a lowest-common-denominator of all the feature sets. So if you chose some technology because you really wanted a differentiating feature then most likely you will find yourself working against such abstractions. Interestingly enough, the dire support for cross-database queries which are perfectly legal in MySQL but not in other vendors is the reason I had to roll my own ORM. But the productivity and maintainability benefits were well worth it. So yeah I guess what I'm saying is: premature optimization is the root of all evil, there are no silver bullets and performance and scalability is about measuring and optimising where needed. And finally: ORMs are not an anti-pattern.
- cwp 15y agoYup, spot on analysis. I do think he missed one thing, however. The few times I've seen ORM layers work well is when they're custom-built for a specific application. It's still not ideal, but it lets you put all the ugliness in one place, and regain efficiency by sacrificing generality.
- JulianMorrison 15y agoI like the iBatis (now renamed mybatis) approach: explicit queries in a separate file that say "input an object of this type reading these fields, output an object of that type setting those fields" and contain raw SQL to be thus parameterized. This avoids the two largest flaws: live proxies with hidden state pretending to be simple data objects, and SQL being generated with no control. It also avoids a mistake I've only seen two ORMs make but they're common ones: defining its own dialect of not-quite-SQL. You still get objects mapped in and out of DB queries, it saves you the pointless grunt work of "copy A, put it in B" and it prevents the as-bad-as-ORM anti pattern of "SQL scattered throughout your code".
- narrator 15y agoI think Ibatis is great for querying data out of the db and ETL operations, especially when I'm using a lot of database features. However, IMHO, Hibernate is better for crud opts on individual records because it takes care of dirty flagging and managing relationships.
- JulianMorrison 15y agoIn my experience Hibernate fails badly at dirty flagging and saves every field in the object rather than updating the changed properties (active record is better). I suspect that's a design decision, to avoid an object in cache being partially stale relative to the DB. But it's an problematic solution to something that didn't need to be a problem. Hibernate does manage relationships - but that is a misfeature. It's doing the wrong thing well. The right thing is not to try to model relations as objects, but to model queries as methods. The relationships exist only in the database - they are not duplicated into the data objects.
- Zvez 15y ago"The relationships exist only in the database" Entity Person contains collection of entity ContactInfo. In DB we have tables PERSON, CONTACT_INFO and foreign key from CONTACT_INFO to PERSON. What the difference in relations in code (between entities) and in DB between tables? In this example.
- encoderer 15y agoSomebody may have already mentioned this, but there's a fantastic essay The Vietnam of Computer Science (2004) on this subject. It's long but so, so worth it. http://blogs.tedneward.com/2006/06/26/The+Vietnam+Of+Computer+Science.aspx http://blogs.tedneward.com/2006/06/26/The+Vietnam+Of+Compute...
- wulczer 15y agoMy first reflex was to write a comment mentioning that essay and then I found this. Wish I could upvote you more...
- JunkDNA 15y agoWow, thanks for contributing this to the thread. I had not seen this before. That has to be the best explanation of the tradeoffs/benefits of ORM I've seen anywhere. Most of the content isn't "new" in the sense that if you've done both a lot of SQL and a lot of ORM work, the issues are pretty obvious. But it does a fantastic job of showing why there's just no simple answer currently.
- cturner 15y agoThe article makes a strong case against ActiveRecord, not against Object Relational Mapping. Under the heading "The problem with ORM" the author writes, The most obvious problem with ORM as an abstraction is that it does not adequately abstract away the implementation details. The documentation of all the major ORM libraries is rife with references to SQL concepts. Some introduce them without indicating their equivalents in SQL, while others treat the library as merely a set of procedural functions for generating SQL. This is true of ActiveRecord, it's untrue of object graph ORMs like Apache Cayenne. I find two patterns to be key to effective ORM: * Data Access Objects. This is for when you have nothing, and want to get an entrypoint into the schema. In this case, you should be able to write near-pure SQL to get what you want. * Entity Objects. This is what the DAO will give you back - either an individual or a list. Each instance represents a row in a table, and has methods that will do lookups to foreign keys. Once you have this, you have an entrypoint into the data graph, and can use foreign keys to crawl around to wherever you need to go. The DAO layer is a simple, centralised place where you can implement permissioning logic. If you need to do something high-performance (usually some sort of report), you create a custom DAO, and have it return custom entities (instances of classes that don't have 1 to 1 association with a table) that fit your need. I've found that after a certain point of complexity in an application, it becomes impractical not to use an ORM. It's like working in a type-unsafe language. You refactor something, and SQL-in-code breaks all over the place. That path leads to the hiring of dedicated DBAs, and abstraction of the schema behind stateless layers of PL/SQL in a doomed attempt to get to grips with the complexity of the problem space. I worked on a system with a very tough customer where they repeatedly demanded major schema changes that were sitting in front of a business logic layer and frontend that had already been written. While the project had lots of problems, those particular refactorings were very straightforward. I was able to modify the ORM, and then just fix complilation problems and a few obvious tentacles from them until the application recompiled, at which point it worked again. Some more criticism: This leads naturally to another problem of ORM: inefficiency. When you fetch an object, which of its properties (columns in the table) do you need? ORM can't know, so it gets all of them (or it requires you to say, breaking the abstraction). I'm rusty but remember that at least in WebObjects EOF at least you can nominate what you want to retrieve, including automatic joins to retrieve stuff over foreign key jumps The author's first suggested alternative "Use objects" offers worse technical debt than ActiveRecord. I anticipate there are a lot of shitty systems being on top of key-value stores. You can get fast results doing it, but it has technical debt and doesn't scale horizontally. The key-value store is becoming the next generation equivalent of "Oh we'll just build it in excel, and worry about the consequences later on". But depends what you're doing. There are situations where foreign keys are good. The second alternative is "Use SQL in the Model" The advice of the heading doesn't match the content of the text that follows. I think the author means to recommend building a service that wraps the model by answering questions. If not, that's the point I think that should be made. It's common for companies to create a database, and then have many entrypoints into it. This is a mistake and creates technical debt. As soon as you have multiple entrypoints like this, you lose ability to refactor your schema (because it's impractical to get multiple stakeholders to make concurrent changes) and your system rots. Instead, you create a model service that wraps the schema, but also has stateful knowledge. For example - it knows the permissions of the user who is talking to it and can tailor its response based on their permissions. Then you return results in a transport format. I can't recommend a good, mainstream mechanism for this. JSON, YAML are fiddly because they're typed, XML is unnecessarily verbose Anyway - there's no reason not to use a good ORM in this business logic layer. For small systems - sure - use SQL in the model. For the larger stuff, you have a more maintainable system if you use an ORM. But if it's a complicated space, steer towards Cayenne or Hibernate, rather than active record patterns.
- Bdennyw 15y agoI have felt the pain of Hibernate. It just sucks. But I think that it is possible to make an ORM like solution work. A great example is NeXT's EOF and now Apple's Core Data. The combo of awesome mapping tools and Objective-C's dynamism make for system that works very well for it's intended purpose. Mind you core data is not a database and likely would not work well for a web service, but I believe that EOF did.
- forgotAgain 15y agoSeems like a "horses for courses argument" but maybe that's just the sign of my being a cranky old man. More important than using an ORM is that everyone consistently uses the same toolset for the project. Personally I can't say I'm a big fan of using an ORM. Just too many bad tastes in my mouth over the years from bad implementations. It's probably improved by now but I long ago developed tools to generate the boiler plate code I need to work with a database. This gives me a generated data layer and a bare business object layer that moves the data out of the data layer. With the metadata available from databases it's fairly simple to automate the generation of the code. Once you have a tool that generates the code then an ORM has much less to offer.
- skittles 15y agoIf a project uses object-oriented programming and a relational database, it will have an ORM. Either one written somewhere else (hibernate, Entity Framework, etc.) or one written in house (whether or not it is thought of as an ORM). An ORM maps object data to relational data and back. That's all.
- andybak 15y agoSometimes 'good enough' really is good enough. For christ's sake, there's a zillion ways to mitigate ORM related performance hits. One of those zillion is 'stop using an ORM' but it's not likely to be your first choice.
- mgkimsal 15y ago"The whole point of an abstraction is that it is supposed to simplify" No, it's supposed to abstract. A simplification is supposed to simplify. Often abstractions have the benefit of simplification, but it's not a requirement. I migrated a project from MySQL to PostgreSQL last summer, and the project was built on Grails with GORM. I had to migrate the data by hand (mostly easy, save for a couple of edge cases like boolean columns), and I had to change the jdbc driver. That was pretty much it. No rewriting of SQL, no changing of escaping logic, etc. I tell a lie - the auto-sequence generation stuff of postgresql wasn't playing nice with some of the GORM identity stuff, and my code had made some assumptions that turned out not to be 100% true. Those likely would have shown up had I written my own stuff rather than relied on GORM, but it was a little bit of a pain to track those down. All in all, using the ORM abstracted away the need to write against specific database commands and syntax. A byproduct of that was simplification of most use cases of the database, but the key use was abstraction.
- kstrauser 15y agoI inherited an incredibly hairy, large, mission-critical database at my current job. While we're slowly phasing in its replacement, we'll be interacting with the current mess for a long time to come. There are seemingly endless little insanities, like "this column references upper(substr(othertable.column,5))". Instead of trying to remember all such idiosyncracies, I defined them all in SQLAlchemy and added a _lot_ of unit tests to make sure I don't accidentally break one later. Now I can use programmer-friendly ORM joins in production code and not have to worry about getting all the weird rules right each time. I'm perfectly comfortable working in SQL. I don't want to write it directly all the time, though, any more than I want to have to write assembler all the time.
- code_duck 15y agoI had a lot of problems working with ORMs when I was 1-2 years into programming. However, I also felt a lot of resistance to learning to use a framework vs. straightforward, procedural code for web apps. It's a matter of wanting to take the time to learn another system, API, DSL, what-have-you just in order to work with something you already know - SQL. The dislike of HQL resonates with me - I was wondering why I would ever work with PHP Doctrine's DQL. Building SQL queries out of a sequence of OO method calls seems absurd, too. As the article and comments note, you shouldn't have to know SQL well to use an ORM. There are definitely issues with the ORM/Framework working against you, too. I love the organization and features in Rails or Django, but I hate when I spend hours working out how to do something that would take 5 minutes in plain PHP. Same with ORMs. Getting them to do the right type of join, not make unnecessary calls, etc. can be a pain. Sometimes it's that I don't know the software well enough, which could either be my own problem or just a reasonable lack of desire to devote my brain to it. Other times it's that the given ORM really does have shortcomings, conceptually and at level of development. The one ORM I've had the most luck with is Django's. It's straightforward, does what I want, is well documented, and doesn't have too many features.
- d4nt 15y agoI'd say the problem is more with OO. The author touches on this a little towards the end, but ORM fulfills the demand to put relational data into objects. The issue is with all these developers insisting that they want to think only in a limited set of types, even when the page they're rendering needs half the properties of one type and a few more from another type. Something like linq2sql can actually be the answer here, if you use it to select an anonymous type containing exactly what you want, the you only need one round trip and you haven't wasted and processing.
- mechanical_fish 15y agoMy conclusion, drawn from the title alone: The term antipattern has apparently jumped the shark. Spend five minutes decoding a particularly hairy regular expression? Regexps are an antipattern. Someone writes an inefficient SQL query? SQL is an antipattern. Stub your toe on a curb? Curbs are an antipattern.
- absconditus 15y agoIt seems like the term was flawed from the start. "Anti" does not mean "bad".
- zizee 15y agoNo, it means opposite. Patterns are "the right way to do things". So, the term anti-pattern means "not the right way to do things".
- Terretta 15y agoPatterns are recurring designs, so anti-patterns might merely be non-recurring designs. This isn't necessarily a value judgment.
- compay 15y agoThe term entered widespread use after the AntiPatterns book was published. The authors were referring to "patterns of failure" commonly seen in software projects.
- StrawberryFrog 15y agoNope, the AntiPatterns book defined it as something like "recurring patterns of failure, which look superficially attractive".
- Terretta 15y agoEdit: yes, I am extremely well aware of the conventional definition. I was deliberately offering an alternative spin in context to the parent. Case in point: ORM can handle 80% of query plans (patterns). Some, it can't handle. They are not patterns, they are the opposite: unique and non replicable. That doesn't make them bad. However, I disagree with headline. Just because ORM doesn't handle non pattern situations doesn't make the concept an anti pattern.
- rjurney 15y agoAgree. Talked about how severe ORM impedance mismatch is here: http://datasyndrome.com/post/3257282059/data-driven-recursive-interfaces-for-graph-data http://datasyndrome.com/post/3257282059/data-driven-recursiv... Your model needs to fit your view.
- perlgeek 15y ago> If your data is objects, stop using a relational database. What does that even mean? My data, is, well, data. Tables and rows are just ways to represent my data, as are the nested hash and array structures of document storage systems. Oh, and tables and rows are also objects. What data is "object" and what data is "non-object"?
- seldo 15y agoYes, I could have clarified that. Here's an attempt: "Relational" data is data whose value stems from its relationships with other data. For instance, if it is statistical data that is viewed in aggregate rather than as individual rows. Or if you need to answer the question "how many rows look like this row?" "Object" data is data that is useful in and of itself, and is largely self contained. A pretty good example is a blog post: each post has a bunch of metadata, including possibly a string of comments. But you seldom if ever run queries across batches of blog posts (other than indexing them by date). It's always bugged me that blog entries -- the staple of the ORM tutorial -- have little to no relational value, which is why they work so well in ORM.
- glenjamin 15y agoSplitting blog posts by date is hardly something thats rare. See also, all blog posts by tag, all posts by author, all posts containing the word "javascript", latest comments by a particular user etc. Well organised blogs have a reasonable amount of relational data.
- joeburke 15y agoAnother mistake the author is making is calling the ORM "an abstraction". ORM is a mapping technology: it takes input from one world and turns it into data suitable to another world. It doesn't abstract anything.
- dazzer 15y agoTechnically, (my interpretation is that) it abstracts your application from your db. The ORM acquires the data, allowing your application code to concentrate on using the data.
- mgkimsal 15y agoOne other thing struck me reading this - it feels like premature optimization. Assuming that every ORM is going to be slow and inefficient to the point where you'll need to override or rewrite all the queries will lead to an inefficient use of developer time, and assumes you know a lot about what will matter under real world use conditions. Yeah, sure, that ORM is adding 200% overhead to the SQL query - it's pulling back 30 columns instead of 4! And... it's taking 38 milliseconds and is run 4 times per day. So what? And when the model changes and you have an extra few columns to represent more data? You've now got to hunt through every SQL query that could possibly reference that table and make sure it's dealing with the new columns appropriately, instead of having an ORM let the computer do what computers do - compute the changes required. Yes, there are other ancedotes that can be trotted out to prove the opposite of my 38ms story above. Then we'll fall back to 'right tool for the right job', and ORMs are currently a good middle ground tool for many of the projects people are developing. Perfect? No. Useful? Yes.
- keithnoizu 15y agoMhh the premature optomization school of thought. Kiss good and all but I think einstein had it right with his "Everything should be made as simple as possible, but not simpler." A modest amount of additional upfront is probably worth it if it saves you time and effort in the long run. It's the same reason why its probably worthwhile to learn a framework rather than rolling your own organic solution, setting up templating, using version control from the get go. They all add upfront complexity but pay off in the long run.
- geebee 15y agoI've used two different formal ORMs, ActiveRecord and JPA (backed by hibernate), and I've never felt completely at ease with them. In fact, I was lining up to agree that ORMs suck, except that I realized I'm probably using one no matter what I do. If I have a model object, and I want to to persist it in a relational database, then I'm going to need to do something that persists and retrieves this object back and forth from the RDBMS, right? And if I want to retain the flexibility to switch to a different database (or different persistence strategy in general), then I'm going to some way to specify the implementation details for each possible approach, and swap them in depending on which approach I take. The java world tends to handle this with an ORM and a DAO tier these days, using DI to swap in the desired implementation (sigh, Java really is a soup of acronyms these days), whereas Rails developers tend to use migrations. But either way, I'm pretty much stuck with an ORM. It may be an ORM that works at a very low level, directly with objects, sql, connections, and transactions rather than through a higher level API, but it's still an ORM... (right?) In spite of the inevitability of an ORM (I really hope I'm defining this correctly), I'm going to agree with a lot of the points made in this blog post. If I'm using SQL, I really don't like having the SQL hidden from me. And I really can't stand languages (like HSQL) that force me to re-learn a variant of SQL. ActiveRecord and Migrations are, without question, very productive, but I like to see the objects and understand very directly how they are being persisted and retrieved. I want to see the fields and methods, and I want to see the SQL. I've found that I almost always end up changing it, and it's easier to do that when it isn't all hidden from me. Rails offers me so much that I can get over this little issue of mine, but I don't feel the same way about Hibernate. My personal experience is that if Rails isn't going to help me, and something has pushed me to use Java, I probably need to write a lot of lowish level code anyway.
- absconditus 15y ago"If your project really does not need any relational data features, then ORM will work perfectly for you, but then you have a different problem: you're using the wrong datastore. The overhead of a relational datastore is enormous; this is a large part of why NoSQL data stores are so much faster." I have never been able to receive a straight answer to this question: Is there a "NoSQL" database that provides the same ACID properties that major RDBMS databases do? Things like "eventual consistency" are entirely unacceptable for the software that I work on.
- smharris65 15y agoI'm not trying to say it's the best fit for you, but Apache CouchDB is fully ACID compliant: http://couchdb.apache.org/docs/overview.html http://couchdb.apache.org/docs/overview.html
- wulczer 15y agohttp://en.wikipedia.org/wiki/CAP_theorem http://en.wikipedia.org/wiki/CAP_theorem
- jtchang 15y agoSo I don't have deep experience with ORMs except for SQLAlchemy. I've worked with others such as Hibernate. All I can say is that I love SQLAlchemy. Projects that use it make my life easier. It is easier to read and troubleshoot than pure SQL. Just because the majority of ORMs such doesn't mean they all suck. It's like saying "web frameworks" are anti-patterns. ORM done right can be a god send. Like all patterns knowing when to apply it is key. Go around creating FactorFactoryFactoryObjectFactory and of course you'd think it is an anti-pattern.
- Swannie 15y agoIn the comments I'm noticing no one ask: when should or shouldn't you use an ORM? Most of the discussions are over the merits of either approach, when to me it seems an ORM has many places it belongs. And a few it doesn't. For most database of record systems, which are a large chunk of your average webapp, an ORM is a god send. When I say DBOR, I mean things like articles, posts, comments, users, products, transaction history. An ORM saves a large amount of work writing SQL, it covers 95% of your queries (particularly insert,update,delete and simple gets) with minimal effort. You create the model, and let it get dealt with by the ORM. Your objects are mainly records. The pain comes when you start wanting to do analytics and interesting reports - but stick with a reporting tool, and keep this out of your application, and you feel less pain. But this breaks down when you move to a database that represents a complex real world system. If you're working on a model that represents, for example, an electrical distribution system, these are not really records. They represent a vast set of complex interrelations, Of course there are still records, but in isolation, away from the complex relationship of say pole->{location,type,maintenance history,conductors,insulator type}, and conductor->{poles traversed,length,a end location,a end join type,b end location, b end join type,material,material batch number,power circuit carried} etc. etc. Then your queries to "find all customers affected by the pole at these coordinates", requires joins through: pole, conductor, circuit, serviced area, customers... we're moving rapidly to lots of complex queries, where hand crafting really is the way to go.
- dasil003 15y agoI started writing a response in a comment, and then I started frothing at the mouth, and pretty soon it ballooned into a whole blog post: http://news.ycombinator.com/item?id=2659442 http://news.ycombinator.com/item?id=2659442
- nathanlrivera 15y agoGeneralizations are an anti-pattern.
- trustfundbaby 15y agoI'd like to learn more about the author's background with ORMs and especially activerecord ... I know SQL pretty well, in fact when I started using Rails, I insisted on still writing my own SQL queries by hand. ActiveRecord might be an anti-pattern, or it might not ... I really couldn't care less, what I do know is that I enjoy dealing with the database using active record far more than error prone dynamically constructed SQL queries I was doing back in my PHP days. It makes my life as a coder easier ... I mean, have you ever tried to construct a really complex search function on a web app using SQL? ... its a pain and a half. ActiveRecord makes stuff like that much easier (named scopes in Rails especially) Yes, if you don't understand databases, you're going to use an ORM in shameful ways, but it works well ... very well, if you know what you're doing and you take the time to learn your craft. I'm glad to have ActiveRecord in my tool belt every morning when I get to work and that ... is what really matters to me.
- jgrahamc 15y agoOne of the problems I frequently see is that people complaining about ORM and SQL are thinking mostly of some object wrapping a row (or set of rows) in a table. Then they get into trouble when they want to wrap something more complex involving joins between tables. All these problems would disappear if people used database views. Then their nice ORM layer (say ActiveRecord) would work perfectly and the nasty joining and updating would be taken care of by the database. I've often wondered if people even realize that database views exist and how powerful they are: http://en.wikipedia.org/wiki/View_(database) http://en.wikipedia.org/wiki/View_(database) Of course, it's only relatively recently that MySQL has started supporting views properly (in 5.0). The other nice thing about views is that it means your code using the ORM is simplified because you aren't indirecting through different objects to get at specific values you need to display. It also means that only the necessary data is retrieved from the database.
- G_Morgan 15y agoThat is the other problem. Every database is a mash of semi intentional subtle incompatibility with the standard and a host of non-standard features. When an ORM comes along many try to expose the non-standard features that make sense in some way but end up needing a custom solution for each DB (see how you do an auto sequence ID for an entity bean between Postgres and MSSQL). So you have a semi portable layer interfacing with a semi portable environment.
- shaydoc 15y agoORM's tend to suck, easy way out. delegating control to a custom ORM says to me, OK give me performance issues. Design your Domain model, keep it simple at the DAL and use Stored Procedures, easy life, ultimate flexibility.
- danssig 15y ago>When you fetch an object, which of its properties (columns in the table) do you need? ORM can't know, so it gets all of them (or it requires you to say, breaking the abstraction). Not true. If you take the nHibernate approach of returning a proxy argument then you can get clients to "tell you" without breaking the abstraction. You normally don't worry about this, though, because pulling 30 properties usually isn't much different then pulling 3.
- tcarnell 15y agoI 100% completely agree with the article. ORM's are dangerous. ORM's solve a problem we dont really have, but introduce a whole load of new problems we never had before (lazy loading does not work, object relations work completley different from table contraints, designing the domain layer to fit our persistance layer, learning new proprietary query languages, inability to control sql queries) I am continuously amazed at how keen developers are to adopt them as a core part of a their product. Thanks for the article - I feel relieved others feel the same way!
- kunley 15y agoit's a pity and a sign of ignorance that people do an implicit assumption that ORM == active record pattern, while the other ORM pattern: data mapper is in fact widely used and superior for many use cases.
- dazzer 15y agoInterpreted pedantically, removing ORM techniques means dealing directly with resultsets or loosly typed structures. This is definitely what I DO NOT want in any of my views. If you're working with an OO language like C#.NET or Java, good luck! Any abstraction on any level is going to add a performance hit no matter what. If this was really an issue, wrap your ORM Framework stuff (differentiating from ORM the pattern) in a DAL layer so that your BLL does not worry about the existence of the ORM. Then as you scale, optimise your DAL with either inbuilt optimisations or when desperate write your own SQL (if you don't even know SQL then you're a poor excuse of a developer) Think of them as like Ikea furniture - they don't look great, and they don't often fit in every household if they have complex requirements. But they're highly modular, and easy to assemble. So when you need something in a jiffy, just bring it home, fix it up and it'll perform its purpose. When it no longer fits the purpose, get something else. And every household has to just start somewhere.