8 ms·
Eva – A distributed entity-attribute-value database in Clojure
- CurrentB 7y agoI'm excited that some of the ideas from Datomic are starting to make it into some open source projects. With crux last month and now Eva, there are now multiple options for modern clojure databases. While there are many excellent ideas embedded in Datomic and these projects, for me just being able to persist the same data structures you're using at a repl and query for them with data is a huge win vs having to start translating types and concepts and query strings to and from SQL is a huge win.
- maximente 7y agocan't agree enough! was really interested in the now defunct mentat project, but was written in Rust - it's a game changer to have a JVM-powered datomic-ish project with time as a first class citizen.
- refset 7y ago> time as a first class citizen It's worth noting that time now has two levels of meaning in this context - transaction time and valid time. Disclosure: working on Crux
- maximente 7y agohey, thanks for the reply. could you highlight the differences between Eva and Crux as you see them, today? i do like passing in a timestamp and getting the state of the world as Crux seems to implement, whereas Eva seems slightly less oriented on time based querying, more so for a sort of log setup (where they have EAV+T+added?)
- refset 7y agoWell, Eva overlaps very closely with Datomic so I would recommend taking a look at the FAQ entry I wrote on the Crux/Datomic comparison (with 8 sections covering the major differences!): https://juxt.pro/crux/docs/faq.html#_comparisons https://juxt.pro/crux/docs/faq.html#_comparisons `EAV+T+added?` is the canonical data representation used by Eva, whereas Crux relies on two discrete "document" and transaction logs as the canonical representation. There are many reasons for this (data eviction, "unbundled" scalability etc.) but bitemporality (i.e. the ability to transact into the past/future) is definitely the primary reason why Crux doesn't follow the `EAV+T+Added?` pattern. You can get a feel for the indexes Crux uses here: https://github.com/juxt/crux/blob/master/src/crux/codec.clj#L26-L64 https://github.com/juxt/crux/blob/master/src/crux/codec.clj#...
- derefr 7y ago> I'm excited that some of the ideas from Datomic are starting to make it into some open source projects Now if only they’d make it into open source projects in languages other than Clojure. I want my Datomic-for-Elixir, darn it! (And, if I wasn’t too busy, I’d be the first person to volunteer to build it!)
- ilikehurdles 7y agoWhich databases are you currently using that are written in Elixir?
- derefr 7y agoOf course, I meant Erlang; these things (Riak, CouchDB, RabbitMQ) always end up getting written in Erlang. (Though I don’t think there’s been a new major infra component project started since Elixir became a viable contender, so maybe that could change.)
- ilikehurdles 7y agoI didn't realize those were written in Erlang! Similarly, on the clojure side, Datomic and Crux I think are both written in Java, despite generally being consumed by clojurey interfaces.
- estsauver 7y agoCrux is definitely written in Clojure and Hickey has said Datomic is written in Clojure in the past.
- skrebbel 7y agoIt's not precisely the same thing, but I wonder why the Elixir community doesn't seem to use mnesia too much. It ships right there with the OTP and you can store any erlang term.
- 7y ago
- eternalban 7y agoI suspect by "ideas from Datomic" you are referring to ideas predating Datomic (and Clojure) that are used and possibly refined in Datomic.
- bitwize 7y agoIn the Clojure community, no new technology officially exists until Rich Hickey invents it.
- fnordsensei 7y agoMy impression is that the Clojure community (including Rich Hickey) is generally transparent and proud about digging up and reusing ideas from the 80s in a modern context.
- abc_lisper 7y agoI think he meant to say they were not mainstream
- deleted 7y ago[deleted]
- emmanueloga_ 7y agoTo be fair, a quick google search for "open source datalog database" does not return much that is usable out-of-the-box, other than datomic: mostly academic projects or adapters for other databases with some different querying API, like SPARQL. Datomic seems to be one of the best ones in terms of execution of the ideas behind it and usage in real world projects, to say the least.
- ngcc_hk 7y agoCan you elaborate a bit more about their relationship?
- marcrosoft 7y agoLast time I saw the EAV pattern used was in Magento. It was an absolute nightmare to query against.
- nerdponx 7y agoFor those of us who have never used such a thing, what didn't you like about it?
- gouggoug 7y agoThe EAV model is great in theory and is probably great in practice _if_ the backing software has been _designed_ with EAV in mind. In the context of Magento, it is a real _nightmare_ and is one of the major contributor of the slowness of the Magento platform (at least for magento < 2.0). The reason for this slowness is that in a relational database, the EAV model makes it so that e-v-e-r-y s-i-n-g-l-e SQL query is one gigantic query made of tons of JOINs. To give you an example, querying a product in Magento may need to join no less than 11 tables! catalog_product_entity, catalog_product_entity_datetime, catalog_product_entity_decimal, catalog_product_entity_int, catalog_product_entity_gallery, catalog_product_entity_group_price, catalog_product_entity_media_gallery, catalog_product_entity_text, etc, etc. In order to fix this issue, the Magento team created what they call "flat tables" which are tables that are created by querying the database with an EAV query (i.e. the query with a million joins) and putting the results in a table with as many columns as there is attributes being returned by the original query. In theory choosing to use EAV was an amazing idea. In practice, this idea did not scale for large Magento stores and it has made Magento hugely complex, slow and hard to use. We use Magento at betabrand.com and I can confidently say 90% of the slowness of our website is due to Magento's EAV tables and we have spent a humongous number of engineering hours optimizing this.
- komon 7y agoAre you on Magento 1 or 2? We've been building a site in Magento 2 and while the number of joins, hassle with re-indexing performance, and lack of null vs zero vs empty string handling in the flat tables are all very annoying, the actual site is quite fast with all the caching turned on on a server with enough gerbils behind it. My main complaint with Magento 2 is with the feature that is it's biggest selling point: its flexibility. The fact that any public method on any class is wrappable/replaceable, and that any class in the system can be replaced wholesale, and any javascript or any template in the system can be wrapped or replaced, from anywhere, by any module, at a distance, just makes the whole thing a huge cluster to deal with once you get any number of 3rd party modules.
- znpy 7y agoIf someone wants to really make money, please clone the Versant Object Database.
- gilbetron 7y agoVersant, while an awesome ORM, taught me that I hated ORMs ;)
- avodonosov 7y agoVersant is not an ORM
- gilbetron 7y agoShit, been so long I forgot. Versant was cool. Too many databases floating around in my head :/ TopLink was the ORM that made me hate ORMs.
- beders 7y agoI've used Versant back in 2001. It was terrific and very fast, but wasn't a good fit for reporting/analysis. Versant was building systems that were updated on the fly without shutdown, which was quite an achievement back then (and probably still is: Imagine updating a running instance of PostgresQL)
- znpy 7y agoThe company I work for runs dozens of Versant instances and it generally is immensely fast and rock solid. The Object-Oriented Database (OODBMS) space is basically just Versant. Actian however seems not to be developing it actively anymore and it's generally not that much advertises (they rebranded it as "Actian NoSQL" apparently). I wonder if anybody else still uses it (besides us).
- emmanueloga_ 7y agoCan someone explains how anybody is willing to even give a chance to a software product that is sold like this? [1] * Generic/"bootstrappy" looking site, mostly marketing speak. * Not even one code example of what it looks like to solve a problem with this DB. * Gigantic "Get a Trial" button that links to a lengthy form that I'm sure has 99% abandonment rate. There are a lot of other software products that are marketed like this, so I'm genuinely curious how this works. 1: https://www.actian.com/data-management/nosql-object-database/ https://www.actian.com/data-management/nosql-object-database...
- lwb 7y ago> Workiva has decided to discontinue closed development on Eva, but sees a great deal of potential value in opening the code base to the OSS community. I'm confused -- does this mean that Workiva themselves are not using Eva? Or are they still using it, but not officially developing it any more? If they were really invested in it, why would they only allow employees to work on it in their 10% time?
- glesica 7y agoIt was open sourced because it was a cool piece of technology and it would have been a shame to keep it back, but it is no longer in use at Workiva. Source: I work there, although I have literally nothing to do with this project.
- lwb 7y agoInteresting! Are you guys switching to another EAV store like Datomic or something else?
- methehack 7y agoQuestion for Eva and Datomic users: How often does your org miss having a separate relational database around for data folks and product folks capable of (read only) SQL? Do you find yourself making a separate store/s to support them?
- valw 7y agoOne thing that's nice with Datomic (don't know about Eva, but probably the same) is that syncing data to other stores is straightforward, thanks in particular to the the easy change detection.
- FpUser 7y agoOh sweet memories I have created almost identical implementation of this as an embeddable library in a year 2000 along with proprietary SQL language. The reason was that my client had inventory of products with some crazy amount of attributes and each product can have its own set and the client kept changing, creating, deleting those. It was in memory but with persistence and atomic transactions. No history though. It was blindingly fast on complex queries. And the schema of the database was kept as a set of entities with some predefined names, values and range of id's . For a while I was contemplating releasing it as a standalone product but as I had enough tasks on my plate decided not to do it. Kinda feel sorry now ;( So all in all very close by idea.
- beders 7y agoGlad to see another contender in the EAVT field! Yay!
- Dowwie 7y agoEAV is generally an anti-pattern. How does Eva break away from its long, dark history?
- perfmode 7y agoBuild Facebook on it, I suppose.
- lukev 7y agoI reject the premise of your statement. It's only a failure when implemented on top of relational databases which are (obviously) not optimized for it. What are the failures of EAV except when applying it to a database optimized for something different?
- foobar_ 7y agoEAV is the best schema. How you can do NoSQL in SQL.
- briandear 7y agoThe name of this is a bit confusing since there is another language by the same name that also is related to databases: https://link.springer.com/chapter/10.1007/978-3-7091-7557-6_61 https://link.springer.com/chapter/10.1007/978-3-7091-7557-6_...
- AtlasBarfed 7y agoAP or CP?
- refset 7y agoDefinitely CP. With Eva and Datomic, transactions are serialised through a single transactor node at a time. The transaction log design is a fundamental design tradeoff that Eva/Datomic/Crux share, which means system throughput is limited by the throughput of a single process. The argument in favour of such a design is that most businesses & business applications don't actually experience transactional data volumes over 10K tx/sec.
- AtlasBarfed 7y agoE-A-V is basically a JSON/BSON document store, perhaps with typing. Anything that distinguishes it from those datastores?
- refset 7y agoIt's the other covering indexes (i.e. AVE, AEV, VAE) that make it interesting because they allow the query engine to perform efficient ad-hoc graph traversals.