3 ms·
Maybe Spark + Cassandra, but Cassandra has definitely shown to be finicky when it comes to 4+ node clusters, meaning, unless you're willing to contribute time a
by j42 11y ago
Maybe Spark + Cassandra, but Cassandra has definitely shown to be finicky when it comes to 4+ node clusters, meaning, unless you're willing to contribute time and resources to devops and dive into Java, it's a point against getting up and running quickly.
That said, this service runs on HBASE which is great, however, queries function similarly to a mapreduce. This has proven consistently slower than SQL-like alternatives and I can think of a few use-cases where you'd definitely want that added speed.
It's really a question of "good enough" and how many components of your stack you're willing to take responsibility for, in exchange for enhanced scalability and IOP/s.
For what its worth though, I think Spark (in light of the recent commitment from IBM) is here to stay, so I'd say it's the unequivocal leader in distributed load/clustering frameworks.
- tstonez 11y agoAlso built-in support for other data store backends e.g., PostgreSQL, MySQL, ...see https://docs.prediction.io/system/anotherdatastore/ https://docs.prediction.io/system/anotherdatastore/ since 0.9.3 release.
- ShirsenduK 11y agoStorage is not what engineers struggle with its the ML setup.