7 ms·
Citus Data (YC S11) Wants To Make Scalable Data Analytics Accessible To Anyone
- johnpmayer 14y agoThis sounds a lot like AsterData Database, which I know has been around for at least a few years. I'm interested to know if you are able to write queries that define explicit parallelism like in the SQL-MapReduce language extension, and also the ability of the query preprocessor for the distributed workload.
- ozgune 14y agoHey, we currently don't have an SQL/MapReduce language extension, but we do have the Map & Reduce execution primitives implemented under the covers (for parallel query processing). For the distributed query processor, we can efficiently parallelize SQL queries that involve look-ups, complex selections, groupings and orderings, analytic functions, and joins between one large and multiple small tables. We also have a lot more coming; are there any queries that you are particularly interested in?
- pella 14y agoany information about PostGis compatibility?
- ozgune 14y agoIn all honesty, I'd have to check. The worker nodes in our architecture will be able to use PostGIS indexes just like regular PostgreSQL instances, and the master node should correctly handle (partition) most PostGIS functions. Still, we'd probably need to implement parallelization for the && operator, which shouldn't be that hard. That said, I don't want to misguide you here before going over PostGIS' documentation more thoroughly. If you could ping us at engage@citusdata.com, we'll send you a reply once we know for sure.
- pella 14y ago"Features Not in v1.0" http://www.citusdata.com/documentation#missing-features http://www.citusdata.com/documentation#missing-features
- bsg75 14y ago"Real-time insert, update, or deletes issued against the master node." Is this a bulk/batch load only system then?
- spathak 14y agoThat is correct. This is primarily a bulk-load system. There are setups (as mentioned in the smilliken's comment above), where it can be used for real-time inserts, but requires more hands-on setup and configuration.
- benbjohnson 14y agoDistributed SQL queries are cool and accessible to people but I feel like projects that apply relational languages to event data don't make much sense. If I have click stream data then I'm more interested in knowing what users are doing after they performed action "A", "B" & "C" than rolling up how many users performed a single action "A". SQL falls flat on its face for this type of analysis. Also, the name also threw me off. I thought it was "Citrus" and not "Citus". [Full disclosure: I am writing an open source, distributed, behavioral database - https://github.com/skylandlabs/sky https://github.com/skylandlabs/sky]
- ozgune 14y agoThere's value in both types of analyses. For knowing what users do after they perform action "A", "B" & "C", many people currently rely on implementing Map/Reduce programs. That can be a bit heavyweight if you want to simply compare people who did action "A" or "B", filter based on complex criteria, or apply simple analytic functions. Also, apart from standard relational algebra operators, SQL provides a lot of convenience functions for math operations, string manipulations, date and time formatting, pattern matching, and so forth. These may come in handy to users who want to quickly gather insights out of their data.
- benbjohnson 14y agoYou're right, there is value from SQL over event data, however, I feel like it's a missed opportunity to simply apply the same paradigms to a different type of data. I'm not suggesting that SQL be thrown out but a new language needs to be available specifically for event data.
- bgilroy26 14y agoI'm working outside the bounds of my understanding here, but why can't that result be formed from a query that pulls from the user_history table and a subquery that pulls from the user_history table with different conditions on each one?
- ozgune 14y ago
- edouard1234567 14y agoCongrats Umur and team. I look forward to trying it for ZeTrip analytics.
- pinarsezer 14y agoLove the video btw!
- smilliken 14y agoWe've been using Citus at MixRank for storing our timeseries data, and it's worked out magnificently well for our use-case. A few points: (i) We can do ad-hoc realtime analytics on hundreds of millions of data points. (ii) We can also do realtime analytics on billions of datapoints as long as we pre-compute along one dimension. (iii) We could do a lot better at (i) and (ii) if we invested more heavily in hardware (and Citus would make this pretty painless, actually). (iv) I'd normally not consider a closed-source solution personally, but since Citus is based so heavily on PostgreSQL (protocol-level compatibility, configuration, codebase), this has been a non-issue for us. We can still lean on the amazing PostgreSQL community, documentation, and for the parts we don't have the source code to, the Citus team has been very helpful in explaining how things work. (v) Fault tolerance is immaculate. At the node level, PostgreSQL is notoriously one of the most reliable and robust databases available. At the cluster level, Citus will magically fall back to a replica mid-query when a server dies. (vi) Although realtime inserts are not supported out of the box, the system is flexible enough that we were able to get this working on our own without help from Citus. (vii) Schema migrations are also not supported out of the box, but we built a schema migration framework that takes care of this for us. (viii) We're not worried about vendor lock-in, since the data is just stored on our servers, in the PostgreSQL serialization format. If we wanted to, we could just give up the features that Citus gives us and build our own data-access layer on top of our cluster. Anyway, it won't be everything to everyone, but it works very well for our OLAP use-case of timeseries ad impression data. I'd definitely recommend looking into it if you're otherwise considering Hadoop, Vertica, Aster, Greenplum, or a sharded MySQL/PostgreSQL setup. Full disclosure: I am extremely biased since I've gotten to know the team very well after using Citus. I'm definitely one of their biggest fans, if for no other reason the amount of time they've saved us at MixRank.
- seboavalin 14y agoScalable data analytics accessible to anyone? Great!
- emre 14y agoCongrats to the Citus team! It looks like a great product, will definitely use it!
- kolistivra 14y agoLooks very promising, I will definitely give it a shot! Good job!
- kt9 14y agoCongratulations on the launch! I've been following this company for a while and its great to see the public launch!