3 ms·
Hi, I'm one of the lead developers on the project. It's a relatively young project, so keep an eye out for lots of substantial releases over the next few months
by jaybaxter 13y ago
Hi, I'm one of the lead developers on the project. It's a relatively young project, so keep an eye out for lots of substantial releases over the next few months.
So far, we have been focusing on smaller dataset sizes, such as 10,000 rows by 100 columns, but everything scales linearly (both query processing and offline inference) so you could use it for larger datasets too. The current version is all in memory, though, so you're limited there.
Right now you must import data from csv, so no images, and you must do all preprocessing of your data before loading it in. We hope to add more and more of this kind of functionality in later releases. I'd love to hear suggestions! I recommend trying out the VM installation if you want to quickly play around with it.
- jzwinck 13y agoHere's a suggestion for importing data: support HDF5. It's like CSV, but a lot faster and better. And the set of people who use HDF5 probably overlaps a fair bit with your target audience.
- tlarkworthy 13y agoThanks for your reply. So its doing something akin to clustering on categorical variables. Might be able to connect to some kind of ontology learning...