4 ms·
R17 is a data mining language that's a cross between SQL and Bash. For example this SELECT username, COUNT(1) AS num FROM users GROUP BY username ORDER BY num
by matthewnourse 15y ago
R17 is a data mining language that's a cross between SQL and Bash. For example this
SELECT username, COUNT(1) AS num FROM users GROUP BY username ORDER BY num;
is roughly equivalent to
io.file.read('users') | rel.select(username) | rel.group(count) | rel.order_by(_count);
The most interesting difference is that each r17 clause executes concurrently :).
Download link is here: http://www.rseventeen.com/#download http://www.rseventeen.com/#download
- pyre 15y agoIs R17 a language and an implementation then?
- matthewnourse 15y agoYes, right now they are one and the same.
- skimbrel 15y agoAt first glance it looks like doing a query involves streaming the entire dataset into memory while selecting and projecting on the fly. If that's true, what happens when you have truly massive rows (i.e., things containing MEDIUMTEXTs or worse)? Okay, reading further down you only get very basic data types. Still, nothing in the spec appears to prohibit very long rows, and I'd imagine performance starts to fall off once you're throwing around tens of kilobytes per row. Any plans to support pushing the projection operation into the read phase so you can work with massive individual records? And where's the source? I want to see exactly how much this differs from a modern SQL engine.
- matthewnourse 15y agoIt streams the dataset into memory 256K (more for longer rows) at a time. It doesn't load the whole dataset into RAM unless it must eg for a join, sort or grouping. I don't currently have plans to push projection into the read phase, but the phases are all pretty close together :) so maybe it wouldn't be required. How massive is "massive" for you? 10s of K? Megs? R17 is not currently open source, but I haven't ruled it out.