Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
matthewnourse
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
R17 1.7.1: easily combine R, Python and r17
(rseventeen.com)
9 points
by
matthewnourse
14y ago
|
0 comments
2.
▲
R17 now open source
(rseventeen.com)
4 points
by
matthewnourse
14y ago
|
0 comments
3.
▲
Hadoop on 5 machines vs r17 on 1 machine
(rseventeen.com)
4 points
by
matthewnourse
14y ago
|
1 comments
4.
▲
R17 1.5.0: floating-point support, ear-pinning performance
(rseventeen.com)
3 points
by
matthewnourse
14y ago
|
0 comments
5.
▲
R17 1.4.3 and (some of) the evils of C's strptime & mktime
(rseventeen.com)
2 points
by
matthewnourse
15y ago
|
0 comments
6.
▲
by
matthewnourse
15y ago
All the best with the stabilizing, please add me to your beta users list whenever you're ready :).
7.
▲
by
matthewnourse
15y ago
Sounds cool, can you share a link/download ?
8.
▲
by
matthewnourse
15y ago
So James & Paul, are your needs covered by Custora or one of the other services? Or is there something more that one or both of you need?
9.
▲
by
matthewnourse
15y ago
Thanks for taking the time to look over the tests in such detail! MySQL's EXPLAIN output says that it's using the indexes. PostgreSQL's EXPLAIN output says that it's _not_ (and from what I can tell so far, this is a deliberate design decis
10.
▲
by
matthewnourse
15y ago
They work as expected for English. They "work" for languages with characters outside the English set, but they order based on ASCII byte values rather than the order expected by native speakers. Proper collation is a TODO for r17. I made
11.
▲
by
matthewnourse
15y ago
It streams the dataset into memory 256K (more for longer rows) at a time. It doesn't load the whole dataset into RAM unless it must eg for a join, sort or grouping. I don't currently have plans to push projection into the read phase, but t
12.
▲
by
matthewnourse
15y ago
Cool, thanks! Will alter and re-run in a few minutes. r17 supports only UTF8 strings but compares with strcmp/strcasecmp/memcmp depending on the situation. Would you like a different collation for PostgreSQL too?
13.
▲
by
matthewnourse
15y ago
r17 doesn't run on Windows, and I don't have any plans to make it work on Windows at this time. Apart from that niggle, I agree...more bakeoffs needed. There's not much statistical power in r17, it's more a brute force thing. If you want
14.
▲
by
matthewnourse
15y ago
I hear you, and I would have preferred to use the SQL (or any other well-understood) syntax. Some of the reasons are: 1) I want the query writer to be in complete control over what happens first, rather than a query optimizer. The r17 syn
15.
▲
by
matthewnourse
15y ago
Sure, here it is: name | current_setting -----------------------+-----------------------------------------------------------
16.
▲
by
matthewnourse
15y ago
[re fail] my main goal is to find out if r17 is "generally useful" for other people doing data mining, _not_ to pull the wool over any eyes. What would you like to see done differently? I am very happy to redo the comparison...or better y
17.
▲
by
matthewnourse
15y ago
I agree that it would be disingenuous to base "faster" on load+index+query. I'm basing it on the query times alone. I would like very much to base it on load+index+query 'cause then I could have said "faster than MySQL and PostgreSQL" :)
18.
▲
by
matthewnourse
15y ago
Yes, right now they are one and the same.
19.
▲
by
matthewnourse
15y ago
R17 is a data mining language that's a cross between SQL and Bash. For example this SELECT username, COUNT(1) AS num FROM users GROUP BY username ORDER BY num; is roughly equivalent to io.file.read('users') | rel.select(username) | rel.gro
20.
▲
R17 on spinning disk faster than PostgreSQL on SSD
(rseventeen.com)
13 points
by
matthewnourse
15y ago
|
24 comments
21.
▲
by
matthewnourse
15y ago
Syntactically it's a cross between UNIX shell script and SQL and it's got some nice performance/scalability characteristics. Am very keen to hear the collective HN wisdom about it.
22.
▲
Show HN: my new data mining language
(rseventeen.com)
3 points
by
matthewnourse
15y ago
|
1 comments
23.
▲
by
matthewnourse
15y ago
Sounds great! Even better if you can squeeze in another run. For me, consistency is the key...whatever I can keep doing consistently is the most helpful. For legs the best things I've found are sprinting in sand and bike riding up long h
24.
▲
by
matthewnourse
15y ago
I'd love to see where the crossover point is on different architectures...especially when the size of the collection is just over the CPU's L1 cache size. I wonder if the vector's 1-or-many cache miss(es) would exceed the cost of the list'