Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kyllo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
kyllo
6y ago
Indexes don't solve the same problem as columnar because with an index you still have to dereference a pointer, read the page (performing I/O if on-disk), and locate each tuple you want within the page--the tuples you need to aggr
62.
▲
by
kyllo
6y ago
Here's a YT video of one of the implementers giving a talk to CMU students about what DuckDB is, why they made it, and how it works: https://www.youtube.com/watch?v=PFUZlNQIndo In a nutshell, it's a file-per-datab
63.
▲
by
kyllo
6y ago
I was floored by this comment yesterday from one of their Developer Relations people: > Did any of you actually read the article? We are passing the Jepsen test suite and it was back in 2017 already. So, no, MongoDB is not losing anythi
64.
▲
by
kyllo
6y ago
I was wondering the same--I guess because it's fairly new (2018) and it came out of a database research group at a European university, rather than a SV tech firm. Therefore, limited marketing budget. By the way, here's a YT video
65.
▲
by
kyllo
6y ago
Basically--it's a struct containing a byte array, weight, scale, and sign, and all the arithmetic operations are implemented in software. So it's really slow, and each RDBMS has to implement this type from scratch, or find a vendo
66.
▲
by
kyllo
6y ago
If you like SQLite for data analysis, you might want to check out DuckDB https://github.com/cwida/duckdb which is billed as "SQLite for analytics." SQLite is a row store, which is best for OLTP (point queries
67.
▲
by
kyllo
6y ago
Most people don't realize that Excel (like PowerBI) has an in-memory, compressed, column store database inside of it. Loading hundreds of millions of rows into it takes a while, but given a commensurate amount of RAM and a reasonable d
68.
▲
by
kyllo
6y ago
Is there any form of distributed query engine in use today that doesn't fit the definition of the MapReduce pattern? Is describing a distributed query engine as MapReduce still a meaningful distinction from some other, non-MapReduce
69.
▲
by
kyllo
6y ago
By contrast, China is the US' third largest trading partner (after Canada and Mexico) at 10% of US trade. Taiwan is much more intertwined with China economically. Which makes sense given the geographic proximity and linguistic and cult
70.
▲
by
kyllo
6y ago
Politically yes, Taiwan is de facto independent, but economically they are deeply intertwined; although they have their own currency, stock exchange etc., a lot of major companies operating in China are under Taiwanese ownership (Foxconn an
71.
▲
by
kyllo
6y ago
Good point, but it's worth noting that "go get a job on autopilot" behavior is entirely rational as it's driven by economic necessity--most people go get jobs because that's the only way they can afford food, shelte
72.
▲
by
kyllo
6y ago
A few people will do that, yes, but there are a couple million software developers in the US, and most of them are not going to suddenly move to sparsely populated interior states. They'll spread out from the SF bay area a bit.
73.
▲
by
kyllo
6y ago
This realistically isn't a major concern because the pool of qualified software engineers in a state like Idaho is far too small to make a noticeable dent in market salary levels.
74.
▲
by
kyllo
6y ago
That's basically what Salesforce did, but they ended up skimming a lot higher % than that
75.
▲
by
kyllo
6y ago
Yes, Filipino spaghetti is definitely intended to be sweet--it uses banana ketchup in the sauce. Bananas were used as a substitute for tomatoes due to the Philippines' relative abundance of the former. https://en.wikipedia.o
76.
▲
by
kyllo
6y ago
Yes and programming languages are human interfaces to machine instructions, so context sensitivity can be manageable and even desirable to human users of the language, even if it makes the implementation of the language interpreter more com
77.
▲
by
kyllo
6y ago
It's still context-free, the reason is because by the time you hit the '=' symbol you already know whether you're in a <Statement> or a <Compare Exp> production rule, based on the preceding symbols (namely th
78.
▲
by
kyllo
6y ago
You're right, I meant that each valid RHS matches one and only one LHS--but that's also true of CSGs.
79.
▲
by
kyllo
6y ago
I have learned both and agree with this statement. I think that Rust is harder to learn if you've only worked with high-level, GC languages, and don't have a background doing lower-level programming in C/C++/Obj-C, as we
80.
▲
by
kyllo
6y ago
Context-free has nothing to do with symbol tables, it just means that in the grammar, the left-hand side of a production rule can only have a single non-terminal symbol, which can always be replaced by the expression on the right-hand side,
81.
▲
by
kyllo
6y ago
This is the reason why ORMs and SQL generators typically quote all column names in statements.
82.
▲
by
kyllo
6y ago
SQL is incredibly verbose compared to dplyr. Modern non-SQL query languages coming out tend to be more more similar to dplyr, based on method chaining or piping data table objects through function calls, like UNIX pipes. It's much more
83.
▲
by
kyllo
6y ago
don't forget the crucial group_by() / summarise()
84.
▲
by
kyllo
6y ago
R's come a long way in the last decade. The tibble and data.table packages both address this issue. data.table ( https://github.com/Rdatatable/data.table ) is the more strongly-typed of the two, by default it fails
85.
▲
by
kyllo
6y ago
SQL connections can be arbitrarily long-lived. ETL processes and other data-intensive jobs often take multiple hours.
86.
▲
by
kyllo
6y ago
TFA isn't even about SQL the language at all though, it's about the scalability and reliability characteristics of databases, especially in a distributed environment.
87.
▲
by
kyllo
6y ago
Yeah, I need to point at the stove and say "stove off" when I turn it off. A few times I have started driving away from my house and suddenly been hit by a sense of dread that I might have left the stove on after cooking breakfast
88.
▲
by
kyllo
7y ago
As a data scientist, most of my projects have been individual--I'm generally the only person writing and reading my code. No one tells me which language I have to use. Python and R are the most popular, and I use either one depending o
89.
▲
by
kyllo
7y ago
RPython is not intended for humans to write programs in, it's for implementing interpreters. If you're after a faster Python, you should use PyPy not RPython. Numba gives you JIT compilation annotations for parallel vector operati
90.
▲
by
kyllo
7y ago
It's an argument that Python being slow / single-threaded isn't the biggest problem with Python in data engineering. The biggest problem is the need to process data that doesn't fit in RAM on any single machine. So you n
More ›