Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
kwillets
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
kwillets
2y ago
Strike the sharding idea as I misunderstood the CDC method.
62.
▲
by
kwillets
2y ago
This paper compares the benefits of lightweight compression and other techniques: https://blog.acolyer.org/2018/09/26/the-design-and-implement...
63.
▲
by
kwillets
2y ago
Does this fragment columns into rowgroups like Parquet, or is it more of a pure columnstore? IME a data warehouse works much better if each column isn't split into thousands of fragments.
64.
▲
by
kwillets
2y ago
One additional thought regarding query performance is that content-defined row groups allow localized joins and aggregations which are much faster than the globally-shuffled kind. If the sharding key matches (or is a subset of) a join or gr
65.
▲
by
kwillets
2y ago
One more: do you prefer the CDC technique over using the rowgroups as chunks (ie using knowledge of the file structure)? Is it worth it to build a parquet-specific diff?
66.
▲
by
kwillets
2y ago
How does this compare to rsync/rdiff?
67.
▲
by
kwillets
2y ago
I'm surprised that Parquet didn't maintain the Arrow practice of using mmap-able relative offsets for everything. Although these could be called relative to the beginning of the file.
68.
▲
by
kwillets
2y ago
OMG I'm that guy -- 3 straight Vertica roles with $B annual revenue. I did learn a lot from watching MDS people try to beat it (in the end I'm also looking for what should come next), but mostly it was confirming the article and R
69.
▲
by
kwillets
2y ago
Microsoft has a seed finder specifically aimed at avoiding a priori bias in experiment groups, but IMO the main effect is pushing whales (which are possibly bots) into different groups until the bias evens out. I find it hard to imagine obt
70.
▲
by
kwillets
2y ago
Once I had tried coca tea in Peru, it was clear what the taste of Coca-Cola comes from. What became funny was the series of "tasters" who claimed they had figured out the formula with other ingredients like cinnamon, etc.
71.
▲
by
kwillets
2y ago
This seems like what I've noticed on MPP systems (a little before cloud): data replicas give a lot more availability than the number of partition events would suggest. I likely need to read the paper linked, but it's common to hav
72.
▲
by
kwillets
2y ago
We came close, as we had architects who knew little besides Snowflake. While there are all kinds of issues it can't handle for CRUD, such as locking and small updates, it's also expensive to do small reads due to the reliance on a
73.
▲
by
kwillets
2y ago
Papers We Love SF has been reconstituting lately; unfortunately I don't know if July has an event yet, but try https://www.meetup.com/papers-we-love-too/ . We've been meeting online, and I'm interested i
74.
▲
by
kwillets
2y ago
SF Papers We Love might be interested in a physical space -- we've been meeting online, but IMO we would benefit by face-to-face.
75.
▲
by
kwillets
2y ago
I'm guessing from spreadsheets riddled with per-cell SQL fetches.
76.
▲
by
kwillets
2y ago
The latency before/after histograms unfortunately use different scales, but it appears that eg the under-200ms bucket is only a few percentage points smaller after the change, maybe 38 before and 33 after. What I'm curious about i
77.
▲
by
kwillets
2y ago
Wouldn't a non-DWH SQL database work as well? Most RAG datasets seem like they would fit in a small DB.
78.
▲
by
kwillets
2y ago
Ironically this model most resembles Teradata, which used to sell their own proprietary hardware/software combination at exorbitant rates. Snowflake compute instance types cost about $.30-.40 an hour on EC2, so it's quite a markup
79.
▲
by
kwillets
2y ago
I've done it both ways. Look into Data Sketches also if you want to see applications. The pros: -- Samples are small and fast most of the time. -- can be used opportunistically, eg in queries against the full dataset. -- can run more c
80.
▲
by
kwillets
2y ago
Back in the 80's and 90's NASA built a National Aerodynamic Simulator, which was a big Cray or similar that could crunch FEA simulations (probably a low-range graphics card nowadays). IIRC they found that the queue for that was as
81.
▲
by
kwillets
2y ago
I believe this problem diminishes as the key count goes up.
82.
▲
by
kwillets
2y ago
I was thinking about an AI to feed you the proper Snowflake sales pitch each time a query runs expensive or fails a benchmark. At my previous org it could replace several headcount.
83.
▲
by
kwillets
2y ago
Yes, I almost added that -- there are several alternatives, but it takes some testing and tuning to be certain.
84.
▲
by
kwillets
2y ago
I'm guessing the next stage will be to keep the raw data and aggregate dynamically. We went through a similar progression with ad audiences -- maintaining pre-aggregated data sketches and summaries was such a PITA that we likely saved
85.
▲
by
kwillets
3y ago
Multiplying digit bitfields by their respective powers of 10 and shift/adding via MUL is a (well?) known technique, see Lemire https://lemire.me/blog/2023/11/28/parsing-8-bit-integers-qui... .
86.
▲
by
kwillets
3y ago
The Nitro chipset claims 100 GB/s encryption, so that doesn't seem to be the reason.
87.
▲
by
kwillets
3y ago
AWS docs and blogs describe the Nitro SSD architecture, which is locally attached with custom firmware. > The Nitro Cards are physically connected to the system main board and its processors via PCIe, but are otherwise logically isolated
88.
▲
by
kwillets
3y ago
EBS is also slower than local NVMe mounts on i3's. Also, both features use Nitro SSD cards, according to AWS docs. The Nitro architecture is all locally attached -- instance storage to the instance, EBS to the EBS server.
89.
▲
by
kwillets
3y ago
You're on the right track with pricing. What I find ubiquitous about MDS people is a lack of old-school ideas like benchmarking and price-performance. Cloud salespeople talk endlessly about your company becoming data-driven until you b
90.
▲
by
kwillets
3y ago
To make up for having a better schema in Terraform than in the database.
More ›