Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
winwang
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
121.
▲
by
winwang
1y ago
Nice! I attended a hackathon by Modular last weekend where we got to play with MI300X (sponsored by AMD and Crusoe). My team made a GPU-"accelerated" BM25 in Mojo, but mostly kind of failed at it, haha. The software stack for AMD
122.
▲
by
winwang
1y ago
Yep, CPU has to transfer data because no RDMA setup on GCP lol. But that's like 16-32 GB/s of transfer per GPU (assuming T4/L4 nodes), which is much more than network bandwidth. And we're not even network bound, even if
123.
▲
by
winwang
1y ago
They have >1PB of data to ETL, with some queries hitting 450TB of pure shuffle. It's very true that most users don't need something like BigQuery or Snowflake. That's why some startups have come up to save Snowflake cost b
124.
▲
by
winwang
1y ago
I'm not sure about the rest of your comment, but we would likely still want GPUs even with highly multicore CPUs. Case in point: the upper-range Threadripper series. It makes sense to have two specialized systems: a low-latency system,
125.
▲
by
winwang
1y ago
I was about to comment that Gluten is only targeting CPU vectorization, but then I found this (very cool!): https://github.com/apache/incubator-gluten/issues/9098 I'm not very familiar with Gluten, but I
126.
▲
by
winwang
1y ago
Since you mentioned non-x86, how are things on the ARM side? I believe I heard AWS's Graviton + Correto combo was a huge increase for JVM efficiency. FPGAs... I somehow highly doubt their efficiency in terms of being the "core&quo
127.
▲
by
winwang
1y ago
Very true. Can't get those numbers even if you get an entire single-tenant CPU VM. Minor note, A100 40G is 1.5TB/s (and much easier to obtain). That being said, ParaQuery mainly uses T4 and L4 GPUs with "just" ~300 GB&#x
128.
▲
by
winwang
1y ago
I'm not too familiar with HeavyDB, but here are the main differences: - We're fully compatible with Spark SQL (and Spark). Meaning little to no migration overhead. - Our focus is on distributed compute first. - That means ParaQuer
129.
▲
by
winwang
1y ago
Is the shuffle the biggest issue? Not too sure about joins but one of the datasets we're currently dealing with has a couple trillion rows. Would love to chat about this!
130.
▲
by
winwang
1y ago
Just checked out Velox. It's awesome that you're reducing duplicate eng effort! What was your talk about?
131.
▲
by
winwang
1y ago
Not sure if Spark is a classical portion for GPU compute ;) Well, outside of HPC and research. SQL on GPUs is definitely a research classic, dating back to 2004 at least: https://gamma.cs.unc.edu/DB/
132.
▲
by
winwang
1y ago
Large scale shuffles: Absolutely. One of the larger queries we ran saw a 450TB shuffle -- this may require more than just deploying the spark-rapids plugin, however (depends on the query itself and specific VMs used). Shuffling was the majo
133.
▲
by
winwang
1y ago
No relationship... yet! Hoping to have a good relationship in the future so I have a business reason to fly to Japan :D Btw, interesting thing they said here: "By utilization of GPU (Graphic Processor Unit) device which has thousands c
134.
▲
by
winwang
1y ago
Still figuring out pricing! For our first customers, we're doing pricing as either bytes scanned or by compute time, similar to BigQuery. Also experimenting with a contract that also gives the minimum of the two potential charges (up t
135.
▲
by
winwang
1y ago
Thanks! I was also told to make a performance-focused demo... didn't do it in time, but was able to go from a 44-minute BigQuery job to a 5.5-minute ParaQuery job, with a similar dataset/query as the video here. 8x faster!
136.
▲
Launch HN: ParaQuery (YC X25) – GPU Accelerated Spark/SQL
135 points
by
winwang
1y ago
|
81 comments
137.
▲
Ask HN: Which function definition keyword do you prefer, def or fn?
1 points
by
winwang
1y ago
|
5 comments
138.
▲
by
winwang
1y ago
This was just a brief moment of thought over a year ago, but I can try to summarize. I was thinking about how to unify variables in certain simple Datalog settings. If we think of a clause as a vector of variables, then simple unifications
139.
▲
by
winwang
1y ago
As someone who doesn't know very much about graphics (ironically), you're welcome and hope it helps!
140.
▲
by
winwang
1y ago
lol, I haven't thought about it like that, true. though of course, I mean compared to CPUs :P I try and use tensor cores for non-obvious things every now and then. The most promising so far seems to be for linear arithmetic in Datalog,
141.
▲
by
winwang
1y ago
I don't -- unfortunately not too well-versed in this field! But I was a bit fascinated with SWAR after I randomly thought of how to prefix-sum with int multiplication, later finding out that it is indeed an old trick as I suspected (I&
142.
▲
by
winwang
1y ago
It's 32 32-bit values which get sorted. I don't think a GPU sort would beat a CPU sort at this scale, even if you don't take kernel launch time into account. CPUs are simply too fast for (super-)small data, especially with AV
143.
▲
by
winwang
1y ago
Oh wow, TIL, thanks. I usually call stuff like that SWAR, and every now-and-then I try to think of a way to (fruitfully) use it. The "SIMD" in this case was just an allusion to warp-wide functions looking like how one might use S
144.
▲
by
winwang
1y ago
Haha, if you're the type to toss out the phrase "well known", then yes!
145.
▲
by
winwang
1y ago
I hope this can help shed the misconception that GPUs are only good at linear algebra and FP arithmetic, which I've been hearing a whole lot! Edit: learned a bunch, but the "uniform" registers and 64-bit (memory) performance
146.
▲
Faster sorting with SIMD CUDA intrinsics (2024)
(winwang.blog)
92 points
by
winwang
1y ago
|
11 comments
147.
▲
by
winwang
1y ago
I'm working on GPU-accelerated SQL+Spark in a zero-hassle package: https://paraquery.com Been prod for a few months, recently ripping through 900TB with ~5x efficiency (customer was on BigQuery). If anyone has any data/
148.
▲
by
winwang
1y ago
In case it's slow to load for others: https://web.archive.org/web/20250426185806/https://www.mail-... I don't quite fully agree with these, but I agree with the general spirit.
149.
▲
by
winwang
1y ago
I agree, it sucks to be lumped in. What's your proposed solution?
150.
▲
by
winwang
1y ago
Fun fact which I'm 50%(?) sure of: a single branch divergence for integer instructions on current nvidia GPUs won't hurt perf, because there are only 16 int32 lanes anyway.
More ›