Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
winwang
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
91.
▲
by
winwang
1y ago
I find it ironic that you talk about distrust but then use ACAB, which I assume means "all cops are bad" (cursory googling).
92.
▲
by
winwang
1y ago
https://en.wikipedia.org/wiki/Commutative_diagram
93.
▲
by
winwang
1y ago
hey chris, I found your posts quite inspiring back then, with very poetic ideas. cool to see you follow up here!
94.
▲
by
winwang
1y ago
I agree about no filter*, I disagree with your reasoning. Scala (2) as a language is quite simple. The complexity came from the incredible power of its building blocks (e.g. implicits and path-dependent types). The lack of filter, as I felt
95.
▲
by
winwang
1y ago
it's quite the coincidence that ideas like friendship and love happen to be good for socialization and reproduction
96.
▲
by
winwang
1y ago
Which questions are you referring to specifically?
97.
▲
by
winwang
1y ago
Oh, my startup isn't about Postgres, but rather a GPU-accelerated Spark: https://news.ycombinator.com/item?id=43964505 What are some bad UX choices you generally dislike in data products?
98.
▲
by
winwang
1y ago
What options do you use? I don't work for Databricks but I am building my own data infra startup, so I'd like to hear what "good" looks like!
99.
▲
by
winwang
1y ago
If cost (or perf) is the issue, we're building a super-efficient, GPU-accelerated, easy-to-use Spark: https://news.ycombinator.com/item?id=43964505
100.
▲
by
winwang
1y ago
I'm building another Spark-based choice now with ParaQuery (GPU-accelerated Spark): https://news.ycombinator.com/item?id=43964505
101.
▲
by
winwang
1y ago
Not a new way like Ray, but a new way to express Spark super-efficiently (GPU-acceleration): https://news.ycombinator.com/item?id=43964505
102.
▲
by
winwang
1y ago
Do you guide the LLM to do this specifically? So it doesn't "waste" time on what can be taken care of by static analysis? Would be interesting if you could also integrate traditional analysis tools pre-LLM.
103.
▲
by
winwang
1y ago
Not the OP but -- I would immediately believe that finding bugs would be a valuable problem to solve. Your questions should probably be answered on an FAQ though, since they are pretty good.
104.
▲
by
winwang
1y ago
One possibility is crafting (somewhat-)minimal reproductions. There's some work in the FP community to do this via traditional techniques, but they seem quite limited.
105.
▲
by
winwang
1y ago
Well, I was rejected 4 times (once for a completely unrelated idea though). The reason I got in was because I actually got traction, but I'm sure my successive failed applications helped a teeny bit at least. So... pretty straightforwa
106.
▲
by
winwang
1y ago
Indeed I have, though I have some reservations about its focus on interaction nets (or rather, its marketing). I ended up making a CUDA-based, data-parallel STLC typechecker (Hindley-Milner)... I want to formally prove its correctness first
107.
▲
by
winwang
1y ago
Depends on the workload. Spark can persist dataframes to the cluster as normal. GPUs can also load certain datasets significantly faster.
108.
▲
by
winwang
1y ago
Gotcha, no problem. > while people like me who see things a little differently seem condemned to struggle in isolation Haha, you should consider founder life then. I think many founders feel like this. The issue with transpiling to GPUs
109.
▲
by
winwang
1y ago
That's a pretty interesting idea, might take a bit to prepare a useful graphic/post. Also, what do you think would be the best way to structure such a post? But, here's a small bit of something perf-y: during large shuffles,
110.
▲
by
winwang
1y ago
If you stand up your compute cluster in the same region as your bucket, there are no egress fees. Otherwise, yes, in general. There are some clouds that don't have egress fees though, i.e. Cloudflare R2.
111.
▲
by
winwang
1y ago
1. Yes! Would love to contribute back to these projects, since I am already using RAPIDS under the hood. My general goal is to bring GPU acceleration to more workloads. Though, as solo founder, I am finding it difficult to have any time for
112.
▲
by
winwang
1y ago
Thanks for the link! Not the first time I came across it, but it's a good one. If I had to bet on the longer term, I think that something like Mojo will win out -- a programming language (mostly) agnostic to the underlying vector proce
113.
▲
by
winwang
1y ago
a) Since being fully-managed is one of my value props, that's probably better for us. b) Of course I'm only accelerating the Spark/query part. Not sure what you mean. And in that case, I took a query which was 44 minutes on B
114.
▲
by
winwang
1y ago
Thanks for the links! I'm planning on contributing kernels back to open source, so will think of a way to be vendor agnostic. As far as I understand, HIP should make that doable.
115.
▲
by
winwang
1y ago
Not sure what your use cases are, but I haven't had too much issue seeing good gains vs bare Spark -- GCS has not been my bottleneck.
116.
▲
by
winwang
1y ago
All the Spark GPU acceleration right now is done via the Spark-RAPIDS plugin, so HIP would somehow have to support that. Since cuDF is the core part and hipDF is a thing, it might be doable in the near future.
117.
▲
by
winwang
1y ago
Enterprises can deploy on their own GCP, and we're planning on releasing on GCP Marketplace. For a similar cost, what if their pipeline were 5x faster, and they don't have to dealing with managing the deployment themselves? Thanks
118.
▲
by
winwang
1y ago
Indeed, Spark-RAPIDS has been around for a while! And it's quite simple to have a setup that works. Most of the issues come after the initial PoC, especially for teams not wanting to manage infra, not to mention GPU infra.
119.
▲
by
winwang
1y ago
Hopefully, we can be another managed solution for those on GCP. And as for your second point, yep, Spark tuning is definitely annoying! BigQuery is a lot more than jusr the engine, and building a simple interface for a complicated, high-per
120.
▲
by
winwang
1y ago
Good question -- it depends. For certain workloads, it might look exactly the same! For others, I found that the memory and VM constraints were creating large inefficiencies. Also, many teams simply don't want to manage that level of d
More ›