Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zackangelo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
zackangelo
1y ago
GPT-OSS will run even faster on Blackwell chips because of its hardware support for fp4. If anyone is working on training or inference in Rust, I'm currently working on adding fp8 and fp4 support to cudarc[0] and candle[1]. This is bei
32.
▲
by
zackangelo
1y ago
Is something like SeaQuery[0] what you're talking about? [0] https://github.com/SeaQL/sea-query/
33.
▲
by
zackangelo
1y ago
Draft model doesn’t degrade quality!
34.
▲
by
zackangelo
1y ago
We also wrote our inference engine in rust for mixlayer, happy to answer any questions from those trying to do the same. Looks like this uses ndarray and mpsgraph (which I did not know about!), we opted to use candle instead.
35.
▲
by
zackangelo
1y ago
Typically a combination of expert level parallelism and tensor level parallelism is used. For the big MLP tensors they would be split across GPUs in a cluster. Then for the MoE parts you would spread the experts across the GPUs and route to
36.
▲
by
zackangelo
1y ago
For anyone more curious about how this works, Fireworks wrote a blog post about it last year (I think): https://fireworks.ai/blog/cursor
37.
▲
by
zackangelo
1y ago
In your forward pass section you give a lot of emphasis to FlashAttention, but it might be worth mentioning Paged Attention as well (which was the paper written by the vLLM authors and I believe was the genesis of the project). PA-style blo
38.
▲
by
zackangelo
1y ago
Yes, this is true. A lot of times labs will hold back necessary infrastructure pieces that allow them to train huge models reliably and on a practical time scale. For example, many have custom alternatives to Nvidia’s NCCL library to do fas
39.
▲
by
zackangelo
1y ago
In the README of the linked library they have a code snippet showing how to have a conversation with the model. Also, even if it were for fine tuning, that would require an implementation of the model’s forward pass (which is all that’s nec
40.
▲
by
zackangelo
1y ago
Your reply adds more confusion, imo. The inference code and model architecture IS open source[0] and there are many other high quality open source implementations of the model (in many cases contributed by Google engineers[1]). To your poin
41.
▲
by
zackangelo
1y ago
Is that input tokens or output tokens/s?
42.
▲
by
zackangelo
1y ago
I think it’s a typo, looks pretty close to their 8xH100 prices.
43.
▲
by
zackangelo
1y ago
There might be a plateau coming but I’m not sure that will be the reason. It seems counterintuitive but there is some research suggesting that using synthetic data might actually be productive.
44.
▲
by
zackangelo
1y ago
H200s are pretty easy to get now. If you switched I'm guessing you'd get a nice bump because the nccl allreduce on the big mlps wouldn't have to cross infiniband.
45.
▲
by
zackangelo
1y ago
This definitely happens, and I'm surprised it's not talked about more often. Some attention kernels are more susceptible to this than others (I've found that paged attention is better than just naive attention, for example).
46.
▲
by
zackangelo
1y ago
I've had excellent experience with several models writing Rust. Wonder if there's just a particular issue with Tauri? I'm primarily writing code on top of the Candle ML framework.
47.
▲
by
zackangelo
1y ago
Yeah, to run the full precision model you need either two 8xH100 nodes connected via Infiniband or one 8xH200 node or one 8xB200 node. Not for the GPU poor, to be sure.
48.
▲
by
zackangelo
1y ago
Are you familiar with min_p sampling? Kind of funny that it was introduced randomly on Reddit a couple of years ago instead of in a journal or something[0]. But I believe it's widely implemented and used now. [0] https://www
49.
▲
by
zackangelo
1y ago
The 3FS chunk engine is written in Rust.
50.
▲
by
zackangelo
2y ago
What codec were you using for the audio data?
51.
▲
by
zackangelo
2y ago
My go to is Programming Massively Parallel Processors by Wen-Mei Hwu, excellent really approachable introduction. [0] [0] https://a.co/d/9fmbZqg
52.
▲
by
zackangelo
2y ago
Allaire Homesite anyone? It was a sad day for me when it got bought and integrated into Dreamweaver.
53.
▲
by
zackangelo
2y ago
Something something Chesterton s fence
54.
▲
by
zackangelo
2y ago
would love for you to test a serverless llm product i'm working on, zack [at] mixlayer.com
55.
▲
by
zackangelo
2y ago
I've been developing on top of wasm (wasmtime, specifically) for several years now. I personally have my doubts about how broadly the components specification (which this cloud platform seems to depend on) will be adopted. Maybe I'
56.
▲
by
zackangelo
2y ago
The authors specifically recommend against using a system prompt in the model card.
57.
▲
by
zackangelo
2y ago
Was it at Strangeloop by any chance? I think I remember that talk too!
58.
▲
by
zackangelo
2y ago
lol good to know! I've had the luxury of only needing the first one :)
59.
▲
by
zackangelo
2y ago
fwiw faiss, although a bit unwieldy, has an optimized full scan search built into it as well
60.
▲
by
zackangelo
2y ago
This is so true. A plain old exhaustive SIMD-optimized similarity search will do just fine in many cases and not have any of the approximation tradeoffs of HNSW.
More ›