Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zackangelo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
by
zackangelo
2y ago
Where did you see that? I thought they used an 8b model for their reward model? > To guide our search strategies, we used RLHFlow/Llama3.1-8B-PRM-Deepseek-Data, an 8B reward model that has been trained using process supervision
62.
▲
LLM Reasoning 101 – Cot and Hidden Tokens
(mixlayer.com)
2 points
by
zackangelo
2y ago
|
0 comments
63.
▲
by
zackangelo
2y ago
Perhaps I don't have such a keen eye for things like font rendering, but the nano texture display has been a pure step forward for me. No negatives at all.
64.
▲
by
zackangelo
2y ago
A few thoughts: * Will this be the start of enshittification of the base ChatGPT offering? * There may also be some complementary products announced this month that make the $200 worth it * Is this the start of a bigger industry trend of pr
65.
▲
by
zackangelo
2y ago
Interestingly, the actual prop was auctioned off in 2022 for more than $18K! [0] [0] https://propstoreauction.com/lot-details/index/catalog/319/l...
66.
▲
by
zackangelo
2y ago
I’ve been writing all of our transformer implementations in Rust using the Candle crate and it’s been great. While dealing with CUDA and GPUs on servers is never a joy, deploying fully contained Rust binaries instead of a morass of python s
67.
▲
by
zackangelo
2y ago
Congrats on the launch Dex!
68.
▲
by
zackangelo
2y ago
So grateful that HSW existed when I was younger. As a teenager, I couldn't afford to get the timing belt and water pump replaced on my car so I had to figure out how to do it myself. I bought the service manual from AutoZone but I need
69.
▲
by
zackangelo
2y ago
Had to reach for a new IPC mechanism recently to implement a multi-GPU LLM inference server. My original implementation just pinned one GPU to its own thread then used message passing between them in the same process but Nvidia's NCCL
70.
▲
by
zackangelo
2y ago
Related recent discussion on twitter: https://x.com/Teknium1/status/1858987850739728635 Looks like other folks get 80 tok/s with max batch size, that's surprising to me but vLLM is definitely more optimi
71.
▲
by
zackangelo
2y ago
> Does the function oscillate over the region of integration? If it does, then make sure that the step size is chosen to be smaller than the wave length of the function. Nyquist limit, but for numerical integration?
72.
▲
by
zackangelo
2y ago
It’s been a minute so my memory might be off but I think when I ran 70b at fp16 it just barely fit on a 2x A100 80GB cluster but quickly OOMed as the context/kv cache grew. So if I had to guess a 96GB H100 could probably run it at fp8
73.
▲
by
zackangelo
2y ago
Ah, makes a lot more sense now.
74.
▲
by
zackangelo
2y ago
Would you care to share your prompts? They posted a haystack benchmark in the blog post that seems too good to be true.
75.
▲
by
zackangelo
2y ago
This is astonishingly fast. I’m struggling to get over 100 tok/s on my own Llama 3.1 70b implementation on an 8x H100 cluster. I’m curious how they’re doing it. Obviously the standard bag of tricks (eg, speculative decoding, flash atte
76.
▲
by
zackangelo
2y ago
I'm surprised this is the case! I've been working on a rope implementation for my own project (needed to account for padding in unique situations) and even an off by one error usually causes the model to produce non-sensical outpu
77.
▲
by
zackangelo
2y ago
I’m working on an inference platform that allows for tokens to be appended to the context after some tokens have been generated. If there’s other sequences in the batch, it means they’ll have to be padded. Currently this means I can’t use F
78.
▲
by
zackangelo
2y ago
Do you explicitly list that a PFAS test was omitted for a municipality? I looked up Austin and didn’t see it listed.
79.
▲
by
zackangelo
2y ago
I'd love it if you checked out what we've been working on. It's still in early stages, but might be usable for something you're trying to build. Here's an example (this buffers the entire JSON object, but you can al
80.
▲
by
zackangelo
2y ago
With mixlayer, because the round trip time to the model is so short, you can alternate between appending known tokens of the JSON output and values you want the model to generate. I think this works better than constraining the sampling in
81.
▲
Show HN: Mixlayer – code and deploy LLM prompts using JavaScript
(mixlayer.com)
5 points
by
zackangelo
2y ago
|
0 comments
82.
▲
ZLUDA's Third Life (Drop-In CUDA on Non-Nvidia GPU)
(vosen.github.io)
5 points
by
zackangelo
2y ago
|
0 comments
83.
▲
by
zackangelo
2y ago
I've only read the abstract but they don't mention quantizing the weights or otherwise trying to shrink the model in any way. They're claiming to be able to efficiently run larger models without loading the entire thing into
84.
▲
by
zackangelo
2y ago
This limitation is common to most implementations of the actor model. In fact, I think a lot of people would consider it a feature, not a limitation because it allows you to reason about your concurrent behavior in a more straightforward wa
85.
▲
by
zackangelo
2y ago
We used candle[0], which uses cudarc and the metal crate under the hood. That means we run on nvidia hardware in production and can test locally on macbooks with smaller models. I would certainly like to use non nvidia hardware but at this
86.
▲
by
zackangelo
2y ago
I'm building this type of functionality on top of Llama models if you're interested: https://docs.mixlayer.com/examples/json-output
87.
▲
by
zackangelo
2y ago
This is incorrect: > With text-only inputs, the Llama 3.2 Vision Models can do tool-calling exactly like their Llama 3.1 Text Model counterparts. You can use either the system or user prompts to provide the function definitions. > Cur
88.
▲
by
zackangelo
2y ago
Just trying to stay focused on launching first ( https://docs.mixlayer.com ) and keeping early customers happy, but would love to open source some of this work. It'd probably be a separate crate from candle. If you haven'
89.
▲
by
zackangelo
2y ago
Yeah, I’ve had to rewrite continuous batching and other scheduling logic. That and multi-GPU inference have been the hardest things to build. I’ll need to get paged attention working as well, but I think I can launch without it.
90.
▲
by
zackangelo
2y ago
Their inference server is written in Rust using huggingface’s Candle crate. One of the Moshi authors is also the primary author of Candle. We’ve also been building our inference stack on top of Candle, I’m really happy with it.
More ›