Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chillee
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
29 ms
·
31.
▲
by
chillee
2y ago
I recommend you read the linked post above: https://gwern.net/story-of-your-life , which argues that the story is not about precognition at all. > and on my first read, I thought [Story of My Life] was downright mediocre—
32.
▲
by
chillee
2y ago
I actually don't think that it's very true to the original - arguably, it loses the entire point of the original. In particular, I agree with Gwern's take ( https://gwern.net/story-of-your-life ) that the origi
33.
▲
by
chillee
2y ago
I'm skeptical of these benchmarks for a number of reasons. 1. They're only comparing against VLLM, which isn't SOTA for latency-focused inference. For example, their vllm benchmark on 2 GPUs sees 102 tokens/s for BS=1, g
34.
▲
by
chillee
2y ago
We actually do track such provenance now ( https://pytorch.org/blog/understanding-gpu-memory-1/ ) - works pretty well I think :)
35.
▲
by
chillee
2y ago
Thanks, I consider that very high praise :)
36.
▲
Matrix Multiplications on GPUs Run Faster When Given Predictable Data
(thonking.ai)
4 points
by
chillee
2y ago
|
0 comments
37.
▲
What shapes do matrix multiplications like?
(thonking.ai)
4 points
by
chillee
3y ago
|
0 comments
38.
▲
by
chillee
3y ago
> maybe this is useful and maybe it isn't, but counting lines of code on top of a huge codebase is not very meaningful. In this case it's pretty reasonable imo, since the kernel itself is fairly independent - the usage of torch
39.
▲
by
chillee
3y ago
I work on PyTorch Compilers at Meta, and I think folks enter ML Systems from all directions :) Some folks start with more familiarity in ML research and dip down as far as they need. Other folks come from a traditional distributed systems&#
40.
▲
Supporting Mixtral in GPT-fast through torch.compile
(thonking.substack.com)
1 points
by
chillee
3y ago
|
0 comments
41.
▲
by
chillee
3y ago
Who is this :think: But no, FlashAttention already solved the memory requirements of attention. RingAttention is primarily useful for parallelizing across the sequence component.
42.
▲
by
chillee
3y ago
https://news.ycombinator.com/item?id=39475528 My source is me :) I work at PyTorch on ML compilers. If you don't believe me, perhaps you'll believe Karpathy's diagram (and the general discussion in the thread
43.
▲
by
chillee
3y ago
> For a sequence with N = 100,000 tokens, it would mean cost dropping by a factor of 100,000× I'm not sure I understand the intended interpretation of this. Concretely speaking, if it cost CoolAI 100k seconds of compute to process a
44.
▲
by
chillee
3y ago
This is extremely wrong. The attention component that's quadratic is a relatively small portion of compute.
45.
▲
by
chillee
3y ago
What’s the issue with getting int8 dynamic quantization to work? As in, you’re unable to get it to quantize or to run with speedups?
46.
▲
by
chillee
3y ago
Just a difference in speed. This repo is primarily showing how you can get really good inference perf with just native pytorch.
47.
▲
by
chillee
3y ago
I wouldn't recommend using it for a batch serving setting today. One crucial optimization for batched serving (which you need if you have a large number of requests) is continual batching, which this implementation doesn't have.
48.
▲
by
chillee
3y ago
This person claims to have compared llama.cpp against gpt-fast on a 4090, and found gpt-fast about 20% faster. https://twitter.com/zeuxcg/status/1730450360895242355
49.
▲
by
chillee
3y ago
Assuming you're talking about the one here: https://pytorch.org/blog/accelerating-generative-ai-2/#start... it's just the pytorch profiler + chrome profiler (chrome://tracing)
50.
▲
by
chillee
3y ago
Unfortunately it's a little bit tricky today. The main issue is that we rely heavily on torch.compile + Triton for performance in this repo, and there isn't an Apple Silicon backend either for torch.compile or Triton. For example,
51.
▲
by
chillee
3y ago
Thanks! I've also written a couple other things along a similar vein you might like at https://horace.io/writing.html (particularly https://horace.io/brrr_intro.html ) and also some of the things I'
52.
▲
by
chillee
3y ago
Yeah it's certainly possible, but it's not the focus of this implementation, which is more latency focused (so BS=1).
53.
▲
by
chillee
3y ago
This is indeed a bit of a dark art. Essentially, you want a balance between "is significantly faster than base model" and "generates similar stuff to the base model". Anecdotally, folks often seem to use say, 70B base +
54.
▲
by
chillee
3y ago
I cover it a bit in the blog post, but unless you have a really long context length (like 32k+), your primary computational cost doesn't come from attention but rather from loading your weights from VRAM into registers. I mean, pract
55.
▲
by
chillee
3y ago
We used an A100-80GB GPU. We didn't compare explicitly to Huggingface TGI but I think you should be able to compare the tokens/s achieved. One note is that this release is optimized for latency , while I think HF TGI might be mor
56.
▲
by
chillee
3y ago
Surprisingly, no. And part of this is that text generation is really expensive. Unlike traditional ML inference (like with, resnets), you don't just pass your data through your model once. You need to pass it over and over again (onc
57.
▲
by
chillee
3y ago
Yeah, for sure. I think for deployment purposes, many times these model conversions are necessary (such as if you don't want to use Python). However, I do think these model conversions are often a significant pain for users. So, in som
58.
▲
by
chillee
3y ago
Hey, author of the blog post here. It's mentioned in the blog post, but one of the intentions of this repo is that it's more of a "tutorial" than it is a library/framework. My hope is that people will copy-paste and
59.
▲
by
chillee
3y ago
I think 40/year is the ARPU globally - the EU number is 71.52. But it’s a bit worse than that, since users willing to pay for adblocking are disproportionately rich users, making them more valuable.
60.
▲
by
chillee
3y ago
It's not about the "unusual tech economics", it's about the profit. Making more money than you spend is not a tech-specific concept.
More ›