Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
srush
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
srush
2y ago
That's a good point. We don't see that in our experiments because it's all in the math domain. However for OAI it's plausible that training for o1 might conflict with standard instruction training, leading to less human
32.
▲
by
srush
2y ago
Thanks for the feedback, and not minor. Sorry about that.
33.
▲
by
srush
2y ago
Full blog is here: https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling... Happy to answer any questions about these methods.
34.
▲
by
srush
2y ago
Nope! Too hard for me. But it would be a great practice for someone who wants to get started in this space. There is a Triton implementation that might be a good starting place.
35.
▲
by
srush
2y ago
I would recommend first learning Numpy or a similar vectorized library. If you have a good sense of those data structures (array broadcasting) it is a good starting point for what you can do in a GPU world.
36.
▲
by
srush
2y ago
Thanks so much!
37.
▲
by
srush
2y ago
Yeah it looks bad in the readme. In the actual code it's cleaner. Font rendering is hard
38.
▲
by
srush
2y ago
Awesome! Here are all of them if anyone else is looking. https://github.com/srush/Triton-puzzles https://github.com/srush/tensor-puzzles https://github.com/srush/autodiff-puzz
39.
▲
by
srush
2y ago
These still hold up, and I think they're a great first step. But they no longer get you to the goal line. Think about it more as conceptual practice, before you enter the jungle.
40.
▲
by
srush
2y ago
I'll take a look. Yeah Pyro is the best thing to do here. But it would be nice to revisit some of these implementationz
41.
▲
by
srush
2y ago
Here is a port without the visualizer: https://twitter.com/srush_nlp/status/1719376959572980094 Here is an amazing in-browser implementation in WebGPU https://www.answer.ai/posts/2024-09-12-gp
42.
▲
by
srush
2y ago
I made these a couple of years ago as a teaching exercise for https://minitorch.github.io/ . At the time the resources for doing anything on GPUs were pretty sparse and the NVidia docs were quite challenging. These days ther
43.
▲
by
srush
2y ago
tweet says the opposite?
44.
▲
by
srush
2y ago
PyTorch is a generationally important project. I've never seen a tool that is so inline with how researchers learn and internalize a subject. Teaching Machine Learning before and after its adoption has been a completely different exper
45.
▲
by
srush
2y ago
These slides from Lucas Beyer are pretty nice. https://docs.google.com/presentation/d/1ZXFIhYczos679r70Yu8v...
46.
▲
by
srush
2y ago
Oh no, yes, please send a stack trace (although if it is in colab I should be able to repro)
47.
▲
by
srush
2y ago
Yup. I often find people learning ML Engineering struggle a lot with shapes and broadcasting. The goal of these puzzles is to force you to really learn the semantics of broadcasting and internalize that data shapes in ML correspond to how m
48.
▲
by
srush
2y ago
Hey, I made these. They're pretty fun. Sometimes people tell me they use them for ML interviews, but they're kind of hard. The motivation was primarily teaching point-free, array programming. I don't think it is a great style
49.
▲
by
srush
2y ago
This book is great. Really mind warping at first read. Fernando Pereira has had an incredible influence across NLP for his whole career. Here is an offhand list of papers to check out. * Conditional random fields: Probabilistic models for
50.
▲
by
srush
3y ago
Hi! Blog author. This was an attempt a couple years ago to understand and write about this paper in a detailed way. Here is a video going through this topic as well: https://youtu.be/dKJEpOtVgXc?si=PDNO0B0qi6ARHaeb Section
51.
▲
by
srush
3y ago
Want to give proper credit to my former student for starting this: Yuntian Deng et al., 2016 ( https://arxiv.org/abs/1609.04938 ). I believe this repo uses the dataset from that paper. Some recent cool work he's bee
52.
▲
by
srush
3y ago
Yup, should work nicely together.
53.
▲
by
srush
3y ago
The model targets the decoder part of the system which is the speed bottleneck. So for tasks like classification it is not likely to be helpful. However a similar method could be used for that use case. (Coauthor)
54.
▲
by
srush
3y ago
When distilling models for speed, you get a better win from removing decoder parameters, since they are run in serial, than encoder parameters. For example see this work https://arxiv.org/abs/2006.10369 - paper co-auth
55.
▲
by
srush
3y ago
This is an awesome notebook. Just want to note that the difference in this paper is that it works without direct access to the embedding models (encoder). So it can't design the embedding space.
56.
▲
by
srush
3y ago
The claim of the paper is that you can store it losslessly! If you assume you have access to an LLM for free, then text is extremely compressible. Storing it in an embedding would be plenty of bits.
57.
▲
by
srush
3y ago
Roughly. It's about learning a decoder for a blackbox decoder. It's much harder in an environment that is not end-to-end trainable.
58.
▲
by
srush
3y ago
Agreed. Embeddings are pretty big > 1024 * 4 bits. And language is really small: ~1 bits per character. So it's not at all crazy that embeddings can be lossless. The paper shows a practical method to recover the text and shows how i
59.
▲
by
srush
3y ago
I got excited about this a couple years ago and ported it to Pyodide in to run in the browser with a pytorch like library: https://srush.github.io/g9py/ It's a little choppier than javascript, but crazy to me that
60.
▲
by
srush
3y ago
Yeah, it's CPU only, and it is using about 38g for 70B and 7g 7B. Guessing that is mostly from the large caches that it keeps around for efficiency. If you wanted to pay some computational cost, you could likely get that down by quanti
More ›