Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lambda-research
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
From bigger models to better intelligence:what NeurIPS25 tells us about progress
(lambda.ai)
2 points
by
lambda-research
10mo ago
|
0 comments
2.
▲
by
lambda-research
2y ago
> What this article misses though is that despite this, each GPU in the distributed cluster still needs to have enough VRAM to load the entire copy of the model to complete the training process. That's not exactly accurate. In the d
3.
▲
Show HN: Text-to-Video Arena
(t2vleaderboard.lambdalabs.com)
4 points
by
lambda-research
2y ago
|
1 comments
4.
▲
by
lambda-research
2y ago
Unlike text generation using LLMs, text-to-video generation brings unique challenges — balancing realism, prompt alignment, and artistic vision is something much more nuanced and intuitive than generated code. But how do we measure the qual
5.
▲
by
lambda-research
2y ago
Something that I always think about when I see discussions about hallucinations or "confidently wrong answers" is that humans have this issue too. For those on tiktok, how many times have you found yourself easily believing some r
6.
▲
by
lambda-research
2y ago
Awesome thank you!
7.
▲
Show HN: Open-Source Python REPL with AI Tutor for Learning and Problem-Solving
(companionai.dev)
5 points
by
lambda-research
2y ago
|
2 comments
8.
▲
by
lambda-research
2y ago
I think the benefit is that SpinQuant had higher throughput and required less memory. At least according to the tables at the bottom of the article. Definitely nice to see them not cherrypick results - makes them more believable that its no
9.
▲
by
lambda-research
2y ago
Hey there are some details about this scattered throughout. The answer really depends on the technique. For DDP you can fairly easily get same throughput as single gpu throughput (we were getting ~80% gpu util for multiple nodes iirc), as l
10.
▲
by
lambda-research
2y ago
The benchmark is matrix multiplcation with the shapes `(6, 1500, 256) X (6, 256, 1500)`, which just aren't that big in the AI world. I think the gap would be larger with much larger matrices. E.g. Llama 3.1 8B which is one of the small
11.
▲
by
lambda-research
2y ago
Let me know if there are any questions or suggestions! Feel free to open issue on github, and contributions are welcome also
12.
▲
Show HN: How to guide on training Llama-405B using PyTorch distributed APIs
(github.com)
3 points
by
lambda-research
2y ago
|
4 comments
13.
▲
by
lambda-research
2y ago
The idea that time is tied to computation makes me wonder if everything we see as 'progress' is just the universe showing us the loading screen percentage of the game of life.