Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
schopra909
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
schopra909
15d ago
One of the reasons we’re focused on training efficiencies is because those are upstream of end user costs. We achieved these specific gains while cutting context windows down 4x. That will have major implications for inference speed and exp
2.
▲
by
schopra909
15d ago
Thanks for the references! I’ll have to take a deeper look, haven’t read these before.
3.
▲
by
schopra909
15d ago
Just skimmed the paper (haven't seen this before). Might need to read more carefully, but at first glance I don't understand their intuition why the "Markovian property limits the model’s ability to fully utilize the generati
4.
▲
by
schopra909
15d ago
Yes, 3.6x cheaper when it comes to training! When it comes to inference it should also be cheaper (since we’ve cut attention sequence length by 4x), but I don’t have a hard number for you how much cheaper for inference
5.
▲
by
schopra909
15d ago
Author here, feel free to drop questions below. Will try to answer to best of my ability!
6.
▲
Training Text-to-Image Models 3.6× Faster
(linum.ai)
55 points
by
schopra909
15d ago
|
10 comments
7.
▲
by
schopra909
23d ago
The idea of getting a speed up in language models from using diffusion is compelling, but it just doesn’t seem like discrete diffusion models work as good at discrete non-diffusion. Which kind of makes sense, tokens don’t have implicit cont
8.
▲
by
schopra909
27d ago
Totally hear you on point a/b! I think ultimately the folks with the purses won’t care enough about a for it to be taken seriously, even if it’s an engineering bottleneck. B definitely has scope but still smaller than I’d expect. When
9.
▲
by
schopra909
27d ago
IMO this will be a blip. There’s a lot of talk in the wake of all the Uber handwringing about token spend. Legacy enterprises want to look innovative to Wall Street without spooking them, so it’s easy to hop on the narrative and “show” that
10.
▲
by
schopra909
29d ago
Totally, RLVR as a concept predates DeepSeek; but they proposed a version that was simple and scalable. Popularizing a specific version of a technique is exactly what I mean by iterations on a theme. It’s only 5% different from what others
11.
▲
by
schopra909
29d ago
Progress is iterative. Everyone is always riffing on other’s ideas and can execute on them given enough support (eg $$). The person to get to an idea first is just 5% away, so it’s possible to catch up. Moreover,I think it’s impossible to k
12.
▲
by
schopra909
1mo ago
That might work! Off the dome, it’s not clear to me whether spatial/depth priors are better/worse than an LLM for this type of task. Only reason I can think why the LLM might still work better here is that it’s trained to solve a
13.
▲
by
schopra909
1mo ago
Aah, for this we're just trying to filter not generate. When it comes to conditioning, you'll still need a model that understands text since the primary control is text. In the original Stable Diffusion, CLIP doubled as part of th
14.
▲
by
schopra909
1mo ago
What would you have in mind for a modern model? Like Dino-V3 or something of that ilk? For the LAION classifier specifically, it's trained on-top of CLIP. The bottleneck for accuracy isn't the linear/non-linear readout, it&#x
15.
▲
by
schopra909
1mo ago
Hi HN, one of the authors here. Lmk if you have any questions, and I'll try my best to answer them!
16.
▲
Getting video models to learn better, faster
(linum.ai)
36 points
by
schopra909
1mo ago
|
11 comments
17.
▲
by
schopra909
1mo ago
Can someone explain the intuition behind the en-gram idea? I know DeepSeek published a paper about it a few months ago and the Gemma models have a lightweight version of it; but it hasn’t clicked for me yet
18.
▲
by
schopra909
1mo ago
Yep on iPhone I just use two dashes —- and it looks like an emdash
19.
▲
by
schopra909
1mo ago
I feel this in my bones, as someone who has been using em dashes in their texts and writing before ChatGPT existed. In the past year, I’ve had 3 or 4 times when someone has “called me out” for using AI when I’m just an em dash organically.T
20.
▲
by
schopra909
1mo ago
Yep checked 3.7
21.
▲
by
schopra909
1mo ago
I think this is more “gray” than this. I feel like I can rip through ideas more quickly then ever before and as a result get a lot better at designing systems and (for my work) get a lot better at designing data/model experiments. But
22.
▲
by
schopra909
2mo ago
From our experiments it’s the best video captioning model in the world by a mile. This was not the case a year ago. When reasoning got introduced a year ago to GPT 5, on average the model performed worse than GPT4-o for short video clip cap
23.
▲
by
schopra909
2mo ago
I’m not entirely sure if local development will lead to Nvidia’s supremacy being challenged. I think a simple reason why it’s been hard to unseat in Nvidia is first mover advantage. A lot more water has flown through Nvidia pipes than TPUs
24.
▲
by
schopra909
2mo ago
100% agreed.
25.
▲
by
schopra909
6mo ago
Honestly never considered the forking use case; but it makes a ton of sense when explained Congrats on the launch. This is cool tech
26.
▲
by
schopra909
7mo ago
Really cool to see innovation in terms of quality of tiny models. Great work!
27.
▲
by
schopra909
7mo ago
Very cool work! We spend a lot of time thinking about "robust representations" in the video space. Are there any alternative ideas to JEPA right now, when it comes to speech encoding that couples meaning and sound? Curious to lear
28.
▲
by
schopra909
7mo ago
Honest question, why were folks posting AI generated comments in the first place? There's such a high inertia to comment. I only comment when I have something to contribute OR find something incredibly interesting. So I'm just baf
29.
▲
We Built an $8/Month GPU-Cluster Monitor
(linum.ai)
3 points
by
schopra909
7mo ago
|
0 comments
30.
▲
by
schopra909
7mo ago
It’s a great question. In terms of pre-training even if they were was enough data at that quality, storing it and either demuxing it into raw frames OR compressing it with a sufficiently powerful encoder likely would cost a lot of $. But th
More ›