Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
roborovskis
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
roborovskis
7mo ago
What would you define as 'distillation' versus 'learning'? How do you know that what a LLM is doing is 'distillation' vs a process closer to a human reading a book? From my perspective, pretraining is pretty cl
2.
▲
Tinker: General Availability and Vision Input
(thinkingmachines.ai)
5 points
by
roborovskis
10mo ago
|
0 comments
3.
▲
by
roborovskis
2y ago
Where are you seeing this? On https://github.com/deepseek-ai/DeepSeek-R1/tree/main?tab=rea... I only see the paper and related figures.
4.
▲
Reviewing Post-Training Techniques: DeepSeek, Qwen-2, and Phi-4
(brianfitzgerald.xyz)
2 points
by
roborovskis
2y ago
|
0 comments
5.
▲
by
roborovskis
2y ago
https://stable-baselines3.readthedocs.io/en/master/ is a great resource for hacking on implementations for RL - many good RL courses out there but https://www.youtube.com/playlist?list=PLwRJQ4m4UJj
6.
▲
Tricks for DPO Tuning a Code LLM, Part 1: Logit Curriculum Learning
(brianfitzgerald.xyz)
2 points
by
roborovskis
2y ago
|
0 comments
7.
▲
ChatGPT for macOS can now work with apps on your desktop
(twitter.com)
1 points
by
roborovskis
2y ago
|
0 comments
8.
▲
Implementation of a Diffusion Transformer and Rectified Flow Sampling in Jax
(github.com)
1 points
by
roborovskis
2y ago
|
0 comments
9.
▲
Stable Diffusion 3 API Now Available
(stability.ai)
246 points
by
roborovskis
2y ago
|
61 comments
10.
▲
by
roborovskis
3y ago
You could definitely use this for upsampling negative prompts, though I haven't tested that much. In theory, future T2I models shouldn't need to be negatively prompted as much; I find it's better to focus on really high quali
11.
▲
by
roborovskis
3y ago
Yup, the model will still forget details sometimes. This is a common issue with prompt upsampling methods, but I'm hoping to improve this with the next version.
12.
▲
by
roborovskis
3y ago
Thanks for the kind words! I started with the 780M param flan-t5-large model, and kept trying smaller and smaller base models - I was shocked at how good the output was at 77M. As you go smaller, though, it's much easier to accidentall
13.
▲
by
roborovskis
3y ago
As Invoke is open-source and already has transformers as a dependency, it should be pretty easy to add.
14.
▲
by
roborovskis
3y ago
I haven't tested extensively with non SDXL based checkpoints but there's nothing really SDXL specific about the model; if you're using a fine-tune that's trained on booru-style tags, it will probably not work as well - b
15.
▲
by
roborovskis
3y ago
will fix these, thanks for the heads up!
16.
▲
SuperPrompt: Better Text to Image Prompts in 77M Parameters
(brianfitzgerald.xyz)
150 points
by
roborovskis
3y ago
|
31 comments
17.
▲
StableLM Zephyr 3B
(stability.ai)
123 points
by
roborovskis
3y ago
|
38 comments
18.
▲
Stable Video Diffusion
(stability.ai)
1330 points
by
roborovskis
3y ago
|
302 comments
19.
▲
by
roborovskis
3y ago
https://dreamstudio.ai/
20.
▲
Comma Three Devkit
(comma.ai)
95 points
by
roborovskis
5y ago
|
60 comments
21.
▲
Elon Musk Called a Tesla Critic’s Boss to Complain About Him
(jalopnik.com)
3 points
by
roborovskis
8y ago
|
1 comments
22.
▲
by
roborovskis
13y ago
The fact that they can see what plugins I have makes uTorrent wanting a browser plugin make total sense. Just thinking.