Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fpgaminer
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
by
fpgaminer
2y ago
I've been consistently surprised by Gemini's OCR capabilities. And yeah, Qwen is climbing the vision ladder _fast_. In my workflows I often have multiple models competing side-by-side, so I get to compare the same task executed on
32.
▲
by
fpgaminer
2y ago
Awww, I was looking forward to seeing some of the leak ;) Oh well. Nice find and breakdown! Somewhat relatedly, it occurred to me recently just how important issues like prompt injection, etc are for LLMs. I've always brushed them o
33.
▲
by
fpgaminer
2y ago
No, it probably exists in the raw LLM and gets both significantly strengthened and has its range extended. Such that it dominates the model's behavior, making it several orders of magnitude more reliable in common usage. Kinda of lik
34.
▲
by
fpgaminer
2y ago
> As the parent says, modern LLMs are finetuned with a different loss function after pretraining. This means that in some strict sense they're no longer autoregressive models – but they do still generate text one word at a time. I t
35.
▲
by
fpgaminer
2y ago
True. See one of Anthropic's researcher's comment for a great example of that. It's likely that "planning" inherently exists in the raw LLM and RL is just bringing it to the forefront. I just think it's helpf
36.
▲
by
fpgaminer
2y ago
Supervised finetuning is only a seed for RL, nothing more. Models that receive supervised finetuning before RL perform better than those that don't, but it is not strictly speaking necessary. Crucially, SFT does not improve the model
37.
▲
by
fpgaminer
2y ago
LORA can be used in RL; it's indifferent to the training scheme. LORA is just a way of lowering the number of trainable parameters.
38.
▲
by
fpgaminer
2y ago
99% of evolution was spent on single cell organisms. Intelligence only took 0.1% of evolution's training compute.
39.
▲
by
fpgaminer
2y ago
All user facing LLMs go through Reinforcement Learning. Contrary to popular belief, RL's _primary_ purpose isn't to "align" them to make them "safe." It's to make them actually usable. LLMs that haven&#x
40.
▲
by
fpgaminer
2y ago
Right, but it leads to too many false conclusions by lay people. User facing LLMs are only trained on next token prediction during initial stages of their training. They have to go through Reinforcement Learning before they become useful
41.
▲
by
fpgaminer
2y ago
> This is powerful evidence that even though models are trained to output one word at a time I find this oversimplification of LLMs to be frequently poisonous to discussions surrounding them. No user facing LLM today is trained on next
42.
▲
by
fpgaminer
2y ago
There are a few different approaches. Meta documents at least one approach quite well in one of their llama papers. The general gist is that you have some kind of adapter layers/model that can take an image and encode it into tokens.
43.
▲
by
fpgaminer
2y ago
It is paramount to a functioning society to have zero tolerance for nazis.
44.
▲
by
fpgaminer
2y ago
The paper basically sums to suggesting (and analyzing) these otpions: * Comparing all possible pair permutations eliminates any bias since all pairs are compared both ways, but is exceedingly computationally expensive. * Using a sorting alg
45.
▲
by
fpgaminer
2y ago
Seems on par with the existing scaling curve. If I had to speculate, this model would have been an internal-only model, but they're releasing it for PR. An optimized version with 99% of the performance for 1/10th the cost will c
46.
▲
by
fpgaminer
2y ago
Because Tesla falls in the "startup" category in most people's minds (whether it _should_ fall in that category is irrelevant). And this is a forum run by startup investors, and frequented by startup investors. QED.
47.
▲
by
fpgaminer
2y ago
I noticed a similar phenomenon in my work on JoyCaption when I began teaching it VQA. JoyCaption was trained on about 800k image-caption pairs, and built from so400m and Llama 3.1 8B Instruct. There's no VQA data in its training. As
48.
▲
by
fpgaminer
2y ago
A lot of problems jump out to me with this article, particularly with the explanation of multi-modal LLMs. I'll say that I _do_ agree with the thrust of the article. Don't trust LLMs. But they probably should have argued legiti
49.
▲
by
fpgaminer
2y ago
DeepSeek is great because: 1) you can run the model locally, 2) the research was openly shared, and 3) the reasoning tokens are open. It is not, in my experience, state of the art. In all of my side by side comparisons thus far in real wo
50.
▲
by
fpgaminer
2y ago
It doesn't look like, it is. e.g. violent coup and Nazi salutes during inauguration. Are people expecting the rise of authoritarianism to look different or something? This is the exact same fascist playbook that's been run every
51.
▲
by
fpgaminer
2y ago
I think "reasoning" models will solve the joke issue (amongst other issues), but not because they're "reasoning". Rather because they help solve the exploration issue and the scaling issue. Having worked with LLMs
52.
▲
by
fpgaminer
2y ago
It does seem like individual prompting styles greatly effects the performance of these models. Which makes sense of course, but the disparity is a lot larger than I would have expected. As an example, I'd say I see far more people in
53.
▲
by
fpgaminer
2y ago
Google's experimental thinking model is similarly casual. Not as casual as QwQ, but more casual than Gemini 1.5 Pro. Flash 2.0 will also go a bit more casual in its responses randomly, and when you tell it to think step by step.
54.
▲
by
fpgaminer
2y ago
I work with LLMs and MLLMs all day (as part of my work on JoyCaption, an open source VLM). Specifically, I spend a lot of time interacting with multiple models at the same time, so I get the chance to very frequently compare models head-to
55.
▲
by
fpgaminer
2y ago
As a long time Rust user, I more or less agree that it's a pain. And yeah, when working on a big project refactoring code, it's difficult to know ahead of time if the borrowing pattern you _think_ will work will actually work. Of
56.
▲
by
fpgaminer
2y ago
Didn't work when I tried it, while other jailbreaks still work.
57.
▲
by
fpgaminer
2y ago
I gave it a try, but when I tried to start a finetune of Llama 3.1 8B it just gave an error every time. I also encountered several server errors just navigating to different pages.
58.
▲
by
fpgaminer
2y ago
> The truth is like 10 or 20 thousand lines of game logic can make a lot of games, and that's really not much code to port to your own game engine compared to the rest of the engine. I intuitively want to agree, but on the other han
59.
▲
by
fpgaminer
2y ago
HuggingFace really is such an amazing resource to the ML community. Not just for storing datasets, but being able to stand up a demo of my models using spaces for anyone to use? It's hard to overstate how useful that is.
60.
▲
by
fpgaminer
2y ago
I found the paper a tad difficult to understand because it spends a lot of time circling around the main thesis instead of directly describing. So, to the best of my understanding: We want to improve LLM's abilities to give correct an
More ›