Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Imnimo
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
121.
▲
by
Imnimo
1y ago
In my experience 4o was already really good at this task. I'd be curious to see an in-depth 4o vs. o3 benchmark.
122.
▲
by
Imnimo
1y ago
The short answer is that they are applying the same defense to audio as to images, and so we should expect that the same attacks will work as well. More specifically, there are a few moving parts here - the GenAI model they're trying
123.
▲
by
Imnimo
1y ago
Any new "defense" that claims to use adversarial perturbations to undermine GenAI training should have to explain why this paper does not apply to their technique: https://arxiv.org/pdf/2406.12027 The answer
124.
▲
by
Imnimo
1y ago
This sounds very cool at a conceptual level, but the article left me in the dark about what they're actually doing with DolphinGemma. The closest to an answer is: >By identifying recurring sound patterns, clusters and reliable seque
125.
▲
by
Imnimo
1y ago
I'm so used to seeing the "fish crawling onto the shore" cartoon of evolution that I assumed the branching always went that way - land creatures are branchoffs of sea creatures. But surely this is oversimplified - are there e
126.
▲
by
Imnimo
2y ago
Here are some example questions that Turing proposed when initially describing the test: >"I have K at my K1, and no other pieces. You have only K at K6 and R at R1. It is your move. What do you play?" >"In the first li
127.
▲
by
Imnimo
2y ago
My favorite part about the original paper is that it was written during a time when "extra-sensory perception" was a big fad, and Turing bought into the idea. He admits that the most likely failure of his test is that humans could
128.
▲
by
Imnimo
2y ago
It gives me a little pause that humans are so much worse than random chance at detecting GPT-4.5. Suppose we reframed the test as: "You interact with 10 witnesses, 5 of which are humans, 5 of which are GPT-4.5. Your task is to separate
129.
▲
by
Imnimo
2y ago
>But while some readers might not subscribe to outlets that give away some of their best journalism for free, it’s just as possible that readers will recognize this sacrifice and reward these outlets with more traffic and subscriptions i
130.
▲
by
Imnimo
2y ago
>The study looked at the medical notes of 21 children aged between two and seven in the UK and Ireland who fell ill after consuming the drinks between 2009 and 2024. Is this just a sample of 21 among a larger set of children, or is that
131.
▲
by
Imnimo
2y ago
Another category of "Lab Play" task I'd be interested in seeing is balancer design. Even small balancers can be quite complicated ( https://factorioprints.com/view/-NopheiSZZ7d8VitIQv9 ), and it would be i
132.
▲
by
Imnimo
2y ago
>Test-time compute/RL on LLMs: >It will not meaningfully generalize beyond domains with easy verification. To me, this is the biggest question mark. If you could get good generalized "thinking" from just training on mat
133.
▲
by
Imnimo
2y ago
Here's an example of language switching: https://gr.inc/question/although-a-few-years-ago-the-fundame... In the dropdown set to DeepSeek-R1, switch to the LIMO model (which apparently has a high frequency of langu
134.
▲
by
Imnimo
2y ago
>To speed up our experiments, we omitted the Kullback–Leibler (KL) divergence penalty, although our training recipe supports it for interested readers. I am very curious whether omitting the KL penalty helps on narrow domains like this,
135.
▲
by
Imnimo
2y ago
I had a similar confusion previously, so maybe I can help. I used to think that a mixture of experts model meant that you had like 8 separate parallel models, and you would decide at inference time which one to route to. This is not the cas
136.
▲
by
Imnimo
2y ago
I wonder if having a big mixture of experts isn't all that valuable for the type of tasks in math and coding benchmarks. Like my intuition is that you need all the extra experts because models store fuzzy knowledge in their feed-forwar
137.
▲
by
Imnimo
2y ago
Yeah, I'm not so much interested in "can you think of the right card name from among thousands?". I just want to see that it can produce a thinking procedure that makes sense. If it ends up not being able to recall the right
138.
▲
by
Imnimo
2y ago
Yeah, that's a great point. While this is evidence that the sort of behavior LeCun predicted is currently displayed by some reasoning models, it would be going too far to say that it's evidence it will always be displayed. In fa
139.
▲
by
Imnimo
2y ago
Just that it's another model where you can read the raw "thinking" tokens, and they sometimes fall into this sort of rut (as opposed to OpenAI's models, for which summarized thinking may be hiding some of this behavior).
140.
▲
by
Imnimo
2y ago
I have become a little more skeptical of LLM "reasoning" after DeepSeek (and now Grok) let us see the raw outputs. Obviously we can't deny the benchmark numbers - it does get the answer right more often given thinking time, a
141.
▲
by
Imnimo
2y ago
A company wants to hire someone to perform tasks X, Y and Z. It's difficult to cleanly evaluate someone's ability to do these tasks in a short amount of time, so they do their best to construct a task A which is easy to test, and
142.
▲
by
Imnimo
2y ago
There is a classic Dave Chapelle bit where his friend gets pulled over for driving recklessly. After being confronted he says to the cop, "I'm sorry officer, I didn't know I couldn't do that." The officer, to Dave&#
143.
▲
by
Imnimo
2y ago
Am I right in understanding that this "declaration" is not a commitment to do anything specific? I don't really understand why it matters who does or does not sign it.
144.
▲
by
Imnimo
2y ago
Is this happening a lot? I definitely don't have any months where I make zero web searches. I'm not even sure I have individual days where I make zero web searches. Are a lot of Kagi customers going on month-long trips into the ra
145.
▲
by
Imnimo
2y ago
My surface-level reading of these two sections is that the 800k samples come from R1-Zero (i.e. "the above RL training") and V3: >We curate reasoning prompts and generate reasoning trajectories by performing rejection sampling
146.
▲
by
Imnimo
2y ago
DeepSeek claims that the cold-start data is from DeepSeekV3, which is the model that has the $5.5M pricetag. If that data were actually the output of o1 (a model that had a much higher training cost, and its own RL post-training), that woul
147.
▲
by
Imnimo
2y ago
I think it would cast doubt on the narrative "you could have trained o1 with much less compute, and r1 is proof of that", if it turned out that in order to train r1 in the first place, you had to have access to bunch of outputs fr
148.
▲
by
Imnimo
2y ago
I think there's two different things going on here: "DeepSeek trained on our outputs and that's not fair because those outputs are ours, and you shouldn't take other peoples' data!" This is obviously extremely
149.
▲
by
Imnimo
2y ago
I feel like this tweet and Altman's tweet ( https://x.com/sama/status/1884066337103962416 ) have a vibe of "our investors are mad and told us we had to say something". Are people getting cold feet ove
150.
▲
by
Imnimo
2y ago
My guess is that OpenAI didn't cheat as blatantly as just training on the test set. If they had, surely they could have gotten themselves an even higher mark than 25%. But I do buy the comment that they soft-cheated by using elements o
More ›