Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aesthesia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
aesthesia
2mo ago
What makes you think that this is actually good PR for the firms involved? Every claim that this is good PR comes from someone who has increased their negative views of OpenAI based on these events. Where are the people coming away with a p
62.
▲
by
aesthesia
2mo ago
I'm not sure we should take it as a given that AI will not also become better than humans at mathematical exposition.
63.
▲
by
aesthesia
2mo ago
Ah, I see, taking "to test A" as an infinitive. But it would be strange to say test B is an alternative without saying what it's an alternative to. The other reading still seems quite unlikely.
64.
▲
by
aesthesia
2mo ago
You can make this kind of claim about living in lots of different places. Some amount of engineering and technology has been necessary to enable humans to live in most parts of the world. What determines the line?
65.
▲
by
aesthesia
2mo ago
In those experiments, the effect only happened between models from the same family (and likely the same weight initialization).
66.
▲
by
aesthesia
2mo ago
How is that ambiguous? The best interpretation I can find where "test" is a verb is an elision: Test [that] B is an alternative to test A. That is an unlikely reading: "test" is a verb in the first instance and a noun in
67.
▲
by
aesthesia
2mo ago
Luna's now cheaper than 5.4-nano (for output tokens). That's a significant improvement.
68.
▲
by
aesthesia
2mo ago
These aren't assertions being made without arguments or evidence. Maybe you've engaged with those and find them unconvincing. But from what you've said in this thread it seems more like you've dismissed them without seri
69.
▲
by
aesthesia
2mo ago
Quines produce similar issues to self-referential sentences without being directly self-referential. e.g. "'Yields falsehood when preceded by its quotation' yields falsehood when preceded by its quotation."
70.
▲
by
aesthesia
2mo ago
I mean, Cervantes was parodying medieval romances. That's what Don Quixote was reading, not modern novels. The point is, it should give you at least some pause when the people most familiar with a technology are among those making the
71.
▲
by
aesthesia
2mo ago
Fun, though as hinted at the end, the point of LLM "truth" probes is to measure the model's internal judgment of truthfulness. There's no reason this judgment, even if measured with 100% accuracy, couldn't be mistak
72.
▲
by
aesthesia
2mo ago
How many of those concerns came from the people building the technologies? John Carmack wasn't up in arms about how Doom was going to destroy society. Charles Dickens didn't think novels rotted the brain. Johannes Gutenberg wasn&#
73.
▲
by
aesthesia
2mo ago
I think the origin of the title formula is Timothy Chow's "You Could Have Invented Spectral Sequences" ( https://www.ams.org/notices/200601/fea-chow.pdf ), which is certainly not aimed at an average r
74.
▲
by
aesthesia
2mo ago
Who pays to build the open weight models? The cost of training a frontier model is orders of magnitude greater than the cost of safety tests.
75.
▲
by
aesthesia
3mo ago
Yeah, there's probably not a lot systematic behind Anthropic's version numbers. Opus 4.5 was a third of the price of Opus 4.1, indicating there was probably a change in underlying architecture. Opus 4.7 changed tokenizers, probabl
76.
▲
by
aesthesia
3mo ago
Unless you believe that OpenAI was trying to get their models to break out of the sandbox and hack into Hugging Face, and aiding them in that goal, I don't see how any of those questions would really contradict that claim. The models d
77.
▲
by
aesthesia
3mo ago
Could there, conceivably, be some third factor at play here? One that both led OpenAI to have concerns about releasing GPT-2 as well as led Microsoft to invest in OpenAI?
78.
▲
by
aesthesia
3mo ago
They could start by releasing the MAI-1 models they announced recently. https://microsoft.ai/models/
79.
▲
by
aesthesia
3mo ago
You ask good questions, and I would like to know the answers to them, but I don't think any likely answers would actually change the conclusion much. Any probable set of circumstances that produces this report from OpenAI represents a
80.
▲
by
aesthesia
3mo ago
Where do you see the claim that "long-horizon goals in real world settings are now effectively settled"? The argument you put in their mouth would be a bad one, but I don't see anyone making it.
81.
▲
by
aesthesia
3mo ago
This is a comment about user-facing responses, which are seldom the thing you're worried about when thinking about token efficiency.
82.
▲
by
aesthesia
3mo ago
It depends a lot on the model. Some models (e.g. Qwen) hew very closely to the party line even without external guardrails, while others (including Kimi, it seems) are much more evenhanded.
83.
▲
by
aesthesia
3mo ago
1. OpenAI being bad at managing risk from misaligned models is not evidence that their models are not misaligned. It's evidence that they're not taking misalignment seriously. 2. Hugging Face did report this incident to law enforc
84.
▲
by
aesthesia
3mo ago
My guess is that RL training being done with particular generation parameters makes models much more brittle to changes in these parameters, and that's why we're seeing changes like this across model providers. But I don't re
85.
▲
by
aesthesia
3mo ago
And what do those who encourage its creation get?
86.
▲
by
aesthesia
3mo ago
Alibaba wrote about a similar but less severe incident during RL training in a paper earlier this year ( https://arxiv.org/abs/2512.24873 ): > When rolling out the instances for the trajectory, we encountered an unant
87.
▲
by
aesthesia
3mo ago
Nah, the point is that if models are commoditized and there's no hope of making significant profits, no one is going to be willing to make the massive investments necessary to continue pushing the scale frontier. How large a training r
88.
▲
by
aesthesia
3mo ago
IDK, Anthropic's public posts are mostly free of Claude-isms. (Their documentation less so...)
89.
▲
by
aesthesia
3mo ago
Oh, interesting. "!" is token id 0 in the Qwen tokenizer; I wonder if there's some tokenizer shenanigans either in inference or training that end up causing this specific behavior.
90.
▲
by
aesthesia
3mo ago
Yeah, I guess that's possible. Anthropic has published some information about systems they use to monitor usage patterns, which indicates they have some related infrastructure set up. https://arxiv.org/abs/2412.136
More ›