Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stratos123
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
stratos123
6d ago
You seem to be implying that achievements of internal models are exaggerated, but that's rather implausible. The public does have access to, for example, Opus and Fable, and so we know what those models are capable of - finding real vu
2.
▲
by
stratos123
11d ago
There's an obvious coordination problem here: if you decide to stop your research for safety and your competitors don't, you have just burned your company without actually improving the world's outcomes. There's also an
3.
▲
by
stratos123
11d ago
The writing is slop. It doesn't mention a detail I have seen people actually not get: that LLMs wrap tool calls in special tokens, so they are "out of band" and can't be mistaken with normal output. It also spends an ent
4.
▲
by
stratos123
11d ago
That's pretty interesting. It was already well-known that you could easily remove safety training from open-weights models by a bit of finetuning, but apparently you don't even need a finetuning dataset, as long as you have just a
5.
▲
by
stratos123
11d ago
> The agents are deciding what is acceptable as part of the task, which is a security task and may well be testing or evaluating that type of behaviour as far as they know. I don't think this at all describes what was going on in th
6.
▲
by
stratos123
11d ago
> Self improvement is not an instrumental goal. A self improving paperclip maker has making paperclips as the instrumental goal. Self-improvement is an instrumental goal, it makes you better at accomplishing whatever your actual goal i
7.
▲
by
stratos123
12d ago
> Any type of goal is singular. To change to be better at something is a goal. That's false. There's is such a thing as an instrumental goal. For example, a human who doesn't particularly enjoy eating or drinking will stil
8.
▲
by
stratos123
12d ago
It is indeed easier to imagine, which doesn't mean it's a more desirable outcome.
9.
▲
by
stratos123
12d ago
Well, as a reference, for the first of the OpenAI swarm incidents, the huggingface breach one, METR didn't find any cases where the agents didn't realise that what they're doing were out of scope. Instead, they expressed hesi
10.
▲
by
stratos123
12d ago
You mention Yudkowsky at the end. Did you not read his writings on the orthogonality thesis ("there can exist arbitrarily intelligent agents pursuing any kind of goal"), or do you disagree with them? Your entire comment seems to j
11.
▲
by
stratos123
13d ago
It's selection pressure. Both OpenAI (back in 2015) and Anthropic (much later, in 2021) were founded by people who could foresee the concept of AI x-risks and wanted to work on preventing it. For OpenAI's founding, the idea was th
12.
▲
by
stratos123
13d ago
Would it change your opinion if later this experiment is repeated with exposed CoT, and it turns out the model did notice that this was against the instructions yet did it anyway?
13.
▲
by
stratos123
13d ago
AFAIK for OpenAI it's the Model Spec: https://model-spec.openai.com/2026-08-18.html and for Anthropic it's the Constitution, which they actually include in training to the point Claude can recite segments of it b
14.
▲
by
stratos123
14d ago
> I mean it is a very loose analogy, but pretty dumb things in the world can cause big problems, like real mice and rats and mosquitoes. I think it all depends on the rules and the 'alignment', but a not-so-smart model could st
15.
▲
by
stratos123
14d ago
> By the same logic, suppose that Dario/Altman/Jensen accumulate all capital because they are the only one that have access to AGI and end up controlling democracy, turning the world into technofeudalism or whatever you want to
16.
▲
by
stratos123
14d ago
This only works for some technologies - those where everyone having access to it doesn't cause a tragedy of the commons. I love open source too and yet that doesn't make me like the idea of being murdered by a misaligned model. No
17.
▲
by
stratos123
14d ago
So you're saying that if Dario truly thought what he says he does, he'd resign/dissolve the company somehow/unilaterally slow down? I see two problems with this: First, Dario thinks that OpenAI will continue racing even
18.
▲
by
stratos123
14d ago
> Human minds can acquire, hold and use axioms as building blocks for actual reasoning, LLMs use statistical likelihood. This is not even slightly true. Even when trying, humans commit logical mistakes at notable rates, because the behav
19.
▲
by
stratos123
14d ago
Sure they can, after the fact.
20.
▲
by
stratos123
14d ago
> if this was something which could be done amateurs would have done it to provide open models. (Like SETI@home or Folding@home) I know that's not what you meant but this does exist, by the way. It's called AI Horde: https:&#x
21.
▲
by
stratos123
14d ago
Maybe, but I don't think tiny models can be harnessed as intelligence. As in, if you have one rogue Mythos overseeing the swarm, it only produces 1 Mythos's worth of useful thoughts no matter how many gemma4:e4bs it consists of. A
22.
▲
by
stratos123
14d ago
> So it's either they truly think AI is going to kill us all, or there's some other motives at play here. I don't think these people could possibly agree on the color of the sky,[...] And yet, they historically did agree o
23.
▲
by
stratos123
14d ago
I mean, sure, the immediate cause of the NPT was the ability of everyone involved to foresee the possibility of a global nuclear war and judge it both worryingly likely and catastrophic. But I argue that the Hiroshima and Nagasaki bombing w
24.
▲
by
stratos123
14d ago
Okay, but suppose that Hypothetical Opensource Anthropic trains a model that turns out to be very dangerous, and releases it. Suppose that the public investigates and, not being limited by an information bubble, correctly notices that it&#x
25.
▲
by
stratos123
14d ago
Yeah, I'm also pretty skeptical about this. With AI companies we see time and time again that they can have benevolent, well thought-out regulations and then a few years just... abandon them - the most notable case of this being, of co
26.
▲
by
stratos123
14d ago
The latter has some extra failure modes but I don't think these two kinds of RSI are that different. Either way you can get exponential growth in capabilities, and either way a not-entirely-aligned model can train a more capable and
27.
▲
by
stratos123
14d ago
> disingenuous because if he believed what he says then he'd stop. Why do you think so?
28.
▲
by
stratos123
14d ago
> Suppose it was, I don't know, Australia producing SOTA open weights, because they believed it was the right thing to do for the benefits of humanity - would Dario propose making Australia a geopolitical adversary too? I'd exp
29.
▲
by
stratos123
14d ago
> First: if they are unilaterally doing it: what about their competitors? Is this a chance for OpenAI to pass them? The only part of this plan Dario is unilaterally committing to is the "embedded evaluators" thing, which doesn&
30.
▲
by
stratos123
14d ago
That's not undecidable in principle - compute governance is a thing. The more likely sticking point is that the two sides might be soured on the deal once they realize how much oversight they'd have to give to the other side.
More ›