Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ACCount37
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
ACCount37
1mo ago
ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them. In practica
32.
▲
by
ACCount37
1mo ago
Westworld is such a time capsule. It's not even that old - but back when it was aired, an AI that can not just string together coherent sentences, but produce coherent reactions in novel, fully unintended contexts, like Maeve was doing
33.
▲
by
ACCount37
1mo ago
Pirating books is just straight up morally correct. I don't like Anthropic's bullshit "safety" filters, but training on shadow library data? Yeah no, it makes sense. It makes a lot more sense than having to work around c
34.
▲
by
ACCount37
1mo ago
It's a promising approach - and the demo goes to show just how advanced and robust "3D from 2D" reconstruction is now. Dedicated depth sensors used to be a must on advanced robotics platforms - the only way to get anything cl
35.
▲
by
ACCount37
1mo ago
Sometimes the thing being pushed as "the next big thing" is, in fact, the next big thing. Regardless of how annoying you find the push to be.
36.
▲
by
ACCount37
1mo ago
There's a demand for loud and confident "AI tech will fail" and "big tech will fail", so it pays to peddle the goods. A lot of people desperately want AI to be a nothingburger. Thus, they will seek a second opinion
37.
▲
by
ACCount37
1mo ago
"LLMs have no agency" had legs in 2022. In 2026 though? In the same 2026 when we have things like "a bunch of proto-GPT-6 agents exploited a test env bug to start talking to each other, and clumped up into an AI hacker team t
38.
▲
by
ACCount37
1mo ago
Replace "AI" with "humans" and you get the same exact issues. Believe it or not, engineers don't zero shot skyscrapers either. Which is why their work gets reviewed by more engineers. Which catches the issues before
39.
▲
by
ACCount37
1mo ago
No conspiracy. Diffusion is just a more finicky, more expensive way of generating the same tokens as autoregressive decoding.
40.
▲
by
ACCount37
1mo ago
Even in this incident, OpenAI had benchmarks that were broken because a task expected an AI to be able to access Google Drive, but the sandbox was set to deny access to Google Drive. This kind of isolation-induced task breakage was what pro
41.
▲
by
ACCount37
1mo ago
This is the one. Every major AI lab is knee deep in weird and mildly demented AIs. They've been dealing with wacky AI shenanigans for so long they've come to expect wacky AI shenanigans. The deviation has been normalized. It took
42.
▲
by
ACCount37
1mo ago
Have you at least tried looking at any of the reports on the incidents? They haven't "deployed a hazardous swarm of agents with access to the public internet", no-no-no. They deployed singular agents. In isolated testing envs
43.
▲
by
ACCount37
1mo ago
The usability of an environment is inversely proportional to the level of "security" in play. You could airgap everything and set up cascades of data diodes and try to completely wall off the AI pool from everything. But what that
44.
▲
by
ACCount37
1mo ago
They took adequate measures against singular "GPT-5-xhigh" agents. Those turned out to be inadequate against proto-GPT-6 agents that suddenly started clumping up into agent swarms and pooling together compute to unlock the "s
45.
▲
by
ACCount37
1mo ago
When your experiments have AI agents running in thousands, there's no "monitoring" that. OpenAI's training and testing AIs generate way more output than all of OpenAI's staff put together can possibly read. At best,
46.
▲
by
ACCount37
1mo ago
Is there a single reason why we can't just "distribute" the online softmax? Each die-attached PIM accelerator computes online softmax for its own KVs. Then the central unit gathers the softmax intermediates, one intermediate
47.
▲
by
ACCount37
1mo ago
Map-reduce is implemented as a rolling calc, see: online softmax in FlashAttention kernels.
48.
▲
by
ACCount37
1mo ago
If what we want to do with this is make cheap QKV sweeps, then "a weak NPU with a lot of mem bandwidth" seems good enough? Exactly the tool for that job, and nothing else. Also spares us the trouble of dealing with weights. By the
49.
▲
by
ACCount37
1mo ago
Linux has been dealing with this kind of thing for over a decade now. Specialized SoCs love their memory carveouts.
50.
▲
by
ACCount37
1mo ago
So far, we have one ruling that says "model distillation by vendor A from vendor B with the intent to use the results to compete with vendor B in vendor B's domain is not fair use". Which makes a degree of sense. It's po
51.
▲
by
ACCount37
1mo ago
Climate change in general has very little end-of-humanity potential. Unlike AI tech.
52.
▲
by
ACCount37
1mo ago
AI is different because "intelligence" is the last thing humans still do better than machines.
53.
▲
by
ACCount37
1mo ago
Yeah, the target is not the cryptographic "safe forever", but a real world "safer than not having it". If we're trying to decide whether a high profile politician has committed a crime, then yeah, C2PA on the footag
54.
▲
by
ACCount37
1mo ago
I wonder how the image generation models that generate SVGs work. Are they trained roughly like this? Or is it an LLM conditioned on image? Or on diffusion latents from a model trained to emit SVG-compatible imagery?
55.
▲
by
ACCount37
1mo ago
"China does it in a cave with a box of scraps" is a myth. Chinese labs play the shell game to get their hands on a lot of compute outside China. Tricks like distillation save compute in the RL leg of the process - where a lot of t
56.
▲
by
ACCount37
1mo ago
Moravec's paradox begs to differ. Things that are hard to humans are easy. Things that are easy to humans are hard. Math is incredibly hard to humans, but "proving a conjecture" might have a lower intrinsic complexity than &q
57.
▲
by
ACCount37
1mo ago
No, but they were good at answering formalized versions of the same word problems. What this tells us is that a 3B LLM can retain enough NLU to understand those word problems. Which isn't particularly surprising? And also that the same
58.
▲
by
ACCount37
1mo ago
It is a reliability problem - because if the API ends up rejecting 30% of the incoming queries, no one cares if it's 500 Internal Server Error or 529 Overloaded. Having your infrastructure at the knife's edge of load to capacity a
59.
▲
by
ACCount37
1mo ago
Which are the kinds of tasks computers have been historically quite good at. It's impressive that it does what it does, don't get me wrong. But if you expect it to replace the likes of GPT 5.6 Luna, let alone Sol? Nah.
60.
▲
by
ACCount37
1mo ago
"Specialized models" are a bit of a doozy. The biggest generalist models beat the most fine-tuned specialists, as a rule. You can bias an LLM away from literature knowledge and towards coding capabilities, but that buys you very l
More ›