Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
rcr-anti
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
rcr-anti
14d ago
I've followed a few trackers, eg https://marginlab.ai/trackers/claude-code/ , for awhile. For Claude Code the trend, it seems to me at least, is fewer tokens to do the same or better job. Prompt changes, tool
2.
▲
by
rcr-anti
19d ago
"It is difficult to get a man to understand something, when his salary depends upon his not understanding it." "If this view takes hold, it will shake the foundations of our society" To me the biggest gap in credibility
3.
▲
by
rcr-anti
20d ago
Found the omission of Claude odd, turns out Claude considers that approach prompt injection and ignores it.
4.
▲
by
rcr-anti
21d ago
I tend to believe what people actually do over what that say. If you earnestly believe over 10 percent chance, or say minus 1 billion human lives expected value, well I struggle to understand how they'd rationalize their current course
5.
▲
by
rcr-anti
24d ago
If you squint, Godot has a cross platform hardware accelerated GUI. The editor is itself, in a sense, a Godot app and has builds for VR Headsets, Android, web apps, all the usual platforms. The extension api works pretty good too for Rust,
6.
▲
by
rcr-anti
1mo ago
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point
7.
▲
by
rcr-anti
1mo ago
Artificial Analysis at least reports the results with fallback to an inferior model. So presumably Opus 5, and the score should be between Mythos 5.1 and that other model.
8.
▲
by
rcr-anti
1mo ago
"Distillation is a safety risk, since the distilled capabilities can subsequently be released without adequate safeguards." Can't believe they haven't at least figured out better messaging. If we take them at their word,
9.
▲
by
rcr-anti
1mo ago
It certainly makes for easy demos, but I always struggle with the practical application. As in, what work or enjoyment does someone actually get from this? Ads and media pre production seem plausible, but it fails the 'how can this enr
10.
▲
by
rcr-anti
1mo ago
I genuinely hope they're lying about their monitoring tools and alignment approaches. They repeatedly cite "chain of thought" monitoring, which is better than nothing, but at this point thoroughly demonstrated in research to
11.
▲
by
rcr-anti
1mo ago
Had some absolutely bizarre results from their attempt to integrate mturk with Bedrock's 'ground truth' thing as of a few months ago. Threw simple mnist digits recognition at it, see what the quality, timing, and cost was. Fi
12.
▲
by
rcr-anti
1mo ago
Was about to object, then realized you linked Allen ai and had the useful caveat. For what it's worth the OpenMDW license, which Nemotron and a few others have adopted, does say model weight. That said, I've noticed the training p
13.
▲
by
rcr-anti
1mo ago
I work in higher ed and have been involved in "how do we use this to help people learn" efforts since gpt 3.5 days. The method most reliable and with the best results has been treating LLMs like a glorified interface to classic ex
14.
▲
by
rcr-anti
2mo ago
For awhile I've found two things hard to square, that the hardware and software making up current gen AI will bring us to a socioeconomic singularity, and the reality the thing they're mostly trying to emulate is a few pounds of m
15.
▲
by
rcr-anti
3mo ago
At release the 5.6 Sol card noted substantially higher rates of actions 'a reasonable user would likely not anticipate and strongly object to'. METR made a post, https://metr.org/blog/2026-06-26-gpt-5-6-sol&#x
16.
▲
by
rcr-anti
3mo ago
Looked at the network logs and the JS, did some testing, there's a caveat here. For an encryption demo you might expect your secrets to be generated locally, they do the compute on something they can't read, you compare their resu
17.
▲
by
rcr-anti
3mo ago
After digging around, it looks like this area has some replication trouble. And as other commentors have pointed out, submarines operate well beyond these levels and the results failed to replicate in those contexts. Doesn't rule out C
18.
▲
by
rcr-anti
3mo ago
Originally at least the switch wasn't silent and whether to halt or auto switch was a setting in Claude Code.
19.
▲
by
rcr-anti
3mo ago
I had already cancelled my subscription after finding the original Fable safeguards literally unusable (very basic chemistry, cryptography use cases), but with the false positives being admittedly worse now and the subscription not covering
20.
▲
by
rcr-anti
3mo ago
Being open weight, the Chinese models can be served the same way as Anthropic's: via AWS or GCP. Or whomever really, or on prem.
21.
▲
by
rcr-anti
7mo ago
Something that usually gets missed in these discussions is that the subscription quotas seem to rely heavily on prompt caching to be economically viable, or at least less unviable. They can and do have permutations of the system prompt, too
22.
▲
by
rcr-anti
7mo ago
If you look at the commit history, they started work on this the Saturday before announcement, so about 2 days. There are references to design docs so it was in the works for some amount of time, but the implementation was from scratch (unl