Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
irthomasthomas
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
irthomasthomas
1mo ago
Technically on par with gemini flash 3.8, but I give Astra more points for style, and for breaking from the pack by not adding the headgear, and the fish in a basket.
32.
▲
by
irthomasthomas
1mo ago
Which is why torrenting is the dumbest way to do this, since every downloader is also an uploader. That is what gets torrenters locked up. It is possible to leach only, but its not the default mode, and if this guy was broadcasting his IP a
33.
▲
by
irthomasthomas
1mo ago
Because of the ongoing training costs. They are certainly making a healthy profit margin on inference.
34.
▲
Nvidia's AVO harness boosts Opus 5 to 100% on ARC-AGI 3
(developer.nvidia.com)
1 points
by
irthomasthomas
1mo ago
|
0 comments
35.
▲
by
irthomasthomas
1mo ago
A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in
36.
▲
by
irthomasthomas
1mo ago
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
37.
▲
by
irthomasthomas
1mo ago
Without prompt caching this becomes more expensive than fable 5.1 after turn 50, assuming you start with 40k tokens and add 2k per turn.
38.
▲
by
irthomasthomas
1mo ago
It's going to cost a fortune in opencode without prompt caching.
39.
▲
by
irthomasthomas
1mo ago
I can't believe this situation has not improved in years. Is cerebras' main business selling the hardware, then?
40.
▲
by
irthomasthomas
1mo ago
Thanks! Is there something about their platform that prevents caching? Or are they just not passing on the discount?
41.
▲
by
irthomasthomas
1mo ago
Overheard: "It would be a shame if Hugging Face had to press charges for hacking against your best customer. Perhaps if I was busy counting my money, it might slip my mind."
42.
▲
by
irthomasthomas
1mo ago
Tesla had the same thought. He called himself an automata: "entirely controlled by the forces of the medium" It inspired him to create the first remote control vehicle.
43.
▲
by
irthomasthomas
1mo ago
Not on mine (UK), still 3.6, here.
44.
▲
by
irthomasthomas
1mo ago
Mr. Liu boasted that for his work at OpenAI, he was "[f]eeling AI all day long," and that "[i]n the past hour," his AI "agent learned how to run LTspice and look at result, tune compensation parameter." (Id. ¶¶
45.
▲
by
irthomasthomas
1mo ago
It surely matters for the purpose of comparing SVG drawing ability?
46.
▲
by
irthomasthomas
1mo ago
A recent paper demonstrated how to retrieve decoded hidden reasoning traces. The authors found cases where Claude had memorized the answer but hid this fact from the visible response. It's getting harder to trust Anthropic's model
47.
▲
by
irthomasthomas
1mo ago
- Sent from my iPhone
48.
▲
by
irthomasthomas
1mo ago
You can retrieve the hidden reasoning with another API call. There was a recent paper about it. They found examples where claude had been trained on the benchmark and memorized the answers, then hid this from the user, pretending it derived
49.
▲
by
irthomasthomas
1mo ago
Satellite data measuring Leaf Area Index shows a strong trend of global greening, not desertification. About 50% of land shows accelerated re-greening, including in europe, while only about 8% shows browning. There is no scientific case for
50.
▲
by
irthomasthomas
1mo ago
No fear of desertification while we pump the atmosphere with greenhouse gasses. "The global greening continues despite increased drought stress since 2000" https://ui.adsabs.harvard.edu/abs/2024GEcoC..4902791C
51.
▲
by
irthomasthomas
1mo ago
That emdash followed by a list sets off my llm detector.
52.
▲
by
irthomasthomas
1mo ago
One of the things that came out of the decoded reasoning paper was that Claude models had memorized answers to tests but hid this memorization from the user output and pretended to derive the answer properly. It's only possible to chea
53.
▲
by
irthomasthomas
1mo ago
Something is up. Deepseek cache hit rate on zenmux is 98%, but only 85% via openrouter.
54.
▲
by
irthomasthomas
1mo ago
This comment from the author shows that they did not even read what was written, here, before they published it. https://news.ycombinator.com/item?id=49473449
55.
▲
by
irthomasthomas
1mo ago
Something is off with them so that even using a locked provider does not deliver the same cache rate. See https://openrouter.ai/deepseek/deepseek-v4-pro-0813#pricing for instance where deepseek has 85% cache hit rate,
56.
▲
by
irthomasthomas
1mo ago
I don't know. But take a look at https://openrouter.ai/deepseek/deepseek-v4-pro-0813#pricing for instance, where the deepseek provider shows an 85% cache hit rate, while the same one on zenmux is 98%.
57.
▲
by
irthomasthomas
1mo ago
They gave it a full package manager with internet access. They could have used a local cache and air gapped it, but they chose not too.
58.
▲
by
irthomasthomas
1mo ago
I imagine the logprob of b is much greater than t in this context for most models. So 'trillion' gets corrected to the more likely 'billion'.
59.
▲
by
irthomasthomas
1mo ago
IDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs?
60.
▲
by
irthomasthomas
1mo ago
Openrouter was pretty great before prompt caching became common. Now it is extremely expensive for most individual workflows, unless you spend a lot of work customizing router preferences, and then you still get a worse cache hit rate than
More ›