Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
GodelNumbering
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
GodelNumbering
5d ago
Every Grok release obscures their cache pricing while highlighting their input/output pricing From their headline comparison: Grok: $2/$6 per million Fable: $10/$50 per million What this doesn't say: Grok co
2.
▲
by
GodelNumbering
5d ago
Yes I am aware and I agree that it is from the IP. The capability of LLMs differ from traditional cookies and location can materially impact your experience which you may not want. If you read the conversation, I had explicit instructions t
3.
▲
There is no way to get ChatGPT to stop injecting your location in context
(chatgpt.com)
2 points
by
GodelNumbering
5d ago
|
2 comments
4.
▲
Ask HN: What is one plausible path to 'AI extinction'?
3 points
by
GodelNumbering
6d ago
|
5 comments
5.
▲
by
GodelNumbering
7d ago
A friend of mine created GoBench[1][2] that evaluates LLMs on 9×9 Go using KataGo opponents as Elo anchors, you see real capability differences there, like Astra Max substantially leading all other models. I think strategy is a generally in
6.
▲
by
GodelNumbering
8d ago
Good point. I would treat this as 'fully specified vs partially specified'. For a fully specified system, my mental model still maintains that the code is the most compact ruleset. I agree that "don't care" is often
7.
▲
by
GodelNumbering
8d ago
The code itself is the most compact representation of the rules you want applied.
8.
▲
by
GodelNumbering
10d ago
A friend created this eval. Currently it is the most unsaturated evals I've seen with gpt-6 Astra leading massively. Results: GPT-6 Astra · Max 2,568 ± 140 $0.15 66s GPT-6 Astra · High 2,227 ± 138 $0.033 10s Claude Opus 5 · High 2,076
9.
▲
GoBench: Evaluating LLMs on 9×9 Go using KataGo opponents as Elo anchors
(rolandgao.com)
2 points
by
GodelNumbering
10d ago
|
1 comments
10.
▲
by
GodelNumbering
15d ago
Some months ago I was evaluating command output compressors to integrate into Dirac[1] as that seemed like an easy win that would compliment and compound with Dirac's other mechanisms. I tested rtk among these and it was actually a net
11.
▲
by
GodelNumbering
16d ago
Tangential to the subject, but this is a bluesky post, containing a screenshot of an X post, which itself starts with "in a detailed Mastodon post"...
12.
▲
by
GodelNumbering
23d ago
> That's quite common with many models Such as? I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
13.
▲
by
GodelNumbering
23d ago
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks: Terminal-Bench 4.0: High (57.9%), Max (56.7%) DeepSW
14.
▲
by
GodelNumbering
23d ago
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
15.
▲
by
GodelNumbering
24d ago
Do you have instructions on how to run a custom harness against this? There are none on the linked page. I want to run Dirac ( https://github.com/dirac-run/dirac ).
16.
▲
by
GodelNumbering
25d ago
Interesting, even if we were to ignore the cache-hits, reads and output, the reasoning cost (aka test time compute) per task should remain a fully comparable metric - it went from $1.25 (Fable5) to $1.48 (+18.4%) for an improvement signific
17.
▲
by
GodelNumbering
25d ago
The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M). This gives a lot of credit to the theory that Anthropic d
18.
▲
by
GodelNumbering
1mo ago
1.2 TB/s bandwidth of M5 Ultra comes from two dies of M5 Max (each 614 GB/s) connected together using 4.4 TB/s inter-die fabric. For a non-quantized Deepseek V4 flash on an ultra, I would estimate about 1000+ tokens per secon
19.
▲
by
GodelNumbering
1mo ago
This is a plug, but relevant. I recently added a 'build native tools on the fly' functionality to Dirac ( https://github.com/dirac-run/dirac ) that works like: 1. You can use the '/new-tool' and
20.
▲
by
GodelNumbering
1mo ago
I agree in principle that the buck has to stop somewhere, but taking full responsibility for the actions of a fundamentally statistical system that you didn't build and have no interpretability of is uncomfortably risky for most partie
21.
▲
If your agent commits a crime, who is responsible?
(signalbloom.ai)
37 points
by
GodelNumbering
1mo ago
|
104 comments
22.
▲
by
GodelNumbering
1mo ago
I think introductory here means more or less permanent but they can't publicly admit there are no takers at a higher price. Anthropic for instance announced a couple of days ago that they are making Sonnet's 'introductory pri
23.
▲
by
GodelNumbering
1mo ago
The corresponding OpenAI post https://openai.com/index/previewing-ultrafast/ There is no pricing info, which could mean it's "if you have to ask..." territory or they are simply gauging interest bef
24.
▲
by
GodelNumbering
1mo ago
> introductory price They should call it 'face saving pricing after we realized just how terribly did we mis-price the flash 3.5' > since google betrayed us with those price hike, people already spent their time making their
25.
▲
by
GodelNumbering
2mo ago
https://xcancel.com/finkd/status/2086755195535413696 "... Soon we'll also release the weights for Muse Spark 1.2, our latest foundation model..." This is bigger news - good for self hosting enthusia
26.
▲
by
GodelNumbering
2mo ago
Am I missing something or the evals do not compare it to the baseline deepseek-v4-flash? Without a baseline comparison, it is hard to tell what works well and what doesn't
27.
▲
by
GodelNumbering
2mo ago
Apologies, I had no idea. I lookup host image and that site came up. I don't think the site itself is of NSFW nature.
28.
▲
by
GodelNumbering
2mo ago
If that was it, wouldn't Claude go directly to the canonical source?
29.
▲
by
GodelNumbering
2mo ago
I just checked Cloudflare for SignalBloom ( https://www.signalbloom.ai , which I own and operate). Over the last 72 hours, Claude-searchbot [1] alone fetched ~205,000 pages. Sent exactly 1 referral. There is a lot of free financia
30.
▲
by
GodelNumbering
2mo ago
You should play Blood on the clocktower
More ›