Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
GodelNumbering
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
by
GodelNumbering
3mo ago
I've playing around in between with Arc-AGI-3 lately. Based on my very quick test prompt, I do not think it will achieve any meaningful score in Arc AGI 3. Not that it was expected to.
62.
▲
by
GodelNumbering
3mo ago
This is not the right thing, this is the tactical thing. If you have an LLM with less than 1% of the share to begin with, you suffer from bad rep and you got caught uploading user data, one of the very few remaining tactical moves to try to
63.
▲
by
GodelNumbering
3mo ago
Interestingly, when opening this page, the first thought I had was not that the benchmarks should be high, but 'I really hope they did not benchmaxx'. I think a model with modest benchmark scores can have much better real world ut
64.
▲
by
GodelNumbering
3mo ago
Sharing this as a datapoint, not a benchmark
65.
▲
GPT-5.6-terra used 48.5% more context than Mimo-2.5-pro
(dirac.run)
3 points
by
GodelNumbering
3mo ago
|
1 comments
66.
▲
by
GodelNumbering
3mo ago
It is not the raw prompt size that matters ultimately, otherwise Pi (and variants) would be the lowest costing agents. What matters is how efficient the prompt it. Prompt minimalism often gets conflated with efficiency. Having said that, CC
67.
▲
by
GodelNumbering
3mo ago
Interesting how all of grep, sed, ls, cp, mv, rm, cat, pwd, chmod etc are well over 50 years old and get used more than ever today. Claude code owes at least some of its success to the well established and solid unix toolchain
68.
▲
by
GodelNumbering
3mo ago
yup it does
69.
▲
by
GodelNumbering
3mo ago
Thanks, I needed to hear that lol. Yes, the site was an afterthought, core work took/takes most my focus. I will look into un-slopping the site soon.
70.
▲
by
GodelNumbering
3mo ago
What that link describes is basically the motivation to go from terminal bench 2.0 to 2.1. The latter simply fixed the common issues/complaints. There is a long github discussion on tbench's about it
71.
▲
by
GodelNumbering
3mo ago
Dirac ( https://github.com/dirac-run/dirac , https://dirac.run/ ) now supports gpt-5.6. This thing does now seem to be on the chatGPT/codex accounts yet. UPDATE: it is now available in chatGPT accoun
72.
▲
by
GodelNumbering
3mo ago
Terminal bench 2 isn't simply about 'somehow' getting a task done, it intends to measure real world behavior of an agent, including environment awareness in a given situation. A few examples from memory: 1. This task [1] asks
73.
▲
by
GodelNumbering
3mo ago
Lot more details in the linked report https://ai.meta.com/static-resource/muse-spark-1-1-evaluatio... From Terminal-bench-2.1 details, > We use a bash-tool-only agent harness to evaluate 89 Terminal-Bench 2.1 tasks
74.
▲
by
GodelNumbering
3mo ago
There are also a lot of fake results out there on Terminal Bench 2 for different reasons (although the great team behind it Ryan/Alex et al, recently cleaned up a lot of dodgy submissions). A lot of labs publish the results by modifyin
75.
▲
by
GodelNumbering
3mo ago
Also, the cache hit pricing is 25% of the input pricing ($2 vs $0.50). Long agentic workflows are dominated by cached input. The US frontier labs typically have this at 10% of the input price, and DeepSeek/Xiaomi etc take it to the ext
76.
▲
Cache hit rate dropping by 20% doubles your agent's bills
(dirac.run)
2 points
by
GodelNumbering
3mo ago
|
0 comments
77.
▲
by
GodelNumbering
3mo ago
Very shallow wrapper around the reuters piece ( https://www.reuters.com/business/zuckerberg-says-ai-agent-de... ), I dont think author adds any tangible value
78.
▲
by
GodelNumbering
3mo ago
Youtube has been incredibly frustrating for many many reasons and is evidently evil in many axes now. We really need competition in video hosting.
79.
▲
by
GodelNumbering
3mo ago
There are infinitely many 3-level hierarchies. My point was about overloading the model sizing with one more unnecessary classification.
80.
▲
by
GodelNumbering
3mo ago
This assumes a perfect problem routing though. Determining the complexity class of an arbitrary problem is generally undecidable or extremely hard (Rice's theorem implication). So, in real use cases, you need to amortize all cases wher
81.
▲
by
GodelNumbering
3mo ago
This would not work in the way that shows any significant genuine benefit IMO. Caching and optimum routing of a single request are at odds with each other. Higher the distinct model count in a conversation, more cache misses you accept. Bas
82.
▲
by
GodelNumbering
3mo ago
I do not like the fact that this forces people to remember one more hierarchy of "Sol vs Terra vs Luna". OpenAI was supposed to simplify their naming since at least 2025.
83.
▲
by
GodelNumbering
4mo ago
I don't see any real point being made in (or point of) the article. The author sort of just...dumped a bunch of links with the noise that is so incredibly mainstream at the moment that I doubt any of it is news to anyone even somewhat
84.
▲
by
GodelNumbering
4mo ago
I have been experimenting with modifying Ghostty lately. It's a well attended codebase and a pleasure to work with, props to Mitchell. Since Ghostty is written in Zig, I ended up adding native Zig AST support in Dirac ( https:/&#x
85.
▲
by
GodelNumbering
4mo ago
yes, one of them is paid.
86.
▲
Don't Conflate 'Minimal' with Minimal Effort
(dirac.run)
2 points
by
GodelNumbering
4mo ago
|
0 comments
87.
▲
by
GodelNumbering
4mo ago
I built this using openrouter data. An obvious asterisk: not everyone routes their proprietary models through openrouter. The underlying assumption is that the percentage distribution of people routing proprietary models through openrouter
88.
▲
OSS models decisively overtook Proprietary models in openrouter market share
(dirac.run)
4 points
by
GodelNumbering
4mo ago
|
1 comments
89.
▲
by
GodelNumbering
4mo ago
As someone that spends all day every day talking to LLMs, I'd say the OSS frontier models + a good harness is already a sufficient combo. For local deployments, we are missing one or two hardware generations (and may not get that soon
90.
▲
by
GodelNumbering
4mo ago
A new CLI for https://github.com/dirac-run/dirac and a paper that may or may not ever publish
More ›