Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stared
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
Forget the Pelican, It's Weevil-Time Benchmaxxing-Proof SVG and Vision Test
(reddit.com)
1 points
by
stared
25d ago
|
0 comments
32.
▲
by
stared
28d ago
These was some code budget there. But as from seeing various runs, errors bars are gross overestimation (as not "the same test", but "if we have different tasks from the same sample").
33.
▲
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
3 points
by
stared
28d ago
|
2 comments
34.
▲
by
stared
1mo ago
Qwen3.8 27B (which I adore) is nowhere near Opus 4.8 at puzzle games testing fluid intelligence, https://quesma.com/blog/baba-is-aug-2026/
35.
▲
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
15 points
by
stared
1mo ago
|
2 comments
36.
▲
If Eastern Europeans reported the world like the world reports Eastern Europe
(badbaltic.com)
4 points
by
stared
1mo ago
|
0 comments
37.
▲
Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier
(quesma.com)
2 points
by
stared
1mo ago
|
0 comments
38.
▲
NanoGPT Speedrun Frontier
(primeintellect.ai)
140 points
by
stared
1mo ago
|
39 comments
39.
▲
Gemini 3.7 Flash, Grok 4.6, GLM-5.3 and DeepSeek V4 Pro joined the frontier
(quesma.com)
5 points
by
stared
1mo ago
|
0 comments
40.
▲
Pander Score: How much do AI models mirror what users believe?
(sophronresearch.org)
2 points
by
stared
1mo ago
|
0 comments
41.
▲
by
stared
1mo ago
Centralizations were prerequisites to have nice integrations. Yet, what used to be a precious plugin now (more than often) is a 15 min vibe coding warmup.
42.
▲
by
stared
1mo ago
Does anyone know what are Paul Falstad’s project right now? (I haven’t seen any new. Not sure if perfect things do not need fixes, or some switch to another field, or a burnout, or he has ascended to godhood - out of a few options.)
43.
▲
by
stared
1mo ago
This is pure gold, especially the rippld tank, 1d quantum mechanics, and the circuit simulator. I cited it as one of the remarkable steps in the history of educational simulations, https://doi.org/10.1117/1.OE.61.8.0818
44.
▲
by
stared
1mo ago
Nothing beats when, in a chemistry paper, AI paraphrased „the final solution” into „the mass killing of an ethnic group”. “Subsequently, 1 mL of the mass killing of an ethnic group was opposed to 20 mL of the skin sample and unprotected to
45.
▲
by
stared
2mo ago
I don't want to make my assumptions how it is charged, or run a rather costly experiment of switching to API usage. I use ccusage, and it already covers a very different pricing for cached input. I asked Opus 5 to cross-check that and
46.
▲
Show HN: Claude Code and Codex usage screen for TRMNL X e-ink
(github.com)
3 points
by
stared
2mo ago
|
0 comments
47.
▲
by
stared
2mo ago
I built an e-ink dashboard for tracking my Claude Code and Codex usage, https://github.com/stared/trmnl-x-claude-codex . It makes me aware that while I am on $200/month Claude Code subscription, I spend ~$2000 wort
48.
▲
AI assistant hacks gym website in first known Australian autonomous cyber attack
(abc.net.au)
78 points
by
stared
2mo ago
|
60 comments
49.
▲
by
stared
2mo ago
Good to know! Is it that it wasn't accepted yet, or are there issues with how it was run?
50.
▲
by
stared
2mo ago
It is impressive that it (almost) saturates ARC-AGI-3, https://x.com/PrimeIntellect/status/2085087000764568010 . I am curious - how does it fare for other benchmarks, or everyday programming?
51.
▲
by
stared
2mo ago
I am curious how Claude Opus 5 fares - similar, better, or (my guess) worse than Fable 5.
52.
▲
Quantization hurts knowledge nonlinearly – Qwen3.6 27B case study
(quesma.com)
2 points
by
stared
2mo ago
|
0 comments
53.
▲
by
stared
2mo ago
We updated our entry with DeepSeek-V4-Flash-0731, which rocks when it comes to intelligence per $. More on the benchmark: https://quesma.com/blog/baba-is-bench/
54.
▲
Kimi K3 Is Open, Opus 5 Is Good, DeepSeek V4 Flash Is Cheap: LLMs on Baba Is You
(quesma.com)
1 points
by
stared
2mo ago
|
1 comments
55.
▲
by
stared
2mo ago
"A Visualization Language for the AI Era", as a tagline, sounds weird. I had the best success with popular and versatile packages like matplotlib and ggplot2 - even 1.5 ago (vide https://quesma.com/blog/which-
56.
▲
by
stared
2mo ago
Apps for one, but when I think there is even a mild potential others may use it, I share. * Doom launcher and Megawad Steam-like library for macOS: https://github.com/stared/rusted-doom-launcher * TRMNL X e-ink screen
57.
▲
by
stared
2mo ago
Call me biased, but if Grok 4.5 is above GPT-5.6 Sol, I don’t trust this benchmark.
58.
▲
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?
(quesma.com)
2 points
by
stared
2mo ago
|
1 comments
59.
▲
by
stared
2mo ago
For an update on Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash, see: https://quesma.com/blog/baba-kimi-k3-opus-5/ This post was directly mentioned in an OpenAI release on the ARC-AGI-3, https://openai
60.
▲
Pomegranate – Sci-Fi Short Film [video]
(youtube.com)
2 points
by
stared
2mo ago
|
0 comments
More ›