Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stared
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
Benchmarking Kimi K3, Opus 5, Grok 4.5, and Gemini 3.6 Flash on Baba Is You
(quesma.com)
4 points
by
stared
2mo ago
|
0 comments
62.
▲
by
stared
2mo ago
Now the friction is close to none, at least for techies. Set a GitHub repo, use pnpm + Astro, and your favourite AI to migrate data there, brew some coffee, and before you finish it, you have your page up. It does not prevent you from posti
63.
▲
Do Qwen 3.6 27B quantizations break the pelican?
(quesma.com)
9 points
by
stared
2mo ago
|
0 comments
64.
▲
Ukraine's Kill Zone
(reuters.com)
6 points
by
stared
2mo ago
|
0 comments
65.
▲
by
stared
2mo ago
It's called frog boiling. We get used to the new level of intelligence so fast, any deviation feels like going back to the stone age. If you don't believe me, create something complex with Opus 5 and then with Opus 4.5, and notice
66.
▲
by
stared
2mo ago
Also top on the freshly released Frontier-Bench, by a large margin: https://www.frontierbench.ai/
67.
▲
Frontier-Bench
(frontierbench.ai)
1 points
by
stared
2mo ago
|
0 comments
68.
▲
by
stared
3mo ago
On another note - who needs WordPress in the age of Astro and LLMs? I mean, it's not a taunt, but a serious question - do people keep WordPress because it used to be the easiest solution to set up years ago, or are there still clear us
69.
▲
by
stared
3mo ago
See the "caveats" section. It was also our initial assumption, but to our surprise there were no signs of models "knowing" the solutions (we investigated all trajectories). Compare and contrast with SWE-Bench Verified, f
70.
▲
by
stared
3mo ago
It is on the way! Sadly, this Kimi K3 API rate limits are devastating, at least on the OpenRouter. So far it solved all The Intro levels (as expected) - I am much more curious how it works with The Lake.
71.
▲
by
stared
3mo ago
I am almost certain it is star or spark. But again, if someone is looking for assholes, they will find them.
72.
▲
by
stared
3mo ago
In a similar vein, some time ago I got curious why the Grafana logo looks like the Zerg emblem see https://www.reddit.com/r/grafana/comments/1o79zxy/grafana_lo... .
73.
▲
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?
(quesma.com)
3 points
by
stared
3mo ago
|
4 comments
74.
▲
by
stared
3mo ago
It was our original assumption. Yet, we went through trajectories and agents did not recall solution. It is with a sharp contrast with task for which agents magically generate solution, e.g. https://openai.com/index/why
75.
▲
by
stared
3mo ago
In the spirit of ARC-AGI-3-like challenges, we just tested if frontier AI models are able to solve a lovely puzzle game, Baba Is You: https://quesma.com/blog/baba-is-bench/ A year ago, Sonnet 4 barely solved the f
76.
▲
Demographic Profiles
(visquill.com)
1 points
by
stared
3mo ago
|
0 comments
77.
▲
Show HN: Tree, truth, druid, dryad, and tar share one Proto-Indo-European root
(p.migdal.pl)
3 points
by
stared
3mo ago
|
0 comments
78.
▲
by
stared
3mo ago
Open-source repo with the benchmark (requires a copy of the lovely game): https://github.com/stared/baba-is-harbor
79.
▲
Baba Is Solved by Fable 5 and GPT-5.6 Sol, but at what cost?
(quesma.com)
7 points
by
stared
3mo ago
|
3 comments
80.
▲
by
stared
3mo ago
I am curious how it works for people with ADHD. I mean, I would love to read more books (and I have a very long backlog), but very rarely I find them stimulating enough to sustain attention. And if they do... well, then time flies and other
81.
▲
Urban heat island effect in Brussels during the late June 2026 heatwave
(eu-space.europa.eu)
5 points
by
stared
3mo ago
|
0 comments
82.
▲
The true cost of saying "Hi" to an AI agent
(quesma.com)
9 points
by
stared
3mo ago
|
4 comments
83.
▲
by
stared
3mo ago
Since when Zuckerberg is an AI expert?
84.
▲
by
stared
3mo ago
I am puzzled (and irritated) why there is „rotate left” without „rotate right”. Does any of you know why?
85.
▲
by
stared
3mo ago
There was "It’s Not Enough to Be Right – You Also Have to Be Kind" https://news.ycombinator.com/item?id=21490714 But I think the core part is WHY we want to be right? To prove something to others, or to ourselves?
86.
▲
by
stared
3mo ago
Nice! If you want to play a hyperbolic minesweeper, Hyperrogue features that https://hyperrogue.fandom.com/wiki/Minefield
87.
▲
by
stared
3mo ago
Yes, it gets really hot really fast. As much as I was tempted to use it on longer projects, I had some reservations about whether it would put too much strain on my MacBook.
88.
▲
by
stared
3mo ago
All experiments with Qwen 3.6 required no more than 48GB Apple Silicon. I believe you can go even further with more aggressive quantizations - one can go down even further. In any cases, from the economic point of view, running models on la
89.
▲
Qwen 3.6 27B is the sweet spot for local development
(quesma.com)
1192 points
by
stared
3mo ago
|
759 comments
90.
▲
by
stared
3mo ago
I didn't know it has a name. But I have been using similar thing, https://github.com/harbor-framework/terminal-bench-science/p...
More ›