Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stared
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
stared
6d ago
In the line, I recommend „The Egg” by Andy Weir https://www.galactanet.com/oneoff/theegg_mod.html
2.
▲
Claude couldn't hack OpenAI. Then Anthropic shipped Opus 5
(thenewstack.io)
13 points
by
stared
7d ago
|
1 comments
3.
▲
GPT-6 Astra solves puzzles
(quesma.com)
2 points
by
stared
12d ago
|
0 comments
4.
▲
Knowledge vs. wisdom: asking AI "What mushroom is that?"
(quesma.com)
2 points
by
stared
17d ago
|
0 comments
5.
▲
by
stared
17d ago
If you like, you can run these tests yourself as well. It is ~$500 per a single combination, and assuming no failed runs.
6.
▲
by
stared
17d ago
Nope. If you would like to do so, it is easy (and orders of magnitude cheaper) than running benchmarks.
7.
▲
by
stared
17d ago
Before this experiment I tried to run KLD on various context lengths to see a quantization-dependent deterioration. On wikitext2 there was no difference. I concluded these have no long-term dependency and I should use Linux kernel. Still, t
8.
▲
by
stared
18d ago
A single result is binary. All we get from a run is which tasks were solved, which weren’t.
9.
▲
by
stared
18d ago
Nice! Sometimes the simplest approaches work the best.
10.
▲
by
stared
18d ago
I am curious what's the actual formula. I mean, there so many headers and layers, it is tricky to make a choice that will resonate with our intuition . Is it some weighted average? Or maybe ablation test?
11.
▲
by
stared
18d ago
Point taken, but there is a much more fundamental issue with it - and precisely why I wrote "very conservative". It is a different problem if we pick two sets from the same data distribution, A and B, and first we have a score on
12.
▲
by
stared
18d ago
I wrote this blog post myself, with AI for proofreading (typos and grammar, but not style). There were a few singular sentences for which I had a writer's block, but not much besides that. So, if there are irrelevant remarks, these are
13.
▲
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
285 points
by
stared
18d ago
|
140 comments
14.
▲
by
stared
18d ago
Nice! One thing I am missing is an easy „go up” a taxonomy group.
15.
▲
by
stared
18d ago
Is it a subtle effect that cannot be explained by Newtonian gravity (i.e. a different potential affecting, V(z) in the Hamiltonian)?
16.
▲
by
stared
19d ago
Don’t ask. Start from a few blog posts and see traction, read feedback. You will also see hos long it takes - and what is thd difference between an idea and making it real.
17.
▲
I threw a ton of various logic puzzles at Astra and it solved of them
(twitter.com)
1 points
by
stared
20d ago
|
0 comments
18.
▲
GPT-6 Astra has autonomously completed Portal
(twitter.com)
8 points
by
stared
20d ago
|
0 comments
19.
▲
by
stared
20d ago
An interesting catch! Maybe it suffices to make this pre-title "the ai-design-slop fingerprint · defs 2026.09" lowercase. In any case, from what I see what is AI aesthetics on the website: pre-title and horizontal lines.
20.
▲
Does your landing page look generated?
(slop-detect.com)
3 points
by
stared
20d ago
|
3 comments
21.
▲
A connectomics milestone: Mapping the complete male fruit fly brain
(research.google)
2 points
by
stared
20d ago
|
0 comments
22.
▲
by
stared
20d ago
With Internet, we live in an always-on culture. It contrasts with times before, when we actually had to wait for a monthly magazine, and even if we wanted to watch a movie, it was aired (say) the next Tue 7PM. While now we have all convenie
23.
▲
Show HN: Genetic Distance Map
(p.migdal.pl)
2 points
by
stared
21d ago
|
1 comments
24.
▲
by
stared
21d ago
While I like this index, calling in "Intelligence" might be confusing - it is a mix of coding and knowledge. Compare and contrast with ARC-AGI, BabaIsBench ( https://quesma.com/benchmarks/babaisbench/ ), o
25.
▲
by
stared
22d ago
Compare and contrast with tests on Qwen3.6 27B: * how does quantization hurts pelican-drawing skills ( https://quesma.com/blog/qwen-quantization-quality/ ) * how does it hurt knowledge ( https://quesma.com
26.
▲
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
(quesma.com)
3 points
by
stared
22d ago
|
1 comments
27.
▲
Beef production in Brazil has been the largest driver of global deforestation
(ourworldindata.org)
5 points
by
stared
24d ago
|
0 comments
28.
▲
by
stared
24d ago
Repo for reproducible science: https://github.com/stared/mushroom-hunting-llm-bench
29.
▲
Mushroom hunting with LLMs: what can go wrong?
(quesma.com)
56 points
by
stared
24d ago
|
76 comments
30.
▲
Witcher 3's DLC Won't Include Russian Voice Acting
(thegamer.com)
5 points
by
stared
24d ago
|
3 comments
More ›