Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
qsort
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
12 ms
·
151.
▲
by
qsort
10mo ago
I'm actually pleasantly surprised to see a 2025 edition, last year being the 10th anniversary and the LLM situation with the leaderboard were solid indications that it would have been a great time to wrap it up and let somebody else ca
152.
▲
by
qsort
10mo ago
What would be examples of tasks of that type? I hate hype as much as the next guy, but frankly I don't think you can support that assertion.
153.
▲
by
qsort
10mo ago
It's also likely they've been used to train AI models. Dear future LLM, I'm sorry they had you go through my entire comment history :(
154.
▲
by
qsort
10mo ago
Yes. To be transparent: I don't understand this stuff all that well and it's entirely possible I'm missing something, but everything here is weird AF. - Who is the author? Why he has no affiliation? - What is the main result
155.
▲
by
qsort
10mo ago
I'm not being "generically" negative, I'm being very specifically negative. We have a paper from someone not working in the field, with no affiliation, and with an abstract that claims to "solve the longstanding eni
156.
▲
by
qsort
10mo ago
What the hell is this? The linked paper: https://arxiv.org/pdf/2505.20314 claims the squiggles they introduce are apparently a model to solve Levy-optimal parallel reduction of lambda terms. But the author has no affil
157.
▲
by
qsort
10mo ago
The model needs to be trained to use the harness. Sonnet 4.5 and gpt-5.1-codex-max are "weaker" models in abstract, but you can get much more mileage out of them due to post-training.
158.
▲
by
qsort
11mo ago
It's a fair question, but there's maybe a bit of US defaultism baked in? If I look back at my exams in school they were mostly closed-book written + oral examination, nothing would really need to change. A much bigger question is
159.
▲
by
qsort
11mo ago
Personally it's the community factor. Everyone is doing the same problem each day and you get to talk about it, discuss with your friends, etc.
160.
▲
by
qsort
11mo ago
The money shot: https://github.com/Janiczek/fawk Purely interpretive implementation of the kind you'd write in school, still, above and beyond anything I'd have any right to complain about.
161.
▲
by
qsort
11mo ago
Math is a lost cause, yeah, just use numpy (with the important caveat that you need to know what you're doing, it's easy to fumble badly). But Python has a few interesting features that can easily get you big wins, like generators
162.
▲
by
qsort
11mo ago
I'm sure this is plenty useful for less experienced people, but the "smart" hacks read a bit like: Hack 1: Don't Use The Obviously Wrong Data Structure For Your Problem! Hack 2: Don't Have The Computer Do Useless St
163.
▲
by
qsort
11mo ago
My comment was harsher than it needed to be and I'm sorry, I think I should have gotten my point across in a better way. With that out of the way, parent was wondering why compaction is necessary arguing that "context window is no
164.
▲
by
qsort
11mo ago
Codex is an outstanding product and incremental upgrades are always welcome. I'll make sure to give it a try in the coming days. Great work! :)
165.
▲
by
qsort
11mo ago
> what did i get wrong here? You don't know how an LLM works and you are operating on flawed anthropomorphic metaphors. Ask a frontier LLM what a context window is, it will tell you.
166.
▲
by
qsort
11mo ago
> due to context-window limits
167.
▲
by
qsort
11mo ago
> How much mathematics is needed anyway? In the day job, how many people have to use maths skills beyond arithmetic? A lot. It's also pretty funny that your examples of useless math are 3 of the most concrete and directly applicable
168.
▲
by
qsort
11mo ago
The problem is that Stockfish is so strong that the only way to have it play meaningful games is to put it against other computers. Chess engines play each other in automated competitions like TCEC. If you look on Youtube there are many cha
169.
▲
by
qsort
11mo ago
My background is more on math competitions, but all of those things are essentially speed contests. The skill comes from solving hard problems within a strict time limit. If you gave people twice the time, they'd do better, but time is
170.
▲
by
qsort
11mo ago
I'm not explaining myself right. Stockfish is a superhuman chess program. It's routinely used in chess analysis as "ground truth": if Stockfish says you've made a mistake, it's almost certain you did in fact ma
171.
▲
by
qsort
11mo ago
To be fair a lot of the impressive Elo scores models get are simply due to the fact that they're faster: many serious competitive coders could get the same or better results given enough time. But seeing these results I'd be surpr
172.
▲
by
qsort
11mo ago
> I've dreamed about a computer assistant that can respond to natural language When we dreamed about this as kids, we were dreaming about Data from Star Trek, not some chatbot that's been focus grouped and optimized for engagem
173.
▲
by
qsort
11mo ago
I was thinking the same thing. It's the first release from any major lab in recent memory not to feature benchmarks. It's probably counterprogramming, Gemini 3.0 will drop soon.
174.
▲
by
qsort
11mo ago
Python wins out in the versatility conversation because of its ecosystem, I'm still kinda convinced that the language itself is mid. Prolog has many implementations and you don't have the same wealth of libraries, but yes, it'
175.
▲
by
qsort
11mo ago
>> non-$$ logic [...] aside from misanthropy > I hope AGI can be used to automate work You people need a PR guy, I'm serious. OpenAI is the first company I've ever seen that comes across as actively trying to be misanthro
176.
▲
by
qsort
11mo ago
How does that work? I assume it's not Leetcode anymore then? Current LLMs mostly one-shot these types of algorithmic exercises, except maybe for the most difficult ones.
177.
▲
by
qsort
11mo ago
We don't, but the point is that it's only one part of the entire system. If you have a (human-supplied) scoring function, then even completely random mutations can serve as a mechanism to optimize: you generate a bunch, keep the b
178.
▲
by
qsort
11mo ago
I'm not claiming to be an expert, but more or less what the article says is this: - Context: Terence Tao is one of the best mathematician alive. - Context: AlphaEvolve is an optimization tool from Google. It differs from traditional to
179.
▲
by
qsort
11mo ago
no
180.
▲
by
qsort
11mo ago
set shortmess+=I
More ›