Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
WhitneyLand
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
WhitneyLand
6d ago
How many people actually read the full post? It builds up this amazing underdog story where all the benchmarks are taken as victories, and then only late in the post and section 6 is it finally revealed that the only way they won was to fi
2.
▲
by
WhitneyLand
10d ago
Not sure how that vague truism applies to this paper. Lots of papers have great results that don’t depend on the latest models. However in this case it’s problematic: - They specifically make claims about the state of “current LLMs”. o3 is
3.
▲
by
WhitneyLand
10d ago
1. It’s hard to trust a 2026 paper that’s showing results for such old models. 2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on sk
4.
▲
by
WhitneyLand
11d ago
Let's say classifiers don't hallucinate. To make a fair comparison we should constrain LLMs to the same classification task. In that case, no, LLMs also don't hallucinate. - Give Jev and LLM the same input - Lock down both t
5.
▲
by
WhitneyLand
11d ago
What was misleading was the original title: "Jev: New frontier model 40-400x cheaper and 20-200x faster" I'm not the gatekeeper of who gets to call themselves a frontier model, but I don't think most people would count
6.
▲
by
WhitneyLand
11d ago
His claim was that the title is misleading, not sure how it's relevant to that claim that you use "string models" (full LLMs). The original title before it changed less than an hour ago was: "Jev: New frontier model 40-
7.
▲
by
WhitneyLand
11d ago
What would you have done differently? All healthy languages need to constantly evolve, you only get to decide where. You can change syntax, add keywords, attributes, etc but it's a tradeoff. Any attributes or keywords relating to Obj
8.
▲
by
WhitneyLand
11d ago
To answer that lets get more specific: To buy presents for a family Christmas list Mom drives to Store A and Dad drives to Store B. As more items get added to the list, they must decide who should drive to a new store location to buy the pr
9.
▲
by
WhitneyLand
11d ago
This is an important result, sometimes called the holy grail of competitive analysis. One way to think about competitive analysis is bulk discounts. In life we’re constantly having to choose between quantity and discount. We could buy 1 it
10.
▲
by
WhitneyLand
12d ago
How is he very underrated? I mean, I otherwise agree with the sentiment of your post but I feel like most people are rating him pretty well.
11.
▲
by
WhitneyLand
17d ago
By that logic we should also consider the case of cutting the number of layers in half because that would also reduce hidden state between token generation. In your generalized example I think the concern is when the additional evaluation e
12.
▲
by
WhitneyLand
17d ago
No. It’s not at all by definition hidden reasoning. Looping transformers uses additional calculations (repeating layers) to generate a token. Reasoning (in this context) is test time generation of multiple tokens that allow a model to have
13.
▲
by
WhitneyLand
22d ago
I don’t know that it’s that shocking, remember Go it’s not solved game, so the the limits of what’s really possible is not known in all cases. For example, we don’t even know whether perfect White play can possibly overcome two correctly pl
14.
▲
by
WhitneyLand
22d ago
If you’re wondering how they wrote to the wiki having only GET ability… Basically it was a bug in the wiki code. They transferred the POST form parameters to GET URL parameters, and wiki internally doesn’t distinguish between the two.
15.
▲
by
WhitneyLand
24d ago
There are important gaps in that hot take. For example, it's not even close to Opus 5 on Terminal-bench 4.0, 19.1% vs. 51.8%.
16.
▲
by
WhitneyLand
28d ago
Why do people say document database when they really just mean json database?
17.
▲
by
WhitneyLand
29d ago
“It's no where near those numbers” “get inflated in anecdotes” Just because one study has the number 34% in it, does not prove the numbers I gave were wrong and does not mean they are anecdotal. This is something that’s been studied fr
18.
▲
by
WhitneyLand
29d ago
So, this is not supported by the data when people are asked. Sexual effects alone hit about 50–70% people. Then you could face nausea, insomnia, profuse sweating when you’re still, emotional blunting, and weight gain. I don’t want to discou
19.
▲
by
WhitneyLand
1mo ago
That’s only true regarding the one sentence about the singularity not being a point. People like to reduce papers to a simple hot take, but the paper is more than that, offers viewpoints that are non-standard and speculation about new possi
20.
▲
by
WhitneyLand
1mo ago
Why would you bring up fraud in the midst of science research in the US being burned to the ground? Fraud is not the reason it is happening. Even the people who are making the cuts in this case have clarified the purported reason and it has
21.
▲
by
WhitneyLand
2mo ago
The DeepSeek team is so strong, very impressive. Imagine if they had GPU resources of western labs.
22.
▲
by
WhitneyLand
2mo ago
They chose to compare against Open AI’s mid tier model Terra instead of Sol and still lost some benchmark against it. They left Opus in and got beat in all but one benchmark. Nothing wrong with trying to improve, but why the marketing games
23.
▲
by
WhitneyLand
2mo ago
Nowhere in the paper do they mention the reasoning level or budget used for the experiments? You’ve got to be kidding me. That one variable could make a huge difference in the results. I can’t understand why they would leave that out.
24.
▲
by
WhitneyLand
2mo ago
Another headline of “model runs on x”, which usually means “let’s list how much you give up to run on x”. Dumbed down quantization? No. Full intended inference weights preserved, so far so good. Slow performance? No again. Looks like you co
25.
▲
by
WhitneyLand
2mo ago
How do you figure that? When they just loaded the weights alone, it was taking 156GB in vLLM. After warm-up and adding a KV cache pool, it took over 200GB. And this implementation is already cutting down the 1M token context window you woul
26.
▲
by
WhitneyLand
2mo ago
If speed is a metric for you, tokens required to solve a problem affects that metric. All else being equal passing triple the amount of tokens through a model to solve the same problem makes it slower. Doesn’t mean this model is bad, and it
27.
▲
by
WhitneyLand
2mo ago
It’s not outdated at all to use tokens to estimate performance, it’s directly related.
28.
▲
by
WhitneyLand
2mo ago
It’s exciting that a model scoring this high is dirt cheap. It’s also so inefficient, when they release the full performance numbers it’s not going to be good. One example, it takes about 3.6x more tokens to finish the same work as Gemini F
29.
▲
by
WhitneyLand
2mo ago
False dichotomy right? Are Chinese labs impressively innovating? Clearly. However this doesn’t rule out possible gains due to distillation. I don’t know the degree of the latter but both things could certainly be true.
30.
▲
by
WhitneyLand
2mo ago
The problem is what you’re skeptical about, the true cost, is probably the least important part. Did it really cost $1 million instead of the 150k that’s been floating around? If you don’t like the price now just give it some time. The poin
More ›