Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thegeomaster
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
thegeomaster
1y ago
Well, Gemini Flash Lite is at least one, or likely two orders of magnitude larger than this model.
32.
▲
by
thegeomaster
1y ago
Gemini 2.5 Pro is severely kneecapped in this evaluation. Limit of 4096 thinking tokens is way too low; I bet o3 is generating significantly more.
33.
▲
by
thegeomaster
1y ago
It just matches the 90% discount that Claude models have had for quite a while. I don't see anything groundbreaking...
34.
▲
by
thegeomaster
1y ago
o3 pricing: $8/Mtok out GPT-5 pricing: $10/Mtok out What am I missing?
35.
▲
by
thegeomaster
1y ago
SWE-Bench Verified score, with thinking, ties Opus 4.1 without thinking. AIME scores do not appear too impressive at first glance. They are downplaying benchmarks heavily in the live stream. This was the lab that has been flexing benchmarks
36.
▲
by
thegeomaster
1y ago
They have worse scores than recent open source releases on a number of agentic and coding benchmarks, so if absolute quality is what you're after and not just cost/efficiency, you'd probably still be running those models. Let
37.
▲
by
thegeomaster
1y ago
GLM-4.5 seems to outperform it on TauBench, too. And it's suspicious OAI is not sharing numbers for quite a few useful benchmarks (nothing related to coding, for example). One positive thing I see is the number of parameters and size -
38.
▲
by
thegeomaster
1y ago
This can be dangerous, because Claude doesn't truly understand why it did something. Whatever it writes a post-hoc justification which may or may not be accurate to the "intent". This is because these are still autoregressi
39.
▲
by
thegeomaster
1y ago
I never understood the hate. Beyond the stranger syntax, it's not terribly different from a language such as Pascal. It's an old imperative language without too much magic (beyond strange syntax sugar).
40.
▲
by
thegeomaster
1y ago
Same here. Reading the article, I could not really relate to the experience of being a single-language developer for 10 years. In my early days, I identified strongly with my chosen programming language, but people way more experienced tha
41.
▲
by
thegeomaster
1y ago
You're assuming that the whole model has to be in SRAM.
42.
▲
by
thegeomaster
1y ago
Yes, there have been multiple (very big) hints dropped by various people that they had no official cooperation.
43.
▲
by
thegeomaster
1y ago
Just curious: why do you prefer/have a requirement of quad-based meshes?
44.
▲
by
thegeomaster
1y ago
Mentioned in TFA as well.
45.
▲
LLMs Are Bad at Being Forced
(morphllm.com)
1 points
by
thegeomaster
1y ago
|
0 comments
46.
▲
by
thegeomaster
1y ago
Post-training allows leveraging the considerable world and language understanding of the underlying pretrained model. Intuition is that this would be a boost to performance.
47.
▲
by
thegeomaster
1y ago
A detail that is not mentioned is that Google models >= Gemini 2.0 are all explicitly post-trained for this task of bounding box detection: https://ai.google.dev/gemini-api/docs/image-understanding Given that t
48.
▲
by
thegeomaster
1y ago
This might be a controversial take, but this approach is just plain old engineering applied to LLMs. Instead of making an LLM perform both an accurate code edit AND follow a strict output schema, you split that into two problems: accurate c
49.
▲
by
thegeomaster
1y ago
LLMs will generate unpredictable, very humanlike code edits in this form. They might use comments like "same function as above", "rest of the function with similar changes", "function ABC is no longer needed",
50.
▲
by
thegeomaster
1y ago
It does require a level of contextual awareness, fuzziness and robustness against crazy inputs that in my mind would be very hard to achieve using classical approaches.
51.
▲
by
thegeomaster
1y ago
You can run more intelligent traditional LLMs at higher speeds than the Google diffusion model. Even then, it runs nowhere near 4500tok/s, and such small models generally suck in terms of accuracy compared to a specialized, fine tuned
52.
▲
by
thegeomaster
1y ago
I watched the video and I don't really understand how this maps to the underlying Git operations and what it can do. What happens if I make changes locally while Cursor is doing something? Is this detected properly? (That might be usef
53.
▲
by
thegeomaster
1y ago
I have been using LLMs for coding a lot during the past year, and I've been writing down my observations by task. I have a lot of tasks where my first entry is thoroughly impressed by how e.g. Claude helped me with a task, and then t
54.
▲
by
thegeomaster
1y ago
There's really no defensible way to call one "learning" and the other not. You can carry a half-full context window (aka prompt) with you at all times. Maybe you can't learn many things at once this way (though you might
55.
▲
by
thegeomaster
1y ago
Hard disagree. Of course technically they didn't do anything explicitly against the public guidance (the checks and balances would never let them), but naming a model with a date very strongly implies immutability. It's the same l
56.
▲
by
thegeomaster
1y ago
I hate to enter this discussion, but learning based on a small number of examples is called few-shot learning, and is something that GPT-3 could already do. It was considered a major breakthrough at the time. The fact that we call gradient
57.
▲
by
thegeomaster
1y ago
Google could at least learn something from this attitude, given their recent 03-25 -> 05-06 model alias switcharoo with 0 notice :)
58.
▲
by
thegeomaster
1y ago
But... is it a real problem? As the author says, the entropy reduction is tiny.
59.
▲
by
thegeomaster
1y ago
If this was the case, regular code review as a practice would be entirely unworkable.
60.
▲
by
thegeomaster
1y ago
I've personally found that I reject around 80% of suggestions with "No, and tell Claude what to do differently". So it requires a lot of babysitting, and it usually means I cannot do another thing effectively while it's
More ›