Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
serjester
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
serjester
1y ago
It's interesting that it seems to the non thinking variant has actually regressed on a quite few benchmarks compared to flash-2.0. They seem to be prioritizing coding above all else. Even the thinking variant only has marginal gains on
32.
▲
by
serjester
1y ago
VLM’s capable of parsing images with high fidelity are 10 - 50X cheaper than the frontier models. Any savings from not parsing, are quickly going to be wiped out if someone has any actual traffic. Not to mention the massive hits to long con
33.
▲
by
serjester
1y ago
There's multiple fundamental problems people need to be aware of. - LLM's are typically pre-trained on 4k text tokens and then extrapolated out to longer context windows (it's easy to go from 4000 text tokens to 4001). This
34.
▲
by
serjester
1y ago
> When I first started at Microsoft, I had a sleeping bag in my office. I coded until 11pm nightly and slept until 3am, at which point I’d code until ~6am, then sleep until my first meeting ~10am. It’s everyone’s pejorative on what they
35.
▲
by
serjester
1y ago
It's smart that they're pivoting to using the user's computer directly - managing passwords, access control and not getting blocked was the biggest issue with their operator release. Especially as the web becomes more and mor
36.
▲
Gemini 2.5 Can Visually Map Text Back to PDFs
(sergey.fyi)
2 points
by
serjester
1y ago
|
0 comments
37.
▲
by
serjester
1y ago
We've submitted tens of millions of requests at a time and never had it take longer than a couple hours - I think the zone you submit to plays a role.
38.
▲
by
serjester
1y ago
Yes, this was completely image-based. Not quite of a point of using it in production since I agree it can be flakey at times. Although I do think there's viable workarounds, like sending the same prompt multiple times, and seeing if t
39.
▲
by
serjester
1y ago
I suspect this is a remnant of how images get tokenized - simplest solution is probably to increase the buffer.
40.
▲
by
serjester
1y ago
I wrote a similar article a couple of months ago, but focusing instead on PDF bounding boxes—specifically, drawing boxes around content excerpts. Gemini is really impressive at these kinds of object detection tasks. https://www.s
41.
▲
by
serjester
1y ago
This article doesn't do the best job explaining the broader picture - stability has been their number one priority up to this point. - Most of the work has just been plumbing. Int/float unboxing, smarter register allocation, free-
42.
▲
by
serjester
1y ago
Tokenizing text is ridiculously small part of the overall computation that goes into serving a request. With that said if you’re doing this on petabytes of data, never hurts to have something faster.
43.
▲
O3-Deep-Research
(platform.openai.com)
4 points
by
serjester
1y ago
|
0 comments
44.
▲
by
serjester
1y ago
Cool stuff! Probably one of the less popular languages, but I noticed that the transcription with Russian is often quite poor. Part of me loves this—no judgement, endless convenience, cheap. But another part mourns, sensing it strips away t
45.
▲
by
serjester
1y ago
I listened to their podcast with Dwarkesh and found them incredibly off putting. They’ve done so little confidence modeling and have completely failed to understand they’re dealing with a chaotic system. The smartest people in the world str
46.
▲
by
serjester
1y ago
I'm completely impartial here - seems like there's only so many ways you can design a schema builder?
47.
▲
by
serjester
1y ago
It crashes with "a problem repeatedly occurred". I think there's some sort of infinite loop - fails on both safari and chrome on my iPhone.
48.
▲
by
serjester
1y ago
Enterprises don't care about faster, but they do care an enormous amount about security. Astral is very well positioned here.
49.
▲
by
serjester
1y ago
Anaconda makes on the order of 100M a year “solving” data science package management. I would argue it has a significantly worse product, attacking a much smaller part of the ecosystem. It seems easy to imagine Astral following a similar pa
50.
▲
by
serjester
1y ago
Congrats on the launch guys, mobile website seems to be broken though.
51.
▲
by
serjester
1y ago
I think you’re straw manning his argument. He explicitly says that both LLMs and traditional software have very important roles to play. LLMs though are incredibly useful when encoding the behavior of the system deterministically is impossi
52.
▲
by
serjester
1y ago
I'm glad that they standardized pricing for the thinking vs non-thinking variant. A couple weeks ago I accidentally spent thousands of extra dollars by forgetting to set the thinking budget to zero. Forgetting a single config parameter
53.
▲
by
serjester
1y ago
I’d be curious to see benchmarks but this kind of query rewriting seems almost guaranteed to already be baked into the model.
54.
▲
by
serjester
1y ago
I found o3 pro to need a paradigm shift, where the latency makes it impossible to use in anything but in async manner. You have a broad question, likely somewhat vague, and you pass it off to o3 with a ton of context. Then maybe 20 minutes
55.
▲
by
serjester
1y ago
I wouldn’t underestimate the impact of having massive communities around a language. Basically any problem you have has likely already been solved by 10 other people. With AI being as frothy as it is, that’s incredibly valuable. Take for ex
56.
▲
by
serjester
1y ago
There’s been plenty of anti missile systems deployed to Ukraine and the tank losses have still been absolutely insane. Fundamentally it’s not about building anti-missile technology, it’s about doing it cheaply and at scale. That is a much,
57.
▲
by
serjester
1y ago
This is just the Peter principle - individuals will keep getting promoted until they're no longer competent.
58.
▲
by
serjester
1y ago
There's the interesting question here - has embedding performance reached saturation? In practice, most people are pulling in 25 to 100 candidates and reranking the results. Does it really matter if a model is 1 - 3% better on pulling
59.
▲
by
serjester
1y ago
Very ambitious but it seems futile if you’re not building the rockets yourself. Personally I’m more bullish on figuring out how to use analog chips to train models.
60.
▲
by
serjester
1y ago
As a counterpoint, I also use cursor as my daily driver and I have been tempted to switch many times because of the endless bugs. Just take a look at their forum.
More ›