Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
redox99
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
61.
▲
by
redox99
21d ago
The problem is not photorealism. The SVG is outright dumb, its on the wrong side of the table, the table has fucked up geometry (its tilted) and many more minor flaws.
62.
▲
by
redox99
21d ago
Images from their X account
63.
▲
by
redox99
21d ago
They realized Astra having the same score as Sol was silly so they rushed to update the index so it fits what people expect. The old index was clearly bad (Astra is way better than Sol) but it's also unscientific to tweak it like this.
64.
▲
by
redox99
22d ago
Those two you mentioned completely demolish opus 4.5. It's not even close. I'd say they are between opus 4.8 and opus 5. And better in some tasks.
65.
▲
by
redox99
22d ago
4.5 became useful for one shotting large features. LLMs were useful for coding ever since GPT3 (copilot), and sonnet 3.5 for agentic coding.
66.
▲
by
redox99
23d ago
Yeah 5 was very underwhelming.
67.
▲
by
redox99
23d ago
The jump from 3.5 to 4 felt gigantic to me back then. GPT 5.0 did feel underwhelming though.
68.
▲
by
redox99
23d ago
I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
69.
▲
by
redox99
24d ago
Definitely not Google, countless horror stories and infamous for killing stuff. OpenAI is still serving GPT 3.5 turbo as far as I remember.
70.
▲
by
redox99
24d ago
You can ask 100 people and they'll all give you a different list. It's subjective. I think a less personal ranking would be, as a business owner, which of those providers is more dependable? As in, you don't care about evil,
71.
▲
by
redox99
24d ago
Google, Zuck, Sama, Elon, Amodei (in no particular order). They all suck. Pick your poison.
72.
▲
by
redox99
24d ago
No because the frontier keeps advancing very fast.
73.
▲
by
redox99
24d ago
Yeah for sure. Easily 10x that. And obviously we're talking just output tokens. I was mostly just pointing out a theoretical lower bound.
74.
▲
by
redox99
24d ago
> I think it's roughly possible for any piece of software. Still not possible for games (many people are attempting it on X). It's definitely getting better with every new model and it's just a matter of time before they c
75.
▲
by
redox99
24d ago
Obviously the agent would need to build a very extensive test setup for Paint.NET, a lot of it with computer use or something similar. Opening the same files, performing the same clicks, etc should output the same on paint.net and the clone
76.
▲
by
redox99
24d ago
700k LoC of human written code is probably like 2M LoC of LLM "slop". That's probably around 8M tokens. Assuming a single fable agent that one shots the new source code at 30 tok/s, that's 74 hours. Of course it nee
77.
▲
by
redox99
24d ago
How much in tokens/Claude subs would it cost to make a paint.net open source clone? Seems like the kind of software that with enough tokens an agent would be able to create from scratch autonomously (because you have a clear goal and r
78.
▲
by
redox99
24d ago
Qwen 3.8 27B absolutely demolishes Mistral best offerings at coding, and you only need a 5090 or 2x3090 to run it.
79.
▲
by
redox99
24d ago
Just wait and see
80.
▲
by
redox99
24d ago
It's just a leak don't take it too seriously, the model releases later this week.
81.
▲
by
redox99
25d ago
Am I 100% sure the author is telling the truth? No Does it matter? No
82.
▲
by
redox99
25d ago
It's not my image
83.
▲
by
redox99
25d ago
Yes
84.
▲
by
redox99
25d ago
Yes https://x.com/lyraxana/status/2093960706051727723
85.
▲
by
redox99
25d ago
>massive PD disaggregated cluster of B300s connected via NVLink. So a headless server. Macs were mentioned because that's what the post is about. It could be a PC (I use a 2x3090 PC). The point is that it's a better experience
86.
▲
by
redox99
26d ago
Only people who live in SF are capable of programming a browser?
87.
▲
by
redox99
26d ago
Firefox has 1.4 billion in reserves and no debt. If they stop paying tens of millions to their management, they could easily afford paying devs in perpetuity just from investing that money.
88.
▲
by
redox99
26d ago
You can, you just need a beefier PC, and it's more annoying in terms of noise and heat vs throwing something on your server closet. Plus you don't need to worry about other software stealing resources and whatnot.
89.
▲
by
redox99
26d ago
Qwen 27B runs very comfortably on a 5090. You need to use Q4 quants and Q8 KV cache. Here's the math https://news.ycombinator.com/item?id=49514141
90.
▲
by
redox99
26d ago
32GB of fast unified memory is enough for Qwen 3.8 27B. - 16GB for the weights at Q4 - 9GB for the full 256K context at Q8 - 7GB spare for overhead and system. The problem is that these Macs have 32GB of slow unified memory. Edit: I
More ›