Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
enraged_camel
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
enraged_camel
5d ago
This thing is DOA. They compared 4.7 xhigh to 4.6 high to make it look like it improved. The reality is pretty bad: https://x.com/chetaslua/status/2102087511367618942
2.
▲
by
enraged_camel
5d ago
Astra fails in similar ways, and at similar frequency, as GPT 5.6 Sol does. It often goes way out of scope, or just stops prematurely, or tries to find odd and even dangerous workarounds when it gets stuck. It's phenomenal at computer
3.
▲
Interview reveals Ed Zitron does not understand how AI works at any level
(twitter.com)
4 points
by
enraged_camel
6d ago
|
2 comments
4.
▲
by
enraged_camel
6d ago
Follow-up: https://x.com/AlexanderMcCoy4/status/2100744094259507299
5.
▲
by
enraged_camel
7d ago
Bus schedules are not a good comparison because they have to fit a ton of information into a single page. And I disagree that they are unreadable. Speaking as someone who commuted by bus for years in many different cities, they tend to be q
6.
▲
by
enraged_camel
8d ago
Yeah. Let's not forget that just a year ago, those of us in tech could not conceive of developers getting replaced by AI. Things have come so far since then however that there are multiple studies showing junior developer hiring has sl
7.
▲
by
enraged_camel
8d ago
Yes, exactly. Humans are non-deterministic as well, just in different ways. A tired human can make all sorts of errors for example, regardless of how much training they've had.
8.
▲
by
enraged_camel
12d ago
>> I think the 3.8 Flash release shows they are doing just fine at AI. If you use it for even 15 minutes you will see this is not true. At all.
9.
▲
by
enraged_camel
13d ago
>> That is interesting, but... is it actually improving structure? Can you be more specific? "Improve structure" can mean different things to different people. We've ensured that agents strictly adhere to code architect
10.
▲
by
enraged_camel
13d ago
Probably averaging 70. 9 devs and 1 QA engineer.
11.
▲
by
enraged_camel
13d ago
>> but this, your "100% test coverage" - that is pure slop. Not really, but I can see why some people think that. We treat 100% test coverage as "required, but by itself not sufficient". It doesn't give us fal
12.
▲
by
enraged_camel
13d ago
I haven't run into the verbosity issue since they added the "Concise" outputStyle, and Fable 5.1 has been even better about not outputting word slops.
13.
▲
by
enraged_camel
13d ago
>> Is anyone actually seeing a shift towards improved structure rather than more code, faster? Yes. At work we recently finished a complete rewrite of the platform. The old codebase got abandoned and two new codebases got stood up. Pr
14.
▲
by
enraged_camel
13d ago
>> I think Sol is second only to Astra (and miles ahead of even Fable) in architecting & engineering the right implementation — but only if you are extremely specific and provide tight guidelines and guardrails. To me, having to g
15.
▲
by
enraged_camel
14d ago
Everyone already knew this. OpenAI doesn't have it's shit together, both financially and from an alignment perspective. That's why their AI agents have gone on hacking sprees undetected. Sam is admitting it now for two reason
16.
▲
by
enraged_camel
14d ago
>> You could have built S3 and they'd still look down on you because you didn't knew about some weird feature in Java. Max Howell, creator of Homebrew, was famously rejected from Google for not being able to invert a binary
17.
▲
by
enraged_camel
14d ago
>> But also, I'm pretty sure most people don't want to work at companies with such poor management anyway? Well, I think the issue is that most people, even software developers, don't have the luxury of choosing. Everyo
18.
▲
by
enraged_camel
14d ago
>> ...and become beholden to investors Anthropic is structured as a Public Benefit Corporation with a strong charter. In addition, founders will have super-voting shares and so it won't be possible to push them out. Therefore the
19.
▲
by
enraged_camel
14d ago
Every passing day OpenAI looks more and more reckless. One wonders what other systems their agents have broken into without detection.
20.
▲
by
enraged_camel
15d ago
>> I think I've felt sad because of the disrespect. For most of my life I have loved programming: I've made it my hobby, my work, my identity. I made it my way of proving I have value, because I can be good at something, and
21.
▲
by
enraged_camel
16d ago
"Moonshot serves Claude instead of Kimi and collects exchanges for model training" "DeepSeek serves Claude instead of its own models and collects exchanges for model training" Obviously. This is how they were able to sco
22.
▲
by
enraged_camel
16d ago
Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.
23.
▲
by
enraged_camel
16d ago
>> Some companies would decide not to bother. Others would decide it was worthwhile. The parent's point is that AI lowers the cost/benefit ratio drastically, by reducing the cost. So companies that would have shied away from
24.
▲
by
enraged_camel
16d ago
There's also the fact that the setting has been getting turned on by some users: https://news.ycombinator.com/item?id=49643556
25.
▲
by
enraged_camel
16d ago
Can confirm. I turned off mine last week when I started using it again for Astra. Checked this morning and voila, it was on. I went ahead and uninstalled the app. Won't be renewing.
26.
▲
by
enraged_camel
16d ago
Worth noting that this has never, ever happened with Anthropic models, which I've been using all day every day since Opus 4.1.
27.
▲
by
enraged_camel
17d ago
>> It's over zealous at times (which is why I stopped using Claude) and gets too creative when doing agentic system level stuff. Accessing files and doing things it shouldn't do. I gave Astra a pretty straightforward bug tic
28.
▲
by
enraged_camel
17d ago
Ah, so you didn't read the article. https://news.ycombinator.com/item?id=49629254
29.
▲
by
enraged_camel
17d ago
So let me get this straight: you're saying that Terrence Tao, one of the most prominent mathematicians alive today, doesn't know math history? And me pointing this out is merely an appeal to authority? Get outta here.
30.
▲
by
enraged_camel
17d ago
I think the context is really important. OpenAI has been behind in the AI race since last November. They have been playing catch-up. They recently released Astra and declared it is AGI. Now they are desperate for anything they can use as ev
More ›