Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
resonious
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
resonious
2mo ago
The game changes when it's not just "Opus vs Sonnet", but "Opus vs GLM". The amount saved is way more than even $0.04. And it's not only money but speed. Some providers can serve GLM crazy fast - I'll ev
32.
▲
by
resonious
2mo ago
This sounds like not the best workflow. If your prompt is that long, you may not be spending your time very efficiently. Splitting problems into smaller problems is huge. You want a concrete idea that you still own. Only tell an agent to do
33.
▲
by
resonious
3mo ago
How much does it cost? I even made an account and I cannot find pricing anywhere...
34.
▲
by
resonious
3mo ago
So it's GLM-5.2 performance for almost twice the price. That said, the speed looks really good. I think it's competitive with Fireworks's GLM 5.2 Fast, although Fireworks is still cheaper.
35.
▲
by
resonious
3mo ago
Why does KV cache matter if they show it's cheaper anyway?
36.
▲
by
resonious
3mo ago
The article shows that it's still cheaper to pay for the switch than it is to let the frontier model do all the work. At least in benchmarks. If you want to interact with plans then I think this technique just isn't for you.
37.
▲
by
resonious
3mo ago
> for frontier work. I'll agree that GPT 5.6 may well be the best given the above contstraint, but for run-of-the-mill dev tasks (real ones, not benchmark ones), GLM 5.2 still blows every other model out of the water. Cost per task
38.
▲
by
resonious
3mo ago
Oh My Pi has it. I'm a big OMP shill right now. Seems not very popular, but it has the stability of Pi with the features Opencode (and more I think; OMP has web browsing and a more advanced edit system too). OMP often outperforms Claud
39.
▲
by
resonious
3mo ago
Just features for the client. The same stuff I was shipping before but much faster and more stable.
40.
▲
by
resonious
3mo ago
Trust me there are plenty of us using cloud AI to actually ship stuff. We just aren't writing blog posts about it.
41.
▲
by
resonious
3mo ago
The artificialanalysis cost per task chart has DeepSeek as the clear winner and Fable as the clear loser. But I would still pick Fable for some tasks, so that also can't be all there is to it. But I agree that price per token figure is
42.
▲
by
resonious
3mo ago
> But artificial intelligence, far more than any tool we’ve ever created, intends us not just to sit forward and behave, but to cease to think critically, to cease to imagine, and, most temptingly, to cease to feel struggle and pain. If
43.
▲
by
resonious
3mo ago
I consider myself an LLM maximalist, at least in the context of software engineering. Even after delegating everything you can to LLMs, there's still work left for you. And it's far more interesting. Even the delegation to LLMs is
44.
▲
by
resonious
3mo ago
Right I saw them saying something along the lines of "they're good at subagents". But this seems true even with third party harnesses. So I'm wondering what Codex is hiding.
45.
▲
by
resonious
3mo ago
I've also seen Google indexing pages with random values in the path that don't get linked to statically (server asks for the URL then redirects to it immediately). I'm pretty sure they index straight out of the Chrome address
46.
▲
by
resonious
3mo ago
I guess this implies that non-Codex harnesses get a little bit worse? In wondering what's so special about their subagents system that they feel the need to hide these messages...
47.
▲
by
resonious
3mo ago
"This is the part where you think/learn" and "this is the part where you just let the LLM figure it out" is a fine format for advice nowadays I think. I am getting a little tired of every single HN comment being abo
48.
▲
by
resonious
3mo ago
This seems like a lot. I told my oh-my-pi agent to build my ios app for me and send it to me remotely. After some chugging, it gave me a tailscale funnel link. And it worked. I didn't have to dig into any of this stuff. I didn't u
49.
▲
by
resonious
3mo ago
I think the moralities of all the big heads in AI are questionable. The training corpus is largely stolen, and they are all in inescapable debt but keep going. But at this point, their products are so useful that almost nobody is willing to
50.
▲
by
resonious
3mo ago
~13 yoe, and I had some nasty WebRTC + CallKit problems that Opus couldn't make a dent on but Fable figured out.
51.
▲
by
resonious
3mo ago
This is an old rumor but I thought Nintendo made a loss on the devices. If that's true, why would they want to sell more?
52.
▲
by
resonious
3mo ago
Okay I hadn't heard of Vending-Bench until reading this and it was quite the ride learning about it through this article. Very fun read. My very native programmer take is that it's not too surprising that their hacker model would
53.
▲
by
resonious
3mo ago
Over optimizing and spending too much time on engineering decisions is also something an incompetent development team would do.
54.
▲
by
resonious
3mo ago
I think it's the tiny chance they it will help humans that makes it so fascinating.
55.
▲
by
resonious
3mo ago
Look, I felt it. I didn't wait for the official apology from Anthropic. I quite before they published that, then felt very vindicated when they did.
56.
▲
by
resonious
3mo ago
Deja Vu... This looks just like the Claude Code performance regression back in April. I just quit my Claude subscription when that happened and went to Codex. Now I'm kinda thinking of trying per token for both, using GLM 5.2 on Firewo
57.
▲
by
resonious
3mo ago
Claude Code downgrades loudly but I'm not sure what happens over API or with other harnesses, OpenRouter, etc.
58.
▲
by
resonious
3mo ago
Yes you pay a big burst right after switching. After that, everything is cached and it's smooth sailing.
59.
▲
by
resonious
3mo ago
I have clients waiting for very gigantic features and the agent harnesses are a godsend.
60.
▲
by
resonious
3mo ago
This lines up with my experience with my mother, though it played out differently. In her case, she would switch doctors every ~5-10 years and each time they'd basically say everything the previous doctor said was wrong. First it was &
More ›