Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sothatsit
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
61.
▲
by
sothatsit
6mo ago
They provide thinking summaries, so I assume they have to call Haiku or some other model to summarise the thinking blocks.
62.
▲
by
sothatsit
6mo ago
People will accept it as a way to build good software. Many are still in denial that you can do work that is as good as before, quicker, using coding agents. A lot of people think there has to be some catch, but there really doesn’t have to
63.
▲
by
sothatsit
6mo ago
I think the anti-AI stance has been reversing on HN as tooling improves and people try it. It’s only been a little over a year since Claude Code was released, and 3 or 4 months since the models got really capable. People need time to adjust
64.
▲
by
sothatsit
7mo ago
Generally I think this happens when people don’t monitor for errors on a regular basis. People only notice if things are actively broken for customers, and tons of small non-fatal bugs slip through and build up over time.
65.
▲
by
sothatsit
7mo ago
It is not just startups or small companies embracing agentic engineering… Stripe published blog posts about their autonomous coding agents. Amazon is blowing up production because they gave their agents access to prod. Google and Microsoft
66.
▲
by
sothatsit
7mo ago
The benchmark is AI making less mistakes than humans, not making no mistakes. Just like autonomous vehicles. And yes, presumably there would be a person who set the firm up, or else our legal system would need to change quite fundamentally.
67.
▲
by
sothatsit
7mo ago
That is why a fully automated firm would be a paradigm shift. Instead of requiring someone to be responsible and to QA things, you just let AI systems be responsible internally, and the company responsible as a whole for legal concerns. Thi
68.
▲
by
sothatsit
7mo ago
You laid out the theoretical limitations well, and I tend to agree with them. I just get frustrated when people downplay how big of an impact filling in the gaps at the frontier of knowledge would have. 99.9% of researchers will never have
69.
▲
by
sothatsit
7mo ago
Fundamentally, I’m more optimistic on how far current approaches can scale. I see no reason why RL could not be used to train models to use memory, and fine-tuning already works, it’s just expensive. The continual learning we get may be a b
70.
▲
by
sothatsit
7mo ago
Memory systems built on top of LLMs could provide continual learning. I do not agree that it is some fundamental limitation. Claude Code already writes its own memory files. And people already finetune models. There is clear potential to us
71.
▲
by
sothatsit
7mo ago
This is what you said: > they are still predicting training set continuations But this is underselling what they do. Probably a large part of what they predict is learnt from their training set, but RL has added a layer on top that does
72.
▲
by
sothatsit
7mo ago
You can’t really say it is just predicting continuations when it is learning to write proofs for Erdos problems, formalise significant math results, or perform automated AI research. Those are far beyond what you get by just being a copying
73.
▲
by
sothatsit
7mo ago
I’d argue this social angle is not very nuanced or effective. Not all people who used Claude Code will be submitting low-effort patches, and bad-faith actors will just lie about their AI-use. For example, someone might have done a lot of in
74.
▲
by
sothatsit
7mo ago
RL on LLMs has changed things. LLMs are not stuck in continuation predicting territory any more. Models build up this big knowledge base by predicting continuations. But then their RL stage gives rewards for completing problems successfully
75.
▲
by
sothatsit
7mo ago
I quite like this direction. Limit new contributors to small contributions, and then relax restrictions as more of their contributions are accepted.
76.
▲
by
sothatsit
7mo ago
The people likely to submit low-effort contributions are also the people most likely to ignore policies restricting AI usage. The people following the policies are the most likely to use AI responsibly and not submit low-effort contribution
77.
▲
by
sothatsit
7mo ago
Trusted contributors using LLMs do not cause this problem though. It is the larger volume of low-effort contributions causing this problem, and those contributors are the most likely to ignore the policies. Therefore, policies restricting A
78.
▲
by
sothatsit
7mo ago
Concerns about the wasting of maintainer’s time, onboarding, or copyright, are of great interest to me from a policy perspective. But I find some of the debate around the quality of AI contributions to be odd. Quality should always be the r
79.
▲
by
sothatsit
7mo ago
I much prefer this, we can choose based on our use-cases, and people who don’t care can still use Auto.
80.
▲
by
sothatsit
7mo ago
This is all manual, so people ask their agent to load Jira issues, edit Confluence pages, etc. Users sign-in using their own accounts using the CLIs, so the agents inherit their own permissions. Then we have the permissions in Claude Code s
81.
▲
by
sothatsit
7mo ago
Yeah, we wrote our own CLIs for Jira/Confluence/Zendesk using their REST APIs instead. Works well, although a bit more work.
82.
▲
by
sothatsit
7mo ago
I'm afraid I can't easily share this, as we have embedded a lot of company-specific information in our setup, particularly for cross-linking between confluence/jira/zendesk and other systems. I can try explain it though,
83.
▲
by
sothatsit
7mo ago
We have been using something similar for editing Confluence pages. Download XML, edit, upload. It is very effective, much better than direct edit commands. It’s a great pattern.
84.
▲
by
sothatsit
7mo ago
> 10% coding and 90% knowing how to do it I think this is the main point where many people’s work differs. Most of my work I know roughly what needs changing and how things are structured but I jump between codebases often enough that I
85.
▲
by
sothatsit
7mo ago
ChatGPT’s instant models are useless, and their thinking models are slow. This makes Claude more pleasant to use, despite them not being SOTA. But ChatGPT is still SOTA in search and hard problem solving. GPT-5.2 Pro is the model people are
86.
▲
by
sothatsit
7mo ago
I strongly disagree about CLI help being a good enough solution. Skills with CLIs backing them is the gold standard right now for a reason. 1. Skills let the agent know the CLI is available because they get an entry in the context window. 2
87.
▲
by
sothatsit
7mo ago
People use UIs for git despite it working so well in the terminal... Many people I knew at uni doing computer science wouldn’t even know what tmux is. I would bet that the demand for these types of UIs is going to be a lot bigger than the d
88.
▲
by
sothatsit
7mo ago
I’ve had a great experience with CLI-related skills at work. We have written CLIs for systems like Jira, along with skills that document the CLIs and describe the organisation of Jira at our company. Claude Code loads these reliably wheneve
89.
▲
by
sothatsit
8mo ago
OpenAI and Anthropic give you a lot of usage/$ through their plans. For the Anthropic Max plans, this can be like a ~90% discount. Copilot does not benefit from this (their pricing model is also different though, it is request-based ra
90.
▲
by
sothatsit
8mo ago
You cannot just directly compare prices like this. It is like comparing share prices, it doesn't really mean much unless you also know how many tokens the models use. For example, GPT-5.2 is even cheaper than Gemini, but in real-world
More ›