Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
theptip
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
theptip
16d ago
State actors already have the resources to do this. The disruption is from actors where the normal deterrence ladders don’t apply.
32.
▲
by
theptip
17d ago
Codex and Claude mobile apps are finally quite usable in the last month or two, it’s been a long time that the experience has been janky. You finally don’t need to do the janky termux ‘Claude —remote-control’ startup dance. Still not as pro
33.
▲
by
theptip
17d ago
It’s a good feature. However I am not convinced that the terminal/ssh abstraction will remain optimal when you orchestrate a swarm of workers across multiple machines. Running via ssh is great when your machines are pets that have loca
34.
▲
by
theptip
17d ago
You are pasting stuff from the open internet into this agent, and it’s crawling the web for you. This is the worst king of hazmat for LLMs, in one of the most adversarially challenging roles (unattended personal agent). Just to be super cle
35.
▲
by
theptip
18d ago
Markets are only zero-sum in any given trade. Allocating capital to assets with higher growth rates (on the marginal dollar) creates value in the long run. So, very simply, if AI can actually do better at picking a better long-term winner t
36.
▲
by
theptip
19d ago
It’s really important to understand the various perspectives in AI discourse; those who view Programing as Art feel a deep affront and threat on many levels. It’s legitimate. If you loved the art of writing code, and also enjoyed getting pa
37.
▲
by
theptip
21d ago
I wouldn’t take the fatalistic stance that it’s fully impossible - but it’s certainly impossible to align a model while racing as fast as any technological paradigm shift has ever raced.
38.
▲
by
theptip
21d ago
OpenAI also found sandbox breaking behavior on a broken biology eval apparently. The evidence suggests it’s more strongly downstream of unsolvable tasks, than the hacking prompt. Anthropic have also observed similar things, so while it seem
39.
▲
by
theptip
22d ago
If the alignment process cannot fix this then we are cooked. The least of our worries is discussions on this forum.
40.
▲
by
theptip
22d ago
Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to conti
41.
▲
by
theptip
23d ago
Concretely, TFA lists a bunch of input validation that is being made more strict in the default configuration.
42.
▲
by
theptip
24d ago
Two problems. One, people mean different things by “coding”. Two, it’s diffused slower than folks predicted. But the core prediction was not crazy for “coding as typing code”, which at most big tech companies is at roughly 100% automation n
43.
▲
by
theptip
29d ago
Why do you need to trust Gates on anything? You can think for yourself and reason to the correct conclusion. He’s not dropping some secret knowledge here. All he is doing is helping move the Overton window to bring “very bad outcomes from A
44.
▲
by
theptip
29d ago
I agree there are some extremely bad outcomes from fast labor replacement. I don’t think it’s inevitable that the economy collapses though. The problem is that this model assumes that capital needs labor, but soon it may not. The easiest mo
45.
▲
by
theptip
29d ago
No, the metric is “p50 task duration”. 7-month doubling time, recently 4-month: https://metr.org/ Plenty of other exponential metrics too, compute built, AI revenue, etc.
46.
▲
by
theptip
29d ago
LLM capability improvement, eg as measured by METR task time.
47.
▲
by
theptip
29d ago
I mostly agree, though there are versions of ASI that I believe leave humans “in charge”. Most of my probability mass for good outcomes is around some variant of “guardian angels” or “benign machine god”. Banks’s Culture novels being on the
48.
▲
by
theptip
29d ago
If you assume democratic access to compute, this may be true. However, the base case is that compute gets hoarded, and as the labor share of profit decreases (capital just needs compute and robots - not human labor - to grow) then why would
49.
▲
by
theptip
29d ago
Agreed, I think AGI->ASI. My happy radical abundance stories are happy ASI stories mostly. And therefore much less likely. BDFL / Benevolent Machine God is definitely one of the ways you can get a happy attractor. Seems scary to rol
50.
▲
by
theptip
29d ago
It looks to me like AI will enable extreme power concentration by default. For example, suppose we have a world with centrally planned economy where humanoid robots outnumber humans. And also suppose the government has taken ultimate contro
51.
▲
by
theptip
29d ago
IMO, this is a highly bimodal probability distribution. I model this as two very strong attractor states (radical abundance and totalitarian power concentration). I believe the latter is way more likely without a concerted effort that we cu
52.
▲
by
theptip
1mo ago
> About four in ten respondents (37 percent) report that AI has contributed positively to their organizations’ EBIT, essentially unchanged from 2025—despite growth in the share of organizations scaling AI technologies. But respondents do
53.
▲
by
theptip
1mo ago
Honestly I have had great success with “I’m worried here about cpu and latency, please rigorously profile and propose fixes”. The models can build micro-benchmarks with a level of rigor that few could muster for a new feature. I agree that
54.
▲
by
theptip
1mo ago
The trick is to set up the harness so that the solution is easy to verify - you’ve profited as long as verification is cheaper than building, but ideally verification is close to automatic (not always achievable of course). Generally you wa
55.
▲
Field Notes from an AI Society
(asteriskmag.substack.com)
1 points
by
theptip
1mo ago
|
0 comments
56.
▲
by
theptip
1mo ago
You play to your outs. Gemini is far behind on quality.
57.
▲
by
theptip
1mo ago
I have my agent write up a summary of the diffs that land each day in my org. If there is something you interesting I’ll ask for an html explainer with code pointers and scan the code in parallel. I wouldn’t say “post reading code” but it’s
58.
▲
by
theptip
1mo ago
It’s a tech tree question. Nuclear weapons have some very specific choke points that can be monopolized or policed, eg highly enriched uranium is largely a weapon input. There are also specific technologies you need to crack, which includes
59.
▲
by
theptip
1mo ago
A related theory here is that Opus is heavily RL’d to be a sub-agent. If Fable is the primary interlocutor then perhaps there is less pushback on the obtuse language. Indeed perhaps the convoluted language acts as a kind of Neuralese betwee
60.
▲
by
theptip
1mo ago
Agents can’t one-shot the problems that boring technology solves. So I think your first premise is mistaken. There is no trivial “just DIY” option for Postgres or Django or Kubernetes.
More ›