Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ddp26
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
Agents Sometimes Catastrophize
(futuresearch.ai)
9 points
by
ddp26
5mo ago
|
2 comments
62.
▲
by
ddp26
5mo ago
Stanford has this policy too. Students get livid when proctoring is proposed, even though cheating is rampant (afaict)
63.
▲
Run Agents Twice
(futuresearch.ai)
6 points
by
ddp26
5mo ago
|
0 comments
64.
▲
by
ddp26
5mo ago
Author here. Great point, and I think this is due to what another commenter points out, that the questions are different. The right test of this is to take the _same_ markets that run for 90+ days, and check accuracy 90 days out vs 30 days
65.
▲
by
ddp26
5mo ago
Author here. Hal Varian pointed me to this 1992 paper, which I think is still considered the canonical empirical piece on what is actually going on in trading behavior that leads to accuracy (or not): https://www.jstor.org/s
66.
▲
by
ddp26
5mo ago
Yeah. People have put together a Prediction Market Database [1] (in a Google sheet), I think it's pretty well sourced and shows a good number of both real money and play money prediction markets from before 2002. DARPA did have a big r
67.
▲
by
ddp26
5mo ago
It's true they are "just" summarizing current knowledge. But there are better and worse summaries of current knowledge! Some summaries, like on some prediction markets, have objective accuracy that is much better than chance.
68.
▲
by
ddp26
5mo ago
Author here. Agree, and I wrote in that section "Absolute accuracy is hard to compare across markets on one platform, and across platforms, because different forecasting questions have different difficulties. I addressed this by tracki
69.
▲
by
ddp26
5mo ago
Yeah, the question in the title can be answered: "by using gpt-4o, a model 2 years behind the frontier, to serve audio responses"
70.
▲
by
ddp26
6mo ago
Training window cutoff is Jan 2026, when Opus 4.6 was Aug 2025. That quite a lot of new world knowledge.
71.
▲
I think Anthropic is worth $100B more than last week
(futuresearch.ai)
9 points
by
ddp26
6mo ago
|
0 comments
72.
▲
by
ddp26
6mo ago
The free open source model does have its competitive advantages!
73.
▲
by
ddp26
6mo ago
The second paragraph starts "Muse Spark is the first step on our scaling ladder and the first product of a ground-up overhaul of our AI efforts. To support further scaling, we are making strategic investments..." This article is a
74.
▲
by
ddp26
6mo ago
Got a source on this? I didn't take into account in this forecast that public markets could be very inefficient in this way.
75.
▲
by
ddp26
6mo ago
There is actually a real bull case for xAI (that I don't endorse), e.g. from people who think that chips & computer is the main determiner of model quality. xAI may plausibly soon have the biggest training apparatus of anyone. I th
76.
▲
by
ddp26
6mo ago
As I wrote in the piece, I'm extremely skeptical that xAI should be valued as if it is a frontier lab. But as you say, going back to the xAI + SpaceX merger, analysts consistently seem to value it as if it is, so I predict the public w
77.
▲
by
ddp26
6mo ago
Yeah, I might have stated this poorly. In the forecast it's just a question of expected value, I don't give almost any probability to "Starship is worthless". My 50% CI on Starship's fair market value at IPO time is
78.
▲
by
ddp26
6mo ago
I read your comment as being glib, but in forecasting this I was really puzzled how much to anchor to how analysts tend to value these businesses. I ended up largely deferring to them, e.g. predicting the public will value xAI at $258 billi
79.
▲
by
ddp26
6mo ago
Yeah, it's wild. But it's not like the P/E should be 30, what do you think would be fair? That's the thing about SpaceX, some businesses are real businesses that can be modeled in normal ways, like the government launch
80.
▲
A forecast of the fair market value of SpaceX's businesses
(futuresearch.ai)
100 points
by
ddp26
6mo ago
|
205 comments
81.
▲
by
ddp26
7mo ago
Yeah, but uvx has this thing where it can automatically build the latest environment, and pull the latest (unpinned) version, right?
82.
▲
by
ddp26
7mo ago
My team was making fun of me for starting all my chats with "Hi Claude"
83.
▲
by
ddp26
7mo ago
Yeah, sharing information across Claude Code sessions really is a problem that needs solving. An urgent hack, where you're using Claude Code to debug and trying to get help from your team, is one such case.
84.
▲
by
ddp26
7mo ago
Yeah, and this is a pattern I saw in the Fancy Bear Goes Fishing book, a lot of discovery of malware is either pure luck, or blunders from the malware developers. https://en.wikipedia.org/wiki/Fancy_Bear_Goes_Phishing
85.
▲
by
ddp26
7mo ago
Agree, lots of hand wringing about us being so vulnerable to supply chain attacks, but this was handled pretty well all things considered
86.
▲
by
ddp26
7mo ago
Sure, but this is a pretty onerous restriction. Do you think supply chain attacks will just get worse? I'm thinking that defensive measures will get better rapidly (especially after this hack)
87.
▲
by
ddp26
7mo ago
Yeah, this was my team at FutureSearch that had the lucky experience of being first to hit this, before the malware was disclosed. One thing not in that writeup is that very little action was needed for my engineer to get pwnd. uvx automati
88.
▲
by
ddp26
7mo ago
I think so, and I've seen other solutions too. The one in the OP is more general, as you say. Have you tried Code Mode?
89.
▲
Ask LLM Agents to Classify Problems Before Starting
(futuresearch.ai)
7 points
by
ddp26
7mo ago
|
1 comments
90.
▲
by
ddp26
7mo ago
I don't understand the CLI vs MCP. In cli's like Claude Code, MCPs give a lot of additional functionality, such as status polling that is hard to get right with raw documentation on what APIs to call.
More ›