Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tedsanders
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
tedsanders
1mo ago
Disagree. Examples: - predict a coinflip: easy to verify, hard to learn - earn $100: easy to verify, hard to learn - increase paid subscriptions in an A/B test: easy to verify, hard to learn I won't get into it, but there are many
32.
▲
by
tedsanders
1mo ago
Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
33.
▲
by
tedsanders
1mo ago
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all. ARC is reporting our score on
34.
▲
by
tedsanders
2mo ago
Not sure, to be honest. I don’t think I have any special insight here. My own approach is conversations with friends and family, and the occasional social media post. Exposure and experience are the best teachers, and that’s one reason I’m
35.
▲
by
tedsanders
2mo ago
Sol is not world-endingly dangerous. I work at OpenAI and I've never heard a single person ever come close to claiming that. I think you're bashing a straw man here. One can simultaneously believe: - GPT-5.6 Sol will not end the w
36.
▲
by
tedsanders
2mo ago
Models already display eval awareness, in which they suspect a question is from an eval and then adjust their behavior. E.g., https://www.anthropic.com/engineering/eval-awareness-browsec...
37.
▲
by
tedsanders
2mo ago
Same spirit as above. API is fixed (though of course things like web search results can change from day to day). ChatGPT and Codex harnesses do change over time, though they never result in name changes. We document the big changes here: h
38.
▲
by
tedsanders
2mo ago
I can explain how we do it at OpenAI. In the API, we keep the models fixed. There are tiny caveats like rare bug fixes or models like `chat-latest`, but this is spiritually true. Suspicions of models changing over time are either human hall
39.
▲
Improving GPT‑5.6 Sol in ChatGPT, expanding GPT‑5.6 Luna access for free users
(openai.com)
315 points
by
tedsanders
2mo ago
|
283 comments
40.
▲
Advancing the price-performance frontier with GPT‑5.6
(openai.com)
610 points
by
tedsanders
2mo ago
|
402 comments
41.
▲
Enabling two settings tripled our scores on the ARC-AGI-3 benchmark
(openai.com)
38 points
by
tedsanders
2mo ago
|
5 comments
42.
▲
by
tedsanders
3mo ago
No, they're different models. Knowledge cutoff has been updated.
43.
▲
by
tedsanders
3mo ago
I work at OpenAI and I can assure you that if we say we don't train on your data, we don't. I acknowledge that if you don't trust OpenAI, then you may not trust me either. But lying about this would be bad for legal liability
44.
▲
Codex / ChatGPT Work has reached 8M active users
(twitter.com)
2 points
by
tedsanders
3mo ago
|
0 comments
45.
▲
by
tedsanders
3mo ago
Arena can definitely be benchmaxxed a bit, if you try. The distribution of prompts there is very different than usage by regular coders. E.g., lots of requests for one-shot games from scratch. So if you fine-tuned your model to be great at
46.
▲
by
tedsanders
3mo ago
Not entirely fixed yet, but should be rarer with 5.6. Don’t have a quantification, unfortunately.
47.
▲
by
tedsanders
3mo ago
> Even worse, it's not a fair comparison: they purposefully just used "adaptive" instead of "max" for Fable. We agree models should be compared on a fair basis. Unfortunately, adaptive was the only publicly avail
48.
▲
by
tedsanders
3mo ago
As usual, even though GPT-5.6 is releasing today, the rollout in ChatGPT and Codex will be gradual over many hours so that we can make sure service remains stable for everyone (same as our previous launches). We usually start with Pro/
49.
▲
by
tedsanders
3mo ago
Pointing out problems (e.g., hidden tests that assume narrow implementation details) is much easier than fixing them (e.g., creating tests that work for any possible choice of implementation).
50.
▲
by
tedsanders
3mo ago
I believe what the post meant to communicate is: - alpha testers will start getting access now - everyone will get access Thursday (barring banned countries / individuals) Historically, some companies and individuals have gotten alpha
51.
▲
by
tedsanders
3mo ago
I work at OpenAI and can confirm that's correct: reasoning tokens are discarded after each new user turn (though not after each message or tool call). Our docs show a diagram here: https://developers.openai.com/api/
52.
▲
by
tedsanders
4mo ago
Unfortunately we're not in a position where we can promise an exact date, but we expect it to take weeks (not days or months). It's the best coding model we've ever trained and we're bummed we can't release it to ev
53.
▲
by
tedsanders
4mo ago
Yeah, we'll share a lot more details and evals when we can release GPT-5.6 widely. We focused on cyber (and bio) here to help explain why it's being held back for now. We would have loved to launch it to everyone - it's the b
54.
▲
by
tedsanders
4mo ago
Makes sense, thanks. I suppose error bars are tricky if trying to handle problem-to-problem variance, rubric-to-rubric variance, and run-to-run variance all at once.
55.
▲
by
tedsanders
4mo ago
The nonprofit (OpenAI Foundation) owns ~26% of the for-profit, plus some extra warrants. The for-profit (OpenAI Group PBC) is what's filing the S-1 Draft. The OpenAI Foundation also exclusively appoints the board of the OpenAI Group PB
56.
▲
by
tedsanders
4mo ago
Very cool! So glad to see people building and sharing evals that are better than SWE bench. I'm curious - any particular reason you didn't put error bars on the graphs? Seems like it could be helpful when there are only 50 unique
57.
▲
An OpenAI model has disproved a central conjecture in discrete geometry
(openai.com)
1429 points
by
tedsanders
5mo ago
|
1055 comments
58.
▲
by
tedsanders
5mo ago
What do you mean by this? We don’t train on evals, and if we did I’d quit on the spot. (The loose version of this that’s true is that there may exist eval data contamination in pretraining. This is a hard problem to fully solve.)
59.
▲
by
tedsanders
5mo ago
Thanks - let me clarify that we don’t switch to lightly quantized models by time of day or when under heavy load either. (I used the adjective heavily because that’s what the original post said. I have no intention of making misleading but
60.
▲
by
tedsanders
5mo ago
For what it's worth, I work at OpenAI and I can guarantee you that we don't switch to heavily quantized models or otherwise nerf them when we're under high load. It's true that the product experience can change over time
More ›