Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tristanj
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
31.
▲
by
tristanj
16d ago
This form is legally binding and has more legal weight than just clicking a toggle. If they still train on my data, they can get sued, and I'll get a payout.
32.
▲
by
tristanj
16d ago
That OpenAI setting helps, but there is a better way to do it. To completely opt-out of training, submit a request via the OpenAI privacy portal. Visit this website https://privacy.openai.com/policies/en/ , click
33.
▲
by
tristanj
16d ago
This post is out of date. OpenAI quietly updated the references on their paper earlier today and added several authors.
34.
▲
by
tristanj
16d ago
Velocity raptor is a free game that accurately simulates the effects of special relativity. It slows down the speed of light to 3m/s and you solve puzzles as a cute dinosaur. https://www.testtubegames.com/velocityraptor
35.
▲
by
tristanj
17d ago
Yes. Pretty much all the models that don't suck are trained on user data, either directly or via derived synthetic data. Many upstart Chinese labs got around the user data issue by just buying copious amounts of Claude and ChatGPT sess
36.
▲
by
tristanj
17d ago
I've wrote about this elsewhere https://news.ycombinator.com/item?id=49621648 and will reproduce my comment. No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set. First,
37.
▲
by
tristanj
17d ago
To explain who solved what (I copied from here: https://x.com/IlinVasily29521/status/2097554700321329393 ) Tristan + Levent: 3D incompressible Euler with forcing OpenAI: 3D incompressible Euler without forci
38.
▲
by
tristanj
17d ago
OpenAI never asked for the removal of another coauthor. The parent comment is spreading misinformation. OpenAI offered to let Buckmaster to write their Millennium Prize paper, so long as Alpoge (who works at Anthropic) was not a coauthor on
39.
▲
by
tristanj
17d ago
Don't fall for this. This is big tech propaganda, trying to convince you that copyright is bad so they can avoid copyright lawsuits and use everyone's data without paying licensing fees. Without copyright, they can use your data f
40.
▲
by
tristanj
17d ago
I think OpenAI desperately wants to make a blanket denial that they didn't look at or train on Buckmaster and Alpöge's chat transcripts, but know they cannot, because the data is anonymized. The fact they can't make a blanket
41.
▲
by
tristanj
17d ago
OpenAI cannot give a definitive answer here, because it is genuinely unknowable if Buckmaster's data is in the training set. OpenAI explicitly uses user feedback (the thumbs up or thumbs down ratings), as RLHF to train models. However,
42.
▲
by
tristanj
17d ago
No, it’s genuinely impossible to know how much of Buckmaster’s Codex data is in OpenAI’s training set. First, the conversations are anonymized, so there's no simple way to inspect the training dataset and identify which specific conver
43.
▲
by
tristanj
17d ago
An OpenAI employee did say so: https://x.com/tszzl/status/2097393423808377173 it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opte
44.
▲
by
tristanj
17d ago
The rumor going around X was that Anthropic had solved a Millennium Prize problem weeks ago and was sitting on the solution, waiting to release it right before their IPO to maximize hype. If I were at OpenAI, I'd naturally want to snip
45.
▲
by
tristanj
17d ago
I'm not an OAI employee and never pretended to be one are you high?
46.
▲
by
tristanj
17d ago
You're making a classic a burden-of-proof fallacy. The burden of proof lies on the person making the claim, not the person questioning it. See Russell's teapot for an explanation https://en.wikipedia.org/wiki/
47.
▲
by
tristanj
17d ago
Incorrect, Buckmaster and Alpöge can comment if they had the ChatGPT "Improve the model for everyone" setting enabled or disabled. If it was enabled, then their work was included in the training dataset.
48.
▲
by
tristanj
18d ago
The burden of proof is on Buckmaster and Alpöge to reveal if they had the "Improve the model for everyone" setting enabled or disabled. OpenAI shouldn't be expected to reveal private user configuration data. You're askin
49.
▲
by
tristanj
18d ago
The flaw with this line of reasoning is that Buckmaster and Alpöge only had a partially completed proof of a weaker version of the Navier-Stokes problem. OpenAI's internal model solved the full, harder problem. This means the key infor
50.
▲
by
tristanj
18d ago
Because the models are trained on hundreds of billions of user conversations, across more than a billion different humans. The conversations are anonymized and not easily traceable back to a specific user. It's unknowable and not possi
51.
▲
by
tristanj
18d ago
The models are trained on the conversations of hundreds of millions of people. ChatGPT has several billion conversations every day. I estimate that the model that solved Navier–Stokes was trained on data from nearly a trillion conversations
52.
▲
by
tristanj
18d ago
The entire drama is that OpenAI sniped a Millennium Prize Problem from an Anthropic-affiliated research team who had been working on the problem for nearly a year. In just 5 days. I don't think that can be understated.
53.
▲
by
tristanj
18d ago
They have to, otherwise people will accuse the OpenAI model of hacking into people's chat logs and stealing the data there. Which is a claim people are already making.
54.
▲
by
tristanj
18d ago
OpenAI finished another pre-train in late August, and they are now building models off that base. He's saying the specific model OpenAI used to solve this problem is currently in post-training, which started on August 28.
55.
▲
by
tristanj
18d ago
Yes, it's a separate problem. That's a mistake in my post.
56.
▲
by
tristanj
18d ago
Thanks, I wasn't aware of this technicality.
57.
▲
by
tristanj
18d ago
Claiming OAI was going to "totally discredit" Buckmaster is baseless. From the article, it seems OAI wanted to continue discussing the situation with Buckmaster and reach a resolution, but Buckmaster did not want to, declined to r
58.
▲
by
tristanj
18d ago
You're taking the phrase too literally. The point is that knowing a solution is possible gives you the conviction to actually find that solution. The hardest part of solving a problem is often a lack of conviction to see it through, an
59.
▲
by
tristanj
18d ago
I find it amusing when people create brand new accounts just to shitpost fake info. Bc if they posted with their real account, they'd lose karma. Since filing for IPO confidentiality, Anthropic hasn't made any public statements ab
60.
▲
by
tristanj
18d ago
He claims to have found a counterexample for NS (see the second paragraph of the article), but the paper is not ready yet. OpenAI claims to already have a full proof (which they produced in the past 5 days after the rumors leaked). Hence th
More ›