3 ms·
evidence that openai trained on the data: they would have denied it if they didn't train on it. did the proof build on the insights: the influence of an indiv
by lukewarm707 25d ago
evidence that openai trained on the data: they would have denied it if they didn't train on it.
did the proof build on the insights:
the influence of an individual text in the training data is deeply weighted by quality, relevance, etc. a high quality proof in advanced mathematics written by a codex user is going to get boosted to the max.
the model is post-trained on prompt material. that is again going to boost it.
the prompt will boost this material specifically. perhaps they even rammed dense maths in particular into the model in post training.
anecdotally i have been able to get near-verbatim copies of original material out of models at inference. the type of work that buckmaster and alpoge fed into openai feels like the exact type of concept that would cause an "aha!" or "but what if?" in chain of thought. in fact i would bet that their work is in the logs.
the likes of astra and fable are thought to be up to 10T parameters in size. i consider it highly plausible that a semantic representation of the euler proof could be pulled out of the model weights in good shape.