4 ms·
The AI Reproducibility Crisis
I've really been struggling of late to replicate recent findings in research that has built on top of GPT-3.5/GPT-4.
This leads me to believe there is growing yet largely unnoticed issue is taking root in recent AI research. I've termed this the "AI Reproducibility Crisis".
The principle is simple, if accessible private models are silently changing in time, then previous results cannot be replicated.
*Key Issues*:
1. Users have reported significant performance shifts post the May release.
2. Beyond community discussions, academic studies are showing differences in performance across time http://arxiv.org/abs/2307.09009.
3. It appears difficult to replicate previous benchmark evals, see our effort here - https://github.com/emrgnt-cmplxty/zero-shot-replication.
4. The centralized approach of major providers amplifies these concerns, underscoring the essential need for research autonomy.
*Proposed Solutions*:
- Lean towards open-source foundational models for transparency.
- Clearly annotate the date of model access when relying on private providers. Push these providers to provide model / inference specifiers to delineate any changes on their end.
- Advocate for continuous third-party benchmarking of LLM providers to monitor model changes. This is an initiative we're currently starting w/ some help from a great academic group.
This extends beyond mere performance dips or changes in GPT. It's a call for long-term transparency, scientific rigor, and sustained progress. Failing to properly address this issue will create major headwinds for the long-term progress in AI research, as our future work will not be able to build upon past work.
- ocolegro 3y ago*Further Reading*: - [GPT-4's decline over time (HackerNews)](https://news.ycombinator.com/item?id=36786407 https://news.ycombinator.com/item?id=36786407) - [GPT-4 downgrade discussions (OpenAI Forums)](https://community.openai.com/t/gpt-4-has-been-severely-downgraded-topic-curation/304946 https://community.openai.com/t/gpt-4-has-been-severely-downg...) - [Behavioral changes in ChatGPT (arXiv)](https://arxiv.org/abs/2307.09009 https://arxiv.org/abs/2307.09009) - [Zero-Shot Replication Effort (Github)](https://github.com/emrgnt-cmplxty/zero-shot-replication https://github.com/emrgnt-cmplxty/zero-shot-replication) - [Inconsistencies in GPT-4 HumanEval (Github)](https://github.com/evalplus/evalplus/issues/15 https://github.com/evalplus/evalplus/issues/15) - [Early experiments with GPT-4 (arXiv)](https://arxiv.org/abs/2303.12712 https://arxiv.org/abs/2303.12712) - [GPT-4 Technical Report (arXiv)](https://arxiv.org/abs/2303.08774 https://arxiv.org/abs/2303.08774)
- ugur2nd 3y agoThis situation happens to me too. I partially solve the problem by writing extra prompts so that it does not recur.
- greatpostman 3y agoI bet they’re A/B testing hundreds or thousands of models
- OhNoNotAgain_99 3y ago[dead]
- deleted 3y ago[deleted]