4 ms·
Is AI Progress Real? Four Independent Metrics Show It
- Khaine 3mo agoIt reads like it was written, or heavily edited by AI
- rbuccigrossi 3mo agoThe text is my transcript of the video formatted by Claude and with sources added at the end. I do use deep research (across Gemini, ChatGPT, and Claude) to gather background and ideas, and Claude for editing. I started machine learning and computer vision research back in 1995, studied NLP, and dove into LLMs with GPT-2, so I wouldn't be surprised if my writing has been deeply influenced by AI.
- Khaine 3mo agoAh, makes sense. Its well written and an interesting piece of work, I just noticed hints of 'claudisms' in places.
- deleted 3mo ago[deleted]
- ffaccount2 3mo ago>METR measures the longest software engineering task a frontier model can complete successfully half the time Longest? Yes, they actually use actual wall clock time. I don't think this metric makes any sense. I don't know about the others, I stopped reading at this point.
- rbuccigrossi 3mo agoThe others are an IQ test (TrackingAI), a test of Ph.D. questions across multiple domains (Humanity's Last Exam), and graphical pattern matching (ARC-AGI-2). What's interesting is that while they are rather different in nature (yes it is odd that METR measures clock time as opposed to iterations etc.) but the behavior of the resulting improvement curves are extremely close. That's the punchline: 4 independent measures point to the same conclusion. That increases the chance that the conclusion is correct.
- rbuccigrossi 3mo agoAuthor here. In short there are 4 different metrics: - METR's time horizon - TrackingAI's offline cognitive test - Humanity's Last Exam, and - ARC-AGI-2 that have lasted longer than 2 years (though ARC-AGI-2 is now saturated). When plotted in the linear domain, they all have an exponential (hockey-shaped) curve, but the interesting thing is that the bend happens right at Q4 of 2025 (right when Gemini 3, Opus 4.5, GPT 5.2 all come out).