13 ms·
Which ones are you claiming have already been achieved? My understanding of the current scorecard is that he's still technically correct, though I agree with y
by merlincorey 9mo ago
Which ones are you claiming have already been achieved?
My understanding of the current scorecard is that he's still technically correct, though I agree with you there is velocity heading towards some of these things being proven wrong by 2029.
For example, in the recent thread about LLMs and solving an Erdos problem I remember reading in the comments that it was confirmed there were multiple LLMs involved as well as an expert mathematician who was deciding what context to shuttle between them and helping formulate things.
Similarly, I've not yet heard of any non-expert Software Engineers creating 10,000+ lines of non-glue code that is bug-free. Even expert Engineers at Cloud Flare failed to create a bug-free OAuth library with Claude at the helm because some things are just extremely difficult to create without bugs even with experts in the loop.
- stingrae 9mo ago1 and 2 have been achieved. 4 is close, the interface needs some work to allow nontechnical people use it. (claude code)
- fxtentacle 9mo agoI strongly disagree. I’ve yet to find an AI that can reliably summarise emails, let alone understand nuance or sarcasm. And I just asked ChatGPT 5.2 to describe an Instagram image. It didn’t even get the easily OCR-able text correct. Plus it completely failed to mention anything sports or stadium related. But it was looking at a cliche baseball photo taken by an fan inside the stadium.
- protocolture 9mo agoI have had ChatGPT read text in an image, give me a 100% accurate result, and then claim not to have the ability and to have guessed the previous result when I ask it to do it again.
- pixl97 9mo ago>let alone understand nuance or sarcasm I'm still trying to find humans that do this reliably too. To add on, 5.2 seems to be kind of lazy when reading text in images by default. Feeding it an image it may give the first word or so. But coming back with a prompt 'read all the text in the image' makes it do a better job. With one in particular that I tested I thought it was hallucinating some of the words, but there was a picture in the picture with small words it saw I missed the first time. I think a lot of AI capabilities are kind of munged to end users because they limit how much GPU is used.
- atomic_reed 9mo ago[dead]
- falloutx 9mo agoI dispute 1 & 2 more than 4. 1) Is it actually watching a movie frame by frame or just searching about it and then giving you the answer? 2) Again can it handle very long novels, context windows are limited and it can easily miss something. Where is the proof for this? 4 is probably solved 4) This is more on predictor because this is easy to game. you can create some gibberish code with LLM today that is 10k lines long without issues. Even a non-technical user can do
- CjHuber 9mo agoI think all of those are terrible indicators, 1 and 2 for example only measure how well LLMs can handle long context sizes. If a movie or novel is famous the training data is already full of commentary and interpretations of them. If its something not in the training data, well I don't know many movies or books that use only motives that no other piece of content before them used, so interpreting based on what is similar in the training data still produces good results. EDIT: With 1 I meant using a transcript of the Audio Description of the movie. If he really meant watch a movie I'd say thats even sillier because well of course we could get another Agent to first generate the Audio Description, which definitely is possible currently.
- zdragnar 9mo agoJust yesterday I saw an article about a police station's AI body cam summarizer mistakenly claim that a police officer turned into a frog during a call. What actually happened was that the cartoon "princess and the frog" was playing in the background. Sure, another model might have gotten it right, but I think the prediction was made less in the sense of "this will happen at least once" and more of "this will not be an uncommon capability". When the quality is this low (or variable depending on model) I'm not too sure I'd qualify it as a larger issue than mere context size.
- CjHuber 9mo agoMy point was not that those video to text models are good like they are used for example in that case, but more generally I was referring to that list of indicators. Like surely when analysing a movie it is alright if some things are misunderstood by it, especially as the amount of misunderstanding can be decreased a lot. That AI body camera surely is optimized on speed and inference cost. but if you give an agent 10 1s images along with the transcript of that period and the full prior transcript, and give it reasoning capabilities, it would take almost endlessy for that movie to process but the result surely will be much better than the body cameras. After all the indicator talks about "AI" in general so judge a model not optimized for capability but something else to measure on that indicator
- deleted 9mo ago[deleted]
- bspammer 9mo agoThe bug-free code one feels unfalsifiable to me. How do you prove that 10,000 lines of code is bug-free, and then there's a million caveats about what a bug actually is and how we define one. The second claim about novels seems obviously achieved to me. I just pasted a random obscure novel from project gutenberg into a file and asked claude questions about the characters, and then asked about the motivations of a random side-character. It gave a good answer, I'd recommend trying it yourself.
- verse 9mo agoI agree with you but I'd point out that unless you've read the book it's difficult to know if the answer you got was accurate or it just kinda made it up. In my experience it makes stuff up. Like, it behaves as if any answer is better than no answer.
- evrydayhustling 9mo agoSo do humans asked to answer tests. The appropriate thing is to compare to human performance at the same task. At most of these comprehension tasks, AI is already superhuman (in part because Gary picked scaled tasks that humans are surprisingly bad at).
- rafaelmn 9mo agoYou can't really compare to human performance because the failure modes and performance characteristics are so different. In some instances you'll get results that are shockingly good (and in no time), in others you'll have a grueling experience going in circles over fundamental reasoning, where you'd probably fire any person on the spot for having that kind of a discussion chain. And there's no learning between sessions or subject area mastery - results on the same topic can vary within same session (with relevant context included). So if something is superhuman and subhuman a large percentage of time but there's no good way of telling which you'll get or how - the result isn't the average if you're trying to use the tool.
- retrac 9mo ago