3 ms·
Just for the right order of history here - this prediction disparity is false - GPT-3 came out in 2020 and a decent slew of researchers and startups had access
by authorfly 2y ago
Just for the right order of history here - this prediction disparity is false - GPT-3 came out in 2020 and a decent slew of researchers and startups had access in 2021. Though unreliable, slow and costly, davinci could mimic chats back then (It was a main feature sample), but OpenAI banned chat applications until IT released ChatGPT in Nov 2022.
For a whole year prior to ChatGPT, "davinci-instruct-002" was at the level of ChatGPT for most purposes, but applications using OpenAIs models had to be approved, and chat apps were disallowed (as were any completions > 300 tokens).
The competitors of that time (GPT-J, GPT-20b-neo) and early BigScience stuff was behind OpenAI, especially on instruction training. So we didn't see applications until OpenAI changed it's application approval process in November 2022. A good reference for where things were is "Machine Learning Street Talks" 4 hour 2020 GPT-3 video - although it expressed doubt about GPT-3, you can see there through many of the examples that the tech was close to ChatGPT relative to the previous iteration (GPT-2, T5, BERT etc).
However, for experts, it was obvious from about Spring 2021 when GPT-3 started taking off in the research zeitgeist with actual research usage, with loads of tweets and recognition of what it meant for the field (both in terms of impact to grant applications, ongoing projects and the future of language models being decoder only for the reasonable short term).
The real gap in prediction disparity was that most experts, possibly due to a bias towards their funding areas researching BERT etc, totally ignored GPT-2 and assumed encoder architectures or other fields (RNNs etc) were still better paths. Established Researchers* would have predicted GPT-3 was 10 years away in 2018 or 2019 which is funny in hindsight. However, even in 2013 (when I was not a researcher but a student), people in computer vision felt that arbitrary tasks/arbitrary recognition was nigh impossible unless using coding schemes etc (which limited to certain image types anyway)
* And I say this because I was researching in 2019, but I was more optimistic after seeing T5 and GPT-2. The field did not ignore CLIP, but there wasn't much obvious apparent research to be done with CLIP initially. Computer vision was all about semantic segmentation and recognition of disease etc in the 10s.