3 ms·
The research for LLMs has been ongoing for years, and all of these company's contributed bleeding edge research using models they internally trained but did not
by filterfiber 3y ago
The research for LLMs has been ongoing for years, and all of these company's contributed bleeding edge research using models they internally trained but did not release.
For a while it was largely academic and not crazy useful (GPT2ish era).
After this they had internal LLM's that could potentially be used to make products. This is where the companies split in their approaches. Google was overly cautious from what I've heard, they were concerned about "safety". Then openai released GPT3.5 and soon after GPT4. During this time Meta took the approach of just releasing their research model outright without first building it into a product. Google then scrambled to build products from their LLMs.
tl;dr - everyone of those are big players in the ML/AI space and had internally trained models for their research. We just recently crossed the line where these models became useful thus this scramble to release products (or the models themself).
EDIT: It's worth noting that the big breakthrough paper about transformers (Attention is All you Need) was released in 2017. People started scaling up the parameter count, optimizing training time, inference cost, etc. since then.