3 ms·
We still don’t know what LLMs are for. By this I don’t mean LLMs aren’t useful. We have known that there was novel behavior happening as early as GPT-2. GPT-3
by throwaway9274 3y ago
We still don’t know what LLMs are for.
By this I don’t mean LLMs aren’t useful. We have known that there was novel behavior happening as early as GPT-2.
GPT-3 represented a clearly novel transformative technology. But we still didn’t know what to do with it.
The reason for this is that the people who come up with products are generally a different set of people from those who innovate novel ML models.
The latter tend to be PhDs or the extremely mathematically talented.
The former tend to be a mix of product-y software engineers & engineer-y product & strategy people.
There were some products, GitHub Copilot being the breakaway. Knowledge needs time to diffuse, and a market demand is necessary to catalyze that quickly.
Out of exasperation, OpenAI decided to take one of the most common prompting use cases on the GPT-3 beta playground, Q&A, and make a chat product, almost as a technology demonstrator.
Then ChatGPT exploded to hundreds of millions of users.
All hell broke loose. Every Fortune 500 promised a “generative AI” rollout. And of course, like feverish corporate-branded “metaverse” stuff it all sucks.
You can’t “corporate partnership” and “internally accelerate” a technological shift of this magnitude. You can’t hire BCG to do it for you. And you can’t tack a model onto your existing product and call it done.
LLMs and other new foundation models require fundamentally new products. They require new middleware to
be deployed like vector databases, RAG frameworks, and agentic systems.
Start ups are starting to crack these problems.
But my fear is that when the bottom drops out of the corporate efforts, the investment attitudes will shift just as there’s the most work to be done.
- dboreham 3y ago> We still don’t know what LLMs are for I have found the LLM concept to be tremendously valuable in illuminating my understanding of the operation of human brains (both mine and others). This mainly arose after reading Wolfram's article, then continued independent thinking along the same lines. Honestly by far the biggest coin dropping moment for me in 40 years, since I first heard John Searle talk. Note I don't actually use LLMs for anything.
- janalsncm 3y agoLLMs have been more useful on the encoder side than the decoder side in my experience. Creating embeddings is useful in all sorts of ways. Specifically, if your business involves any sort of recommendation system, embeddings are useful. On the decoder side, the use cases are more subtle. Rarely do you want your product to be the raw output of a statistical language model. More often the output is an enabler of other things your business is doing. For example, you can use it to develop “doc queries”, which are queries that a doc might be surfaced under. This can help with cold start issues by supplementing existing doc info.