4 ms·
I’m not as inclined to grade Gary Marcus’ predictions leniently as Gary Marcus seems to be. • 7-10 GPT-4 level models He only gets to claim this is true if yo
by mordymoop 2y ago
I’m not as inclined to grade Gary Marcus’ predictions leniently as Gary Marcus seems to be.
• 7-10 GPT-4 level models
He only gets to claim this is true if you count a bunch of different versions of the same base model, or if you’re willing to say that some models that outperform GPT-4 on some benchmarks count as being GPT-4 level. I don’t think Marcus was right in spirit, here.
• No massive advance (no GPT-5, or disappointing GPT-5)
Seems too unquantifiable to judge, I would call o1 a massive advance over 4o, but I’m sure Marcus would not.
• Price wars
I guess so? From what I’ve read, the frontier models companies are still profitable, and OpenAI now has a $200/mo commercial model, hardly the action of a company deciding its prices purely to undercut the competition.
• Very little moat for anyone
It still seems like the only companies who have pulled off frontier model capabilities have spent many millions of dollars doing it. I think this might become true next year but I don’t think this can be judged as correct based on what we saw in 2024 alone.
• No robust solution to hallucinations
You only use words like “robust” in a prediction like this so you have room to weasel out of it later when the hallucinations diminish greatly but don’t quite go extinct.
• Modest lasting corporate adoption
My industry is oil and gas. A pretty hidebound and conservative industry. Adoption of LLMs has been massive.
• Modest profits, split 7-10 ways
Define modest.
I score Marcus at 0/7, at best 2/7.
- fooblaster 2y agoHow is the oil and gas industry gaining value from llms?
- deleted 2y ago[deleted]
- dartos 2y agoI’d bet the same way any office worker gains value from LLMs. Summarizing things they don’t want to read. Generating things they don’t want to write, And RAG tools for surfacing documents, which imo is the killer app for LLMs.
- mordymoop 2y agoThe RAG application gets the most attention and is probably most widely used, but I use it for everything.
- dartos 2y agoYeah. I’m a software engineer, and don’t find much use for LLMs for coding. But when I worked at a big (4k+ employees) org, the RAG service we used was very helpful. Everything from surfacing documentation, to announcements on confluence, to spicy internal blog posts were much easier to find.
- Mistletoe 2y agoI was talking to a kid at our NYE party last night and he is using machine learning at his company to determine which natural gas wells will produce the most.
- dartos 2y agoI think just about everyone company uses machine learning (read: statistical analysis) somewhere tbf. I think colloquially, when people say AI they mean LLMs.
- var_cw 2y agoyeah well his 2024 "predictions" were nothing but regression to mean
- torginus 2y ago> No robust solution to hallucinations ChatGPT/Claude still sucks at providing factual info. You can see how in some areas (like coding), massive amounts of training has taken place, and the model is almost always right when recalling existing knowledge, particularly if you ask about the major frameworks. Then there are somewhat more obscure topics, where the models are mostly right, but can be convincigly wrong, and it's very hard to tell which is which. If you ask about something that wasn't deemed to have economic value, such as book recommendations (I thought a model that was advertised to have been trained on the entirety of human literaly works would be good at that), what you get is almost always obviously wrong info.
- mordymoop 2y agoI agree with what you say here, but I still don’t feel like this gives Marcus the edge here, because Perplexity does a fantastic job at finding correct information and, in my experience, never hallucinates. I would call Perplexity and whatever they’re doing a pretty “robust” solution to hallucination. So there are solutions to hallucination, but the product has to prioritize that to the extent that it reduces the model’s effectiveness as a general tool.
- scotty79 2y ago> if you’re willing to say that some models that outperform GPT-4 on some benchmarks count as being GPT-4 level How else woild you degine being on the same level? You are better at some things and worse at others.