3 ms·
People somehow have expectations that are both too high and too low at the same time. They expect (demand) current language models completely replace a human en
by tmnvdb 2y ago
People somehow have expectations that are both too high and too low at the same time. They expect (demand) current language models completely replace a human engineer in any field without making mistakes (this is obviously way too optimistic) while at the same time they are ignoring how rapid the progress has been and how much these models can now do that seemed impossible just 2 years ago, delivering huge value when used well, and they assume no further progress (this seems too pessimistic, even if progres is not guaranteed to continue at the same rate).
- mirsadm 2y agoChatGPT 4 was released 2 years ago. Personally I don't think things have moved on significantly since then.
- yojat661 2y agoExactly. I have been waiting for gpt5 to see the delta, but after gpt4 things seemed to have stalled.
- tmnvdb 2y agoThis seems like a bizarre claim on the surface, see also my other message above. https://epoch.ai/data/ai-benchmarking-dashboard https://epoch.ai/data/ai-benchmarking-dashboard
- tmnvdb 2y agoReally now. I think that deserves a bit more explaination, given the cost per token has dropped by several orders of magnitude, we have seen large changes on all benchmarks (including entirely new capabilities), multimodality is now a fact since 4o, test time compute with reasoning models is making big strides since o1.... It seems on the surface a lot is happening. In fact, I wanted to share one of the benchmark overviews, but none include ChatGPT 4 anymore since it is totally not competitive anymore..
- chrz 2y agoits bigger, shinier, faster, but still doesnt fly
- morsecodist 2y agoBenchmarks are meaningless in and of themselves, they are supposed to be a proxy for usefulness. I have used Sonnet 3.5, ChatGPT-3, ChatGPT-3.5, ChatGPT-4, ChatGPT-4o, o1, o3-mini, o3-mini-high nearly daily for software development. I am not saying AI isn't cool or useful but I am experiencing diminishing returns in model quality (I do appreciate the cost reductions). The sorts of things I can have AI do really haven't changed that much since I got access to my first model. The delta between having no LLM to an LLM feels an order of magnitude bigger at least than the delta between the first LLM and now.