5 ms·
Good article. Agree that general unreliability will continue to be an issue since it's fundamental to how LLMs work. However, it would surprise me if there was
by thorum 1y ago
Good article. Agree that general unreliability will continue to be an issue since it's fundamental to how LLMs work. However, it would surprise me if there was still a significant gap between single-turn and multi-turn performance in 18 months. Judging by improvements in the last few frontier model releases, I think the top AI labs have finally figured out how to train for multi-turn and agentic capabilities (likely RL) and just need to scale this up.
- karn97 1y agoReasoning is just the worst kind of stop gap measure. The state that should emerge internally is forced through automating prompts. And you can clearly see this because the models rarely follow their own "reasoning". Its just auto self prompting
- koakuma-chan 1y agoThey’re reliable enough for many use cases
- bluefirebrand 1y agoWhat this should be doing is exposing how those use cases are faulty, if they can accept such inconsistent and poorly defined outputs