4 ms·
This is one of the things that make me uncomfortable about proprietary llm. They get task performance by doing a lot more than just feeding a prompt straight t
by shubb 3y ago
This is one of the things that make me uncomfortable about proprietary llm.
They get task performance by doing a lot more than just feeding a prompt straight to an llm, and then we performance compare them to raw local options.
The problem is, as this secret sauce changes, your use case performance is also going to vary in ways that are impossible for you to fix. What if it can do math this month and next month the hidden component that recognizes math problems and feeds them to a real calculator is removed? Now your use case is broken.
Feels like building on sand.
- BoorishBears 3y agoI'm not sure you realize how proprietary LLMs are being built on. No one is doing secret math in the backend people are building on. The OpenAI API allows you to call functions now, but even that is just a formalized way of passing tokens into the "raw LLM". All the features in the comment you replied to only apply to the web interface, and here you're being given an open interface you can introspect.
- edgyquant 3y agoIt was a contrived example to make a point, one that seems to have flown over your head.
- BoorishBears 3y agoNo it was a bad (straight up wrong) example because you don't understand how people are building applications on proprietary LLMs. If you did you'd also know what evals are.
- shubb 3y agoThank you for pointing that out - I had assumed that things were not how they are. Although performance has varied over time https://arxiv.org/pdf/2307.09009.pdf https://arxiv.org/pdf/2307.09009.pdf I also notice that the API allows you to use a frozen version of the model which avoids the worries I mentioned.
- BoorishBears 3y agoThat was a pretty deeply flawed paper, one of the largest drops recorded was simple parsing errors in their testing: https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-time https://www.aisnakeoil.com/p/is-gpt-4-getting-worse-over-tim... Overall evals and pinning against checkpoints are how you avoid those worries, but in general, if you solve a problem robustly, it's going to be rare for changes in the LLM to suddenly break what you're doing. Investing in handling a wide range of inputs gracefully also pays off on handling changes to the underlying model.
- rightbyte 3y ago> No one is doing secret math in the backend people are building on. How do you know that? With SaaS you are at the mercy of the vendor.