3 ms·
> Running a remote cloud LLM costs about nothing in dollars to the user. The user doesn't care where it runs, because the user interacts with my product, not m
by usrbinbash 3y ago
> Running a remote cloud LLM costs about nothing in dollars to the user.
The user doesn't care where it runs, because the user interacts with my product, not my backend. He also doesn't pay my cloud provider, he pays me.
I think we don't need to argue the fact that an on premise solution is cheaper than a cloud solution for a lot of tasks, especially when talking about bounded resources. There is some convenience in setup, and some maintenance tasks are easier, but this comes at significant costs, especially as projects get larger.
> Hopefully being based on more than predictions of local LLM performance.
What else should they be based on, pray, given the fact that smaller models suitable for on-premise and even on-machine use are improving rapidly? https://arxiv.org/pdf/2303.16199.pdf https://arxiv.org/pdf/2303.16199.pdf
Are they at the performance levels of cloud based very large LLMs? Not yet. But their turnover times are measured in weeks, not months. And it's not a question if there will be a higher quality base model, only when that will happen.
- yyyk 3y ago>I think we don't need to argue the fact that an on premise solution is cheaper than a cloud solution for a lot of tasks In the absolute sense where we look at the total cost of running the model (and not care how it's distributed or include profits), you may be right even with scale efficiencies - making a determination requires data about cloud server costs we do not have. But I can make an informed guess about the dollar cost the user sees. The cost the user sees is influenced by the factor called 'Microsoft and Google (etc.) have a lot of money, and seem to be perfectly willing to absorb costs to control the market and get user data', and that's enough to get user costs very low when calling to Cloud LLMs. >>>I can easily see scenarious where big tech fail to pivot into AI properly and go down. >>I don't think you have much to base scenarios where big tech goes down upon, but I'll be glad to hear an argument. Hopefully being based on more than predictions of local LLM performance. >What else should they be based on, pray, given the fact that smaller models suitable for on-premise and even on-machine use are improving rapidly? My consistent point is that technical performance is not enough. While the small open models have a tendency to overclaim[0], they'll get to GPT-4 level in time. Success, however, is not determined by technical performance alone. There are some very big hurdles ahead. Why should big tech fall when they pivoted ahead in time, and maintain some very useful moats? [0] https://arxiv.org/abs/2305.15717 https://arxiv.org/abs/2305.15717
- usrbinbash 3y ago> and seem to be perfectly willing to absorb costs to control the market That only works when the competing product incurs a cost that can be undercut, and can be pushed off market in the process. Self-hosten open source LLMs don't incur any cost beyond the utilities. They also cannot be pushed off the market. Trying this tactic would be like trying to replace Linux as the dominant server OS by lowering the licensing costs for Windows Server. > and maintain some very useful moats? https://www.semianalysis.com/p/google-we-have-no-moat-and-neither https://www.semianalysis.com/p/google-we-have-no-moat-and-ne... But even acknowledging the fact that larger models still have advantages in performance, how shall that moat be maintained over time? Even in their current state, smaller models are useful for specialised tasks. And I know I'm repeating myself, but they are also cheaper, work offline, and can be run on a laptop. And it isn't a question if there will be better open source base models, and better training data for RLHF it's only a question of when that happens. To wit, we are still waiting for the 15B and 30B checkpoints of stableLM. And other than with the giant models developed behind closed doors, development turnover for smaller LLMs happens in weeks, not months. Which isn't surprising, because the talent pool open source development can draw from, is basically limitless.
- yyyk 3y ago>That only works when the competing product incurs a cost that can be undercut You're right here, they can't destroy open source (unless they lobby/scare legislators, but that's an entirely different matter). But this subsidization means that the smaller OSS models don't appear cheaper to the client-side users. The companies can absorb the costs - it's a typical data for free service deal, and we already know users can be receptive to these deals. subsidization is also useful to scare off commercial competition. >... The Google 'leak' was a dumb spin and/or an example for why Google Research failed at converting its lead because it doesn't understand business. The important moats are not in raw performance following the initial training runs. That metric is secondary. The moats are in data and access, and both require productization. All of the specialized and local models need access to user data for their task, and BigCorp already has access and data from its products. Lots of telemetry! Everyone else are likely to get a scary user prompt for 'security' which they try to access user data. In the LLM world, data => performance, having better data could mean BigCorp keeps improving beyond OSS. Everything needs to be deployed, and BigCorp can just push it as an OS update. OSS needs word of mouth. BigCorp can aggregate data from multiple users on its remote end for retraining. Local models are likely to be intermittently updated (who's going to pay for that? And based on what data?) and have access only to local user data and what it saw in the original training. OSS catching up to GPT-4 performance will eventually happen, but by that time, BigCorp could achieve a strong product moat and improve its own performance beyond GPT-4. Right now, OSS is behind where it matters, and there's no guarantee this would change. One could hope...