5 ms·
I am having trouble understanding what the complaint is here. The docs still mention bigger models with 128k tokens and smaller models with 8k tokens. It seems
by azeemba 2y ago
I am having trouble understanding what the complaint is here.
The docs still mention bigger models with 128k tokens and smaller models with 8k tokens. It seems reasonable to optimize for big and small use cases differently? I don't see how we are being "robbed".
- xena 2y agoThe main limit is that you can have 128k tokens of input, but only 4k tokens of output per run. gpt4-32k lets you have up to 32k tokens of output per run. Some applications need that much output. Especially for token dense things like code and JSON.
- seeknotfind 2y agoSo it's a price concern? Because you could run for 4k output 8 times to get 32K? Or does the RLHF stuff prevent you from feeding the output back in as more input and still get a decent result? The underlying transformers shouldn't care because they'll be doing that already effectively.
- xena 2y agoI'd say it's less a price concern and more a consistency of output concern. It doesn't make much sense to continue incomplete JSON like that I don't think. I need to do some more research.
- peab 2y agoyou can just feed that output into another call, to have the next call continue it, since you have more than 28k extra context. The output per token is faster anyways right, so speed isn't an issue. It's just slightly more dev work (really only a couple lines of code)
- DelightOne 2y agoHow do you know it will have the same state of mind? And how much does that cost.
- jhgg 2y agoBecause the state of mind is derived from the input tokens.
- DelightOne 2y agoIs there a study or anything that that is guaranteed adding an incomplete assistant: response as the input and the API taking off exactly the same way on the same position?
- sshumaker 2y agoIt’s how LLMs work - they are effectively recursive at inference time, after each token is sampled, you feed it back in. You will end up with the same model state (not including noise) as if that had been the original input prompt.
- DelightOne 2y agoLLMs sure. My question is whether it is the same in practice for LLMs behind said API. I found no official documentation that we will get exactly the same result as far as I can tell. And no one here touched how high a multiple the cost is, so I assume its pretty high.
- t-writescode 2y ago> I am having trouble understanding what the complaint is here. The appropriate level of due diligence for each LLM model transition is to run your various prompts into the new model and make sure they still produce the correct output; and, if they don't produce good output, to update the prompts so that they continue to produce good output. Just yesterday, I was experimenting with 4o and assumed I could do a flat migration for some work. 4o actually provided worse results - results I explicitly asked to *not* have in my 4 output (and that I didn't have in my 4 output). It's tedious to have to change models after you've already done a proper validation suite against one model. That would be (at least my) complaint. I've even version-stamped the models I use on purpose to avoid surprises.
- hit8run 2y agoAt that point why not use your own hosted open source model that is more reproducible for you?
- t-writescode 2y agoCost of maintaining architecture, cost of complexity of internal infrastructure, knowledge level required for self-hosting, complexity of local one-boxing, and a slew of other reasons.
- xena 2y agoEverything's a tradeoff, but it seems that part of the tradeoffs include access to tools critical for your product to function correctly being taken away without a way to get it back. Maybe that can be an acceptable tradeoff, but I'd personally not like living with that.
- t-writescode 2y agoAt present, my startup isn't making money - in fact, it's not even released. As a result, I'm trying to prioritize getting it out of the door while still being affordable for myself enough to bootstrap it. To do this, I've made, and continue to make tradeoffs. Among many of the tradeoffs I'm currently making that I intend to resolve ASAP is having OpenAI as a single source of failure. I intend to have some of the other hosted solutions as other options for LLM processing. One of the many, many options that will be considered at that time is self-hosting it, as well, as another option. I've already spent more time than I should perfecting various smaller pieces, increasing reliability, etc. Each time I choose perfection, I lose more time, more runway, more potential market share; and, something I've recently had to learn: Each time I lock myself in a previous step to get that step perfect, I miss the lessons I'm about to have to learn in the next stage of the process, including new issues I'll run into that increase the next step's complexity above my initial estimates. Everything is a tradeoff. Choosing to use a commercially available solution with known and relatively set costs while accepting it may slowly change underfoot (while also knowing I have alternatives I can swap to if an emergency comes up that should only take a little bit to transfer to) is one I've made.
- matsemann 2y agoIf you've spent ages fine tuning your prompt/context to have it work for your integration, it's not a given it will work similarly on a model of a different size. Might have to essentially start from scratch.