3 ms·
it costs 0.001$ per 1K which is slightly cheaper than GPT-3.5-turbo. I have just tested it and it shows extremely worse results on the tasks in my pipelines. No
by behindai 3y ago
it costs 0.001$ per 1K which is slightly cheaper than GPT-3.5-turbo. I have just tested it and it shows extremely worse results on the tasks in my pipelines. Not a game change, unfortunately.
- phillipcarter 3y agoThis seems to be the general pattern so far. A particular benchmark shows better or equal performance for a non-openai model, but then someone else tries a different task and it's just not even close. I think that's really significant. It's arguably a case for fine-tuning models for specific tasks. That's great for teams with ML experience. But for product engineering teams without ML engineers, they can just use a foundational model and get great performance for a low cost.
- ldjkfkdsjnv 3y agoThe research is dishonest, they know they arent showing the full picture
- phillipcarter 3y agoI think even if it's dishonest, it's still representative of a bullish case for fine-tuning. I'm excited to see people get involved in that to make it easier, because...whew, it's not easy.
- waleedk 3y ago[Author] Can you be specific about which part of the research is dishonest? We've shared the code openly, you can reproduce it yourself if you want.
- BoorishBears 3y agoI think in a vacuum it's not dishonest, but it is half-baked in a way a lot of these benchmarks are. Time and time again we see benchmarks like this where 0 effort into actually trying to get anything good out of the model. Of course you can make extremely specific changes that essentially game the test, so to me a fair way to test this is to establish a realistic token budget across all the models, maybe even adjusted for cost depending on the goals of your testing. Then for each model, use your budget to get the highest quality result possible. _ You're asking the GPT models for example, in the worst possible way. You've moved detailed instructions into the user prompt, despite their newest updates focusing on steerability via the system. Meanwhile from tinkering LLAMA 2 pretty much doesn't care about the system prompt vs user prompts. You're also not allowing for any form of chain-of-thought. Giving the models a few hundred tokens to form a conclusion instead of spending those tokens trying to force out a mathematically induced order bias (which I wouldn't expect to do much) would have been much better. There's also no way that GPT 4 should have struggled to give you a well defined output: All of the models would have probably benefited from a well defined output format in terms of a schema, rather than asking for a single letter response. The model rarely has to produce a single letter answer and will struggle with that. Something like asking for { "output" : "A_IS_MORE_FACTUAL" | "B_IS_MORE_FACTUAL" } would have been better Overall I think there's no open dishonesty, and obviously there's no objective correct amount of effort to put into the test setup here... but there's a conflict of interest that would have made me want to see more effort put into it. I think there were trivially low hanging fruit ignored here that I'd expect people selling LLMOps to have pick up on.
- skybrian 3y agoThey didn’t test it on other tasks and yes, it seems there’s no particular reason to believe that the results generalize. If someone wants to try it, they can, but I think it’s based on some unwarranted assumptions that LLM’s are likely to have balanced performance, just because there are some that do seem to be pretty balanced?
- Vetch 3y agoSomething to realize is that different models require different prompting styles. You can't prompt non-gpt4 models with GPT4 tuned stylistic ticks and expect similar results. I've gotten great performance from llama2 derivatives. Out of the box performance is not near GPT4 but it is still very strong in its own right. And, if you are able to break down your problem so precise logit control coupled with guidance from forward or backwards chaining on knowledge graphs is applicable, you can easily exceed gpt4's reasoning ability for your domain. No fine-tuning necessary either. I've been getting useful things out of LLMs since the days of roberta and raw T5, when Large stood for hundreds of millions of parameters. I am flabbergasted when people say a 7B parameter model is no good for them.
- sthatipamala 3y ago> precise logit control coupled with guidance from forward or backwards chaining on knowledge graphs What do you mean by this?
- thewataccount 3y agoI'm not 100% sure they're talking about this specifically but logit control/manipulation is often used to conform to a specific schema. https://github.com/guidance-ai/guidance https://github.com/guidance-ai/guidance I'm going to butcher this explanation - after you've generated your selection of logits but before you sample from them, you check which ones conform to your schema. If you want the only two options to be "true" or "false", then you take any of the logits that would provide invalid answers and lower their probabilities manually. Another example is structures like JSON can be validated so when your sample is "{'name':'Carl'" you lower the probability of "{" since that would invalidate the json. In fact the only valid ones you'd likely have left would be ",", " ", and "}"
- waleedk 3y ago[Author] Can you cite some examples of this so we're not talking in the void? I've actually tested Llama 2 for summarization but haven't blogged about that yet, and across multiple domains, Llama 2 is pretty good. I do see some differences -- Llama 2 doesn't follow instructions as cleanly, and is more verbose. It's not a plugin replacement -- in the article itself I point it out. GPT 3.5 largely followed instructions to return A or B. Llama 2 didn't so I had to use another LLM to post-process. I'm not saying you should always use Llama 2 or ChatGPT. It's that for some use cases, you can save a lot of money by using an open source alternative.
- waleedk 3y ago[Author] Can you share more details? How does it fail? What type of domain? Perhaps there is some tweaking to prompts that's required? It would help to understand what you're seeing.
- 19h 3y agoWhat are those “tasks in your pipeline”? We dedicated an entire team of 6 for two months on evaluating LLMs and while Claude 2 was the final choice, we found Llama 2 70b to be absolutely great (for summaries and structured data generation from free text). We chose Claude because running Llama on an A100 or H100 comes with a baseline cost that doesn’t go to zero when you don’t need it (you could spawn a new instance but right now GPUs are so rare everywhere except for expensive cloud providers that it’s possible you don’t get one). That said, we found the smaller Llama models to be so hilariously bad we have an internal slack channel where Llama 7b writes jokes and the funny part isn’t the jokes but how utterly stupid and random they are.