5 ms·
Stuff like this shows how much better the commercial models are than local models. I’ve been playing around with fairly simple structured information extraction
by beoberha 2y ago
Stuff like this shows how much better the commercial models are than local models. I’ve been playing around with fairly simple structured information extraction from news articles and fail to get any kind of consistent behavior from llama3.1:8b. Claude and chatGPT do exactly what I want without fail.
- minimaxir 2y agoThe Berkeley Function-Calling Leaderboard tracks function calling/structured data performance from multiple models: https://gorilla.cs.berkeley.edu/leaderboard.html https://gorilla.cs.berkeley.edu/leaderboard.html Llama isn't on there but a few finetunes of it (Hermes) are OSS.
- lolinder 2y agoLlama 3 70B is on there, ranked 20.
- thatcat 2y agoI mean, those aren't comparable models. I wonder how the 405b version compares.
- Tiberium 2y agoYou raise a valid point, but 4o is way smaller than 405B. And 4o mini that's described in the article is highly likely <30B (if we're talking dense models).
- A4ET8a8uTh0 2y ago<< Stuff like this shows how much better the commercial models are than local models. I did not reach the same conclusion so I would be curious if you could provide rationale/basis for your assessment in the link. I am playing with humble llama3 8b here and results for federal register type stuff ( without going into details ) was good for what I was expecting to be.. not great. edit: Since you mentioned llama explicitly, could you talk a little about the data/source you are using for your resutls. You got me curious and I want to dig a little deeper.
- 0tfoaij 2y agoOpenAI stopped releasing information about their models after gpt-3, which was 175b, but the leaks and rumours that gpt-4 is an 8x220 billion parameter model are most certainly correct. 4o is likely a distilled 220b model. Other commercial offerings are going to be in the same ballpark. Comparing these to llama 3 8b is like comparing a bicycle or a car to a train or cruise ship when you need to transport a few dozen passengers at best. There are local models in the 70-240b range that are more than capable of competing with commercial offerings if you're willing to look at anything that isn't bleeding edge state of the art.
- Baeocystin 2y agoAny pointers on where we can check the best local models per amount of VRAM available? I only have consumer level cards available, but I would think something that just fits in to a 24Gb card should noticably outperform something scaled for an 8Gb card, yes?
- fnord77 2y agolm studio tells you what models fit in your available RAM, with or without quantization
- kgeist 2y agoIn my tests, Llama 3.1 8b was way worse than Llama 2 13b or Solar 13b.
- gdiamos 2y agoI usually come to a different conclusion using the JSON output on Lamini, e.g. even with Llama 3.2 3B https://lamini-ai.github.io/inference/json_output https://lamini-ai.github.io/inference/json_output Most of these models can read. If the relevant facts are in the prompt, they can almost always extract them correctly. Of course bigger models do better on more complex tasks and reasoning unless you use finetuning or memory tuning.
- dcreater 2y agoYou should probably disclose you're the founder of lamini. Do you have any publicly available validation data demonstrating 100% json compliance?
- gdiamos 2y agoI am a founder. It’s not meant to be a secret. Obviously I’m biased, but I also spend every day using tools like this. Regarding json compliance, we have a formal grammar and a test suite. If you find a bug please report it. I’d appreciate having more test coverage.
- int_19h 2y agoYour problem isn't that you're using a local model. It's that you're using an 8b model. The stuff you're comparing it to is two orders of magnitude larger.
- tpm 2y agoIn my experience the Qwen2-VL models are great at this.