4 ms·
M2 macs can do it: https://twitter.com/junrushao/status/1681828325923389440 https://twitter.com/junrushao/status/1681828325923389440 in practice 10 tokens per
by frozenport 2y ago
M2 macs can do it: https://twitter.com/junrushao/status/1681828325923389440 https://twitter.com/junrushao/status/1681828325923389440
in practice 10 tokens per second is kinda annoyingly slow
most local people would opt for a smaller 7b model
- zarzavat 2y agoHave been playing around with Llama3 7b today, it’s not very good. I’m sure that Facebook put everything they could into making it good, but 7B is apparently just not enough parameters.
- pennomi 2y agoI assume you mean 8B? There is no Llama 3 7B.
- sieszpak 2y agoLlama 3 8B seems sad to answer... this is the first model in a long time that has had trouble telling me how much is 3! - (factorial)
- mistrial9 2y agollava-v1.5-7b-q4.llamafile yes agree that the impression is poor overall
- d-z-m 2y agonot very good compared to what? Hard to reconcile your comment with its outsize performance on arena/glowing praise from others comparing it to much larger models.
- zarzavat 2y agoTrying the 8B on translation gives some hilariously garbled results. It’s a small model so perhaps not unexpected but it’s definitely nowhere close to GPT-3. I’d be cautious of synthetic benchmarks, you never know how much of the scores are due to contamination or survivorship bias.
- anon373839 2y agoLanguages other than English are “out of scope” according to the model card, so I wouldn’t expect strong translation performance. In English, though, it’s incredibly capable for its size.
- deleted 2y ago[deleted]
- d-z-m 2y agoI didn't think arena was a synthetic benchmark?
- frozenport 2y agoYeah thats why OpenAI, etc don't run 7B But for scrapping, and tool calling it represents a substantial gain.