2 ms·
Have you tried Ling-3.0-tiny? It runs fine on CPU-- on a 14700KF gets 40tg/s and 250pp/s and on a ordinary gpu (RTX 4070) does over 200tg/s with no MTP and 718
by nullc 10d ago
Have you tried Ling-3.0-tiny? It runs fine on CPU-- on a 14700KF gets 40tg/s and 250pp/s and on a ordinary gpu (RTX 4070) does over 200tg/s with no MTP and 7185pp/s.
It's certainly not as capable as something that needs a high memory gpu for quick performance, but I was quite impressed with it for what it is.
(and fwiw, I had it translate your last paragraph to German, then used google translate back to english: "I wish they had handled this clearly and transparently via an opt-in mechanism—not enabled by default—that explains what Mistral is (not a major American cloud company, but a relatively small French startup) and that your prompts and LLM activities are sent to their servers. I also wish there were documentation explaining how the data is handled and stored in a way that inspires trust.").
- walrus01 10d agoHow much RAM does it take up in total? I'll have to give that a try on one of my test systems. Looking at a somewhat randomly chose GGUF quantization of it, looks like just under 5GB on disk in Q4, so RAM usage somewhere around 5-6GB? https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF https://huggingface.co/bartowski/Ling-3.0-tiny-GGUF
- nullc 10d agoThat sounds about right for Q4. it's an extremely sparse MOE, so there is some odds of acceptable performance using a smaller in-memory cache and the rest on flash. ... I don't have a setup to test that right now. (Of course, if translation is all you want much smaller models will work. Ling-tiny can do summarization, dom manipulation, scripting, etc. too).
- hugodan 10d agoAndreessen Horowitz led Mistral's €385M Series A in December 2023.
- u8080 9d agoIs that like 2-4 RTX9000 cards price rn?