4 ms·
API only model, yet trying to compete with only open models in their benchmark image. Of course it'd be a complete embarrassment to see how hard it gets trounc
by Jackson__ 2y ago
API only model, yet trying to compete with only open models in their benchmark image.
Of course it'd be a complete embarrassment to see how hard it gets trounced by GPT4o and Claude 3.5, but that's par for the course if you don't want to release model weights, at least in my opinion.
- GaggiX 2y agoYes, I agree, for these small models it's wasted potential to be closed source, they can only be used effectively if they are open. EDIT: HN is rate-limiting me so I will reply here: In my opinion 1B and 3B truly shine on edge devices, if not than it's not worth the effort, you can have much better models for already dirt cheap using an API.
- zozbot234 2y agoThere are small proprietary models such as Claude Haiku and GPT 4o-mini.
- GaggiX 2y agoThey are way bigger than 1B or 3B.
- k__ 2y agoWhile I'm all for open models; why can't the small models not be used effectively? Wouldn't they lower the costs compared to big models drastically?
- Bilal_io 2y agoI think what the parent means is that small models are more useful locally on mobile, IoT devices etc. so it defeats the purpose to have to call an API.
- echelon 2y agoThese aren't the "small" models I'm thinking of. I want an LLM, STT, or TTS model to run efficiently on a Raspberry Pi with no GPU and no network. There is huge opportunity for LLM-based toys, tools, sensors, and the like. But they need to work sans internet.
- thebiss 2y agoYou may be interested in this tread regarding whisper.cpp on an Rpi4: https://github.com/ggerganov/whisper.cpp/discussions/166 https://github.com/ggerganov/whisper.cpp/discussions/166
- derefr 2y agoBig models take up more VRAM just to have the weights sitting around hot in memory, yes. But running two concurrent inferences on the same hot model, doesn't require that you have two full copies of the model in memory. You only need two full copies of the model's "state" (the vector that serves as the output of layer N and the input of layer N+1, and the pool of active low-cardinality matrix-temporaries used to batchwise-compute that vector.) It's just like spawning two copies of the same program, doesn't require that you have two copies of the program's text and data sections sitting in your physical RAM (as those get mmap'ed to the same shared physical RAM); it only requires that each process have its own copy of the program's writable globals (bss section), and have its own stack and heap. Which means there are economies of scale here. It is increasingly less expensive (in OpEx-per-inference-call terms) to run larger models, as your call concurrency goes up. Which doesn't matter to individuals just doing one thing at a time; but it does matter to Inference-as-a-Service providers, as they can arbitrarily "pack" many concurrent inference requests from many users, onto the nodes of their GPU cluster, to optimize OpEx-per-inference-call. This is the whole reason Inference-aaS providers have high valuations: these economies of scale make Inference-aaS a good business model. The same query, run in some inference cloud rather than on your device, will always achieve a higher-quality result for the same marginal cost [in watts per FLOP, and in wall-clock time]; and/or a same-quality result for a lower marginal cost.) Further, one major difference between CPU processes and model inference on a GPU, is that each inference step of a model is always computing an entirely-new state; and so compute (which you can think of as "number of compute cores reserved" x "amount of time they're reserved") scales in proportion to the state size. And, in fact, with current Transformer-architecture models, compute scales quadratically with state size. For both of these reasons, you want to design models to minimize 1. absolute state size overhead, and 2. state size growth in proportion to input size. The desire to minimize absolute state-size overhead, is why you see Inference-as-a-Service providers training such large versions of their models (OpenAI's 405b models, etc.) The hosted Inference-aaS providers aren't just attempting to make their models "smarter"; they're also attempting to trade off "state size" for "model size." (If you're familiar with information theory: they're attempting to make a "smart compressor" that minimizes the message-length of the compressed message [i.e. the state] by increasing the information embedded in the compressor itself [i.e. the model.]) And this seems to work! These bigger models can do more with less state, thereby allowing many more "cheap" inferences to run on single nodes. The particular newly-released model under discussion in this comments section, also has much slower state-size (and so compute) growth in proportion to its input size. Which means that there's even more of an economy-of-scale in running nodes with the larger versions of this model; and therefore much less of a reason to care about smaller versions of this model.
- lumost 2y agoAn open small model means I can experiment with it. I can put it on an edge device and scale to billions of users, I can use it with private resources that I can't send externally. When it's behind an API its just a standard margin/speed/cost discussion.
- Jackson__ 2y agoI'd also like to point out that they omit Qwen2.5 14B from the benchmark because it doesn't fit their narrative(MMLU Pro score of 63.7[0]). This kind of listing-only-models-you-beat feels extremely shady to me. [0] https://qwenlm.github.io/blog/qwen2.5/ https://qwenlm.github.io/blog/qwen2.5/