3 ms·
The Flash version is 284B A13B in mixed FP8 / FP4 and the full native precision weights total approximately 154 GB. KV cache is said to take 10% as much space a
by zargon 6mo ago
The Flash version is 284B A13B in mixed FP8 / FP4 and the full native precision weights total approximately 154 GB. KV cache is said to take 10% as much space as V3. This looks very accessible for people running "large" local models. It's a nice follow up to the Gemma 4 and Qwen3.5 small local models.
- sbinnee 6mo agoPrice is appealing to me. I have been using gemini 3 flash mainly for chat. I may give it a try. input: $0.14/$0.28 (whereas gemini $0.5/$3) Does anyone know why output prices have such a big gap?
- girvo 6mo agoOutput is what the compute is used for above all else; costs more hardware time basically than prompt processing (input) which is a lot faster
- tokenmaxxinej 6mo agoinput tokens are processed at 10-50 times the speed of output tokens since you can process then in batches and not one at a time like output tokens
- regularfry 6mo agoI'm going to blow my bandwidth allowance again this month, aren't I.