3 ms·
FYI you should have used llama.cpp to do the benchmarks. It performs almost 20x faster than ollama for the gpt-oss-120b model. Here are some samples results on
by ggerganov 1y ago
FYI you should have used llama.cpp to do the benchmarks. It performs almost 20x faster than ollama for the gpt-oss-120b model. Here are some samples results on my spark:
ggml_cuda_init: found 1 CUDA devices:
Device 0: NVIDIA GB10, compute capability 12.1, VMM: yes
| model | size | params | backend | ngl | n_ubatch | fa | test | t/s |
| ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | -: | --------------: | -------------------: |
| gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | CUDA | 99 | 2048 | 1 | pp4096 | 3564.31 ± 9.91 |
| gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | CUDA | 99 | 2048 | 1 | tg32 | 53.93 ± 1.71 |
| gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | CUDA | 99 | 2048 | 1 | pp4096 | 1792.32 ± 34.74 |
| gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | CUDA | 99 | 2048 | 1 | tg32 | 38.54 ± 3.10 |
- yvbbrjdr 1y agoI see! Do you know what's causing the slowdown for ollama? They should be using the same backend..
- __mharrison__ 1y agoCurious to how this compares to running on a Mac.
- deleted 1y ago[deleted]
- xs83 1y agoTTFT on a Mac is terrible and only increases as the context increases, thats why many are selling their M3 Ultra 512GB
- Eggpants 1y agoSo so many… eBay search shows only 15 results, 6 of them being ads for new systems… https://www.ebay.com/sch/i.html?_nkw=mac+studio+m3+ultra+512gb+ram&_sacat=0&_from=R40&_oaa=1&Processor=Apple%2520M3&RAM%2520Size=512%2520GB&_dcat=111418&rt=nc&LH_All=1 https://www.ebay.com/sch/i.html?_nkw=mac+studio+m3+ultra+512...
- rajatgupta314 1y agoIs this the full weight model or quantized version? The GGUFs distributed on Hugging Face labeled as MXFP4 quantization have layers that are quantized to int8 (q8_0) instead of bf16 as suggested by OpenAI. Example looking at blk.0.attn_k.weight, it's q8_0 amongst other layers: https://huggingface.co/ggml-org/gpt-oss-20b-GGUF/tree/main?show_file_info=gpt-oss-20b-mxfp4.gguf https://huggingface.co/ggml-org/gpt-oss-20b-GGUF/tree/main?s... Example looking at the same weight on Ollama is BF16: https://ollama.com/library/gpt-oss:20b/blobs/e7b273f96360 https://ollama.com/library/gpt-oss:20b/blobs/e7b273f96360
- xs83 1y agoNow this looks much more interesting! Is the top one input tokens and the second one output tokens? So 38.54 t/s on 120B? Have you tested filling the context too?
- ggerganov 1y agoYes, I provided detailed numbers here: https://github.com/ggml-org/llama.cpp/discussions/16578 https://github.com/ggml-org/llama.cpp/discussions/16578
- nialse 1y agoMakes sense you have one of the boxes. What's your take on it? [Respecting any NDAs/etc/etc of course]