4 ms·
Agreed. I also wonder why they chose to test against a Mac Studio with only 64GB instead of 128GB.
by newman314 1y ago
Agreed. I also wonder why they chose to test against a Mac Studio with only 64GB instead of 128GB.
- yvbbrjdr 1y agoHi, author here. I crowd-sourced the devices for benchmarking from my friends. It just happened that one of my friend has this device.
- ggerganov 1y agoFYI you should have used llama.cpp to do the benchmarks. It performs almost 20x faster than ollama for the gpt-oss-120b model. Here are some samples results on my spark: ggml_cuda_init: found 1 CUDA devices: Device 0: NVIDIA GB10, compute capability 12.1, VMM: yes | model | size | params | backend | ngl | n_ubatch | fa | test | t/s | | ------------------------------ | ---------: | ---------: | ---------- | --: | -------: | -: | --------------: | -------------------: | | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | CUDA | 99 | 2048 | 1 | pp4096 | 3564.31 ± 9.91 | | gpt-oss 20B MXFP4 MoE | 11.27 GiB | 20.91 B | CUDA | 99 | 2048 | 1 | tg32 | 53.93 ± 1.71 | | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | CUDA | 99 | 2048 | 1 | pp4096 | 1792.32 ± 34.74 | | gpt-oss 120B MXFP4 MoE | 59.02 GiB | 116.83 B | CUDA | 99 | 2048 | 1 | tg32 | 38.54 ± 3.10 |
- yvbbrjdr 1y agoI see! Do you know what's causing the slowdown for ollama? They should be using the same backend..
- __mharrison__ 1y agoCurious to how this compares to running on a Mac.
- deleted 1y ago[deleted]
- xs83 1y agoTTFT on a Mac is terrible and only increases as the context increases, thats why many are selling their M3 Ultra 512GB
- Eggpants 1y agoSo so many… eBay search shows only 15 results, 6 of them being ads for new systems… https://www.ebay.com/sch/i.html?_nkw=mac+studio+m3+ultra+512gb+ram&_sacat=0&_from=R40&_oaa=1&Processor=Apple%2520M3&RAM%2520Size=512%2520GB&_dcat=111418&rt=nc&LH_All=1 https://www.ebay.com/sch/i.html?_nkw=mac+studio+m3+ultra+512...
- rajatgupta314 1y agoIs this the full weight model or quantized version? The GGUFs distributed on Hugging Face labeled as MXFP4 quantization have layers that are quantized to int8 (q8_0) instead of bf16 as suggested by OpenAI. Example looking at blk.0.attn_k.weight, it's q8_0 amongst other layers: https://huggingface.co/ggml-org/gpt-oss-20b-GGUF/tree/main?show_file_info=gpt-oss-20b-mxfp4.gguf https://huggingface.co/ggml-org/gpt-oss-20b-GGUF/tree/main?s... Example looking at the same weight on Ollama is BF16: https://ollama.com/library/gpt-oss:20b/blobs/e7b273f96360 https://ollama.com/library/gpt-oss:20b/blobs/e7b273f96360
- xs83 1y agoNow this looks much more interesting! Is the top one input tokens and the second one output tokens? So 38.54 t/s on 120B? Have you tested filling the context too?
- ggerganov 1y agoYes, I provided detailed numbers here: https://github.com/ggml-org/llama.cpp/discussions/16578 https://github.com/ggml-org/llama.cpp/discussions/16578
- nialse 1y agoMakes sense you have one of the boxes. What's your take on it? [Respecting any NDAs/etc/etc of course]