6 ms·
Anyone tried running on a Mac M1 with 16GB RAM yet? I've never run higher than an 8GB model, but apparently this one is specifically designed to work well with
by pamelafox 1y ago
Anyone tried running on a Mac M1 with 16GB RAM yet? I've never run higher than an 8GB model, but apparently this one is specifically designed to work well with 16 GB of RAM.
- thimabi 1y agoIt works fine, although with a bit more latency than non-local models. However, swap usage goes way beyond what I’m comfortable with, so I’ll continue to use smaller models for the foreseeable future. Hopefully other quantizations of these OpenAI models will be available soon.
- pamelafox 1y agoUpdate: I tried it out. It took about 8 seconds per token, and didn't seem to be using much of my GPU (MPU), but was using a lot of RAM. Not a model that I could use practically on my machine.
- steinvakt2 1y agoDid you run it the best way possible? im no expert, but I understand it can affect inference time greatly (which format/engine is used)
- pamelafox 1y agoI ran it via Ollama, which I assume uses the best way. Screenshot in my post here: https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol3jfb2r https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol3... I'm still wondering why my MPU usage was so low.. maybe Ollama isn't optimized for running it yet?
- wahnfrieden 1y agoMight need to wait on MLX
- turnsout 1y agoTo clarify, this was the 20B model?
- pamelafox 1y agoYep, 20B model, via Ollama: ollama run gpt-oss:20b Screenshot here with Ollama running and asitop in other terminal: https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol3jfb2r https://bsky.app/profile/pamelafox.bsky.social/post/3lvobol3...
- roboyoshi 1y agoM2 with 16GB: It's slow for me. ~13GB RAM usage, not locking up my mac, but took a very long time thinking and slowly outputting tokens.. I'd not consider this usable for everyday usage.