4 ms·
Running it in a MacBook Pro entirely locally is possible via Ollama. Even running the full model (680B) is possible distributed across multiple M2 ultras, appar
by keheliya 2y ago
Running it in a MacBook Pro entirely locally is possible via Ollama. Even running the full model (680B) is possible distributed across multiple M2 ultras, apparently: https://x.com/awnihannun/status/1881412271236346233 https://x.com/awnihannun/status/1881412271236346233
- rsanek 2y agothe 70B distilled version that you can run locally is pretty underwhelming though
- vessenes 2y agoThat’s a 3 bit quant. I don’t think there’s a theoretical reason you couldnt run it fp16, but it would be more than two M2 Ultras. 10 or 11 maybe!
- bildung 2y agoWell there's the practical reason of the model natively being fp8 ;) One of the innovative ideas making it so much less compute-intensive, apparently.