2 ms·
will be great fun if one M5 Ultra with 512GB memory at 1.2T bandwidth capable of doing 3x smallish local model inferencing each at Opus 4.5 level of intelligenc
by tw1984 1mo ago
will be great fun if one M5 Ultra with 512GB memory at 1.2T bandwidth capable of doing 3x smallish local model inferencing each at Opus 4.5 level of intelligence.
- brianwawok 1mo agoThe memory bandwidth and size seems to be there, but what is the tokens per sec on like a qwen model? And you can basically do 3x opus 4.5 on the $100 a month claude plan. Your payback will be near infinity years after electricity.
- tw1984 1mo agoyou are assuming the idea, the data, the code being touched by opus worths nothing. how about measuring returns in the sense that I no longer need to share my flagship idea with some random 3rd party just because it hosts the LLM I am using?