4 ms·
You could not get 500GB of VRAM for less than $100k at the time. It was the absolute best bang for the buck as far as local inference and fine tuning. I am not
by srslack 1mo ago
You could not get 500GB of VRAM for less than $100k at the time. It was the absolute best bang for the buck as far as local inference and fine tuning. I am not sure why you are so hostile when I specifically called out “latency sensitivity” as far as prompt preprocessing and various other bottlenecks. I run Linux.
- ActorNightly 1mo agoThe reason Im so hostile is because it personally irks me when a technical forum like HN is filled with Apple-slop being passed off as technical discussion. Its fine to like Apple for what it is, the ecosystem, the industrial design, the advantages of dedicated hardware with specialized software to extend battery life. But instead, there is Apple-slop, which is basically starting with the assumption that whatever you can do on Mac is the standard, and that Macs are great because they fit that standard. Being able to run frontier models at <10 tok/sec is just an excersize in showing that you can afford a Mac. Its useless for any real tasks. Gemma4 is on par with a lot of the older frontier models like Gemini Flash, and you can run that at 100+ tok/sec on a 2x3090 rig which costs way less than even top of the line Mac Studio.