11 ms·
I just ran a local Qwen 3.6 on cheap local hardware ( 12GB VRAM ) and it performed BETTER ( okay, 10x slower, but acceptable ) than Copilot. Big AI is DEAD.
by damnitbuilds 5mo ago
I just ran a local Qwen 3.6 on cheap local hardware ( 12GB VRAM ) and it performed BETTER ( okay, 10x slower, but acceptable ) than Copilot.
Big AI is DEAD.
- newaccountman2 5mo ago10x slower is a deal breaker lol. I think max I could tolerate personally is 2x slower for same quality.
- damnitbuilds 5mo agoWorking on it ! Well, other people are working on it, I'm just trying to get what they do running on my system. MTP looks good at the moment.
- newaccountman2 5mo agoMTP?
- damnitbuilds 5mo agoThis explains it ( not sure about the diagram though ): https://www.hardware-corner.net/multi-token-prediction-llm-speed/ https://www.hardware-corner.net/multi-token-prediction-llm-s...
- newaccountman2 5mo agothnx
- _aavaa_ 5mo agoHow many active chats can you have going in parallel?
- damnitbuilds 5mo agoI only did coding with a MoE model, one prompt at a time, which maxed out my GPU.