Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
EntityDeletr
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
2 ms
·
1.
▲
by
EntityDeletr
5mo ago
The only model I have seen like that is GPT OSS, natively quantized to MXFP4.
2.
▲
by
EntityDeletr
5mo ago
I would disagree. I have 8 GB of VRAM and 32 GB of RAM. I can either run a 4B BF16 dense model fully on GPU at around 30 t/s or Qwen3.6 35B A3B Q5_K_M at 20 t/s with GPU offload. Which one would I choose?
3.
▲
by
EntityDeletr
5mo ago
Then go with Qwen3.6 35B A3B. It's way faster (up to 5x) and it is 80% as capable as the 27B. The 27B is for serious people looking for one shot coding. The 35B is for iterative and quicker coding. I am in the same situation as you (ma