2 ms·
I've done pretty decent local prose->json extraction using Qwen and Phi and Gemma. I'm sure most of it comes down to prompts, and all of them run over 100tps o
by jermaustin1 29d ago
I've done pretty decent local prose->json extraction using Qwen and Phi and Gemma.
I'm sure most of it comes down to prompts, and all of them run over 100tps on a 3090. Smaller cards will likely be slower, but Qwen3.5 9B is small enough to fit on most consumer cards.