3 ms·
It’s because it’s practically useful and enabled things that were impractical previously.
by Tycho 5d ago
It’s because it’s practically useful and enabled things that were impractical previously.
- petesergeant 5d ago> and enabled things that were impractical previously I think that there are not _that_ many use-cases that have been opened up by this that tool-calling on other models didn't solve already. Really depends what benchmark you're looking at. This one against BANKING77[0] has many issues, but suggests it's really not far off DeepSeek 4.1 Flash. This one against BoolQ[1] shows marginal improvement over Qwen3.6. This one against MMLU-Pro[2] (same author as the previous) shows significant improvements over two Qwen models. So there's definitely _some_ alpha there, but I don't think it's the sea-change that the hype would suggest; that is to say, yes, some things that weren't practical before are now, but many things were already very practical with the existing tools. 0: https://sanand0.github.io/llmevals/jev/ https://sanand0.github.io/llmevals/jev/ 1: https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26d5d7b3b2184a524fc61779dbcc1/evals/results/boolq-2026-09-18/comparison/report.md https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26... 2: https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26d5d7b3b2184a524fc61779dbcc1/evals/results/qwen38-27b/report.md https://github.com/ekzhang/openjev-sglang/blob/a3554ed9e9c26...
- Fabricio20 5d agoThe part about "tool-calling on other models didn't solve already" is what gets you, sure I could tool call deepseek, glm or any other model, but the latency is huge and you get no confidence score. I gave JEV a shot via OpenRouter and it has a reply in less than 400ms, it's fast enough and cheap enough that you can hook it up to a game loop for example (so highly state dependant) and it can do decisions in real time.
- petesergeant 5d agoPlease refer to the latency and price figures for DeepSeek V4.1 Flash in the first link I shared.