4 ms·
Any diffusion model is potentially a Jev in disguise: https://github.com/vllm-project/vllm/pull/57250 https://github.com/vllm-project/vllm/pull/57250 Runs ~0.2
by mmastrac 18d ago
Any diffusion model is potentially a Jev in disguise: https://github.com/vllm-project/vllm/pull/57250 https://github.com/vllm-project/vllm/pull/57250
Runs ~0.2s per decision on my DGX Spark.
10/10 programming language detection
9/10 human language detection
10/12 unit magnitude comparison
All incorrect answers are marked with low-P.
It (DiffusionGemma with the Jev mode) can also solve an ASCII maze.
- budro 18d agoYour maze demo gave me the idea to try out the "speculative fan-out" pattern [1]. It seemed interesting to try solving mazes in one-shot. Unfortunately it seems like Jev can't reliably solve basic mazes even with a step count of 1! [2] I was very surprised. Could you point me in the direction of your maze solving code so I can see if it's a skill issue? The only other explanation I can come up with is that Jev was not trained on spatial reasoning tasks at all, and on the other hand DiffusionGemma has a vision tower and significantly more spatial data in its training set. [1] https://docs.typesafe.ai/patterns/fan-out https://docs.typesafe.ai/patterns/fan-out [2] https://github.com/Bud-ro/jev-demos/tree/master/packages/maze_lookahead https://github.com/Bud-ro/jev-demos/tree/master/packages/maz...