3 ms·
This should be much more efficient in theory, right? Why dont we see more leading labs adopt this?
by pwmglenn 1mo ago
This should be much more efficient in theory, right? Why dont we see more leading labs adopt this?
- ux266478 1mo agoWeak "chain-of-thought" abilities, high error rates and a very bad ability to recover from errors. If they produce a non-sequitur somewhere (which they are highly prone to doing), that global refinement spreads it everywhere like a blood infection superhighway. Diffusion text models are cool, but they're functionally much less reliable than autoregressive transformers... and man that's really saying something. Right now most research on them is trying to figure out what complementary systems they need to be reasonably useful.
- vikramkr 1mo agoGoogle was at least trying. Wouldn't be surprised if the others were experimenting with it too. The bar is going to be a lot higher now for for any diffusion model to go from experiment to product since it needs to compete with stuff like glm 5.3 flash and Luna on cost/efficiency for a given quality of output, which is not going to be easy. If it was easy Gemini diffusion would have landed - if it requires a bunch of money and effort it has a much higher bar to make it to market, if it requires some clever breakthrough you have no way of knowing where that's going to come from or what it'll look like/if it'll even seem important when it happens