3 ms·
Most of the responses here seem to imply that the author doesn't understand that physics can be complicated (in the sense of being hard to learn or having big e
by maxwells-daemon 5y ago
Most of the responses here seem to imply that the author doesn't understand that physics can be complicated (in the sense of being hard to learn or having big equations). He studies theoretical physics at MIT [1], so I expect he does.
On the content: it's pretty weird that our best models don't use much of the world's underlying structure at all. State-of-the-art vision models like vision transformers and MLP-mixer do just fine when you shuffle the pixels. You could argue that modern image datasets are so big that any relevant structure could be learned by attention, but it still feels like we're doing _something_ wrong when pixel order doesn't matter at all ¯\_(ツ)_/¯
[1] https://danintheory.com/ https://danintheory.com/
- deleted 5y ago[deleted]
- erostrate 5y agoVision models need the pixel ordering to match the one they have been trained on, in order to work. They won't generalize after training to transformations of the data that they haven't been trained on, even simple ones such as rotations, whereas humans will. So I would argue that vision model do use the "underlying structure", and even that one of their problems is that they make use of some of the "underlying structures" that are not actually important, such as image luminosity, rotations etc. I think people usually augments the data with these transformations beforehand during preprocessing to enforce invariance.