4 ms·
My biggest woah moment wrt transformers was when I saw a paper that took a pretrained Roberta model and made a slight modification to it's embedding layer to fe
by rdedev 3y ago
My biggest woah moment wrt transformers was when I saw a paper that took a pretrained Roberta model and made a slight modification to it's embedding layer to feed in image and audio data and it worked. Granted the performance was probably not on par with actual multi model transformers but the fact that it could figure out how to incorporate multimodal data was not something I ever expected
- uoaei 3y ago> figure out It's statistics. If you can cast your data to a common format (relatively trivial in cases where they are stored as binary representing numerical values) you can learn from patterns. It is not surprising.
- lawrenceyan 3y agoAll information is the same, just represented differently?
- tnecniv 3y agoIn information theory, information is a property of the distribution generating your data. Informally, for a given distribution, information is defined in terms of how much you learn from observing a sample from that distribution on average. If your distribution just puts all the probability mass on the number 3, you learn nothing new by gaining a sample. If your probability mass is really spread out, you gain a lot of information from observing a sample. I haven’t been keeping up with the LLM papers because it’s not really my academic interest but I do find them impressive, so maybe this has been figured out but there are two reasons they could accommodate new data modalities really easily: either the sequential data we generate in the real world is more similar than we would have guessed or the hard part isn’t in handling the domain-specific data, but learning to process and predict future signals really well. The former case would be more surprising to me but it is certainly a possibility — most domains have a “language” of sorts, e.g., visual motifs or licks that get passed between musicians. The network could be picking up on those “linguistic” features born out in data it is fed and just needs to alter its vocabulary from words to pixels or whatever. The second case would be my guess. If you have an algorithm that is good at predicting the future based on the recent past, the hard part is done. The rest is just optimizing it for the task (language, sound, video) at hand.
- Legend2440 3y agoDon't be so dismissive, this was an impossible goal for decades. It's extremely impressive that the same algorithm can handle many different types of data with no changes.
- uoaei 3y agoBut that's the point. If you think they are "different types of data" you are confusing yourself. It's all just bits representing floating-point values. If they're ints you can trivially make them floating-point using routines that have existed since those data formats were invented. If they're discrete values you can encode them in ways that make them legible to the models. One of the more recent interesting developments was token embeddings, admittedly, but again this is just an example of taking slightly more abstract representations and turning them into bits representing floating-point values, which has been an established paradigm known as one of the pillars of "feature engineering" since the beginning of the ML field. One-hot encoding is just a special case of token embeddings. It's amazing that humans figured out how to store data in useful ways, not that models can "figure out" what to do with things they are already capable of ingesting and processing.