6 ms·
This is interesting with the recent work showing that models using different architectures, but the same dataset converge to similar capabilities. It makes the
by TrueDuality 3y ago
This is interesting with the recent work showing that models using different architectures, but the same dataset converge to similar capabilities. It makes the case that the data set itself and the compute over it is the secret sauce for these models.
It'll be especially interesting with the ruling indicating that the output of these models isn't copyrightable as that and internally generated (from scratch) data will be the only bits omitted from this reporting. I'll be curious if the gap gets bridged on including the disclosure of data generated by a model containing copyrighted data.
- Zetobal 3y agoThe dataset is really 75% of it. Just look how runway "fixed" stable diffusion 1.4.
- visarga 3y ago> This is interesting with the recent work showing that models using different architectures, but the same dataset converge to similar capabilities. It makes the case that the data set itself and the compute over it is the secret sauce for these models. Yes, I go as far as putting 99% the merit on the dataset, given a compute budget. Humans with similar cultural exposure are also remarkably close in intelligence, even though brains are very different at micro level. If language data is actually the source of most of our acquired intelligence, if intelligence is a collective process, if it has an evolutionary drive, then isn't it silly to discuss so much about models or brains while forgetting about its crystallized form - language and text. It is a repository of past experience spanning millennia, and a self replication medium for ideas. All our important knowledge is encoded in language. We (and more recently LLMs) draw heavily on this recorded experience. It cost a lot to be earned in the first place. If we lost all this knowledge, it would take us millennia to recover. It's smarter than all of us.