5 ms·
This looks really promising! Other than this sentence: > We curated a large dataset of videos and languages from public book and video datasets, consisting of
by jerpint 3y ago
This looks really promising!
Other than this sentence:
> We curated a large dataset of videos and languages from public book and video datasets, consisting of videos of diverse activities and long-form books.
I didn’t see any other mention of datasets used, is this on intentional?
- catchnear4321 3y agothe sentence itself should be alarming. the claim is that a dataset was created. of words and of videos, and that it was created from public datasets of books and videos, those datasets containing books, and videos. it takes too many words to say almost nothing. nothing to see here. if that isn’t the intent, then the authors need to do better.
- pk-protect-ai 3y agoThe information is in the model card though: Books3 dataset 700B text-image pairs from Laion-2B-en, filtered to only keep images with at least 256 resolution 400M text-image pairs from COYO-700M, filtered to only keep images with at least 256 resolution 10M text-video pairs from WebVid10M 3M text-video pairs from a subset of InternVid10M 73K text-video chat pairs from Valley-Instruct-73K 100K text-video chat pairs from Video-ChatGPT
- catchnear4321 3y agowouldn’t that have been a cleaner explanation than the sentence provided? books and videos, see model card. the redundant language is a smell whether the emitter wishes to acknowledge or not. the point still stands, the hot mess of a sentence didn’t need to be that way. > …so petty and pedantic… if nothing else, think of the language models that need to digest this. sure you can send in gobbledygook and get out plausibly sense, but why? llms will push pedantry to the forefront. or suffer from it. who knows. have fun.
- lern_too_spel 3y agoYou've never decided to rewrite a sentence and forgot to check the entire sentence again after an incomplete refactoring? I'd say you're in the minority. This is a v1 draft on Arxiv. I don't expect the final paper to have that sentence.
- catchnear4321 3y agohence the feedback.
- brucethemoose2 3y agoThey have at least some of the dataset uploaded: https://huggingface.co/LargeWorldModel https://huggingface.co/LargeWorldModel And the model page specifically mention Books3
- nickpsecurity 3y agoWhile I’m not sure about this one, many AI’s do hide their training data because it’s illegally obtained (ie file sharing of copyrighted works). That’s half of why I dropped AI. The “Proving Wrongdoing” part of my article has specific examples of it: http://gethisword.com/tech/exploringai/ http://gethisword.com/tech/exploringai/
- candiodari 3y agoSame is true for "training data" of most/all humans.
- nickpsecurity 3y agoNo it’s not. Pre-Web, humans were mostly trained by our parents, our schools/colleges, places we go, and things they had access to (eg cable TV). Whether free or paid, they had legal access to that data. It would only be illegal if they started distributing extra copies or doing their own performances of the band. These companies do the very thing that file sharing cases already ruled was illegal. They also scrape all kinds of material whose licenses often say they can’t use it without citations, commercially, etc. The authors asked for some benefit in return for free goods they shared. After not giving them that, the AI suppliers have the nerve to both sell the results and put legal restrictions on them, including terms for sharing. So, they ignore their training suppliers’ legal rights while asserting the same kinds of legal rights for themselves for profit. How humans are trained has nothing to do with AI’s unless you were raised by theft, cons, and hypocrisy. There’s certainly people like that. It says more about the sinful nature of humanity than training AI’s, though.
- candiodari 3y agoHumans produce new works based on their experiences, which is a nice way of saying: "based on others' works they have seen". This is considered original work unless it's too blatantly copied, despite those humans never having a license to create derivative works. In other words it's legally treated as if no other works contributed to it (again, unless it's too blatantly copied) Note: this is law working like this. Not a license, not a contract. Authors do not have any power under copyright to prevent this, nor do they have power to demand something in return. Not even in cases where it damages then, like parodies or reviews destroying a work's appeal/reputation/sales. In practice "blatant" has to be pretty damn blatant. Almost always only exact copies are found to be violating and even then (e.g. Google summaries do not violate copyright despite copying portions of the source material) Hence human works are the same as AI works. Assuming not too blatantly copied, why shouldn't they be treated as original works?
- agnosticmantis 3y agoAccording to [0]: - Books3 dataset - 700B text-image pairs from Laion-2B-en, filtered to only keep images with at least 256 resolution - 400M text-image pairs from COYO-700M, filtered to only keep images with at least 256 resolution - 10M text-video pairs from WebVid10M - 3M text-video pairs from a subset of InternVid10M - 73K text-video chat pairs from Valley-Instruct-73K - 100K text-video chat pairs from Video-ChatGPT 0: https://huggingface.co/LargeWorldModel/LWM-Chat-1M-Jax#training-dataset https://huggingface.co/LargeWorldModel/LWM-Chat-1M-Jax#train...