4 ms·
That's not how the scaling laws work. The number of samples required to reach a given quality level reduces exponentially over time. Most researchers use small
by NavinF 2y ago
That's not how the scaling laws work. The number of samples required to reach a given quality level reduces exponentially over time. Most researchers use small datasets.
- dangus 2y agoInteresting, because in The New York Times' lawsuit there is a very large block of text repeated verbatim. Page 30: https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20... How much of a copyrighted work do I have to copy and reproduce/redistribute to violate copyright law? Am I allowed to sell my handheld recording of the last two minutes of Gladiator 2 for $1.99 at the flea market?
- NavinF 2y agoThat has nothing to do with my comment. My point is that generative ai models can achieve good perf without using large datasets
- dangus 2y agoIn theory. In practice, it’s a plagiarism machine. If they didn’t need the large dataset why did they use it? And honestly I think you need to provide more proof of your assertion since it’s so far off of the status quo lived experience of AI. I ask ChatGPT about product specifications and other specific search engine-like queries like that and I get accurate answers. If there’s no big dataset where is that information coming from? Nothing is being generated, I’m getting regurgitated product specs and reviews when I ask for them.