25 ms·
Any data scraped would be instantly deduplicated after the fact by whatever semantic dedupe engine they've cooked up.
by ShamelessC 2y ago
Any data scraped would be instantly deduplicated after the fact by whatever semantic dedupe engine they've cooked up.
- anonymousDan 2y agoWhat has it got to do with deduplication? I'm talking about crafting some kind of alternative (not necessarily duplicate) data. I agree some kind of post data collection cleaning/filtering of the data before training could potentially catch it. But maybe not!
- ShamelessC 2y agoAh fair enough. The OP here mentioned having highly similar content on each of the many domains.
- trescenzi 2y agoThe funny way to do this would be to use an LLM to generate the content you respond with. Have 2 smallish LLMs talk to each other about topics chosen at random and generate infinite nonsense pages that have a few hundred words.