4 ms·
That's only because LLMs haven't been a target until now. Search worked great back before everything became algorithmically optimised to high hell. Over time, t
by stuartjohnson12 2y ago
That's only because LLMs haven't been a target until now. Search worked great back before everything became algorithmically optimised to high hell. Over time, the quality of information degrades because as metric manipulation becomes more effective, every quality signal becomes weaker.
Right now, automated knowledge gathering absolutely wipes the floor with automated bias. Cloudflare has an AI blocker which still can't stop residential proxies with suitably configured crawlers. The technology for LLM crawling/training is still mostly unknown, even to engineers, so no SEO wranglers have been able to game training data filters successfully. All LLMs have access to the same dataset - the internet.
Once you:
1. Publicly reveal how training data is pre-processed
2. Roll out a reputation score that makes it hard for bots to operate
3. Begin training on non-public data, such as synthetic datasets
4. Give manipulated data a few more years to accumulate and find its way into training data
It becomes a lot harder.