5 ms·
There's so much bullshit on the internet how do they make sure they're not training on nonsense?
by breakyerself 1y ago
There's so much bullshit on the internet how do they make sure they're not training on nonsense?
- bgwalter 1y agoMuch of it is not training. The LLMs fetch webpages for answering current questions, summarize or translate a page at the user's request etc. Any bot that answers daily political questions like Grok has many web accesses per prompt.
- 8organicbits 1y agoIs an AI chatbot fetching a web page to answer a prompt a 'web scraping bot'? If there is a user actively promoting the LLM, isn't it more of a user agent? My mental model, even before LLMs, was that a human being present changes a bot into a user agent. I'm curious if others agree.
- bgwalter 1y agoThe Register calls them "fetchers". They still reproduce the content of the original website without the website gaining anything but additional high load. I'm not sure how many websites are searched and discarded per query. Since it's the remote, proprietary LLM that initiates the search I would hesitate to call them agents. Maybe "fetcher" is the best term.
- ronsor 1y ago> The Register calls them "fetchers". They still reproduce the content of the original website without the website gaining anything but additional high load. So does my browser when I have uBlock Origin enabled.
- danaris 1y agoBut they're (generally speaking) not being asked for the contents of one specific webpage, fetching that, and summarizing it for the user. They're going out and scraping everything, so that when they're asked a question, they can pull a plausible answer from their dataset and summarize the page they found it on. Even the ones that actively go out and search/scrape in response to queries aren't just scraping a single site. At best, they're scraping some subset of the entire internet that they have tagged as being somehow related to the query. So even if what they present to the user is a summary of a single webpage, that is rarely going to be the product of a single request to that single webpage. That request is going to be just one of many, most of which are entirely fruitless for that specific query: purely extra load for their servers, with no gain whatsoever.
- snowwrestler 1y agoWhile it’s true that chatbots fetch information from websites in response to requests, the load from those requests is tiny compared to the volume of requests indexing content to build training corpuses. The reason is that user requests are similar to other web traffic because they reflect user interest. So those requests will mostly hit content that is already popular, and therefore well-cached. Corpus-building crawlers do not reflect current user interest and try to hit every URL available. As a result these hit URLs that are mostly uncached. That is a much heavier load.
- shikon7 1y agoBut surely there aren't thousands of new corpuses built every minute.
- bgwalter 1y agoWhy would the Register point out Meta and OpenAI as the worst offenders? I'm sure they do not continuously build new corpuses every day. It is probably the search function, as mentioned in the top comments.
- snowwrestler 1y agoIt says in the first sentence of the article that it is 80% bots (crawlers) and only 20% fetchers. Of course they are crawling every day to improve their training data. The goal is LLMs that know everything, but “everything” changes on a daily basis. Meta and OpenAI are simply the largest after Google, but Google has had ~20 more years to learn how to politely operate crawlers at full-Internet scale.
- prasadjoglekar 1y agoBy paying a pretty penny for non bullshit data (Scale Ai). That along with Nvidia are the shovels in this gold rush.
- danny_codes 1y agoMaking a lot of assumptions about the quality of scale AI.
- danaris 1y agoI mean...they don't. That's part of the problem with "AI answers" and such.