3 ms·
The article addresses this point with the following: > It would of course be possible to sic Chatty-Jeeps on the raw markup and have it extract all of this stu
by peterlk 2y ago
The article addresses this point with the following:
> It would of course be possible to sic Chatty-Jeeps on the raw markup and have it extract all of this stuff automatically. But there are some good reasons why not.
>
> The first is that large language models (LLMs) routinely get stuff wrong. If you want bots to get it right, provide the metadata to ensure that they do.
>
> The second is that requiring an LLM to read the web is throughly disproportionate and exclusionary. Everyone parsing the web would need to be paying for pricy GPU time to parse out the meaning of the web. It would feel bizarre if "technological progress" meant that fat GPUs were required for computers to read web pages.
- tsimionescu 2y agoThe first point is moot, because human annotation would also have some amount of error, either through mistakes (interns being paid nothing to add it) or maliciously (SEO). Plus, human annotation would be multi-lingual, which leads to a host of other problems that LLMs don't have to the same extent. The second point is silly, because there is no reason for everyone to train their own LLMs on the raw web. You'd have a few companies or projects that handle the LLM training, and everyone else uses those LLMs. I'm not a big fan of LLMs, and not even a big believer in their future, but I still think they have a much better chance of being useful for these types of tasks than the semantic web. Semantic web is a dead idea, people should really allow it to rest.
- tossandthrow 2y agoWhile both of these points a valid today they are likely going to be invalidated going forward - assume that what you can conceive is technically possible will become technically possible. In 5 years resource price is likely negligible and accuracy is high enough that you just trust it.