4 ms·
One of the main factors that makes LLMs popular today is that scaling up the models is a simple and (relatively) inexpensive matter of buying compute capacity a
by RodgerTheGreat 2y ago
One of the main factors that makes LLMs popular today is that scaling up the models is a simple and (relatively) inexpensive matter of buying compute capacity and scraping together more raw text to train them. Without large and highly diverse training datasets to construct base models, LLMs cannot produce even the superficial appearance of good results.
Manually curating "tidy", properly-licensed and verified datasets is immensely more difficult, expensive, and time-consuming than stealing whatever you can find on the open internet. Wolfram Alpha is one of the more successful attempts in that curation-based direction (using good-old-fashioned heuristic techniques instead of opaque ML models), and while it is very useful and contains a great deal of factual information, it does not conjure appealing fantasies of magical capabilities springing up from thin air and hands-off exponential improvement.
- totetsu 2y agoIt’s not unethical if people in positions of privilege and power do it to maintain their rightful position of privilege and power.
- dang 2y agoPlease don't post in the flamewar style to HN. It degrades discussion and we're trying to go in the opposite direction here, to the extent that is possible on the internet. https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- threeseed 2y ago> properly-licensed and verified datasets is immensely more difficult, expensive Arguably the bigger problem is that many of those datasets e.g. WSJ articles are proprietary and can be exclusively licensed like we've seen recently with OpenAI. So we end up with in a situation where competition is simply not possible.
- piva00 2y ago> Arguably the bigger problem is that many of those datasets e.g. WSJ articles are proprietary and can be exclusively licensed like we've seen recently with OpenAI. > So we end up with in a situation where competition is simply not possible. Exactly, and Technofeudalism advances a little more into a new feud. OpenAI is trying to create its moat by shoring up training data, probably attempting to not allow competitors to train on the same datasets they've been licencing, at least for a while. Training data is the only possible moat for LLMs, models seem to be advancing quite well between different companies but as mentioned here a tidy training dataset is the actual gold.
- gessha 2y agoIt's landmines no matter how you approach the problem. If you treat the web as a free-for-all and you scrape freely, you get sued by the content platforms for copyright or term of service violation. If you license the content, you let the highest bidder get the content. No matter what happens, capital wins.
- motohagiography 2y agothe irony is that if large media providers aren't represented in the training sets, my comments on internet forums over the decades will be over-represented, which is kind of great, really.
- pennomi 2y agoRight? I always ask people - Hypothetically if someone created a superintelligent AI that took over the world, wouldn’t you WANT it to share your opinions and morals? Every tiny bit of text you write is a vote in the election of our future AI overlords.