5 ms·
To be fair, this isn't a ladder OpenAI used themselves, right? (The seemingly-extreme shortcut of training from an external LLM model) Their ladder (using publ
by jomohke 3y ago
To be fair, this isn't a ladder OpenAI used themselves, right? (The seemingly-extreme shortcut of training from an external LLM model)
Their ladder (using public data, and hiring humans to classify to taste) is still available I believe.
- matwood 3y ago> Their ladder (using public data, and hiring humans to classify) is still available I believe. Not really. Once chatGPT came out, many sites changed their terms and/or significantly increased their API access costs to prevent/limit/make cost prohibitive future scraping.
- jomohke 3y agoGood point. Though that affects OpenAI too for new data. I had assumed most of their web content was from Common Crawl, and the older pre-ChatGPT Common Crawl datasets used would still be available. But it looks like Twitter, for one, was not in Common Crawl.
- nonethewiser 3y ago> Once chatGPT came out, many sites changed their terms Which is not openAI “pulling up the ladder behind them”
- lelanthran 3y agoNo, it is them spoiling the pitch. There is literally no way for them to avoid looking like assholes once they take enact barriers that they themselves did not have to overcome.
- richardw 3y agoIt’s not like OpenAI trained their model using someone else’s and now won’t allow it done to them. This seems more like saying “get your own content and do the work like everyone else”.
- withinboredom 3y agoExcept they made it nearly impossible to do so.
- moogly 3y ago> using public data As far as we know.