3 ms·
That's a good point, though i already thought the OpenAI team had been very agressive in sweeping both reddit, usenet, + various illegal megatorrents of books,
by MyFirstSass 2y ago
That's a good point, though i already thought the OpenAI team had been very agressive in sweeping both reddit, usenet, + various illegal megatorrents of books, forum dumps etc. I remember there were some controversy around it on twitter a few months ago.
One thing though is books/content/media from other language spheres though that could probably at least 10x the size of the data, and as far as i know translation starts to work rather well in these larger models so it would probably just plug right into the knowledgegraph for all languages?
- vidarh 2y agoThere are still vast amounts of data locked up behind login screens etc., though. E.g. to the foreign language data, a lot of national libraries around the world are either not even fully digitized yet or have lots of locked-down content. The Norwegian one is pretty open, but there's still huge amounts (like most newspapers newer than a century or so) that is either only available based on geolocation (I have my VPN for genealogy because of that - I'm Norwegian but live in the UK, and it's a nuisance), or only in a physical library in Norway. Similarly I was looking for something from the British Library at one point and it was behind a paywall (a copying fee). I have no idea how to even start to estimate how much data is locked down like that, and it's harder yet to try to figure out which parts of that it'd be possible to negotiate access to for various players, and what they can circumvent (e.g. say by buying book collections and the like - OpenAI is large enough by market cap it could afford to buy some of the largest extant publishers, for example, if they thought it gave them sufficient benefits).
- MyFirstSass 2y agoThat is incredibly interesting to me because i've heard both historians, linguists and people "just not from the anglosphere" complain about just how isolated and limited our language, cultural perspectives are. In other words if LLM's could somehow bridge that gap through both regions and time i'm pretty sure something magical could happen, different than the already a bit tired and conformist "echo chamber" like quality to LLM's mostly trained on reddit, corporate speak, and anglo pop culture, or even just western thought in general.
- vidarh 2y agoIt's been pretty fascinating. ChatGPT clearly has a very small Norwegian corpus, but it can not only translate to and from both Norwegian written languages (they're more like dialects - they're mutually intelligible) but can make a passable attempt at translating into at least some highly localized dialects, and can explain key differences between certain sociolets. And it took only a slight explanation to get it to give a plausible translation into this weird sociolect constructed by the circle around a fringe Maoist group from the 70's that wanted to sound more working class and adopted a bunch of affections that does not match any "natural" Norwegian dialect (several members ended up as prominent authors, and so it spread wider than the size of the group otherwise would have allowed for). That said, English and Norwegian are pretty close. How well it will handle languages with more significant differences without larger amounts of tokens is another matter. Even for pretty small language groups there ought to be enough, though.