4 ms·
I wonder if that would even generate good outcomes given how much of the data in some disciplines is conflicting, and if you don’t have a good way to get it to
by techdragon 4y ago
I wonder if that would even generate good outcomes given how much of the data in some disciplines is conflicting, and if you don’t have a good way to get it to ignore bad subsets of that data it could be more likely to hallucinate convincing scientific sounding explanations based on a preponderance of bad poorly researched evidence.
There’s a lot of papers that discredit older papers or work or entire subfields of study and the current large language models would just all all of it and go “these are words” with no analytical reasoning applied.
These are predictive mechanisms not analytical ones and until we get a better handle on how we can make them “stop and think”… I’m just not sure how much benefit each larger dataset will add beyond “sounds more human” while it’s core problems of “still hallucinates” and “can’t do basic reasoning” remain.
Also one question I’ve got on this front is how much duplicate data is processed out of the input sets? Using LibGen as an example there’s a lot of books that get reprinted and uploaded multiple times in different formats… does having these remain in the data provide a “valuable bias” or is it something that needs removing? Do the people building large language models pre-process their data to de-duplicate out these kinds of things from their current data sources?