3 ms·
I question whether a 3B model can have “a lot of knowledge”.
by tomComb 2y ago
I question whether a 3B model can have “a lot of knowledge”.
- foxhop 2y agoMy guess is it uses the same vocabulary size as llama 3.1 which is 128,000 different tokens (words) to support many languages. Parameter count is less of an indicator of fitness than previously thought.
- lolinder 2y agoThat doesn't address the thing they're skeptical about, which is how much knowledge can be encoded in 3B parameters. 3B models are great for text manipulation, but I've found them to be pretty bad at having a broad understanding of pragmatics or any given subject. The larger models encode a lot more than just language in those 70B+ parameters.
- cyanydeez 2y agoOk, but what we are probably debating is knowledge versus wisdom. Like, if I know 1+1 = 2, and I know the numbers 1 through 10, my knowledge is just 11, but my wisdom is infinite in the scope of integer addition. I can find any number, given enough time. I'm pretty sure the AI guys are well aware of which types of models they want to produce. Models that can intake knowledge and intelligently manipulate it would mean general intelligence. Models that can intake knowledge and only produce subsets of it's training data have a use but wouldn't be general intelligence.
- BoorishBears 2y agoI don't think this is right. Usually the problem is much simpler with small models: they have less factual information, period. So they'll do great at manipulating text, like extraction and summarization... but they'll get factual questions wrong. And to add to the concern above, the more coherent the smaller models are, the more likely they very competently tell you wrong information. Without the usual telltale degraded output of a smaller model it might be harder to pick out the inaccuracies.
- wongarsu 2y agoFrom quizzing it a bit it has good knowledge but limited reasoning. For example it will tell you all about the life and death of Ho Chi Minh (and as far as I can verify factual and with more detail than what's in English Wikipedia), but when quizzed whether 2kg of feathers are heavier than 1kg of lead it will get it wrong. Though I wouldn't treat it as a domain expert on anything. For example when I asked about the safety advantages of Rust over Python it oversold Rust a bit and claimed Python had issues it doesn't actually have
- ravetcofx 2y agoI wonder if spelling out the weight would work better. two kilogram for wider token input.
- dotnet00 2y agoIt still confidently said that the feathers were lighter than the lead. It did correct itself when I asked it to check again though.
- apitman 2y ago> it oversold Rust a bit and claimed Python had issues it doesn't actually have So exactly like a human
- fennecfoxy 2y agoWell the feathers heavier than lead thing is definitely somewhere in training data. Imo we should be testing reasoning for these models by presenting things or situations that neither the human or machine has seen or experienced. Think; how often do humans have a truly new experience with no basis on past ones? Very rarely - even learning to ride a bike it could be presumed that it has a link to walking/running and movement in general. Even human "creativity" (much ado about nothing) is creating drama in the AI space...but I find this a super interesting topic as essentially 99.9999% of all human "creativity" is just us rehashing and borrowing heavily from stuff we've seen or encountered in nature. What are elves, dwarves, etc than people with slightly unusual features. Even aliens we create are based on: humans/bipedal, squid/sea creature, dragon/reptile, etc. How often does human creativity really, _really_ come up with something novel? Almost never! Edit: I think my overarching point is that we need to come up with better exercises to test these models, but it's almost impossible for us to do this because most of us are incapable of creating purely novel concepts and ideas. AGI perhaps isn't that far off given that humans have been the stochastic parrots all along.
- ac29 2y agoAs a point of comparison, the Llama 3.2 3B model is 6.5GB. The entirety of English wikipedia text is 19GB (as compressed with an algorithm from 1996, newer compression formats might do better). Its not a perfect comparison and Llama does a lot more than English, but I would say 6.5GB of data can certainly contain a lot of knowledge.