5 ms·
> I don't care who King Charles is every single time Always curious when someone will figure out how we can elide most of the data from an LLM (but retain the
by swiftcoder 2mo ago
> I don't care who King Charles is every single time
Always curious when someone will figure out how we can elide most of the data from an LLM (but retain the logical ability). I don't actually need an LLM to have a very big internal knowledge base to be useful, so long as it can invoke a search tool...
- kanbankaren 2mo ago> Always curious when someone will figure out how we can elide most of the data from an LLM (but retain the logical ability). I don't actually need an LLM to have a very big internal knowledge base to be useful, so long as it can invoke a search tool... I think this can be achieved already. Take a base model and train only on source code. In fact, the very early Granite models from IBM were like that though it didn't support reasoning which limited its performance. You can do it too. I don't know how much it will cost to train on just source code repos. $10K in total? Not sure.
- kccqzy 2mo agoI’m a bit skeptical. Without instructional materials from textbooks, programming language reference manuals and guides, as well as general knowledge about logic and discrete mathematics, I doubt a model could work very well.
- kanbankaren 2mo agoYes. Model would be limited in its performance, but it would perform well in the limited domain because LLM interpolate from training data. They don't think like humans. We might think that knowledge from logc and discrete math would spill over to coding. Unfortunately, it doesn't seem to work like that. Even 1T parameter LLM fail on tasks if there are no variants of it in the training data.
- tyromaniac 2mo agoI'm skeptical that the "logical ability" is much more then the elided data. Obviously some things get fully memorized and other things don't, but I don't think there's anything like functional circuits.
- plandis 2mo agoIsn’t this essentially what MoE partially solves with varying levels of accuracy?
- musebox35 2mo agoSadly no. Despite the name, the experts are not routed per concept or topic but per token. So for the same sentence you might activate multiple experts for different tokens. What it solves is the distributed training and inference problem. As long as each expert fits a single gpu, coordinating the model evaluation is much easier and it is faster. It does not buy as much for running on a single device though still less costly than a dense version.
- rufo 2mo agoApple’s new Foundation model for the 27 OS releases does some interesting things in exactly this area: https://machinelearning.apple.com/research/introducing-third-generation-of-apple-foundation-models https://machinelearning.apple.com/research/introducing-third...
- anentropic 2mo agoI suspect there's a conceptual problem here to what extent is "retain the logical ability" meaningful without attaching it to some knowledge
- swiftcoder 2mo ago> to what extent is "retain the logical ability" meaningful without attaching it to some knowledge Some knowledge is obviously required, I'm just less sure that a specific task like coding benefits all that much from having Shakespeare in the training set...
- apothegm 2mo agoThat’s the thing. LLMs don’t have any logical ability. Only predictive ability. They’re not the same. And that’s why LLMs are a) unreliable and b) not a path to AGI.