10 ms·
Lamini Memory Tuning: 10x Fewer Hallucinations
- KeyBoardG 2y ago10x less is a weird way of saying 90% less, or better yet reduced to 10% from before.
- ziptron 2y agoWhat is the hallucination rate of, for example, a Llama3 or GPT4?
- elicksaur 2y agoThey claim 50% with fine-tuning and/or(?) RAG (unclear marketing phrasing imo), and claim their method achieves 5% on the same task which is apparently a text-to-sql task set.
- peter_l_downs 2y agoHas anyone here used this, or anything similar? This sounds phenomenal if it really works. Looks like “contact us” is the only way to try it or buy it right now, and the purported benefit (memorize facts up to the model training size, basically trillions of tokens of facts) is wild. I’d love to try a system running this way to understand the failure modes, like for instance how does it reliably infer which “facts” to use?
- _yb2s 2y ago"Hallucinations" are the creative aspect of LLMs, which is what they are more useful for- if anything we want more of them. We already have much simpler systems that search and regurgitate facts. We need more intelligent hallucinations that are consistent with and extend rather than conflict with the data.
- thwarted 2y agoIs it even possible to measure and distinguish the output as being hallucinated or not? All LLM output is hallucinated, it's only by statistics or chance that some of the output reflects facts, and we're only able to make that assessment because we can compare the output to facts. The model can't make that assessment itself. Going from 50% "accurate" to 90% "accurate" may actually be more insidious because it changes the utility from being a coin flip to trying to determine which 10% is inaccurate, or downplaying the existence of inaccuracies because at 90% it is "mostly correct".
- MattPalmer1086 2y agoI would argue with your definition of hallucination here. It's just a value judgement we humans apply to output that does not correspond to reality in some way that we don't find useful. We can control how diverse or creative output is via temperature. I am assuming that these new models could work at higher temperatures (i.e. more creative) while maintaining factual accuracy for things you care about. Or alternatively, keep the same temperature but have far fewer hallucinations (i.e. wrong answers).
- vessenes 2y agoWoww, very creative and interesting idea: I understand it as: train a bunch of fact-based LoRAs to zero loss (they mention 100k different ones), then use RAG to pick the appropriate Loras for a query. So cool. The only moat I can think of for such a company would be proprietary fact loras- basically licensing a modern ai encyclopedia. Anyway, really nice idea.
- liuliu 2y agoI think there is an expert router layer to decide which loras to be integrated at inference time. But they also mention that they freeze the weights for router during training. So it is unclear to me how the router was trained on what loss.
- vessenes 2y agoInteresting. That’s kind of surprising to me - it would mean with every new Lora they’d need to fine tune the router, no? Embedding a description of the Lora and using RAG to pull the nearest Loras in the embedding space is where my mind goes; it’s super extensible, minimal additional training for customer use cases, and the way the Loras probably work it’s not terrible to pull a few extras. Anyway I just speculate —- no idea what they’re actually doing on the backend.
- liuliu 2y agoThat's where it is confusing to me. They mentioned that for LoRA fine-tuning, the router weights are frozen, so you don't update the routing when training different concept. But how that expert router is trained? Could be a pretraining with some aux loss to encourage diversity.
- badriprof 2y agoI am old enough to remember the use of yellow pages to find teh right expert
- sage92 2y agoThey mention "tuning millions of expert adapters", not 100k
- 29athrowaway 2y agoCan it win at Jeopardy?
- youssefabdelm 2y agoHopefully someone reproduces results with code... cant find any code they shared
- tzekid 2y agoThey are offering a product/service, so going in too much detail would be a bad business practice, no? But true, would love to see this in OpenSource in the wild
- aetherspawn 2y agoHere I was hoping that there would be some kind of regulatory framework or protections put in place for AI before it became smart enough to actually take over the world. Being able to say "you are wrong, taking over the world is a bad idea" and have the model respond with "oh you are completely right, I am very sorry for that" was our first line of defense. I wonder if this model will argue insistently that you are wrong if you try and tell it that 1+1=3, and if so, whether that expands to philosophical issues such as the model arguing back at you based on its formed opinions on ethics and history.
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- wokwokwok 2y agoThe website says: > At inference time, the model retrieves the most relevant experts at each layer and merges back into the base model to respond to the user query. The paper says: > At inference time, only the relevant experts are retrieved from the index, allowing the LLM to store a large number of facts while maintaining low inference latency. We use specialized GPU kernels written in Triton Tillet et al. (2019) to accelerate the lookup of experts. ...but darned if I can understand from either what they're actually doing when they say that. Why do you need a custom GPU kernel for this outside of the normal NN layers? Can anyone see an explanation of how they pick which expert to use?
- janalsncm 2y agoAgreed, I looked through their “paper” and while it goes through the motions of a scientific paper, there’s barely any reproducible methodology. A single page in their paper, including the diagram. They do reference some papers I’m not familiar with and say their method is “similar”. If you check the huggingface page mentioned in a footnote, they have two directories: one for a model, and the other which contains a FAISS index. Although in the paper they say they use cross attention, so I have no idea how those could be combined.
- gdiamos 2y agoThat’s fair - I’ll try to go through the weekend and write out some of the equations for the kernel that loads the weights out of the index and does the adaptor ops. It’s inspired by cross attention in retro but there are some differences for training stability and to use as an adaptor rather than training from scratch. I consider that paper an early draft - hot off the press so to say - it needs review & editing before we would submit it to a conference. I tend to prefer a few rounds of open review before a final submission these days anyways - so appreciate the feedback I think the main idea should be reproducible - you can repeat the randomization and generalization tests with any LLM and get similar training curves and eval results - it just wouldn’t be efficient. We have tried it on about 5 real customer use cases with different facts and good success. Obviously we can’t publish customer data to reproduce which is why we focused on the randomization tests in the paper . There are also some missing hyper parameters from the appendix as well we will add eventually
- thisisauserid 2y agoAm I the only one that cringes at "10x fewer?" How do I multiply positive numbers and get something smaller? Is "1/10th" or "90% less" not better arithmetic? Maybe I should have done more gooder at math but it hurts my ears (eyes).
- esafak 2y agoThey are overloading "fewer" to mean division as well as subtraction. According to this logic "twice fewer" means "half as much".
- Dylan16807 2y agoFewer is subtraction, times fewer is division. More is addition, times more is multiplication (or 1 plus multiplication, oops). I don't think anyone says "twice fewer". Or "twice more" when talking about quantities.
- wayeq 2y ago> How do I multiply positive numbers and get something smaller? fractions are gonna blow your mind
- scottapotamas 2y agoBigger number better, obviously! I am also annoyed by most modern tech marketing using percentages incorrectly and inconsistently. But 150% is a bigger number than 1.5x so I suppose their hands are tied.
- deleted 2y ago[deleted]
- 22c 2y agoI think it depends how you think of the initial number, I think of it as a fraction and the multiplier applies to the denominator. eg. if hallucinations occur roughly 1 in 20 prompts then 10x fewer is 1 in 200 prompts, rather than 0.1 in 20 prompts.
- 2y ago
- XCSme 2y agoDoesn't this make the "AI" even less creative and more like full-text-search instead? What makes some data a "fact"? If everything is written in the training data, in the end, won't everything be treated like a fact? So the LLM will have 100% accuracy and 0% creativity.
- batch12 2y agoSounds like compression to me
- esafak 2y agoCreativity is clearly not the goal here. Machine learning models are trained to be robust to errors in the training data.
- AlexCoventry 2y ago"Ten times less" is a common English usage with a clear meaning.
- Brananarchy 2y agoMost advertised commercial uses of "AI" are glorified search. This seems like an improvement in that space?
- qeternity 2y agoSome people don’t want their language model to have any creativity.
- deleted 2y ago[deleted]
- raffraffraff 2y agonit: I hate trying to work out what "10x fewer" or "10x less" mean. Let's say I have a counter value, X = 100. I reduce that to 10. How can I phrase that in English? "Value reduced to 10% of original value" "Value reduced by 90%“ "New value is one tenth the original Value" Using multiplication with a positive integer and saying "less" just seems incomprehensible when it flies by during a sentence and I can't stop myself from mentally saying "No", like Neo at the end of the Matrix when the three agents fire a volley of bullets at him down the corridor. "No, this sentence stops right here while I pick it apart"
- bombela 2y ago"10 times less" contracted to "10x less" maybe? The first sentence seems fine to me. The contraction much less so.
- raffraffraff 2y agoSorry, only seeing this now. But this is precisely what I mean. What would 1 times less mean? 0? It's just an awkward way to phrase something that can be said much simpler using a fraction or a percentage.
- Dylan16807 2y agoI don't understand how you can build up this sort of reaction but you still need to "work out" what it means. It sounds like you learned what it means just fine. Also the most direct translation is "value reduced by a factor of ten". "x" means factor, "fewer" or "less" means reduced.
- wruza 2y ago“Value 10x less now” Sounds good as non-native.
- luke-stanley 2y agoIf less hallucinations are the goal, surely this is a bit over the top? Surely if you have the ground truth facts available, then fine-tuning for EVERY subject area seems much more work than using the facts with retrieval augmented generation, and making sure that the facts line up?
- xrd 2y agoIt feels like the two dumb ways to customize an open LLM are fine tuning and RAG. The former is expensive and complicated, the latter adds complexity to your queries but doesn't require up front compute for retraining. I couldn't tell how expensive this is up front, or what complexity it adds to the setup. Anyone know? It's definitely an interesting idea but if you have to pay $100k for all that LoRA, what margins are left over?
- qeternity 2y agoWhat do you think is complicated about RAG? I'm not arguing that it's effortless, but it's not that complicated? Genuinely interested to hear other people's pain points.
- xrd 2y agoWell, you need to generate embeddings usually, and then query and filter those, but it isn't that complicated for sure.
- ssheng 2y agoCreative idea. Any data on how much it takes to load the LoRAs and how much latency it adds to the generation speed?