5 ms·
LoRA Without Regret
- Yenrabbit 1y agoThinking Machines have put out a string of incredibly high-quality posts lately. Hard to oversell how much cred it's buying them with the AI research community! Keep up the great work folks
- sudohalt 1y ago[flagged]
- dang 1y ago"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- mijoharas 1y agoWhat else has there been. I've only seen this one (which is great!)
- joloooo 1y agoTheir Defeating Nondeterminism in LLM Inference was interesting for me. Worth reading their others!
- _def 1y agoTook me a moment to realize this is not about LoRa.
- ellisv 1y agoI also mistook it to be about LoRa and not about LoRA
- chrystalkey 1y agoI too fell victim to mistaking LoRa for LoRa
- logannyeMD 1y agoMissed opportunity to title this "Lo-RAgrets"
- HumblyTossed 1y agoThe name gets me every single time. Always think it’s going to be about radio LoRa
- frostyel 1y ago[dead]
- dannyfritz07 1y agoDang it! Got me too! I've been wanting to hop into Meshtastic lately.
- ijustlovemath 1y agoSet up a node! Bare boards that work with the app are like $50 and take a few clicks to flash and setup. The basic antenna with no amp makes contacts up to 50mi away if the conditions are right. I have one in a window and one in a backpack at all times.
- jacquesm 1y agoIt's insane how far you can go between hops, really most impressive. Where I live the mesh density is fairly high but I've also tried it in places where it was vanishingly low and yet I never completely lost contact. LoRa is very much an underappreciated technology.
- wkjagt 1y agoI have a couple of nodes up, but not seeing a lot of traffic
- mrandish 1y agoYeah, kinda disappointed it's just more AI stuff...
- canadiantim 1y agoI thought it was Lora the CRTD implementation, but then realized that Loro
- eagsalazar2 1y agostupid website hijackes cmd-back-arrow.
- markisus 1y agoCan someone explain the bit counting argument in the reinforcement learning part? I don’t get why a trajectory would provide only one bit of information. Each step of the trajectory is at least giving information about what state transitions are possible. An infinitely long trajectory can explore the whole state space if there are no absorbing states. Such a trajectory would provide a massive amount of information about the system, even if we ignored the final reward.
- mountainriver 1y agoA fair amount of research has shown that RL doesn’t add knowledge to the base model it just optimizes paths that already exist. Now ProRL from Nvidia showed there are ways of adding knowledge, mostly through progressive merging. I’m still not fully convinced of the 1bit claim, they made other mistakes in the blog post
- navar 1y agoI believe it's because the way you measure things in RL, each episode only tells you whether it was good (say reward +1) or bad (say 0 or negative reward), it does not tell you anything about the trace that was produced to get the outcome. This reward is the only thing measured to produce your gradients. Hence why the amount of info in it is O(1). This is in contrast to more "supervised" forms of learning where you could get a loss for each token produced (e.g. cross entropy loss), and where you'd get, as a consequence O(number of tokens) information into your gradients.
- deleted 1y ago[deleted]
- mountainriver 1y ago> LoRA works well when not capacity constrained, i.e., the number of trainable parameters exceeds the amount of information to be learned, which can be estimated in terms of dataset size I’m shocked they didn’t look at progressive merging of LoRAs. Research shows that’s the best way of improving its ability to model higher level features. Seems like a massive miss, not to mention there is other research that contradicts a lot of their findings. This feels a bit like a researchers first pass at learning LoRA
- yenepho 1y agoI am curious, would you mind sharing a citation?
- Mkengin 1y agohttps://arxiv.org/abs/2311.13600 https://arxiv.org/abs/2311.13600 https://arxiv.org/abs/2410.22911 https://arxiv.org/abs/2410.22911 https://arxiv.org/abs/2409.16167 https://arxiv.org/abs/2409.16167
- mountainriver 1y agoDon’t forget ReLoRA! https://arxiv.org/abs/2307.05695 https://arxiv.org/abs/2307.05695
- let_tim_cook_ 1y agoI'm not sure why progressive LoRa merging needs to be addressed here. They show there is a regime of problem where LoRa performs equivalently to FFT. Progressive merging of LoRa is somewhere inbetween and categorically more complex than just LoRa so would be dominated by standard LoRa in that case. While progressive merging could train faster as fewer params are trainable at any given time, it results in very larger adapter diffs OTO the size of the original model and doesn't retain the benefits of being able to deploy multiple adapters over the same base model idt.
- raaron773 1y agoThe amount of people who mistook this for long range radio and were disappointed when it isnt about it is way too damn high. (This is including me)
- ineedasername 1y agoIt might be useful to use this thread in a dataset to train a LoRa so that LLM agents can more easily disambiguate the great LoRa acronym collision of ‘25. No longer will future generations suffer the indignity of either/or/both confusions.
- kouteiheika 1y ago> However, the literature is unclear on how well LoRA performs relative to FullFT. I think the literature is clear on that? "LoRA vs Full Fine-tuning: An Illusion of Equivalence" -- https://arxiv.org/abs/2410.21228v1 https://arxiv.org/abs/2410.21228v1 Quoting from the conclusions: > The paper describes the finding that LoRA and full fine-tuning, with equal performance on the fine-tuning task, can have solutions with very different generalization behaviors outside the fine-tuning task distribution. We found that LoRA and full fine-tuning yield models with significant differences spectral properties of their weight matrices: LoRA models often containing “intruder dimensions”, high-ranking singular vectors approximately orthogonal to the singular vectors of pre-trained weight matrices. The existence of intruder dimensions correlates with the fine-tuned model forgetting more of the pre-training distribution as well as forgetting more when trained on tasks sequentially in a continual learning setup. I'm surprised they didn't cite this; it's a well known paper.
- adhi01 1y agoTo say that the 'literature is clear on that' while citing a single paper, which has been rejected from ICLR, is a bit of an overstatement.
- muragekibicho 1y agoThanks for this comment.
- kouteiheika 1y ago> which has been rejected from ICLR Oh, you mean rejected just like these papers? Efficient Estimation of Word Representations in Vector Space[1], one of the most influential papers in the space with tens of thousands of citations[2]? Or the RoBERTa[3] paper (dramatically improved upon BERT; RoBERTa and derived models currently have tens of millions of downloads on HF and still serve as a reliable industry workhorse)? Or the Mamba paper[4] (pretty much the only alternative to transformers that actually gets used)? Do you want me to keep going? Honestly, I find that whether a paper gets rejected or not means diddly squat considering how broken the review system is, and through how much honestly terrible papers I have to wade through every time I'm looking through the conference submissions for anything good. [1] -- https://openreview.net/forum?id=idpCdOWtqXd60 https://openreview.net/forum?id=idpCdOWtqXd60 [2] -- https://scholar.google.com/scholar?cites=7447715766504981253 https://scholar.google.com/scholar?cites=7447715766504981253 [3] -- https://openreview.net/forum?id=SyxS0T4tvS https://openreview.net/forum?id=SyxS0T4tvS [4] -- https://openreview.net/forum?id=AL1fq05o7H https://openreview.net/forum?id=AL1fq05o7H
- rco8786 1y agoI've been curious about LoRA and find a lot of these articles interesting. But I've been unable to find a good "LoRA for idiots" kind of starting point that gets me started actually doing some training with my data. Anybody know of a more practical guide I could use for that?
- CaptainOfCoit 1y agoUnsloths documentation probably gets as close to practical as it can get: https://docs.unsloth.ai/get-started/fine-tuning-llms-guide https://docs.unsloth.ai/get-started/fine-tuning-llms-guide Be sure to validate everything you're reading though as of late I've come across more and more things that don't seem 100% accurate in their docs, seems to heavily depend on what section.
- ijk 1y agoMy sense is they need to go back and update previous docs; they release a lot of software updates and a lot of notebooks showing how to use the features, but the two might fall out of sync. Would that match your observations?
- sgt101 1y agoQuestion for dudes building modern nn's... what's the thinking on estimating structural capacity for real world problem? How should I estimate how many parameters to choose for the model?
- p1esk 1y agoYou test different models on your real world problem, and pick the smallest one that works.
- sgt101 1y agoI just think that there has to be some heuristic..
- BoorishBears 1y agoClosest thing to a heuristic is trying the task with non fine-tuned models and building an intuition for how far off each model is, what directions it's off in, and how easily you can improve that direction via fine-tuning. For example, for classification, if is hallucinating semantically similar, but not technically valid classes, you can probably fine-tune your way out of the gap with a smaller model. But if your task requires world knowledge, you likely need a larger model. It's not cheap, efficient, or generally useful to fine-tune for additional world knowledge directly.
- _spduchamp 1y agoWell since we all thought this was about Meshtastic stuff, let's just give in and make this that radio/Meshtastic comment thread. Stumbled on this today... https://hackerpager.net/ https://hackerpager.net/ I really want something like this with flip out keyboard and could do Signal on LTE/WiFi.
- lewtun 1y agoFor those interested in playing with an implementation of these ideas, my colleagues at HF made some recipes here: https://github.com/huggingface/trl/blob/main/docs/source/lora_without_regret.md https://github.com/huggingface/trl/blob/main/docs/source/lor...