4 ms·
Is there any credence to the view that these startups are basically dspy wrappers
by nextworddev 1y ago
Is there any credence to the view that these startups are basically dspy wrappers
- -_- 1y agoDSPy is great for prompt optimization but not so much for RL fine-tuning (their support is "extremely EXPERIMENTAL"). The nice thing about RL is that the exact prompts don't matter so much. You don't need to spell out every edge case, since the model will get an intuition for how to do its job well via the training process.
- nextworddev 1y agoIsn’t the latest trend in RL mostly about prompt optimization as opposed to full fine tuning
- ag8 1y agoprompt optimization is very cool, and we use it for certain problems! The main goal with this launch is to democratize access to "the real thing"; in many cases, full RL allows you to get the last few percent in reliability for things like complex agentic workflows where prompt optimization doesn't quite get you far enough. There's also lots of interesting possibilities such as RLing a model on a bunch of environments and then prompt optimizing it on each specific one, which seems way better than, like, training and hot-swapping many LoRAs. In any case, _someone's_ ought to provide a full RL api, and we're here to do that well!
- nextworddev 1y agoThanks. Is this mainly for verifiable tasks or any general task
- -_- 1y agoThere needs to be some way of automatically assessing performance on the task, though this could be with a Python function or another LLM as a judge (or a combination!)
- ag8 1y agoIt's for any task that has an "eval", which is often verifiable tasks or ones that can be judged by LLMs (e.g. see [0]). There's also been recent work such as BRPO [1] and similar approaches to make more and more "non-verifiable" tasks have verifiable rewards! [0]: https://runrl.com/blog/funniest-joke https://runrl.com/blog/funniest-joke [1]: https://arxiv.org/abs/2506.00103 https://arxiv.org/abs/2506.00103
- omneity 1y agoPerhaps less about DSPy, and rather about this: https://github.com/OpenPipe/ART https://github.com/OpenPipe/ART
- -_- 1y agoART is also great, though since it's built on top of Unsloth it's geared towards single GPU QLoRA training. We use 8 H100s as a standard, so we can handle larger models and full-parameter fine-tunes.
- omneity 1y agoInteresting, do you have benchmarks on FFT vs QLoRA for RL?
- ag8 1y agowe should publish some; the high-order effect seems to be that LoRAs significantly hurt small model performance vs FFT, with less of an effect for large models. This is maybe because large models have more built-in skills and thus a LoRA suffices to elicit the existing skill, whereas for small models you need to do more actual learning (holding # parameter updates constant). In general I think it's better to get a performant small model with FFT than a performant large model with a large LoRA, which is why we default to FFT, but I agree that we should publish more details here.
- omneity 1y agoThanks! Personally I found FFT is not necessarily a strict improvement over (Q)LoRA as it can sometimes more easily lead to instability in the model, hence the bit of extra scrutiny. Curious to see your thoughts and results whenever you get something out.