4 ms·
Somewhat off topic, but I've started wondering if we can actually make local LLMs feel smarter than the frontier closed models, by post-training it for your spe
by runeks 1mo ago
Somewhat off topic, but I've started wondering if we can actually make local LLMs feel smarter than the frontier closed models, by post-training it for your specific use case.
Say Company X has a software product which consists of a million lines of code, including a ticket for every bug and new feature for this piece of software. Then wouldn't it make sense to try using an open weight model but post-train it on that specific code base while using the tickets to teach the model about past bugs and features. Not sure exactly how, but it could involve doing reinforcement learning solving a past bug on that historic version of the code base, and rewarding the model if it comes up with the correct solution (as defined by the linked PR which fixed the bug).
So essentially post-training your local open-weight model using reinforcement Learning with Verifiable Rewards (RLVR) on your software products history of bug reports and their ultimate solution. And the same for new features.
- rovr138 1mo agoInstead, I'd build tooling for the model to be able to query the tickets, pr, commits, diff, etc This is something I can reuse better.
- runeks 1mo agoI'd like to have a model that just "knows" our code base and past issues, instead of having to query them for the same reason I don't want the model to have to query a dictionary to speak proper English. I could be mistaken here, but I'm curious to see if it would improve both output speed and quality. Currently, every new prompt I start it basically spends 10-15 minutes "getting familiar with the code" which is both annoying to wait for and wasteful money-wise.
- Zylokloto 1mo agoYou could do this for sure but its quite a lot of effort and a normal software stack is quite well known due to the massive amount of github projects and other opensource code. You can also get a lot out of a harness in this case or your Agents.md or Claude.md file by just enhancing the context. You might even question yourself if you are doing something wrong if a modern LLM really struggles with your code. We do the finetuning only on small semantic data were it helps a lot. I'm still wondering when we see smaller models (faster and cheaper) for more specific stacks like spring boot + java + angular + english only or so. Interstingly enough, i assumed LLMs are really good in any language but it seems that non english languages do reduce the ooutput quality of an LLM. At least last years GTC there was a talk about it. There are companies though which ahve this exact problem with programming languages you normally don't see. ABAP for example is a very well known SAP language.
- runeks 1mo ago> We do the finetuning only on small semantic data were it helps a lot. This sounds interesting. Would you care to expand a bit on how you do this? What is this semantic data? > There are companies though which ahve this exact problem with programming languages you normally don't see. This applies to my company. We have a substantial code base in our own dialect of APL. You think post-training would help substantially here?
- Zylokloto 1mo agoSo we have free text paragraphs which describe machines. Our model detects if the machine in question is of the main category we are looking for, then our finetuned model extracts from the free form text semantic information about the machine. Like color, features, horsepower etc. This would normally take quite a long time to do manually but we already had the semantic version of these texts because the company was doing this for a while. We now use gemma or Qwen (we regularly re-finetune the newest models to just see if they get better and they actually do) and then use these finetuned models to save us a lot of time. ---- If I had a coding language which isn't available much online and coding LLMs are bad on it, I would definitly try to finetune this but its defintily a lot more work than just doing what I explained above. Depending on what your usage of this APL Dialect is, it might be easier to fine tune it to migrate from your APL to something a lot more common. If this is not an option at all: You need to start creating data for the finetuning. You need a few hundred up to a few thousand of them in different formats like Q & A pairs. Documentation, syntax, a lot of diverse small examples. The intersting thing about this data generation: You can either leverage, to a certain degree , what you already have, or you start collecting them through your team/work collegues or you do them by hand. You can put the text into RAG and experiment with context engineering until a LLM is 'good enough' in it to be able to help you generating examples for and with you. Like you give an LLM all the relevant context for it, then you let it generate pairs: { "instruction": "Write an expression to find the maximum along the rows of a 2D array.", "input": "Array matrix: A", "dialect_notes": "Custom dialect uses ⌆ (max-reduce) and ⌥ for axis specification instead of /[1].", "output": "⌆ ⌥2 A" } (I have no clue about APL this is just a random example I asked an LLM to generate). You might have luck and finding communities with the same issue you face.
- nonethewiser 1mo agoI wonder the same thing. Code is pretty open ended though so I wonder if it’s not the best example. On the one hand it would definitely be useful for something like classifying support requests into priority. But would it be worth it to just train your own model? I guess one advantage is you could give it some well informed guidelines without training something on lots of data. There is probably a better example between discrete labeling and code though.