4 ms·
I mostly use local models when the data has personal information. Earlier this year, I felt the coding quality was still not as good as Claude Code. One thing
by Alephinitesimal 1mo ago
I mostly use local models when the data has personal information. Earlier this year, I felt the coding quality was still not as good as Claude Code.
One thing that works for me is to ask the local model to make some fake data with the same format, let Claude Code work on the fake data, and then bring the code back and run it locally on the real data.
This way the real data never leaves my machine, but I can still use a stronger model for most of the coding.
- latentsea 1mo agoQwen3.8-27B has been the turning point for me. It's not as strong as the absolute frontier, but it's the first time I feel local coding models are actually functionally useable as daily drivers.
- Forgeties79 1mo agoMan I am having a hell of a time trying to optimize 3.8 over 3.6. I don’t have a particularly powerful setup but I can usually push 20-30tok/s on 3.6 and I can barely get to 10 on 3.8. Both unsloth same VRAM/RAM distribution more or less. My 3.6 is still producing consistently better results and faster
- Alephinitesimal 1mo agoThat's interesting since both models are dense. I wonder if this is more of an optimization issue with 3.8 rather than something inherent to the architecture.
- johnnyApplePRNG 1mo agoCould have sworn I read these were the same architectures the other day .... 3.6 and 3.8 at this size.
- Forgeties79 1mo agoI’m also not an engineer/coder so it’s equally possible I’m just doing something wrong.
- julianlam 1mo agoYou're likely using 3.6-35B-A3B, the 3.8 is currently a 27 billion parameter dense model.
- Forgeties79 1mo agoI have definitely use that to great success, and I do think it colors some of my memory here. I need to check if the 3.6 27B I was using previously was also a dense model. Good suggestion appreciate it
- latentsea 1mo agoMaybe you're holding it wrong because it's the same architecture between the models, assuming you're using the dense 27B model in both cases. And 3.8 is a significant improvement on 3.6.
- illusive4080 1mo agoHave you tried on personal finance analysis? That is what I most want to do but haven’t gotten around to it.
- Alephinitesimal 1mo agoI did try some finance analysis earlier this year. I was using a DGX Spark, so I could run some relatively large models, but the results were pretty mixed at the time. I honestly can't remember which models I used anymore. Might be worth trying again now though.
- jacquesm 1mo agoI'm having a really hard time doing on twin DGX spark what I could do on my quad 3090 rig (which is a scaled down version of what I was using before, the power requirements and the noise were really an issue but I loved the speed and the amount of VRAM). The results tend to be inconsistent, there is lots of looping, far more tokens generated for the same job and lower quality output. I suspect there is some kind of regression in the B12X kernels or something to that effect because none of that should happen, the exact same model on both machines gives wildly different results. Probably this will sort itself out over time. If I may ask, what model / software combo were you using?
- Alephinitesimal 1mo agoI was only using a single DGX Spark, and this was earlier in the year, so I was running some pretty aggressively quantized models — probably in the 1–3 bit range. My main issue at the time was that my financial data had lots of messy notes, comments, and irregular annotations. The quantized models often failed to process all of that context consistently and would miss things. So I ended up generating a fake dataset with the same structure, asking Claude Code to work out the analysis on that, and then bringing the result back to the local model for the final pass. I was mainly using llama.cpp at the time, before B12X support was integrated into vLLM, so I think I wasn't using it then.