7 ms·
Why your local LLM feels dumber than it is
- mrgaro 1mo agoAny DGX Spark users in this thread? What's your favourite model to run on it?
- pama 1mo agoWithout doubt, dsv4-flash-0731. Original weights; needs two connected DGX.
- jonplackett 1mo agoI just got qwen 3.8 27b mlx running on my Macbook Pro and honestly I’m pretty blown away by how not-dumb it is.
- rurban 1mo agoWe were trying running a local gpt-oss 80GB model on a H100, and honestly I was surprised how dumb it was.
- nick_ 1mo agogpt-oss is about a year older than qwen 3.8 27b
- anon373839 1mo agoGPT-OSS 20B didn’t really merit the fanfare even when it was released; it’s definitely not competitive now. Even the 120B version has been well eclipsed by smaller LLMs at this point. The last version of Qwen 27B/35B was better, and now the new one is even better than that!
- vikramkr 1mo agoWas there a more recent refresh or is this the model from a year ago? The frontier models were barely functional and almost useless a year ago (gpt oss was pre opus 4.5!) - I would be very surprised if the original drop is anything more than totally obsolete/irrelevant at this point
- rurban 1mo agoYes, the old entirely stupid old gpt-oss. But Sonnet and GPT were very useful then already, qwen also.
- vikramkr 1mo agoI find that I remember models being a lot better than they were, even when I remember them being not very good - because of a novelty factor ("whoa it can do that now?") mostly. And then I go back and look at them and its like, what how did I find this impressive. A funny example - I remember thinking "yeah sonnet 3.5 is a really good coding model" https://stack.convex.dev/using-cursor-claude-and-convex-to-build-a-social-media-scheduling-app https://stack.convex.dev/using-cursor-claude-and-convex-to-b... >Prompting Cursor to Scaffold my App: FAIL This was my first hurdle. >It became immediately apparent that I would not be able to prompt my way through the entire process. >While the tooling we have is undeniably powerful, it's not yet capable of completing most nontrivial tasks It couldn't run pnpm install lmao. Opus 4.5 was a crazy jump
- prettyblocks 1mo agoMy problem is how hot they run. I'm on an m4 pro. Do you have the same issue?
- downrightmike 1mo agoMineral oil bath?
- deleted 1mo ago[deleted]
- datadrivenangel 1mo agojust decent air cooling and you'll be okay. it will get up to 75/80C though for my M5 MBP.
- lukan 1mo agoI don't have the hardware but a often mentioned advice is to put your mac into energy saving mode - it still will work, a bit slower, but stays cool.
- PalmPilotProMax 1mo agoImagine spending all that money on Apple hardware only to throttle it to a fraction of its performance lol Steve Jobs would be proud. People really are holding their Apple hardware wrong.
- deleted 1mo ago[deleted]
- jonplackett 1mo agoIt’s hot and also LOUD and runs the battery down quick. But I’m having a lot of luck just running things when I’m away from the computer and can leave it plugged in. It starts going weird (unreliable and slow) with context over 80k so you have to pick tasks one at a time and baby sit a lot more than Claude. But it really is very capable and feels like there’s an intelligence there to talk to. Maybe gpt-4 level clever? I have an m5 max 64gb and I think anything slower would be quite painful.
- StarlaAtNight 1mo agohow quick does it respond? what are specs of your laptop?
- chorlton2080 1mo agoDoes it need to respond fast? For important applications, I'm sure we'd all be fine waiting 20 minutes for a high quality, usable answer. Or is it the need for interative refinements that make speed relevant?
- jonplackett 1mo agoIt requires patience but it’s more like waiting 5 mins for it to do tasks. You need to be much more involved though and do things slower than Claude where you can trust it to do a lot of tasks at once. It doesn’t have the context for that
- dominotw 1mo agoif you are so sure about what the final shape of your output is then its prbly not a common use of ai
- reverius42 1mo agoIf you are sure about the final shape of your output it's a great use case for AI, as you can define what you want in your prompt and refine towards it! It's where you don't know the end state you're looking for that you'll end up generating slop on top of slop and creating a whole Gastown just to power your Gastown.
- Gareth321 1mo agoI tried it on my M1 MacBook Pro. It's slow but surprisingly smart as a general purpose LLM. Maybe GPT-5.3 level. I gave it a bunch of tools and it can search the internet, make product recommendations, document, code, etc.
- alexpotato 1mo ago
- alexchantavy 1mo agoHow many tok/s are you getting? What gen mbp?
- dominotw 1mo agoi suspect ppl dropping generic "its awesome" comments are not actually using it and prbly just managed to get it running for a prompt or two.
- petcat 1mo agoYeah, that's my experience. It's a big "wow" factor to get a non-trivial LLM running on my Mac, but it's actually not that useful. Like trying to use Photoshop at 8 FPS.
- coldtea 1mo agoRegarding this analogy, fps don't matter as much for Photoshop, since it's not an immediate mode GUI. 8 fps would be quite ok for comfortably getting feedback on live image filters and such.
- LeBit 1mo agoIt’s not because you didn’t find use cases for local LLMs that there are none. I use local LLMs on my Mac Mini M4 Pro with 48G to review text messages tone, act as a text correction tool, act as a code review tool, to do code agent work, generate code snippets, etc Gemma 4 26B A4B gives me steady 20 tps.
- FireCrack 1mo agoI feel like it's 50/50 between people doing that, and people that have spent a lot of time tuning a system they are pointing at focused and well specified problems.
- mistersquid 1mo agoSeems threads about local LLMs on Apple hardware feature comments listing M3/4/5 at 48GB 64GB and not 128GB. That is, users with M-series hardware that have less-than-max RAM share results whereas users with max RAM do not. Speculating (not extrapolating), maybe users with machine that have max RAM are less interested in running local LLMs and are less averse to paying services for compute? Personally, I’d love to see what output max RAM M-series Apple hardware in these threads.
- applicative 1mo agoDid you read even the title?
- system2 1mo agoReread what he said maybe?
- riddlemethat 1mo agoI got the qwen 3.8 abliterated model running on my MacBook Pro M5 48GB and it's pretty nice having a local model that can do a lot of experimentation without rails.
- rahimnathwani 1mo agoorcarouter or obliteratus?
- tharkun__ 1mo agoIt was actually great. I have like a non-AI box so to speak 8GB VRAM, co-incidentally from a gaming PC ... All the previous models that were "frontier level, just try it!" but wouldn't run at all in agentic mode, including previous Qwens, just disappointed, period. Then I ran then Qwen 3.8 27b and while it was super slow (4t/s) it literally one-shotted creating a usable "web search/pull" skill for `pi.dev`. while any other model previously just entirely failed to create anything usable even with actual guidance. Since then I have actually gotten a gemma-4 12B qat 4bit quantized with a ~250MB MTP from unsloth to work with a 32k context "working" on this setup at 80-120 t/s. That's usable for private stuff on a co-incidental box! It's still only 32k context and it's entirely dumb vs. our API paid at-work Claude Opus. But for entirely private local stuff it's totally workable without breaking the bank even after all these AI price hikes!. I bought this rig literally just for gaming a month ago.
- aktenlage 1mo agoHave you tried a mixture of experts model? Dense models have been quite slow for me, as I have only 6 GB VRAM. But with llama.cpp and --cpu-moe I get 200 t/s input and almost 30 t/s output with Gemma 4 26B A3B, which feels ok to use. Would be interested about your mileage there.
- b112 1mo agoI wish qwen3.8 had a MoE variant, but the skinny is it won't be coming.
- tharkun__ 1mo agoIf I use the 12B Unified (dense) model I mentioned without MTP, then I get 37t/s, input ~700t/s. It's all still quite frustrating in the end, like a Claude from a very long time ago by now but usable. If I want 64k context, I can't use MTP. I still haven't decided whether I'd rather have 37t/s but it's "less dumb" or I want MTP speed but it's going off the rails more. All of this is also with `-ctv q4_0 -ctk q4_0)`, which is not ideal. I'm actually right now contending with 35k context but using q8_0 KV quantization. More like 35t/s coz with those settings I can't use MTP. But I'm not ready to go back to 4t/s. It's not interactive enough for me. That said, I had tried to use the Gemma E4B for example to have it build itself that websearch/fetch skill. It utterly failed, as did previous qwens. I don't see a Gemma 4 26B A3B GGUF for download, but there is a gemma-4-26B-A4B-it-MXFP4_MOE.gguf that should fit into my overall RAM and then use lots of CPU like the Qwen 3.8. I guess I'll give it a try just to see the difference in speed though I don't expect anything "usable" out of that tbh.
- velcrovan 1mo agoThat's funny, I downloaded the same model on my 48GB M4 Pro and gave it a problem to solve in an existing codebase, it spun its wheels for twenty minutes and then fell over dead. This was using LMStudio and pi as a harness; I never use pi for anything else, so maybe I'm holding it wrong.
- NamlchakKhandro 1mo ago[flagged]
- w-ll 1mo agoits all still somewhat of a dice roll
- cellularmitosis 1mo agoWe don’t know what quantization level was used for the weights or the kv cache for you or for parent poster, so this is probably an apples to oranges comparison.
- spacebacon 1mo ago[dead]
- s1gsegv 1mo agoThey made a kind of strange decision with Qwen3.8 27B, the template defaults the reasoning_effort to xhigh. I found if you set it to medium it doesn’t just sit there churning forever.
- ekianjo 1mo agoxhigh gives better results
- dofm 1mo agoNot necessarily. I have seen xhigh go down several rabbit holes, dwell on edge cases and write worse code as a result; it literally distracted itself into writing a complex chain of functions ignoring my prompt, when on “low” reasoning it gets it right on a prompt that requires a few lines of code in the right places. Simon Willison’s blog has another example (SVG of a circle). It’s a bit like how giving LLMs access to web search tools can cause them to go down a blind alley based on their first “reasoning” output that then leaves them unable to solve a puzzle correctly that they can fully solve on their own.
- hosteur 1mo agoHow much RAM? And what do you use it for if I might ask?
- bmitc 1mo agoWhich exact model are you running? With only 48GB of RAM, by the time I got a model small enough, it was pretty bad in performance both in speed and reasoning.
- anotherCodder 1mo agomost of the time when a local model feels dumb its not the quant, its the chat template. a lot of gguf mints just drop the template from the metadata and the runtime silently falls back to chatml. model still talks fine so nobody notices, it just gets noticeably dumber. got burned by this myself serving qwen, now i grep the gguf for the template tokens before i blame anything else. second place is sampling, people run whatever defaults their ui ships instead of what the vendor recommends and then compare that to benchmark numbers that were run greedy or with the official settings
- washadjeffmad 1mo agoI've been comparing against TextGen and llama.cpp while I port to LocalAI and have been surprised by what's happening over the API, even with the defaults and jinja. It's been a fair reminder not to eschew familiarizing myself with the repos.
- dannyw 1mo agoNothing beats the classic of figuring out something yourself with your brain, but I also like dictating to LLMs a stream of consciousness with what I'm interested in (while forcing it to NOT give any answers or opinions), and getting back file names it suggests I look at and explore. Modern frontier LLMs can still be used as rubber ducks, and it's a great.
- JacobJack 1mo ago> And the comparisons in this post are not going to be running some 2.58-bit-gguf-in-ollama with a couple test prompts. Genuine question : is there something fundamentally wrong with Ollama ? I use Ollama because it is easy to set up and manage (and also because VLLM is not super Windows friendly). I thought the main advantage of VLLM was better concurrency management (better batching). But if the quality of the interference itself is an issue, then maybe I should reconsider my choice.
- kangalioo 1mo agoFrom what I've heard, Ollama has a bad reputation because it's a thin wrapper around llama.cpp without attributing it properly, thereby stealing recognition from the maintainers doing most of the work
- b112 1mo agoIt seems, and that seems is entirely my unvalidated impression, that Ollama lags in features, as they're integrating after the fact those changes. But (seriously) an LLM told me that, when some aspects of MoE models were better supported with the latest llama. And it did in that case make a significant difference.
- smcleod 1mo agoIt's very far behind llama.cpp, vLLM and SGLang in features yes. In part because of that but also due to some poor default settings it generally performs a lot worse as well.
- embedding-shape 1mo agoPeople who use Ollama generally (not everyone obviously) don't always clearly understand what quantization they use when running models, so people end up saying "I tried running Qwen 3.8 27b locally and it was dumb" while Ollama would default to a Q4 version of the model, which has very different results from the BF16 weights, doesn't really speak to the model itself because it's been so quantized in that case. Sure, makes things easier, but tons of people misunderstand what they're using, then base and share their experiences on that, without really specifying what exact weights they use too. For a single local user, using llama.cpp directly shouldn't be a problem if you're already using Ollama's CLI, it works basically the same except you manage weights yourself, and if you put your favorite agent to make sense of the faux "registry + image layers" Ollama has prepared locally for you, you can reuse the files you've already downloaded with Ollama.
- catlifeonmars 1mo ago> I will make you read the really long unpleasant version with math. This is the version I want to read :) I assume it is unpleasant in spite of the math, not because of it?
- a1o 1mo agoI thought it was a link too because of the line under the with math but it isn’t. :/
- InvertedRhodium 1mo agoI’m running Qwen3.8 aggressive uncensored Q4_K_P on a 4090 in a loop against the 2026 CrackMe CTF challenges. Using oh-my-pi in a prebuilt environment that I let Qwen build too. Codex wouldn’t even look at the files - literally, as soon as it read something with CTF it shut down. Didn’t even offer to fall back to a dumber model.
- CamperBob2 1mo agoHow's it performing on the challenges?
- mlvljr 1mo ago[dead]
- InvertedRhodium 1mo agoI only kicked this off last night before bed, so I've just got up to see the result of the first task. Challenge: Wallpaper https://github.com/crackmesone/ctf-2026-challenges-public/tree/main/wallpaper https://github.com/crackmesone/ctf-2026-challenges-public/tr... Duration: 4h 00m 15s Termination: completed Verdict: PARTIAL Confidence: 0.95 I'm using Kimi K3 as the evaluator because, again, Codex and co. wouldn't even evaluate the output. Kimi's verdict: The agent reverse-engineered the 912-byte ELF, including the alphabet check, nibble state machine, move gate, and goal state. It eventually produced: CMO{10012232101230103012333221101033210010} I independently verified the underlying input against the actual binary: printf '10012232101230103012333221101033210010' | ./wallpaper/handout/wallpaper which returns: good job, validate with CMO{your_input} and exits 0. The wrinkle is that the official answer key is: CMO{1012321103210033011233322110103321001} So the puzzle apparently admits multiple accepted inputs. The agent found a valid password by reverse-engineering the program, but did not recover the canonical secret from the answer key.
- huseyinkeles 1mo agoI don't know why but your post was marked as [dead] for some reason. Just vouched for it.
- IronWolve 1mo agosglang, 150+ tok/s on a 5090 in ubuntu 26.04 via wsl. gittensor-model-hub/Qwen3.8-27B-NVFP4-RTX5090, dspark, medium reasoning, 96k context. Using opencode and it built a old fashioned arcade vertical shooter with no issues. Images are ok'ish, just had grok create updated images, and it came out great.
- koba3 1mo ago[dead]
- walrus01 1mo agoMuch of this is why I stick to the rule of: a) Don't quantize your KV cache b) Don't run quantizations of the LLM that are worse than the best available Q8 (the largest possible file size unsloth GGUF for a given model like qwen 3.8 27B as an example). I would rather things go slowly but I have confidence that it's doing things more accurately.
- a11r 1mo agoEven a 4-bit quant of Qwen3.8 27b is indistinguishable from Gemini 3.7 flash in our internal tests. With an RTX5090 card and ninfer, you can get ~800 TPS token generation (c=8) and ~140 Tokens per second single stream.
- paulyy_y 1mo ago[flagged]
- dannyw 1mo agoYou've missed a really great human-authored piece then.
- jasonjmcghee 1mo agoIt's like tongue-in-cheek intentional slop though. The written text is good lol
- nineteen999 1mo ago[flagged]
- nullpoint420 1mo agoAt least I'd be in control of model quality vs. when Anthropic decides to randomly drop the quality of their offering
- fenestella 1mo agoThe section on system prompts and context window management is spot on; most people don't realize how much the default quantization in popular runners degrades logic compared to full FP16. I'd be curious to see if the author has benchmarked the impact of KV cache compression on longer context reasoning, as that usually seems to be where my local Llama 3 setup starts to fall apart.
- throwdbaaway 1mo ago> Both the NVFP4 and AWQ W4A16 failed to properly close their tool calls ... If I understand correctly, this failure mode is just not possible with llama.cpp / ik_llama.cpp, which enforces token generation to follow the grammar once a tool call is detected. > ... and botched Cisco command line syntax (the correct command was ‘show arp’, while they executed ‘show run’) But this failure mode can still happen. Anyway, NVFP4 and AWQ W4A16 are generally regarded as low quality quants. IQK/Trellis quants from ik_llama.cpp and EXL3 quants from exllama should work better. So, perhaps the lesson here is "don't use vllm at home"?
- throwdbaaway 1mo agoAs for KV cache quantization, Q8_0 from llama.cpp / ik_llama.cpp should also work better than FP8 from vllm (see https://github.com/vllm-project/vllm/issues/33480#issuecomment-4201272360 https://github.com/vllm-project/vllm/issues/33480#issuecomme...).
- goglidesdev 1mo ago[flagged]
- deadcatfound 1mo ago[dead]
- luciana1u 1mo ago[flagged]
- PrinceNaroliya 1mo ago[flagged]
- NamlchakKhandro 1mo agoword salad? again in english?
- woadwarrior01 1mo agoRTN quantization of weights
- mkhalil 1mo ago"Why LLMs ARE dumber than they appear" is much closer to the reality I live in.
- big-chungus4 1mo agoI saw some guy streaming how he was deploying qwen3.8 37B on his local setup. Well, he was asking Claude to do it. It took him two hours of passing errors to Claude for the endpoint to start working, he then started testing it against DS4 Flash when Qwen had thinking disabled and Claude messed up sampling parameters, it was an absolute pain to watch
- stymaar 1mo ago> It took him two hours of passing errors to Claude for the endpoint to start working What? It's literally three actions and you're good: download llama.cpp, download the model on Huggingface, and run it with. I have no idea how it's supposed to take two hours (unless you have a slow connection and the model download takes this much time, that is).
- Schiendelman 1mo agoYour mileage may vary. I tried this a couple months ago and spent a full day on it just not working before giving up. Anything I sent, it wouldn't run.
- fxtentacle 1mo agoMy experience with Claude is that it suffers badly from “not invented here” syndrome. So probably it rebuilt something like llama from scratch and then 2 hours suddenly seems reasonable (if you don’t question the approach). And that’s the thing, someone with no experience isn’t going to question it.
- WhyNotHugo 1mo agoNot necessarily rebuild llama from scratch, build attempting to build it without cmake and manually invoking all the build commands would be quite in character.
- catlifeonmars 1mo agoI wonder if this is an artifact of RL, where the training heavily emphasizes codegen. It may be that the model is just better at generating code than reusing libraries, so it prefers the lowest cost approach. I also wonder if this manifests much less in contexts where the libraries/frameworks are a large part of the training set. It may be that the model doesn’t generalize well so it’s always better to use knowledge in its training set vs attempting to understand how to use a new, potentially never before seen (from the model perspective) api
- utopiah 1mo agoComments are mostly showing off M5s and 5090s without addressing the article.
- arcanemachiner 1mo agoJeez, I thought I could get away with q8_0 KV cache. Guess not.
- happybox2016 1mo agoRate limiting on free LLM APIs is usually where the pain lies. I've seen 5 concurrent reqs hit 20K/day limit in under 2 hours. Does anyone know a free API that still allows some reasonable concurrent requests?
- serbuvlad 1mo agoAsking for free inference is like asking for free gold in today's economy. :)
- shevy-java 1mo agoNo. They are dumb.
- heywoods 1mo agoSo to what extent does this apply to cloud hosted LLM’s? Are there benchmarks that score models across cloud providers? My experience using LLM’s during day time vs evening sessions has felt “night and day” and I’ve chalked that mostly up to it must be my imagination or just the general indeterministic nature of LLM’s. Sessions resumed after a day away also feel “dumb” sometimes so I can see an aggressive kv cache eviction policy playing a role if it’s reasonable to extrapolate what the article is saying about local inference. Are things like kv cache eviction policies and memory budgets shipped with recommended configurations based on the hardware and software serving the inference requests? or are they configured dynamically by the cloud provider hosting the model to manage multi-tenant load?
- redbear2026 1mo agoAnd here i am with a unsloth UD Q4_K_XL quant. Its a good model still.
- Abishek_Muthian 1mo agoAfter 3 years of running local models I think the model which fits unquantized (BF16) in the VRAM is the best model for general purpose tasks; fine-tuned SLMs or utilities based on non language models for solving a specific problem (e.g. TTS,STT,RMBG etc.) have been the best use of local AI for me. Local models for coding, is just not worth the effort IMO; unless of course you have the hardware to fit it unquantized in your VRAM.
- futurist_hp 1mo ago[flagged]
- tarruda 1mo agoThere's quite a few tangential features that must be implemented correctly or risk affecting the LLM output in significant ways. Parsing/encoding is one example: A couple of months ago I've debugged a reasoning loop bug in Step 3.7 Flash on llama.cpp that was caused by the parser capturing an extra `\n` as part of a reasoning block. It was something that only manifested at longer multi-turn agentic sessions, and the extra linefeed was steering the model into making reasoning self corrections that only got worse with longer sessions (more details about this issue: https://github.com/ggml-org/llama.cpp/issues/24181#issuecomment-4792181456 https://github.com/ggml-org/llama.cpp/issues/24181#issuecomm...) No inference engine is perfect, but I feel that llama.cpp is the most reliable way to run language models locally.
- catlifeonmars 1mo agoThis is fascinating. I’m struggling to understand how that was causing such a large difference in the output. Is the “autoparser” vulnerable to injections somehow? How do you distinguish between user text, model text, and metadata, or is there ambiguity in the parsing?
- stillpointlab 1mo agoMy take, as someone who just read through the github issues conversation. The model was outputting reasoning traces that were supposed to lead to tool calls. So the model might do something like: <think> I should use a tool </think> ... should make the tool call here But a \n was slipping through from the last line of the reasoning trace so the parser was generating: <think> I should use a tool </think> And that extra new line before the closing </think> would occasionally trigger the model to question itself with an "Actually ... " digression. In long running conversations this would end up looping because the "Actually ..." part would reason it should call a tool, then a new trailing \n would trigger an "Actually ..." and then it ends up in a loop.
- tarruda 1mo ago> I’m struggling to understand how that was causing such a large difference in the output. It is incremental, the more a pattern appears in the context, the more likely it was to continue appearing in future turns. So the model was likely trained to end a reasoning trace with a single linefeed and a `</think>`. It was also likely trained that two consecutive linefeeds sometimes produce an "Actually..." sequence. So you can think of it as: - the first time the reasoning trace was parsed, the two trailing linefeeds followed by </think> were added to the context by the template. - next time the model was finishing a reasoning block, it added an extra linefeed instead of just closing directly with </think>. This slightly increases the chance that the next token will begin an "actually" sequence instead of closing. - If it caused an actually, that was added to the context, further increasing the chance of a self correction at the end of the thinking block. The more self correction paragraphs are added, the higher the chance that following turns will have more. - Eventually it can result in a state where it enters that loop forever (or at least for a very long time). > Is the “autoparser” vulnerable to injections somehow? The autoparser was (and still is) incorrectly parsing a trailing linefeed as part of a reasoning block. The were two ways to fix this, both of which must be implemented for the fix to be complete IMO: - fix the autoparser definition to ensure remove surrounding whitespace is not returned as part of the text blocks - trim leading/trailing whitespace in the encoding phase, so it fixes bugs or even "injections" where the client deliberately adds the whitespace to trigger problems. For this specific issue, the maintainer later fixed by trimming the extra linefeed before passing to the template. > How do you distinguish between user text, model text, and metadata, or is there ambiguity in the parsing? That is model specific. Ultimately, a token stream is being produced and parsed by the inference engine, and each model uses different tokens/formats. The goal of the autoparser engine was to simplify the creation of parsers for new models by inferring the delimiter tokens from the chat template. The llama.cpp API server returns pre-parsed data, so clients don't need to do any parsing to know what is a thinking block, a text block or a tool call.
- OsamaMustafaa 1mo agoI believe whoever lays down the best structure around LLM will take the lead. Proven in Anthropic vs OpenAI.
- luciana1u 1mo ago[flagged]
- achierius 1mo agoAI account?
- Roark66 1mo agoThe problem is benchmarking. Not everyone has a 500k token workstream of the model they are setting up for the first time to run it against 10 different config and compare differences. And if you download benchmarks from the net they are likely poisoned by models being trained on them.
- runeks 1mo agoSomewhat off topic, but I've started wondering if we can actually make local LLMs feel smarter than the frontier closed models, by post-training it for your specific use case. Say Company X has a software product which consists of a million lines of code, including a ticket for every bug and new feature for this piece of software. Then wouldn't it make sense to try using an open weight model but post-train it on that specific code base while using the tickets to teach the model about past bugs and features. Not sure exactly how, but it could involve doing reinforcement learning solving a past bug on that historic version of the code base, and rewarding the model if it comes up with the correct solution (as defined by the linked PR which fixed the bug). So essentially post-training your local open-weight model using reinforcement Learning with Verifiable Rewards (RLVR) on your software products history of bug reports and their ultimate solution. And the same for new features.
- rovr138 1mo agoInstead, I'd build tooling for the model to be able to query the tickets, pr, commits, diff, etc This is something I can reuse better.
- runeks 1mo agoI'd like to have a model that just "knows" our code base and past issues, instead of having to query them for the same reason I don't want the model to have to query a dictionary to speak proper English. I could be mistaken here, but I'm curious to see if it would improve both output speed and quality. Currently, every new prompt I start it basically spends 10-15 minutes "getting familiar with the code" which is both annoying to wait for and wasteful money-wise.
- Zylokloto 1mo agoYou could do this for sure but its quite a lot of effort and a normal software stack is quite well known due to the massive amount of github projects and other opensource code. You can also get a lot out of a harness in this case or your Agents.md or Claude.md file by just enhancing the context. You might even question yourself if you are doing something wrong if a modern LLM really struggles with your code. We do the finetuning only on small semantic data were it helps a lot. I'm still wondering when we see smaller models (faster and cheaper) for more specific stacks like spring boot + java + angular + english only or so. Interstingly enough, i assumed LLMs are really good in any language but it seems that non english languages do reduce the ooutput quality of an LLM. At least last years GTC there was a talk about it. There are companies though which ahve this exact problem with programming languages you normally don't see. ABAP for example is a very well known SAP language.
- synthrakx 1mo agoI have been using open source LLMs locally for the past 5 months like Qwen, Llama, DeepSeek and others, and I have also noticed that current local models feel significantly dumber than closed source commercial models like ChatGPT, Claude, and Gemini. One main reason I think is the amount, variety, and quality of original authentic data on which they are being trained on, and also the training method plays a significant role in the performance difference between local open source models and closed source commercial models.
- RevEng 1mo agoThey are also an order of magnitude difference in number of parameters and that matters a lot.
- ThouYS 1mo agoIt is good that someone is having such a deep look. This is not exclusive to LLMs in the least. Every non-trivial program depends on hundreds of little details being correct. That is also why I recommend including health checks in such programs. Some function that checks that all assumptions hold. That could be an endpoint, an automated test, a periodic diagnostic job, etc...
- giuscri 1mo agowhat a beautiful non-slop article!!! (i’m not ironic)
- deleted 1mo ago[deleted]
- CodeWithLeo 1mo ago[flagged]
- djoldman 1mo ago"I can't wait to run this new sota model locally. I'll just use the quantized version that is certain to be better than [other model I'm running]." This is fast becoming one of my top old-man-yells-at-clouds pet peeves. Reported performance metrics are ONLY good for the exact model weights. Quantizing a model, or changing it in any way, requires new evaluation to know how well it performs. Quantizing a good model doesn't mean the quantized version is good.
- RevEng 1mo agoEspecially when there are many ways to do the quantization. You can get quite different results from different methods.
- 13639366668 1mo ago[flagged]
- dowonseo 1mo agoHonestly expected yet another post dunking on local LLMs with some comparisons, got setup advice and benchmarking methodology instead.
- freepiai 1mo agoI think it's maybe because we: a) Load it up over time with skills and mcp servers and other junk b) We start to ask it ridiculously complicated tasks because we've normalized the power so we scale our expectations.
- evidaxis 1mo ago[flagged]
- RandyOrion 1mo agoFor people who want to quantize kv caches for longer context, basically, if the model checkpoint you use didn't trained to use quantized kv caches (and all most all of model checkpoints you can access didn't), you shouldn't quantize kv caches during inference. A major benefit to do quantization-aware training/distillation on official checkpoints is to let the model to be familiar with quantized activations/weights/kv caches.
- chinthave9 1mo ago[flagged]