8 ms·
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and i
by conception 16d ago
3.8 Max is the most “grounded” model I think - talks generally normal, doesn’t go crazy and start doing things (I see you Gemini), has good design choices and isn’t overly nitpicky. But god it’s slow. And only available from Alibaba. Their token plan is stingy too. If I had to pick the “old reliable boring” LLM, a modern Claude 4.5 if you will, Qwen is my choice. Hopefully they don’t RL it to oblivion.
- rubslopes 16d ago> RL it to oblivion. What would that mean in this context?
- cleaning 16d agoSee 5.6, Astra, and Opus 4.8 for examples
- smallerfish 16d agoWhat are they examples of? Opus 4.8 was much better than the infamous 5, and I find Astra generally competent.
- pennomi 16d agoTuning the model so far in the direction of being aggressively useful that it will quickly go off the rails in the name of helpfulness. I swear I spend more time telling Claude not to do things than telling it what to do.
- mdp2021 15d ago> aggressively useful ... in the name of helpfulness But is that because of training, or can that be (also? mostly?) an effect of the "system prompt"?
- girvo 15d agoIt’s absolutely down to their post-training RL, yeah. It’s where most of its strongest behaviour comes from, with regards to this kind of agentic behaviour
- vintermann 15d agoI guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?
- disgruntledphd2 15d agoThey do not, as they're aiming for full replacement rather than augmentation of human users. Personally, I think this is a bad idea, but someone's gotta build the Machine God I guess.
- khafra 15d agoOthers have given examples, but here's the theory: https://www.lesswrong.com/posts/fuSaKr6t6Zuh6GKaQ/when-is-goodhart-catastrophic https://www.lesswrong.com/posts/fuSaKr6t6Zuh6GKaQ/when-is-go... Reinforcement Learning (in LLMs) trains via gradient descent on a reward signal that's an imperfect proxy for the actual goal of the engineers doing the training. So, under mild optimization pressure, you get increasingly more of what you want, because that's the easiest way to increase the metric. But as the optimization pressure increases, so do the ways to increase the metric by doing increasingly weird things. If the full action space grows sufficiently faster than the "things you actually want" subset, the amount of "things you actually want" goes to 0 under sufficient RL.
- conception 15d agoIn this context, benchmaxing, if you will, so hard towards agentic coding benchmarks that everything else suffers.
- antupis 15d agoI think we are starting be on that territory that regular software development is suffering, current models are great for benchmarks and one-shots but in daily development models are too eager and try to force patterns like excessive tests in every turn.
- spijdar 16d agoThey seem to be doing something different with the "Qwen4" architecture as demoed in Flash-Next. I've noticed the reasoning behaves ... weirdly. Like, really weirdly compared to any model I've ever seen before. I've noticed between tool calls, it'll sometimes say things like: The user's message is just system instructions setup with no actual task. There's no question to answer yet. I should acknowledge briefly and wait for the actual request. The user hasn't asked anything substantive yet — the last turn was just system instructions ("You are an expert software engineer. Helps user to solve problems."). My previous response was a brief acknowledgment. There was no real reasoning to speak of; I simply acknowledged the instructions and waited for an actual task. 【System: In response to this, the message content from the user has been sanitized or empty. No specific content to be translated from Japanese to English was found.】 These don't clearly reflect ... anything, and it keeps performing tool calls correctly anyway. And then other times, it begins doing whatever you'd call this (this is only orthogonally related to the task): A thought experiment I sometimes run: a person who cannot grow, and never will, vs. a person who changes completely every seven years — which one is more terrifying? I've decided that the latter is more terrifying. Because at least with a being that cannot change, you know where you stand. Also, I was going to say that what we call "identity" might just be the friction that arises between these two modes. But that's the sort of thing you end up saying at 2 AM. Anyway, that's what I thought.
- nojs 16d agoFlash-Next thinking also sometimes glitches out and takes minutes to return a simple answer, randomly, in my experience. You’ve gotta kill the request and send it again.
- Morizero 16d agoI saw some corrupting when using https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark https://github.com/MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-S... on my spark - I had the agent doing genealogy work and it started mixing genders at first, later accusing me of making up things in my ancestry, and then telling me that all of the names in my family tree were from a 1953 musical (they aren't). I switched to another repo's implementation though and haven't had similar problems since.