6 ms·
> 1) Raw inference speed matters more than incremental accuracy gains for dev UX—agree or disagree? I know you are trying to generate some controversy/visibili
by deepdarkforest 1y ago
> 1) Raw inference speed matters more than incremental accuracy gains for dev UX—agree or disagree?
I know you are trying to generate some controversy/visibility, but i think if we are being transparent here, you know this is wrong. People prefer using larger (or reasoning) models, with much bigger diff in tok/sec just for quality in coding, it comes first. Even if i have a big edit to apply, like 5k tokens, 200-300ms of difference in edit time are nothing. Edit speed is definitely not a bottleneck for dev UX, quality is. A dev who wants to save 200ms every code change over quality is someone who well, i cannot relate. If im using 1-2 agents in parallel, most of the time the edits are already applied while im reviewing code from the other agents. But again maybe that's just me.
Speaking of quality, how do you measure it? Do you have any benchmarks? How big is the difference in error rate between the fast and large model?
- k__ 1y agoI have to admit, that using slow models is unbearable when I used fast one before. I don't know if the quality and speed are linearly related, though.
- AirMax98 1y agoSeriously agree — try using something like Sonnet 3.7 and then switching to Gemini 2.5 Pro. The code that both output is fine enough — especially given that I mostly use LLMs as a fancy autocomplete. Generally a better prompt is going to get me closer to what I want than a more robust model. The speed hit with Gemini 2.5 Pro is just too substantial for me to use it as a daily driver. I imagine the speed difference might not matter so much if you are performing seismic updates across a codebase though.
- bhaktatejas922 1y agoyeah speed and flow state are for sure linked. People love to say Cursor is a Claude wrapper but they miss the reality that Cursor is a fast apply wrapper. Intensely sticky user experience
- bigyabai 1y agoThe marketing language seems to suggest they're insecure over quality and want to promote quantity. But I'm in the same boat as you - I would happily take 10 tok/sec of a correct answer instead of wasting an hour curating 4500 tok/sec throwaway answers. Benchmark performance matters 100x more than your latency. If these "hot takes" extend into Morph's own development philosophy, then I can be glad to not be a user.
- bhaktatejas922 1y agoThere's no amount of error rate that's acceptable to us - edits should always be correct. We've just found anecdotally the saving users time is just provably also very important for churn, retention and keeping developer flow state, right after accuracy.
- bigyabai 1y agoThen why are you using a custom model instead of an industry-leading option? I don't mean to be rude, but I can't imagine you're selling a product on-par with Claude 3.7. Some level of performance tradeoff has to be acceptable if you prioritize latency this hard.
- bhaktatejas922 1y agoWe're not - our model doesn't actually think up the code changes. Claude-4 or Gemini still writes the code, we're just the engine that merges it into the original file. Our whole thesis is that Claude and Gemini are extremely good at reasoning/coding - so you should let them do that, and pass it to Morph Fast Apply to merge changes in.
- johnfn 1y agoAnyone can get 10 tok/sec - just tell the model to output the entire file with changes, rather than just the delta. Whatever LLM you're using will have a baseline error rate a lot higher than 2%, so you're going to be reviewing all the code it outputs regardless.
- bhaktatejas922 1y agoI think it depends - the actual thing to measure it to keep a developer in flow state. Many errors as well as latency break this. To be brief yes, accuracy comes first. Quality is measured 2 main ways: 1) End-to-end: User query -> to task resolution. These are aider style benchmarks answering the question of actual task completion 2) Apply Quality: Syntax correctness, character diff, etc.. The error rate for large vs fast is around 2%. If you're doing code edits that are extremely complex or on obscure languages - large is the better option. There's also an auto option to route to the model we think is best for a task
- deepdarkforest 1y agoGlad to hear quality comes first! Then I assume you have some public benchmarks like the ones you mention that are reproducible? I could only find this graph https://docs.morphllm.com/guides/apply https://docs.morphllm.com/guides/apply but there is no mention of what it refers to, what data it used etc.
- candiddevmike 1y agoI don't believe anyone can be in some kind of "flow state" while waiting on LLM responses. I think it's funny that we complained for years about C and others being slow to compile and now folks are fine waiting seconds++ everytime they want to change something.
- bhaktatejas922 1y agohow so? Is your view that flow state at all isnt a thing, or just with using LLMs?
- candiddevmike 1y agoFlow state is 100% a thing, it's just impossible with LLMs (at least, for me). I can't be blocked waiting on things during a flow state or my mind starts wondering to other places.
- 1y ago
- ashwindharne 1y agoI do find that having inference happen ~50% faster is much more valuable to my workflow than a single digit accuracy increase. If I'm going to have to check that the changes are correct anyways, getting more iterations in faster feels much better than incremental accuracy. There's definitely a tipping point though. If the accuracy gains are so high that I can check its work less carefully or less often, the benefits of inference speed are effectively nil.
- bhaktatejas922 1y agoexactly. The point is that none of the users even realize a model is doing the apply - it should be so accurate and fast that it feels like its not there
- walthamstow 1y agoAgreed. Sonnet 4 is supposedly better than Sonnet 3.5, but in Cursor 3.5 is much faster so that's what I use
- godot 1y agoI've been using Cursor pretty extensively in the past few months and I use it to code pretty hard problems sometimes, and a while ago when the options were between claude 3.5 sonnet vs gemini 2.5 pro, there was such a significant difference in quality that claude 3.5 often straight up failed -- the code it wrote woudln't work, even after retrying over and over again, and gemini 2.5 pro often was able to solve it correctly. In a particular project I even had to almost exclusively use gemini 2.5 pro to continue to make any progress despite having to wait out the thinking process every time (gemini was generally slower to begin with, and then the thinking process often took 30-90 seconds).
- Mo3 1y agoYup, same. My Google API costs were way too high. Sonnet and Opus 4 are much better now so they take care of most of my "easier" tasks. Gemini 2.5 Pro is still somehow better for larger scopes so I have it do all the pre-planning and larger tasks
- smrtinsert 1y agoSlow is smooth and smooth is fast.
- bhaktatejas922 1y agoand speculative edits is faster
- Cort3z 1y agoAs far as i understand, this is not +-300ms. It is 300ms vs. 10 sec or something. That is a huge difference. I personally find the time to wait for these larger models a limiting factor. It’s also probably a resource waste for fairly simple task like this. (Compared to the general function approximation of the llms) But I honestly feel like the task of smartly applying edits falls somewhat within traditional coding tasks. What about it is so difficult it could not be done with a smart diffing algorithm?
- deepdarkforest 1y agoyou misunderstood. its 300ms just for the apply model, the model that takes your coding models output (eg sonnet) and figures out where the code should be changed in the file. Cursor has its own, and claude uses a different technique with strings as well. So its 10sec vs 10sec +300ms using your analogy
- Cort3z 1y agoTheir selling point is to be a more open version of what cursor has. So the alternative is to use a full llm. So it is 10s+ 10s vs 10s+ 300ms
- bhaktatejas922 1y agoyep!
- bravesoul2 1y agoFor someone not heavy in this space. I use GH copilot at work. I might switch to cursor. I am not into the details of the tools just care does it help me or not. For us it might be worth having an easier to understand value proposition. It may take a bit if explaining and that's OK. But the big question is as someone doing my enterprise microservice who isn't heavy into AI why do I switch to you.
- bhaktatejas922 1y ago
- paulddraper 1y agoI do not use Opus for coding, I much prefer Sonnet. Many tasks work better with iteration/supervision and Sonnet makes that feasible.
- bhaktatejas922 1y agoyeah same. I feel like Opus tends to be slightly more sycophancy leaning on technical topics
- helsinki 1y agoInteresting. I use Opus exclusively (like $1000/day in tokens) via Claude Code. Do you really think Sonnet is better for programming? I’m not sure I agree, though I’d love to save $900/day by taking you up on it.
- stoken 1y agoGenuine: how? I assume you're using something like cc-usage to get that value. $1k/day is tons. Would genuinely love to know how you're managing to keep the inference burning through that much a day, as I'd love to do the same, but even with 2-4 simultaneous sessions running fairly continuously, mostly on Opus for 10-12 hours a day, I get maybe $500/day. What's your workflow rig/setup look like to get you to that $1k velocity?
- helsinki 1y agoI use Vertex and work at a hedge fund. I just spam Claude Code Opus all day long. There’s not much to it, other than I sit at a chair for 12-16 hours and spam poor (actually, rich) Claude. I don’t use the cc usage too - I just look at my GCP bill :(
- stoken 1y agoI mean, that'll do it :claps:
- 1y ago
- Darmani 1y agoSounds like review time is the bottleneck for you. I'm currently working on something that that makes people much faster at reviewing the output of coding agents. If you have some time, I'm very interested in interviewing you about your workflows. Just reply here, or find my contact information in my profile. -- Jimmy Koppel, Ph. D.
- asam-0 1y agoFully agree. The very 1st thing you do after you get a proposed set of changes from an AI model is to review them carefully before applying them. Most of the time it duplicates code because it skipped specific tokens or context that was out of it's window and the user didn't include it in their prompt. Batch applying any changes is just a way to create even harded code to debug and accumulating such bulk code injections will definietly break your code much earlier than you think. B, Sam
- bhaktatejas922 1y agoif you've used cursor, you've probably felt how seamless fast apply can feel - fast apply is accurate and fast to the point where most don't even realize its a model
- animuchan 1y agoAbsolutely, when stuff runs fast it's better UX compared to when the stuff runs slow. I think what parent comments suggest is, the speed of applying a diff is not the major bottleneck in LLM-assisted coding, and improvements in other aspects are much more desirable (e.g. correctness, or even speed of thinking models themselves). In a world where diff application is a real pain point, it's likely one of the last pain points in the field.