10 ms·
We rewrote our Rust WASM parser in TypeScript and it got faster
- rpodraza 7mo agoPress x to doubt
- DaleBiagio 7mo ago[dead]
- neuropacabra 7mo agoThis is very unusual statement :-D
- blundergoat 7mo agoThe real win here isn't TS over Rust, it's the O(N²) -> O(N) streaming fix via statement-level caching. That's a 3.3x improvement on its own, independent of language choice. The WASM boundary elimination is 2-4x, but the algorithmic fix is what actually matters for user-perceived latency during streaming. Title undersells the more interesting engineering imo.
- shmerl 7mo agoMore like a misleading clickbait.
- sroussey 7mo agoYeah, though the n^2 is overstating things. One thing I noticed was that they time each call and then use a median. Sigh. In a browser. :/ With timing attack defenses build into the JS engine.
- Aurornis 7mo ago> Title undersells the more interesting engineering imo. Thanks for cutting through the clickbait. The post is interesting, but I'm so tired of being unnecessarily clickbaited into reading articles.
- socalgal2 7mo agosame for uv but no one takes that message. They just think "rust rulez!" and ignore that all of uv's benefits are algo, not lang.
- estebank 7mo agoSome architectures are made easier by the choice of implementation language.
- crubier 7mo agoIn my experience Rust typically makes it a little bit harder to write the most efficient algo actually.
- catlifeonmars 7mo agoThat’s usually ok bc in most code your N is small and compiler optimizations dominate.
- Defletter 7mo agoWould you be willing to give an example of this?
- lukeweston1234 7mo agoNot OP, but one example where it is a bit harder to do something in Rust that in C, C++, Zig, etc. is mutability on disjoint slices of an array. Rust offers a few utilities, like chunks_by, split_at, etc. but for certain data structures and algorithms it can be a bit annoying. It's also worth noting that unsafe Rust != C, and you are still battling these rules. With enough experience you gain an understanding of these patterns and it goes away, and you also have these realy solid tools like Miri for finding undefined behavior, but it can be a bit of a hastle.
- catlifeonmars 7mo ago
- azakai 7mo agoO(N²) -> O(N) was 3.3x faster, but before that, eliminating the boundary (replacing wasm with JS) led to speedups of 2.2x, 4.6x, 3.0x (see one table back). It looks like neither is the "real win". both the language and the algorithm made a big difference, as you can see in the first column in the last table - going to wasm was a big speedup, and improving the algorithm on top of that was another big speedup.
- hrmtst93837 7mo ago[flagged]
- nulltrace 7mo agoYeah the algorithmic fix is doing most of the work here. But call that parser hundreds of times on tiny streaming chunks and the WASM boundary cost per call adds up fast. Same thing would happen with C++ compiled to WASM.
- hrmtst93837 7mo ago[flagged]
- catlifeonmars 7mo agoYou’re not wrong, but that win would not get as many views. It’s not clickbaity enough
- adastra22 7mo agoNo AI generated comments on HN please.
- wolvesechoes 7mo ago> The real win here isn't TS over Rust Kinda is. We came up with abstractions to help reason about what really matters. The more you need to deal with auxillary stuff (allocations, lifetimes), more likely you will miss the big issue.
- coldtea 7mo agoThe opposite: the more you rely on abstractions the more you miss the lower level optimization opportunities and loose understanding of algorithms and hardware.
- wolvesechoes 7mo ago> of algorithms Yes, sprinkling your code logic with malloc, .clone() or lifetime annotations on the other hand brings algorithmic enlightenment.
- coldtea 7mo agoDealing and having to think about the cost of malloc, clone() and lifetimes, brings algorithmic enlightenment more than working on an high abstraction ivory tower where things "magically happen". Is your argument that the average Python or Typescript dev gets to think and care more about algorithms than the average C/C++/Rust dev?
- zahrevsky 7mo agoThey even directly conclude at the end of the article that improvements in algorithm are more important than the choice of language: > Algorithmic complexity improvements dominate language-level optimisations. Going from O(N²) to O(N) in the streaming case had a larger practical impact than switching from WASM to TypeScript. Yet they still have chosen to put the “Rust rewrite” part in the title. I almost think it's a click bait.
- deleted 7mo ago[deleted]
- dmix 7mo agoThat blog post design is very nice. I like the 'scrollspy' sidebar which highlights all visible headings. Claude tells me this is https://www.fumadocs.dev/ https://www.fumadocs.dev/
- nine_k 7mo ago"We rewrote this code from language L to language M, and the result is better!" No wonder: it was a chance to rectify everything that was tangled or crooked, avoid every known bad decision, and apply newly-invented better approaches. So this holds even for L = M. The speedup is not in the language, but in the rewriting and rethinking.
- MiddleEndian 7mo agoNow they just need a third party who's never seen the original to rewrite their TypeScript solution in Rust for even more gains.
- nine_k 7mo agoIndeed! But only after a year or so of using it in production, so that the drawbacks would be discovered.
- baranul 7mo agoTruth. You can see improvement, even rewriting code in the same language.
- azakai 7mo agoYou're generally right - rewrites let you improve the code - but they do have an actual reason the new language was better: avoiding copies on the boundary. They say they measured that cost, and it was most of the runtime in the old version (though they don't give exact numbers). That cost does not exist at all in the new version, simply because of the language.
- necovek 7mo agoIt's doing copies and (de)serialization on both sides into native data types. If they used raw byte structures, implemented the caching improvements on the wasm side, the copies might not be as bad. But they still have an issue with multi-language stack: complexity also has a cost. Python/C combo does not have this issue because you can work with Python types natively in C, but otherwise, this is a cross-language conversion issue, and not a Rust issue at all.
- spankalee 7mo agoI was wondering why I hadn't heard of Open UI doing anything with WASM. This new company chose a very confusing name that has been used by the Open UI W3C Community Group for over 5 years. https://open-ui.org/ https://open-ui.org/ Open UI is the standards group responsible for HTML having popovers, customizable select, invoker commands, and accordions. They're doing great work.
- caderosche 7mo agoWhat is the purpose of the Rust WASM parser? Didn't understand that easily from the article. Would love a better explanation.
- joshuanapoli 7mo agoThey use a bespoke language to define LLM-generated UI components. I think that this is supposed to prevent exfiltration if the LLM is prompt-injected. In any case, the parser compiles chunks streaming from the LLM to build a live UI. The WASM parser restarted from the beginning upon each chunk received. Fixing this algorithm to work more incrementally (while porting from Rust to TypeScript) improved performance a lot.
- evmar 7mo agoBy the way, I did a deeper dive on the problem of serializing objects across the Rust/JS boundary, noticed the approach used by serde wasn’t great for performance, and explored improving it here: https://neugierig.org/software/blog/2024/04/rust-wasm-to-js.html https://neugierig.org/software/blog/2024/04/rust-wasm-to-js....
- slopinthebag 7mo agoDid you try something like msgpack or bebop?
- SCLeo 7mo agoThey should rewrite it in rust again to get another 3x performance increase /s
- slowhadoken 7mo agoAm I mistaken or isn’t TypeScript just Golang under the hood these days?
- iainmerrick 7mo agoHmm, there's an in-progress rewrite of the TypeScript compiler in Go; is that what you mean? I don't think that's actually out yet, and more importantly, it doesn't change anything at runtime -- your code still runs in a JS engine (V8, JSC etc).
- koakuma-chan 7mo agonpm i -D @typescript/native-preview You can use it today.
- slowhadoken 7mo agoYou get a gold star.
- jeremyjh 7mo agoThere is too much wrong here to call it a mistake.
- wiseowise 7mo agoYes, you've uncovered grand conspiracy.
- slowhadoken 7mo agoIt’s funny because a decade ago people said I was crazy for thinking that Oracle owning JS was going to become an issue in the future.
- szmarczak 7mo ago> Attempted Fix: Skip the JSON Round-Trip > We integrated serde-wasm-bindgen So you're reinventing JSON but binary? V8 JSON nowadays is highly optimized [1] and can process gigabytes per second [2], I doubt it is a bottleneck here. [1] https://v8.dev/blog/json-stringify https://v8.dev/blog/json-stringify [2] https://github.com/simdjson/simdjson https://github.com/simdjson/simdjson
- kam 7mo agoNo, serde-wasm-bindgen implements the serde Serializer interface by calling into JS to directly construct the JS objects on the JS heap without an intermediate serialization/deserialization. You pay the cost of one or more FFI calls for every object though. https://docs.rs/serde-wasm-bindgen/ https://docs.rs/serde-wasm-bindgen/
- szmarczak 7mo agoIndeed, you're right. However, it still needs to encode and decode strings. WASM just needs native interop.
- nallana 7mo agoWhy not a shared buffer? Serializing into JSON on this hot path should be entirely avoidable
- devnotes77 7mo ago[dead]
- mavdol04 7mo agoI think a shared array just avoids the copy, not the serialization which is the main problem as they showed with serde-wasm-bindgen test
- notnullorvoid 7mo agoYou can avoid the serialization in WASM by pushing structured bytes to the SharedArrayBuffer, then do serialization in JS which should be relatively cheap compared to pushing JSON strings across the boundary.
- ivanjermakov 7mo agoGood software is usually written on 2nd+ try.
- joaohaas 7mo agoGod I hate AI writing. That final summary benchmark means nothing. It mentions 'baseline' value for the 'Full-stream total' for the rust implementation, and then says the `serde-wasm-bindgen` is '+9-29% slower', but it never gives us the baseline value, because clearly the only benchmark it did against the Rust codebase was the per-call one. Then it mentions: "End result: 2.2-4.6x faster per call and 2.6-3.3x lower total streaming cost." But the "2.6-3.3x" is by their own definition a comparison against the naive TS implementation. I really think the guy just prompted claude to "get this shit fast and then publish a blog post".
- chvish 7mo agoThis. It’s so annoying to read these types of blogs now where the writer clearly didn’t put the effort to understand things fully or atleast review the blog their LLM wrote. Who is this useful for?
- JimDabell 7mo agoThe article as a whole makes no sense. They are generating UI with an LLM. How fast the UI appears to the user is going to be completely dictated by the speed of the LLM, not the speed of the serialisation.
- rabisg 7mo agoas an author of the blog - ouch did a little bit more than prompt claude but a lot of claude prompting was definitely involved I understand your frustration with AI writing though. We are a small team and given our roadmap it was either use LLMs to help collate all the internal benchmark results file into a blog or never write it so we chose the former. This was a genuinely surprising and counterintuitive result for us, which is why we wanted to share it. Happy to clarify any of the numbers if helpful.
- patapim 7mo ago[flagged]
- nssnsjsjsjs 7mo agoRewrite bias. Yoy want to also rewrite the Rust one in Rust for comparison.
- jeremyjh 7mo agoIt would be surprising if rewriting in Rust could change the WASM boundary tax that the article identified as the actual problem.
- rabisg 7mo ago(author here) We'd be really surprised if a rewrite could fix the boundary tax but if it does, we'd happily move over to it. People (including me) really underestimate how insanely fast browser's JSON.parse is
- rented_mule 7mo agoSomething not unlike this happened to me when moving some batch processing code from C++ to Python 1.4 (this was 1997). The batch started finishing about 10x faster. We refused to believe it at first and started looking to make sure the work was actually being done. It was. The port had been done in a weekend just to see if we could use Python in production. The C++ code had taken a few months to write. The port was pretty direct, function for function. It was even line for line where language and library differences didn't offer an easier way. A couple of us worked together for a day to find the reason for the speedup. Just looking at the code didn't give us any clues, so we started profiling both versions. We found out that the port had accidentally fixed a previously unknown bug in some code that built and compared cache keys. After identifying the small misbehaving function, we had to study the C++ code pretty hard to even understand what the problem was. I don't remember the exact nature of the bug, but I do remember thinking that particular type of bug would be hard to express in Python, and that's exactly why it was accidentally fixed. We immediately started moving the rest of our back end to Python. Most things were slower, but not by much because most of our back end was i/o bound. We soon found out that we could make algorithmic improvements so much more quickly, so a lot of the slowest things got a lot faster than they had ever been. And, most importantly, we (the software developers) got quite a bit faster.
- envguard 7mo ago[flagged]
- sincerely 7mo agoAI account
- DaleBiagio 7mo ago[dead]
- apitman 7mo agoI don't think the better software part is playing out
- slopinthebag 7mo agoThis article is obviously AI generated and besides being jarring to read, it makes me really doubt its validity. You can get substantially faster parsing versus `JSON.parse()` by parsing structured binary data, and it's also faster to pass a byte array compared to a JSON string from wasm to the browser. My guess is not only this article was AI generated, but also their benchmarks, and perhaps the implementation as well.
- StilesCrisis 7mo agoIt's vibe code all the way down!
- jeremyjh 7mo ago> The openui-lang parser converts a custom DSL emitted by an LLM into a React component tree. > converts internal AST into the public OutputNode format consumed by the React renderer Why not just have the LLM emit the JSON for OutputNode ? Why is a custom "language" and parser needed at all? And yes, there is a cost for marshaling data, so you should avoid doing it where possible, and do it in large chunks when its not possible to avoid. This is not an unknown phenomenon.
- kennykartman 7mo agoI dream of the day in which there is no need to pass by JS and Wasm can do all the job by itself. Meanwhile, we are stuck.
- envguard 7mo agoThe WASM story is interesting from a security angle too. WASM modules inheriting the host's memory model means any parsing bugs that trigger buffer overreads in the Rust code could surface in ways that are harder to audit at the JS boundary. Moving to native TS at least keeps the attack surface in one runtime, even if the theoretical memory safety guarantees go down.
- marcosdumay 7mo agoIt would be great if people stopped dismissing the problem that WASM not being a first-class runtime for the web causes.
- dualblocksgame 7mo ago[flagged]
- vmsp 7mo agoNot directly related to the post but what does OpenUI do? I'm finding it interesting but hard to understand. Is it an intermediate layer that makes LLMs generate better UI?
- rabisg 7mo agoIts the library that bridges the gap between LLMs and live UI. Best example would be to imagine you want to build interactive charts within your AI agent (like Claude) The most obvious approach would be to let LLMs generate code and render it but that introduces problems like safety, UI consistency and speed. OpenUI solves those problems and provides a safe, consistent and token optimized runtime for the LLMs to render live UI
- aquariusDue 7mo agoIs it kinda similar to the new GenUI SDK for Flutter in that sense? https://docs.flutter.dev/ai/genui https://docs.flutter.dev/ai/genui
- rabisg 7mo agoHaven't looked in depth but yes it feels like they are solving the same problem. This is an alternative to json-render by Vercel or A2UI by Google which I'm guessing the flutter implementation is based on
- ConanRus 7mo ago[dead]
- owenpalmer 7mo agoSo this is an issue with WASM/JS interop, not with Rust per se?
- aimarketintel 7mo ago[flagged]
- derodero24 7mo ago[flagged]
- measurablefunc 7mo agoI tried a similar experiment recently w/ FFT transform for wav files in the browser and javascript was faster than wasm. It was mostly vibe coded Rust to wasm but FFT is a well-known algorithm so I don't think there were any low hanging performance improvements left to pick.
- wintermute4282 7mo agoIt looks like FFTW3 is working on wasm support: https://github.com/FFTW/fftw3/issues/293 https://github.com/FFTW/fftw3/issues/293 You could also try pretty fast fft: https://github.com/JorenSix/pffft.wasm https://github.com/JorenSix/pffft.wasm
- measurablefunc 7mo agoIt was just an experiment in vibe coding. It's easy enough to try different architectures w/ AI coding but what I wanted to see was whether naive numeric calculations were faster w/ wasm or javascript & it turned out that javascript was faster so the performance trade-off between wasm & javascript is not as simple as between a high-level language like python & SIMD optimized C/assembly.
- simonbw 7mo agoYeah if you're serializing and deserializing data across the JS-WASM boundary (or actually between web workers in general whether they're WASM or not) the data marshaling costs can add up. There is a way of sharing memory across the boundary though without any marshaling: TypedArrays and SharedArrayBuffers. TypedArrays let you transfer ownership of the underlying memory from one worker (or the main thread) to another without any copying. SharedArrayBuffers allow multiple workers to read and write to the same contiguous chunk of memory. The downside is that you lose all the niceties of any JavaScript types and you're basically stuck working with raw bytes. You still do get some latency from the event loop, because postMessage gets queued as a MacroTask, which is probably on the order of 10μs. But this is the price you have to pay if you want to run some code in a non-blocking way.
- jesse__ 7mo agoThis should be the top comment
- osullivj 7mo agoStrongly agree from an Emscripten C++ wasm pov: it's key to minimise emscripten::val roundtrips. Caches must be designed for rectilinear data geometry, and SharedArrayBuffers are the way for bulk data. But only JS allows us to express asynchrony, so we need an on_completion callback design at the lang boundary.
- tankenmate 7mo agoIndeed a whole class of issues become moot if you just don't use javascript anywhere. In the browser world this is obviously difficult/impossible; I look forward to the day when WASM can run natively in a browser and doesn't need javascript at all, DOM, network, etc, etc. On the server side? Just steer clear of the javascript ecosystem altogether.
- fHr 7mo agoSo the actual processing is faster in rust/c/c++ but the marshaling costs are so big so ts is faster in this case? No vlue how something like swc does this but there it's way faster then babel.
- sakesun 7mo agoI heard a lot of similar stories in the past when I started using Python 20+ years ago. A number of people claimed their solutions got faster when develop in Python, mainly because Python make it easier to quickly pivot to experiment with various alternative methods, hence finally yield at more efficient outcome at the end.
- horacemorace 7mo agoI’m more of a dabbler dev/script guy than a dev but Every. single. thing I ever write in javascript ends up being incredibly fast. It forces me to think in callbacks and events and promises. Python and C (or async!) seem easy and sorta lazy in comparison.
- jesse__ 7mo agoThis somehow reminds me of the days when the fastest way to deep copy an object in javascript was to round trip through toString. I thought that was gross then, and I think this is gross now
- athrowaway3z 7mo agoIts also worth underlining that it's not just "The parsing computation is fast enough that V8's JIT eliminates any Rust advantage", but specifically that this kind of straight-forward well-defined data structures and mutation, without any strange eval paths or global access is going to be JITed to near native speed relatively easily.
- deleted 7mo ago[deleted]
- mwcampbell 7mo agoI hope we can still get to a point where wasm modules can directly access the web platform APIs and get JS out of the picture entirely. After all, those APIs themselves are implemented in C++ (and maybe some Rust now).
- shevy-java 7mo agoSo ... Rust. WASM. TypeScript. I am slowly beginning to understand why WASM did not really succeed.
- bulbar 7mo agoIs this an outlier or has Rust started to be part of the establishment and being 'old' so that people want to share their "moving away from Rust" stories? I didn't mind reading articles that are not about how Rust is great in theory (and maybe practice).
- quotemstr 7mo agoThere's a certain segment of the industry that's always chasing the newest thing. Many of them like Zig for some ghastly reason. That said, Rust does have real problems. Manual memory management sucks. People think GC is expensive? Well, keep in mind malloc() and free() take global locks! People just have totally bogus mental models of what drives performance. These models lead them to technical nonsense.
- zozbot234 7mo agoThis story is about moving away from WASM for an application that's unsuitable for it. It's not really about Rust.
- notnullorvoid 7mo agoIt's not an unsuitable application for WASM. They could've drastically reduced the WASM boundary impact if instead of mapping to JSON in Rust they streamed out structured bytes to JS then mapped to JSON there. And the streaming fix was language independent. So it's more so a story about architectural mistakes.
- Yanko_11 7mo ago[dead]
- Dwedit 7mo agoJS and WASM share the main arraybuffer. It's just very not-javascript-like to try to use an arraybuffer heap, because then you don't have strings or objects, just index,size pairs into that arraybuffer. Anyway, Javascript is no stranger to breaking changes. Compare Chromium 47 to today. Just add actual integers as another breaking change, then WASM becomes almost unnecessary.
- fHr 7mo agoI almost can't believe this swc for example is 80x faster then babeljs.
- gettingoverit 7mo agoIn ye olden days of WASM just added to the browser, the difference between native JS and boost::spirit in WASM was x200. In their worst case it was just x5. We clearly have some progress here.
- deleted 7mo ago[deleted]
- pjmlp 7mo agoThis is why, when a programming language already has tooling for compilers, being it ahead of time, or dynamic, it pays off to first go around validating algorithms and data structures before a full rewrite. Additionally even after those options are exhausted, only a key parts might need a rewrite, not the whole thing. However, I wonder how many care about actually learning about algorithms, data structures and mechanical sympathy in the age of Electron apps. It feels quite often that a rewrite is chosen, because knowing how to actually apply those skills is the CS stuff many think isn't worthwhile learning about.
- coldtea 7mo ago>However, I wonder how many care about actually learning about algorithms, data structures and mechanical sympathy in the age of Electron apps. Never mind the age of Electron apps, even fewer care about those in the age of agents.
- pjmlp 7mo agoAgreed, however I would assert that in the age of agents, programming languages will become irrelevant to most, other those lucky enough druids to write AI runtime stack, at the AI overlords. And those will still care about CS.
- moomin 7mo ago“We saw huge speed-ups when changing technology.” Looks inside “The old implementation had some really inappropriate choices.” Every time.
- LunaSea 7mo agoThis has been known by Node.js developers for a while with many C++ core and NPM modules being rewritten in JavaScript to improve performance.
- bluelightning2k 7mo agoGreat write up. It feels like craft in the age of slop. Not sold about the fundamental idea of OpenUI though. XML is a great fit for DSLs and UI snippets.
- twoodfin 7mo agoAre you kidding? To the extent this was “crafted” it was by an LLM from somebody’s notes in a prompt. The other day, someone linked back to this 2018 post on finding a cache coherency bug in the Xbox 360 CPU: https://randomascii.wordpress.com/2018/01/07/finding-a-cpu-design-bug-in-the-xbox-360/ https://randomascii.wordpress.com/2018/01/07/finding-a-cpu-d... So much more genuinely engaging than any of the AI-“enhanced” sloppy, confused, trite writing that gets to the front page here daily because it’s been hyper-optimized for upvotes.
- rabisg 7mo agoWe tried all formats - XML, json, jsonl, even toon - before deciding that we need to invest in OpenUI Lang The primary motivation was speed and schema cohesion. We were running a JSON based format, Thesys C1, in production for a year before we realized we cannot add features fast enough because we were fighting the LLMs at multiple levels. It's probably too much to write in a comment but we'd like to write about the motivation and all the things we tried ona a separate blog soon
- arthurjean 7mo ago[dead]
- wangnaihe 7mo ago[dead]
- diablevv 7mo ago[flagged]
- abitabovebytes 7mo ago[dead]
- ryguz 7mo ago[flagged]
- ata-sesli 7mo ago[dead]
- mohsen1 7mo agoWhen there is a solid test harness, AI Coding can do magic! It was able to beat XZ on its own game by a good margin: https://github.com/mohsen1/fesh https://github.com/mohsen1/fesh
- applfanboysbgon 7mo ago> I had no idea how any of this works. This is apparent. xz's own game is not "a specialized compression pre-processor for x86_64 ELF binaries.". xz's own game is a general-purpose compression utility suited for a range of tasks, not optimized for one ridiculously specific domain. Also, any compression benchmark really ought to include speed of de/compression, not only compression ratio, as compression algorithms occupy along a scale trying to maximize one trade-off or another.
- mohsen1 7mo agoI never claimed to beat xz as a general-purpose compressor. .tar.xz is the dominant format for Linux source tarballs and distro packages. So optimizing for ELF + x86_64 is optimizing for a very real and common case, not some toy benchmark. btw goal of the project was not building a production ready solution. It was curious case of black box software development. Compression is great because input and output are precise bits. As for speed, I think it's comparable since it's using most of XZ infra anyways.
- Yanko_11 7mo ago[dead]
- leontloveless 7mo ago[dead]
- gavinray 7mo agoWhy weren't you able to use WASM shared heaps to get zero-copy behavior? AFAIK, you can create a shared memory block between WASM <-> JS: https://developer.mozilla.org/en-US/docs/WebAssembly/Reference/JavaScript_interface/Memory#creating_a_shared_memory https://developer.mozilla.org/en-US/docs/WebAssembly/Referen... Then you'd only need to parse the SharedArrayBuffer at the end on the JS side
- hackwaly_new 7mo agoYou don't need rewriting if you are using MoonBit. It gives you wasm, wasm-gc, js at once.
- mpajares 6mo agoHad the opposite experience. Our JS FEM solver (~550ms per load case) was rewritten in Rust and dropped to ~270ms. But we compile to native.exe, not WASM — we call it via stdin/stdout with JSON from a Node.js compute engine. Tried the WASM route first but the serialization overhead for large stiffness matrices ate the gains, exactly like this article describes. Native binary + stdin/stdout turned out to be the sweet spot: no boundary tax, no FFI, and you get full native SIMD. The sparse solver variant (sprs crate, COO/CSC assembly) scales even better for larger models.