7 ms·
> Users don't care if you wrote it in Rust Totally agree! But they do care if Word corrupts a document that’s been open for 10 days (a bug that a principal eng
by briantkelley 4y ago
> Users don't care if you wrote it in Rust
Totally agree! But they do care if Word corrupts a document that’s been open for 10 days (a bug that a principal engineer spent many, many weeks hunting down).
> If you take 20% longer to write it in Rust, that's short term cost for long term gain, but users don't think that way.
There are many steps between writing code and delivering a product to users. Sure, I can commit some C++ in less wall time than Rust (though I’m not actually sure that’s true!). But when we have to delay a major release by 2 weeks because there’s a heap corruption bug in the flagship feature, everyone loses.
> the pressure is always on
Sigh, yeah. It’s difficult for organizations to act in their own best interests. If a codebase primarily has architecture debt and little implementation quality debt, it’s much easier to keep taking the shortcuts to ship now. But with implementation quality debt, UB is lurking behind every corner and can cause the project to go off the rails at any time. This is why I would like to see more research about the value of Rust with respect to the predominant cost of engineering in extremely large systems: maintenance.
> Users don't care if you need 30% fewer CI servers
Sure, but DevOps does. When we have to drop jobs due to capacity constraints but mark the commit as stable anyway, it’s just a matter of time until a major regression sneaks in.
> The CEO doesn't care if one additional dev has to spend his time making sure asan, ubsan, etc are running properly
Haha, yeah, in practice it‘s always one dev maintaining the asan, ubsan, etc. loops! But, it’s up to engineers writing code to make sure it’s exercised by those systems, so in reality the loops are highly under utilized and eventually hard-to-find UB bug slips in.
> Meanwhile the devs are wrangling the complexity of the software itself, which requires mental resources.
Yes. And encoding ownership and lifetimes into the type system is a huge reduction in the mental tax when working through code that hasn’t been touched in years, or wiring up some feature to an interface another team just landed that’s not documented at all. Not having to reverse engineer the code to discern ownership and concurrency patterns is a huge time saver.
> Any additional complexity introduced by Rust slows down development and frustrates the team
I agree artificial complexity due to limitations of Rust’s soundness checks are frustrating, but a cost I’m willing to pay given the long term maintenance benefits. But the cost of system design complexity has to be paid at some point, and I’d rather pay as much as possible at compile time vs. discovering problems at run time.
> And the bugs are still there.
Certainly. But essentially eliminating classes of UB bugs that are hard to repro saves all the engineering teams time.
> Operating efficiency becomes important only after the product is out and stable, with paying customers and product-market fit
Yes. Total cost of ownership is what I want to see more research in.
And I think this is the source of our deferring perspectives. My original comment wanted more research on TCO, but I don’t think that’s particularly applicable before product-market fit.
> but by then, why boil the ocean with a rewrite?
Who said anything about a rewrite? Use Rust for new development. And, as the system grows, components will be rewritten (whether due to new business needs or to improve maintainability). Don’t keep playing with fire when it’s no longer necessary!
> Sometimes leadership can be sold on a new and improved v2
Perhaps I’m jaded, but my motto is “Rewrites always fail.”
> we need languages that make it less mentally taxing for developers to write product software
For extremely large systems whose codebases have been written over decades, we need languages that make it less mentally taxing for developers to maintain product software. The complexity of maintaining these systems is why FAANG employs tens of thousands of engineers who seem to move slower than an aircraft carrier. The difficulty in reasoning about system behavior greatly impedes progress.
- pjmlp 4y ago> Totally agree! But they do care if Word corrupts a document that’s been open for 10 days (a bug that a principal engineer spent many, many weeks hunting down). Caused by memory corruption, or the remaining 30% of logical bugs present in any safe language?
- briantkelley 4y agoIIRC, that specific bug was due to a few issues: 1. Mismatched type definitions in different translation units caused an implicit 32 to 16 bit conversion. 2. An addition to the 16 bit value overflowed, causing it to become negative. (The “run for 10 days” thing helped the value get to the point of overflowing.) 3. The negative value led to an out of bounds write, corrupting a key bookkeeping data structure. Rust would have failed to compile at step 1, and would have panicked at 3 (and potentially 2, depending on compilation settings). I’m not sure what other safe languages are candidates for this domain, but I would suspect these issues would likely be similarly identified. That bug, unfortunately was identified after release. The release blocking heap corruption bug I mentioned was due to a C++ object being deleted when it called back to an event handler and wrote to its member variables after the handler returned. Ownership and lifetimes would have prevented this design error (which is surprisingly not that uncommon).
- humanrebar 4y ago> Mismatched type definitions in different translation units caused an implicit 32 to 16 bit conversion. Rust would not have failed to compile at step one necessarily. Or were you assuming a full revamping of the build and dependency management systems in addition to the rewrite of the program itself? If we need to clean up the build and dependency management system to start using Rust correctly, isn't that two major migrations?
- briantkelley 4y ago> Rust would not have failed to compile at step one necessarily. True. In this case there was old code and new code. I assumed in a hypothetical hybrid Rust codebase, bindgen would be used to bring the types in old code to Rust, so the compiler would identify the type mismatch in perhaps the impl From. > If we need to clean up the build and dependency management system to start using Rust correctly, isn't that two major migrations? At the scale of Office, managing the build system is an evergreen project staffed by dozens of engineers. Adding support for Rust integration is a typical deliverable for such an org.