7 ms·
Saving 100 terabytes of memory by optimizing 1.1.1.1's DNS cache
- eviks 1mo ago> Once we store a DNS response in the cache, however, we never modify it again. The capacity field serves no purpose, but still costs 8 bytes per Vec Were there no design discussions/reviews when the system was setup to catch trivial things like this?
- micromacrofoot 1mo agoit was working so no one thought to check
- mhitza 1mo agoPremature optimization argument fits right in. Now that memory is up to 10x more expensive it is worth considering optimizing programs with large memory footprint.
- eviks 1mo agoHow does that fit? What would be the evil of not wasting memory for many years at 1x?
- jgrahamc 1mo agoOne of the "evils" of premature optimization is how much time you spend on the optimization vs. the benefit you get from it. If your goal is correctness and shipping fast and you're not memory constrained then spending time using the least amount of memory is a waste of time specifically because you want to ship fast. Another interesting thing that happens is you don't necessarily know what form your actual optimizations will need to take. Later when your systems grow you discover the suboptimal parts you hadn't optimized for. Very early on at Cloudflare I worked on part of the DNS infrastructure that took DNS records from the UI and got them in a state for actual authoritative serving. The system had been constructed anticipating Cloudflare having millions of customers with unique domains, but it had not been constructed for a single customer with a single domain with millions of records. This caused a periodic slow down in DNS record updating while the system churned on that one customer. In a different job I worked on a piece of optimization software that needed to keep track of "node" A is reachable from node "B". This had been implemented as a matrix (literally a malloced NxN matrix of ints storing 0 or 1) which worked really well for small systems. But you'd be out of memory really fast on a large project. I replaced the matrix with a hash table and all was good because the matrix was actually really sparse.
- stickfigure 1mo agoAbsolutely true, but I will say that LLMs have changed the equation somewhat. With a rather short prompt, claude/codex will take your code, write a harness, profile it, build experiments, profile those, and give some pretty solid advice which one to pick. Then integrate the changes. It's the kind of goal-directed, bite-sized job that LLMs excel at. Extremely low-commitment. Except for the whole "making changes in production at scale" problem, of course.
- gbear605 1mo agoEngineers are expensive, especially good system engineers who are trained in your code base. Very possible that this just hadn't gotten to the top of the priority list.
- eviks 1mo agoI don't understand why you need training on your code base to design a cache format for read only vs rw workloads, but anyway yours is a comment about neglect, not the "evil" that would happen if you did that design
- win311fwg 1mo ago> I don't understand why you need training on your code base to design a cache format Because anyone willing to come in just to design your cache format is going to expect payment that is many multiples more than the engineers you already cannot afford? Long-term employees cost less, which brings them closer to being affordable, but you have to be able to keep them busy for long periods of time to realize that reduction in cost. A engineer who doesn't understand your codebase isn't going to be useful for very long.
- eviks 1mo agoYou explained why it's beneficial for other workloads, but the original point was about this specific design
- Spooky23 1mo agoI see your point but disagree. Engineering is about constraints. Time, materials, labor, scope. The “evil” of premature optimization is that it’s a misapplication of priority. If I have an acute medical problem that needs attention, it’s not the right time to talk about chloresterol and statins, get my broken leg set. There’s always a tension between engineering management who needs to deliver a solution to the business and engineers who want to deliver a beautiful object.
- toast0 1mo agoUsing obviously better data structures the first time isn't premature optimization.
- lbriner 1mo agoIt is often not worth optimising in the early days. You don't know how popular it will become, you might not know how many DNS records you will hold, it was possibly written in an earlier language and ported as-is. At the point someone queries the 100TB of RAM, then maybe it is worth revisiting but even that has risks. You have to design the migration path, have fallback mechanisms etc.
- eviks 1mo agoIt's also often that you can avoid all those future migration/fallback risks and pains if you invest a little bit of design thinking upfront. So how would you decide which path to take in situations like this?
- suriyaG 1mo agoIt only looks super obvious in hindsight and the well explained blog post. when a team of 5 is tasked with getting a completely new DNS up at the scale and integrate well with cloudflare. if you spend cycles on nitty gritty opinions like this time to market goes out further and further out. some napkin math, 130 gen13 servers cost "only" ~$2.6M. relative to the importance of the 1.1.1.1 and the market at the time. that is nothing to cloudflare. this is not to say good system design does not matter. it very much does, but making that call at that time would've butchered the prodcut very much similar to google+, youtube etc.
- eviks 1mo agoThis one also looks pretty obvious "in foresight" (using the same tools that existed back then. Maybe owner dedupe might be less obvious and require a bit of knowledge and probing into actual data, but for rw vs ro you are fine knowing nothing?) and you forgot the napkin math re. how much your precious "time to market" would have been delayed by. It's also not nothing, otherwise it would never be optimized away now, but left as is. After all, wasting time on optimization delays "time to market" for other useful features. I also don't get the reference to YouTube, it's a very successful product, how was it butchered by good system design???
- scott_meyer 1mo agoDiscussing trivial optimizations is a waste of valuable design time. You're never going to "forget" an optimization. The running system will remind you when the optimization is actually needed.
- r3trohack3r 1mo agoRob Pikes 5 Rules of Programming: Rule 1. You can't tell where a program is going to spend its time. Bottlenecks occur in surprising places, so don't try to second guess and put in a speed hack until you've proven that's where the bottleneck is. Rule 2. Measure. Don't tune for speed until you've measured, and even then don't unless one part of the code overwhelms the rest. Rule 3. Fancy algorithms are slow when n is small, and n is usually small. Fancy algorithms have big constants. Until you know that n is frequently going to be big, don't get fancy. (Even if n does get big, use Rule 2 first.) Rule 4. Fancy algorithms are buggier than simple ones, and they're much harder to implement. Use simple algorithms as well as simple data structures. Rule 5. Data dominates. If you've chosen the right data structures and organized things well, the algorithms will almost always be self-evident. Data structures, not algorithms, are central to programming. https://web.archive.org/web/20260314210910/https://users.ece.utexas.edu/~adnan/pike.html https://web.archive.org/web/20260314210910/https://users.ece...
- eviks 1mo ago> Data structures, not algorithms, are central to programming So you agree that they should've designed the system to use the appropriate data structure from the beginning?
- ecnahc515 1mo agoNotice rules are ordered. You don't optimize until you know you need it. They started with a data structure they though would be fine. Clearly it was fine since it worked and they decided it was later worth optimizing.
- deleted 1mo ago[deleted]
- win311fwg 1mo agoThe existence of 1.1.1.1 speaks to a much larger design problem. If you want to talk about what should have been done, you need to step much, much further back.
- ratmice 1mo agoBoxed slice isn't really the most well known type/optimization, There usually aren't that many vec's that it makes a big difference.
- irdc 1mo agoThis is why system programming still matters. Looks like they're missing the obvious optimisation of putting the record data right after the CacheEntry members instead of allocating memory separately though. But that might just be me as a C-programmer talking and not be all that easy in Rust.
- deleted 1mo ago[deleted]
- mkeeter 1mo agoFor the curious, this is technically possible in Rust using a dynamically sized type [1], but in practice is difficult and doesn't really play nice with the rest of the language. The nomicon entry concludes with "Yes, custom DSTs are a largely half-baked feature for now." [2] [1] https://doc.rust-lang.org/reference/dynamically-sized-types.html#r-dynamic-sized.struct-field https://doc.rust-lang.org/reference/dynamically-sized-types.... [2] https://doc.rust-lang.org/nomicon/exotic-sizes.html https://doc.rust-lang.org/nomicon/exotic-sizes.html
- cobalt 1mo agoless ergonomic, but still totally doable
- sdcfgy 1mo agoSystem programming always matters. Things are cheap until they aren't one day.
- strenholme 1mo agoWith my own MaraDNS, I aggressively optimized the memory usage of blacklist entries by having a single really big malloc() to allocate the memory for the entries, then traversing that memory block for potentially blacklisted entries. When I was using one malloc() per entry, a large blacklist took up 237 megabytes of memory. The same blacklist, once optimized to be loaded with a single malloc() call, only took up 9.5 megabytes of memory. https://samboy.github.io/blog/entries/MaraDNS.html#BlogEntry-2022-12-28 https://samboy.github.io/blog/entries/MaraDNS.html#BlogEntry...
- badatnames 1mo agoWhy do I always find interesting new Twitter accounts just as the person is leaving :)
- strenholme 1mo agoTwitter has become a cesspool, and there’s a lot of reasons why people are leaving it in droves. My personal issue is the misogynists who have created a hateful completely false narrative that 80% of the women sleep with 20% of the men (including the very demeaning and hurtful notion that all women are sexually promiscuous, but only if you’re one of the 20% of supposedly “Alpha” men) [1] Twitter is also full of—let’s call a spade a spade—racists who constantly post some video from years before showing some random Black person doing a criminal act, and then a bunch of racists comment that that’s how all Black people are and it’s the “evil left wing media” suppressing this supposed “truth”. Just as Twitter has become a right-wing cesspool, Reddit has become a leftist cesspool, so I also avoid Reddit, which, like Twitter, is also becoming a closed walled garden—they just this month started clamping down on people reading old.reddit.com anonymously, so now you have to log in to have a usable interface with Reddit. Excuse me, no. [1] This annoyed me to the point I researched the claims to verify it’s a bunch of bullshit. https://samboy.github.io/blog/80-20-myth.html https://samboy.github.io/blog/80-20-myth.html
- ignoramous 1mo agoMight be of interest: - TigerBeetle: A database without dynamic memory allocation, https://news.ycombinator.com/item?id=33192288 https://news.ycombinator.com/item?id=33192288 (2022). - Succinct Data Structures: Cramming 80,000 words into a Javascript file, https://news.ycombinator.com/item?id=2348619 https://news.ycombinator.com/item?id=2348619 (2011).
- OptionOfT 1mo ago> we store the records as a single Box<[u8]> containing each record encoded as a 2-byte length prefix followed by its raw bytes. Interestingly this is exactly how netlink works-ish: https://manpages.ubuntu.com/manpages/focal/man3/netlink.3.html https://manpages.ubuntu.com/manpages/focal/man3/netlink.3.ht... You start, get the type & length, and then that is how many bytes you read. Some issues with that when you deserialize, from a raw stream in to `[u8; 4096]` buffer, the alignment is only guaranteed to be on 1 byte, not 4 bytes. In practice it is 4 bytes, but if you run those tests with Miri, you'll get yelled at. So the fix there is to declare the buffer with a type that mandates the alignment of the largest type that you're going to be deserializing. So then you start your buffer as follows: `[u32; 1024]`, and with `slice::from_raw_parts` you get to turn that into `[u8; 4096]` with the expected alignment. As an exercise I wrote a streaming parser for netlink, the current existing package serializes everything, all at once.
- pocksuppet 1mo agoIt's called TLV encoding - tag/length/value. It's very common in all sorts of network protocols and serialisation formats. It allows you to skip unidentified tags. Sometimes, like in the PNG file format, there's a fixed bit in the tag that tells you whether it's safe to skip or if you have to reject the whole thing because you don't understand this tag. Hey dang can I get my rate limit turned off pretty please?
- jandrewrogers 1mo agoThis kind of encoding[0] is ubiquitous in networking protocols. It scales down to small silicon well and enables the receiver to estimate resource requirements or skip parts of a serial byte stream without storing it in memory first. These encodings usually aren't aligned by design. [0] https://en.wikipedia.org/wiki/Type–length–value https://en.wikipedia.org/wiki/Type–length–value
- mannyv 1mo agoOne question the article doesn't answer is: why are they cacheing at all? If your cache is that big it isn't a cache. How much bigger is the dataset in question? There are 250 billion entries. Assuming 80/20, that implies 1.25 trillion records? What's the speed of service/response time relative to the data source? At that point it might be enough to replace your multiple caches with fewer in-RAM databases? It's an interesting problem.
- eggnet 1mo agoThey’re adding the cache consumed across all of their servers. It’s not one giant deep cache.
- bastawhiz 1mo agoMaybe I'm misunderstanding, but this powers 1.1.1.1, it doesn't front an internal dataset. A cache miss hits a nameserver. Which is to say, the dataset is "every DNS record in the world"
- auspiv 1mo agoI think the question is probably more along the lines of - why not do a database with 100 TB of storage/records instead of a cache? tomato / tomato.. especially with smart caching in front of database. 100TB of flash is a good bit cheaper than 100TB of memory
- robotresearcher 1mo agoThis is smart, task-specific caching in front of database.
- ecnahc515 1mo agoBecause it would be slower and have different scaling requirements than the ones they want.
- fc417fc802 1mo agoI'm no expert but presumably all of throughout, latency, and churn. DNS is approximately a giant KV store where the typical record has a TTL of ~5 minutes.
- dshat 1mo agoI'll buys some spare RAM you now have. I only need 64GB.
- vinkelhake 1mo agoThese seem like some fairly standard approaches for reducing memory usage. I can't help to think that the approach of joining several distinct list into a single one in some way undercuts Rust's safety guarantees. If you previous had three distinct Vec objects, then Rust would guarantee that you can't index out of bounds. If you now put all those objects into a single Vec and rely on offsets, then you now open the door to indexing out of range of these sub-slices without any panics. It's a minor point, and it doesn't really invalidate the optimization, but I'm surprised the article didn't mention it.
- vsgherzi 1mo agoyou could always do a .get into the vector and handle the error, it doesn't necessarily need to panic. Thank being said in this case it should be impossible to index out of bounds so maybe a panic is warented.
- FpUser 1mo agoTools exist to serve us, not the other way around.
- asgraham 1mo agoSure, and usually one of the ways Rust serves us is with safety guarantees. Which isn’t to say this optimization is a bad idea, just to say it’s sort of a straw man to imply coding in Rust to take advantage of safety guarantees is “serving Rust”
- ratorx 1mo agoI think it’s more of a time vs code tradeoff, if done properly. For example in the Vec case, you could theoretically build an alternative which encodes the “three sections” property internally, and ensures correctness at construction time for the pointers. Not as completely safe as a Vec, but you can still get similar benefits for the “business logic”. But I agree, just having a custom structure that does not provide a safe wrapper around this would be sacrificing standard guarantees.
- afdbcreid 1mo ago
- 9bot 1mo agoThe most interesting result to me is that the richer parsed representation was not necessarily the faster one. If the hot path is mostly “read from cache and serialize back to DNS,” parsing everything upfront only to serialize it again can become unnecessary work and hurt locality....
- bhouston 1mo agoI've run into issues with using public wifi when I override my MacBook's DNS server to 1.1.1.1 or 8.8.8.8. I believe this is because captive portals require custom resolution of the name captive.apple.com. And external DNS servers will not resolve that correctly to the local gateway's authorization page.
- brians 1mo agoThat’s a Mac bug if so—it should be always using dumb udp/53 for captive detection, not some fancy DoH thing.
- MayeulC 1mo agoAFAIK (at least it worked like that some 10 years ago) the captive portal just intercepts the HTTP page load and inserts its own content (most often a 302). So it just has to be a http web page. Firefox uses http://detectportal.firefox.com/canonical.html http://detectportal.firefox.com/canonical.html Relevant support page, though light in details: https://support.mozilla.org/en-US/kb/captive-portal https://support.mozilla.org/en-US/kb/captive-portal Edit: ah, yes, DNS can be hijacked too (requires intercepting outgoing traffic on port 53 therefore incompatible with DoH), that may require fewer computing resources. Still need http otherwise the server cannot use the correct cert chain. Edit 2: Wikipedia says both methods are used: https://en.wikipedia.org/wiki/Captive_portal https://en.wikipedia.org/wiki/Captive_portal and also mentions RFC 8910. I suspected something like that existed, hence my initial disclaimer. My point was: that domain is not treated any differently from other domains.
- fc417fc802 1mo agoCan we take a minute to appreciate how utterly broken this state of affairs is? The dogged over centralization of DNS is an endless source of problems.
- comprev 1mo agoI've had reliable success by using http://neverssl.com http://neverssl.com to force a basic HTTP connection for kickstarting a public WiFi portal login, although I have to disable NextDNS (iOS) too.
- 0xAstro 1mo agoIt's weird that it took so long for these trivial optimizations but it might just be that they were working on optimizing other stuff.
- sergq 1mo agothis applies to more than DNS caches. In 1998 I mailed Microsoft a proposal to replace search engine crawlers with a push-based filesystem monitor (detect change → extract → compress → push to index). Got a 5-line rejection letter. They built the same thing 20 years later as IndexNow. Full story with the original letter: https://dev.to/andrew_vl/in-1998-i-proposed-push-based-search-indexing-to-microsoft-they-rejected-it-1lpi https://dev.to/andrew_vl/in-1998-i-proposed-push-based-searc...
- MagicMoonlight 1mo ago[dead]
- rfgplk 1mo agoFrankly weird that they were resorting to high level containers for this in the first place. Also, this line struck me as odd > Big Pineapple uses jemalloc, an allocator designed for multithreaded, allocation-heavy workloads. jemalloc multithreaded performance is actually poor(ish) compared to other modern allocators, which makes it a weird choice. But even weirder is why they're even using an allocator in the first place compared to a va MAP_ANON | MAP_NORESERVE arena carveout approach? You can also do punning that way too, which I'm not even certain if Rust supports?
- senderista 1mo agoI would also have instinctively reached for a large VM reservation to exploit demand paging. I have used that pattern a lot in C++ but not in Rust, so I don't know how difficult it would be to implement there.
- cobalt 1mo agoRust supports punning via pointer casting, but you'll want to use #[repr(C)] on any data types used
- jandrese 1mo agoAn approach like that would be at constant war with the borrow checker in Rust. Apparently it is possible but there is enough friction that these guys went a different route.
- kevincox 1mo agoIt sounds like you are just writing your own allocator? That sounds great, but why is it going to be better than jemalloc? There are many reasons a specialized allocator can be better, but just saying "write your own" doesn't really add much value to the discussion.
- foltik 1mo agoIf you squint, even their TLV encoding in a Box<[u8]> is kind of a specialized arena allocator.
- edflsafoiewq 1mo agoGeneral theme: A programming language's native in-memory object format is typically optimized for random access, uniformity, and mutability (fields at fixed offsets, etc). Serialization formats for network or disk tend to be designed explicitly to be more compact. But you can design your own in-memory representation too, with the properties you need.
- kccqzy 1mo agoThat’s the old school of thought. These days, designers of newer serialization formats realize that designing a more compact format doesn’t really buy much on modern CPUs and modern networks. See for example Cap’n Proto (whose inventor, kentonv, also works at Cloudflare) and flatbuffers.
- inigyou 1mo agoThat's also the ancient school of thought, before compaction was viable and before portability was needed.
- lpapez 1mo agoThis is the right way to deliver software. Produce working product first, validate the idea, stabilize the business, start generating profit, and then you can start optimizing your costs. In fact optimization is by far the easiest part of the process because there are many system programming experts on this HN thread who consider these optimizations to be trivial.
- vanviegen 1mo agoOr optimize a bit earlier and prevent having to scale out to a bazillion systems.
- deleted 1mo ago[deleted]
- bcrosby95 1mo agoThe way I usually prevent having to scale out to a bazillion systems is never getting more than 10 users.
- bch 1mo agoThe Art of Production
- froh 1mo agopro move. made my evening.
- steve_adams_86 1mo agoI wonder why Cloudflare didn’t think of this
- kevin_thibedeau 1mo agoThis is Broadcom's business model
- 1mo ago
- cpriest 1mo ago[dead]
- cristaloleg 1mo agoObvious question: why wasn’t this done earlier? It looks like all the data was already available. At THAT scale, reducing memory usage is a must-have, not a nice-to-have. Weird.
- yieldcrv 1mo agoprobably agents going through tech debt or finding wins every dept knows what they could do with more budget, the budget for those things just never comes now agents have utilized budget more effeftively, unbottlenecking many things, including engineering blogs
- sophacles 1mo agoCloudflare talks about having datacenters in 300+ cities. Presumably they have at least a few servers per datacenter. They saved 130 servers worth of memory... not even the minimum number of servers they have (seriously though, they probably have a LOT of servers)... a few GBs of memory per server running the service. At that scale this is a nice-to-have.
- pocksuppet 1mo ago> 56% A records, 25% AAAA, and 19% TXT And they say nobody uses IPV6.
- varispeed 1mo agoNow put the 100 terabytes of memory back to the market. Stop hoarding RAM.
- dumbotronz 1mo ago[dead]
- mu54 1mo agoThe Art of Production.
- 1saadcodes 1mo agoWe're finally seeing more appreciation for this kind of engineering. Not everything needs to be solved by throwing more hardware at the problem
- xyst 1mo agoWith memory costs soaring due to AI bubble, this type of engineering has become "profitable".
- ManBeardPc 1mo agoThe Record struct contains rtype and data where RecordData is a tagged union. Aren’t those two always in sync? Not a DNS expert, just wondering if this is redundant or there is a reason both are there. Doesn’t matter anymore if they store it already serialized but I would be interested why it was this way.
- kevinbaiv 1mo ago[dead]
- superze 1mo agoHow much is this in euro or do we measure money in ram now?
- inigyou 1mo agoCurrently $15 per GB, he saved Cloudflare $1,500,000 and got exactly $0 bonus. He must really believe in cloudflare's vision (global enshittification). In related news, three times today Cloudflare told me that I'm a bot and shall not pass - not that it needs to check if I'm a bot before it lets me pass.
- winrid 1mo agoI saved an employer $4m/yr, entirely myself. I think I got a 10k bonus.
- BikiniPrince 1mo agoFunny thing about cloudflare. I have a dns warming script that uses their top 1k or 10k addresses. Then when my master starts up it warms the entire cache. Everything else uses memcache so the cluster is nice and toasty. As far as I can tell no one else releases domain statistics like them.
- aaron695 1mo ago[dead]
- squirrellous 1mo agoI wonder at their scale, why wouldn’t it make sense to store the entries lightly compressed in memory?
- winrid 1mo agoYeah with snappy they probably could get like 30% savings (I bet domains don't compress well) without much more cpu usage. Alternatively a btree somehow they can take advantage of prefix compression
- squirrellous 1mo agoAh yes, something like prefix tree would work well storing reverse domain names like com.abc.www (although this is such a well known thing I feel like I must be missing things).
- akoboldfrying 1mo agoDefinitely. Even a custom compression/encoding that knows to treat the different fields differently -- just a 4-byte binary IP address for A records, normal LZ-based text compression of the domain name (perhaps using a custom starting dictionary and/or Huffman table), etc.
- ww520 1mo agoNot sure what they use to hold the cache key and entry. If a hashmap is used, then a radix tree (adaptive radix tree) would be better in saving memory space. Most of content of the qname field of the CacheKey is hostname, like www.site.com. The reverse version com.site.www fits nicely in navigation path of a radix tree. The common prefixes like "com." are shared and compressed in the parent nodes of the tree. Even a BTree with compressed prefix keys can save space in the qname.
- Tuna-Fish 1mo agoAt this scale, moving from one pointer chase to multiple is almost certainly a huge loss, even if radix tree would save a lot of memory.
- ww520 1mo agoUnless the keys are completely random, compressed keys shorten the tree height and cause fewer pointer jumps. Hostnames are highly compressible. Plus the root and the upper levels of the tree are always hot, most likely in L1/L2/L3 all the times. OTOH collisions in hash table cause pointer chase as well.
- vismit2000 1mo agoExactly: https://github.com/pytries/marisa-trie https://github.com/pytries/marisa-trie
- didgetmaster 1mo agoWhy do people seem to think that optimization is something you only have to deal with once the software scales so much that 100s of TB of memory or disk space (or thousands of hours of processing time) are being wasted. It is almost like nobody even thought during the design phase about what might happen down the road. This is why so much software is bloated and often buggy. Just gets something that half-way works out the door ASAP and worry about the rest later (too often, never).
- jacquesm 1mo agoIt can be quite hard to predict where particular usage patterns will take a piece of software under extreme load, especially with things that have lots of internal state. Obviously when you get to spend 100 T or more the pay off of an optimization is much larger than what it is in the case of 1T or less, and your typical developer is not going to have that kind of memory even in aggregate to play with. I tend to be forgiving when it comes to watching software bloat that I did not cause myself (and yet, I'm frustrated that Ubuntu's start-up greeting message takes a whopping 500 M). In the case of internet infrastructure I don't think there was anybody even up to the year 2000 who had any idea of how bit this was going to be. And even now we have IPV4 and lots of legacy to deal with. Cloudflare is not my favorite company, let's put it like that, but in this case they show how the sausage is made and I think that should be applauded. Much better than 'why were down again for X hours'.
- philippta 1mo agoI recently heard a great analogy for this exact problem under the premise of "Make it work, make it right, make it fast". Assume a sorting algorithm as a metaphor for your whole program. You make it work by implementing the simplest thing you know how to write: bubble sort. Works. At the end you notice it's way too slow and replace with something much better: quicksort. Now, how much of your program is surviving? Almost nothing, perhaps except for the "greater than" comparison. If you apply that idea to a real world program, we're pretty much talking about a full rewrite.
- jacquesm 1mo ago
- Dylan16807 1mo agoSo they optimized from Vec to Box, but they're still using Box all over and spending 16 bytes on it? The things they're boxing need 2 bytes for length, and their memory use is low enough that they could cram the pointers into 4 bytes. Trying to pack that into 6 bytes is probably too much fuss for the benefit, but I see no reason to use more than 8 bytes.
- fulafel 1mo agoWhere are their users coming from? Besides the few manually putting 1.1.1.1 in their settings.
- kibwen 1mo agoFirefox's default DNS-over-HTTPS provider is Cloudflare.
- grep_it 1mo agoThis reminds me how you can save a bunch of bytes just by making sure your structs are aligned. In go for example: type Wasteful struct { a int16 b int c byte } type Aligned struct { b int a int16 c byte } Will have sizes of 24bytes and 16bytes (on a 64bit system). Same data 8bytes more. If you are storing millions of those objects, then it adds up.
- masklinn 1mo agoRust does that automatically unless you switch to the C layout. In langages that don’t there’s a tension between memory use and human readability / consistency of the layout. There are also other domains which can be affected e.g. databases, it’s a concern / issue when using postgres for instance as it uses aligned columns and stores them in schema order.
- jordiburgos 1mo agoWhy this is not done automatically by the compiler? That seems something quite easy to calculate to me.
- JeromeLon 1mo agoThere is no way in C to express that you don't care about the orde. When you express a struct in C, you list what you want in the struct and (sometimes without wanting it) exactly in what order you want it. Interestingly, there is also no way to write a loop on i for all the values between 0 and 99 without specifying the order. Luckily, in this case, the compiler is allowed to prove that the order has no impact (because it's local), and to decide that it will scan the values in a different order for optimisation purposes. So the compiler could do it on a structure as well, as soon as it's able to prove that the structure is not exposed in any way to any code that it doesn't control, but that's much more difficult than proving that variable i is not visible outside of a tight loop.
- cestith 1mo agoIt could be a new keyword rather than counting on the compiler to prove certain access patterns don't exist. That's a bit of a messy tradeoff. Maybe something like 'unordered struct' or 'packed struct' works, but it would be a nonstandard extension for some time.
- zamalek 1mo agoThe intermediate level Rust dogma is to try your hardest to avoid the heap, and to tear your hair out at the throne of monomorphization. While both are broadly true, it's articles like this that show that a single pointer (or call) indirection can sometimes be better.
- kibwen 1mo agoI'd say that boxing large enum variants is itself an intermediate level Rust topic, and a well-accepted practice. Clippy will even point out places where you might benefit from boxing an enum variant: https://rust-lang.github.io/rust-clippy/master/index.html?search=variant+large#large_enum_variant https://rust-lang.github.io/rust-clippy/master/index.html?se...
- Agentlien 1mo agoOne of my proudest professional moments was when me and three others managed to reduce memory load of the game Wavetale from 20+GiB to under 3GiB so we could port it to Nintendo Switch. The 100 TiB number almost gives me vertigo. Though in this context it was "just" 50%
- adzm 1mo agoI'd love to hear what was taking up that 17GiB if you can share, even if it is commonly optimized things like packing, still fun to hear about
- Agentlien 1mo agoIt's an indie open world game originally released for Stadia and in its original form simply loaded the whole world into memory at boot. We had to implement a streaming system and figure out a good way of chunking the world. This was a challenge because everything was on water and you could see nearby islands quite far away. An intern called Tommi did a great job identifying a good strategy and writing the system. We also had to reduce the density and model complexity of a lot of environmental details such as rocks, vegetation, and stuff like pots and clotheslines. This was done largely by sorting things by memory size and frequency of use and identifying outliers. One of the biggest issues was actually really silly: the journal fetched Portrait images and names of characters by referencing the actual NPC and having them embedded there. Meaning the journal, which was always loaded, would pull every NPC involved in a quest into memory including their behaviors, textures, and models. On my end I also found a lot of silly details wasting hundreds of megabytes. Special render passes using huge textures and render targets, poor structuring of the render pipeline caused memory increases, several key shaders referenced huge textures which weren't necessary, ... My blog posts about the project[0] mainly focus on rendering performance because that's where I spent more time and it contains more interesting content for discussion. But reducing memory was an ongoing concern with countless little improvements over the 1.5 year porting process [0] https://agentlien.github.io https://agentlien.github.io
- NahImG 1mo agosees cloudflare leaves
- makach 1mo agothis is so interesting, memory and storage is cheap until it isn't and then you optimize.
- gxdbox 1mo ago[flagged]
- empiricus 1mo agomy first thought was if is this will impact the memory market prices :)
- marsx-dev 1mo ago[dead]
- iwontberude 1mo ago[dead]
- ricudis 1mo agoREWRITE IT IN C!
- jedisct1 1mo agoEdgeDNS was rewritten and is now EtchDNS https://etchdns.dnscrypt.info https://etchdns.dnscrypt.info
- ozereray1 1mo ago[flagged]
- CyberDildonics 1mo agoGreat article, but I'm surprised they waited until they were using $2 million USD of memory before shaving off all the unused bytes at the end of a vector.
- parasxos 1mo ago[dead]
- myshapeprotocol 1mo ago[dead]
- fudgy73 1mo agoAre they selling that RAM?
- soricus 1mo ago[flagged]