25 ms·
What every programmer should know about SSDs
- effnorwood 5y agothey do not spin
- dang 5y agoWhat someone else said about that in 2014: What every programmer should know about solid-state drives - https://news.ycombinator.com/item?id=9049630 https://news.ycombinator.com/item?id=9049630 - Feb 2015 (31 comments)
- cottsak 5y agohaha! very similar sections too .. almost looked copied for a brief moment as i skimmed there
- personjerry 5y agoHow big is the write cache usually and how does it work? Typically I've seen the write caches be something like 32MB in size, but the "top speed" seems to be sustained for files much bigger than 32MB, which doesn't make sense to me if that top speed is supposedly from writing to the cache. How does that work?
- bserge 5y agoOn SSDs? 32 is way off, the Samsung 470 had 256MB RAM cache and the 860 Pro a whopping 4GB for the top model. Although they started removing it entirely for NVMe SSDs, I guess the direct transfer speed is enough to not need a cache at all.
- mastax 5y agoNVMe drives can access system memory over the PCIe bus.
- wtallis 5y agoThe DRAM you're referring to is for the most part not a write cache for user data. Most of that DRAM is a read cache for the FTL's logical to physical address mapping table. When the FTL is working with the typical granularity of 4kB, you get a requirement of approximately 1GB of DRAM per 1TB of NAND. Drives that include less than this amount of DRAM show reduced performance, usually in the form of lower random read performance because the physical address of the requested data cannot be quickly found by consulting a table in DRAM and must be located by first performing at least one slow NAND read.
- opencl 5y agoIt varies quite a bit. There are two different types of caches: SLC and DRAM. Most drives use SLC caching, higher end drives often use both. Typically the SSDs with DRAM have a ratio of 1GB DRAM per TB of flash. SLC caching is using a portion of the flash in SLC mode, where it stores 1 bit per cell rather than the typical 2-4 (2 for MLC, 3 for TLC, 4 for QLC) in exchange for higher performance. SLC cache size varies wildly. Some SSDs allocate a fixed size cache, some allocate it dynamically based on how much free space is available. It can potentially be 10s of GBs on larger SSDs.
- igg 5y agoThe 1 GB DRAM per 1 TB Flash is to store the Flash Translation Layer mapping from logical addresses of the host system to the physical address in Flash. The write cache is separate and much more limited in size.
- wtallis 5y agoGetting full throughput from the SSD is less about file size and more about how much work is in the SSD's queue at any given moment. If the host system only issues commands one at a time (as would often result from using synchronous IO APIs), then the SSD will experience some idle time between finishing one command and receiving the next from the host system. If the host ensures there are 2+ commands in the SSD's queue, it won't have that idle time. Then there's the matter of how much data is in the queue, rather than how many commands are queued. Imagine a 4 TB SSD using 512Gbit TLC dies, and an 8-channel controller. That's 64 dies with 2 or 4 planes per die. A single page is 16kB for current NAND, so we need 2 or 4 MB of data to write if we want to light up the whole drive at once, and that much again waiting in the queue to ensure the drive can begin the next write as soon as the first batch completes. But you can often hit a bottleneck elsewhere (either the PCIe link, or the channels between the controller and NAND) before you have every plane of every die 100% busy. If you're working with small files, then your filesystem will be producing several small IOs for each chunk of file contents you read or write from the application layer, and many of those small metadata/fs IOs will be in the critical path, blocking your data IOs. So even though you can absolutely hit speeds in excess of 3 GB/s by issuing 2MB write commands one at a time to a suitably high-end SSD, you may have more difficulty hitting 3 GB/s by writing 2MB files one at a time.
- dataflow 5y agoWhat's the flash translation layer made of? Is the flash technology used for that more durable than the rest of the SSD itself? (like say MLC vs. QLC?)
- SeanCline 5y agoYou're right that the FTL has some durability concerns which, in addition to performance, is why it's typically cached in DRAM. Older DRAM-less SSDs were unreliable in the long-term but that's been improving with the adoption of HMB, which lets the SSD controller carve out some system RAM to store FTL data.
- pkaye 5y agoThe FTL is like a virtual memory manager. It is firmware/hardware to manage things like the logical to physical mapping table, garbage collection, error correction, bad block management. Yes there will be a lot of FTL data structures stored on the flash. It can be made durable by redundant copies, writing in SLC mode or having recovery algorithms. I used to develop SSD firmware in the past if you have further questions.
- jng 5y agoHey that's very interesting! How much of the FTL logic is done with regular MCU code vs custom hardware? Is there any open source SSD firmware out there that one could look at to start experimenting in this field, or at least something pointing in that direction, be it open or affordable software, firmware, FPGA gateway or even IC IP? I believe there is value in integrating that part of the stack with the higher level software, but it seems quite difficult to experiment unless one is in the right circles / close to the right companies. Thanks!
- pkaye 5y agoTypically the Host and NAND interface have custom hardware. When the host issues a command, the hardware might validate it and queue up data to a buffer. On the NAND interface there might be a similar queue for NAND commands. You might have multiple queues for different priorities of operation. The error correct will also be in hardware. When you issues NAND reads and writes, the ECC will be checked or encoded. The rest of the FTL might all be in firmware. Perhaps a single core does everything. Or maybe its partitioned between two cores, one for the host related code and the other for the FTL related. Some companies have tried lots of cores, each with a dedicated state machine to handle some part of the operation. These can be complex to coordinate their operation and to debug. Some companies convert some of these state machines into custom hardware. The only open SSD platform I've read about is http://openssd.io/ http://openssd.io/ but I've never played with it. One of the challenges is the NAND manufacturers a lot of the critical documentation under an NDA these days. You really need that information to make a reliable SSD. When you learn how the internals of an SSD work, its a wonder that it retains data at all! In terms of integrating SSD with the higher software level, I believe FusionIO was doing this in the past. They put the whole logical to physical mapping into the host memory.
- CoolGuySteve 5y agoThe claim about parallelism isn't true. Most benchmarks and my own experience show that sequential reads are still significantly faster than random reads on most NVME drives. However, random read performance is only somewhere between a 3rd to half as fast as sequential compared to a magnetic disk where it's often 1/10th as fast.
- pkaye 5y agoWhat kind of queue depth do you test the read performance? The sequential can be made fast at low queue depth by the SSD controller doing prefetch reads internally. I've worked on such algorithms myself.
- CoolGuySteve 5y agoShow me a benchmark at any queue depth where random reads are as fast as the fastest sequential rate for that drive. It's simply not true. I suspect it has something to do with prediction on the controller but I'm also not confidently spewing a bunch of bullshit about drive architecture unlike this article.
- 1_player 5y agoA lot of talk about pages, but no mention about how big these pages are. From a quick look on Google, most SSDs have 4kB pages, with some reaching 8kB or even 16kB.
- wtallis 5y agoSSDs mostly tell the host system that they have 512-byte sectors or sometimes 4kB sectors, and the typical flash translation layer works in 4kB sectors because that's a good fit for the kind of workloads coming from a host system that usually prefers to do things (eg. virtual memory) in 4kB chunks. But the underlying NAND flash page size has been 16kB for years.
- cbsmith 5y ago...and all that cruft, and the logic to try to make handling of it not so bad, makes for a lot of complexity and unintended consequences.
- wtallis 5y agoEmulating 4kB or 512B sectors when the underlying media has a 16kB native page size really doesn't add much more complexity on top of the stuff that was already required to handle the fact that erase blocks are multiple megabytes.
- cbsmith 5y agoThe complexity doesn't come from the emulation. It comes from trying to do the emulation efficiently based on assumptions about the behaviour of the other moving parts... which are also doing the same thing. So, you've got firmware that is pretending you've got 512B/4kB chunks when really you have 16kB, and anticipating how the other layers might be doing things in order to maximize performance. Then you have a filesystem/VFS layer, which tries to optimize its access patterns anticipating how the underlying solid state storage might be really doing things in 16kB sizes and how it might be optimizing 512KB & 4kB accesses to fit that. Both those layers are dealing with filesystem journaling and how that might impact performance. Then you might have a database, which is now trying to anticipate how the filesystem and the underlying firmware might be optimizing access patterns, and so it's trying to optimize to fit all that. You also potentially have application logic that is trying to anticipate how the database might do things... What you tend to end up with are many layers of redundant caching that are all working against each other in a very inefficient manner.
- andrewmcwatters 5y agoMy opinion is probably... not technically correct... until you have to deal with drive reliability and write guarantees, but I don't think programmers actually have to know anything about SSDs in the same way that developers had to know particular things about HDDs. This is out of pure speculation, but there had to be a period of time during the mass transition to SSDs that engineers said, OK, how do we get the hardware to be compatible with software that is, for the most part, expecting that hard disk drives are being used, and just behave like really fast HDDs. So, there's almost certainly some non-zero amount of code out there in the wild that is or was doing some very specific write optimized routine that one day was just performing 10 to 100 times faster, and maybe just because of the nature of software is still out there today doing that same routine. I don't know what that would look like, but my guess would be that it would have something to do with average sized write caches, and those caches look entirely different today or something. And today, there's probably some SSD specific code doing something out there now, too.
- rzzzt 5y agoYou can optimize for less/shorter drive seeks on rotational media by reordering requests: https://en.wikipedia.org/wiki/Elevator_algorithm https://en.wikipedia.org/wiki/Elevator_algorithm
- forrestthewoods 5y agoGames used to spend a lot time optimizing CD/DVD layout. Because reading from that is REALLY slow. Optimize mostly meant keep data contiguous. But sometimes it meant duplicate data to avoid seeks. The canonical case is minimize time to load a level. Keep that level’s assets contiguous. And maybe duplicate data that is shared across levels. It’s a trade off between disc space and load time. I’m not familiar with major tricks for improving after a disc is installed to drive. (PS4 games always streamed data from HDD, not disc.) Even consoles use different HDD manufacturers. So it’d be pretty difficult to safely optimize for that. I’m sure a few games do. But it’s rare enough I’ve never heard of it.
- andrewmcwatters 5y ago
- jedberg 5y agoThis page tells me a lot about SSDs, but it doesn't tell me why I need to know these things. It doesn't really give me any indication about how I should change my behavior if I know that I'll be running on SSD vs spinning disk. I've always been told, "just treat SSDs like slow, permanent memory".
- deleted 5y ago[deleted]
- danbst 5y agoyeah, article should talk about periodic TRIMming, though this is more an admin advice
- smt88 5y agoDon't modern OSes transparently TRIM periodically anyway?
- snazz 5y agoYes, although you have to set it up manually if you’re using a more bare-bones Linux distribution or something like that.
- PaulKeeble 5y agoI have found trim is not sufficient at least on Windows, we still need to rarely defragment SSDs from what I can tell. On a Windows server we were having SSD performance issues where sequential reads were often down to 100MB/s, it was kind of confusing but we tried all sorts of ways to copy it with the same result. I eventually tested the drive with a fragmentation tool and it was really high at 80% but most importantly the problem files had so many fragments that they were tending towards 4k IO reads. What I did was remove all the files to another drive, force trimmed the drive and gave it several hours to sort itself out and then copied them back and performance was restored to 550MB/s as would be expected. I wrote a quick go program to test sequential read speed of all files across all the drives and I found plenty of files where performance was degraded. This was across a range of SSDs I had, SATA and NVMe from differing vendors. I suspect this is a bigger problem than most people realise, normal use absolutely can get the drive into a bad performing state and trim wont fix it. Very few people expect that the drive will degrade down to its 4K IO speed on a sequential copy but it apparently can.
- kortilla 5y agoThe title should be “why SSDs mean programmers no longer have to think about hard drives”. These are all reasons SSDs are much more pleasant to work with than old platter disks.
- cbsmith 5y agoWell, they no longer need to think about hard disks, but there are a lot assumptions from the world of hard disks that play out very differently in the SSD world.
- formerly_proven 5y agoI don't think there's any optimization for hard drives that is going to hurt on SSDs, and unoptimized workloads are always going to work better on SSDs. I'm inclined to agree with GP that SSDs are quite close to random-access storage and so there is little to worry about.
- cbsmith 5y agoSure there are. If nothing else, hard disks have much more consistent latency characteristics for reads and writes. So, for example, you might trade some extra write IOs to ensure data is organized efficiently on disk, reducing the number of read IOs you will subsequently have. With an SSD it's largely a waste of time, because the random reads are so much cheaper and the "contiguous" blocks you think you are seeing are mapped all over the drive anyway. You want to organize things reasonably efficiently when you write, and then rewrite as little as possible, ideally never. LSM's tend to fit the SSD paradigm so much better than say... balanced trees for this reason. Similar story with clustered indexes in databases. If you use a clustered index on an SSD, usually it's for an index on something like time where new records are invariably going to go near the end of the index; anything else will have bad write performance on a hard disk, but it might be worth it for the read performance... with the SSD, it is just an unmitigated disaster. There was a time where people thought of hard drives as "just random access storage" and consequently "there is little to worry about" and "unoptimized workloads are always going to work better on SSDs". Yup, SSDs are way faster than what came before them, but that if anything tends to mean that data structures & algorithms that used to make sense might not make much sense any more.
- teddyh 5y agoWhat everyone should know is that flash drives can lose their data when left unpowered for as little as three months.
- mercora 5y agoif that is true disks should come with a very visible note stating this... seriously, 3 months would be nothing. i doubt it is true because 3 months is a time frame which should be surpassed quite often making this more known.
- teddyh 5y agoDepending on manufacturer, and storage conditions, it can be up to about ten years. But the “three months” number is real: https://web.archive.org/web/20210502042514/http://www.dell.com/downloads/global/products/pvaul/en/Solid-State-Drive-FAQ-us.pdf https://web.archive.org/web/20210502042514/http://www.dell.c...
- crazygringo 5y agoThat's a document from nine and a half years ago, and it states: > It depends on the how much the flash has been used (P/E cycle used), type of flash, and storage temperature. In MLC and SLC, this can be as low as 3 months and best case can be more than 10 years. The retention is highly dependent on temperature and workload. Are there any modern sources provide more accurate stats? "3 months to 10 years" is so vague as to be useless.
- adrian_b 5y agoConsumer SSDs (unlike enterprise SSDs) must have a retention time of at least 1 year at the end of their life. To achieve that target, when they are new they must have a retention time of a few years, but you should better not count on that.
- anticensor 5y agoYep, they are semivolatile limited write memory modules, not disks. Everyone should use that SV-LWMM acronym.
- FpUser 5y agoIt is really puzzling why "every programmer" should burden their already overloaded brains with this. If they're reading/writing some config/data files this knowledge would not help one bit. If they're using database then it falls to the database vendor's to optimize for this scenario. So I think that unless this "every programmer" is a database storage engine developer (not too many of them I guess) their only concern would be mostly - how close my SSD to that magical point where it has to be cloned and replaced before shit hits the fan.
- bob1029 5y agoThings I have learned about SSDs: If you want to go fast & save NAND lifetime, use append-only log structures. If you want to go even faster & save even more NAND lifetime, batch your writes in software (i.e. some ring buffer with natural back-pressure mechanism) and then serialize them with a single writer into an append-only log structure. Many newer devices have something like this at the hardware level, but your block size is still a constraint when working in hardware. If you batch in software, you can hypothetically write multiple logical business transactions per block I/O. When you physical block size is 4k and your logical transactions are averaging 512b of data, you would be leaving a lot of throughput on the table. Going down 1 level of abstraction seems important if you want to extract the most performance from an SSD. Unsurprisingly, the above ideas also make ordinary magnetic disk drives more performant & potentially last longer.
- pclmulqdq 5y agoI used to think the same thing, but now that I work on SSD-based storage systems, I'm not sure this holds up in today's storage stacks. Log structuring really helped with HDDs since it meant fewer seeks. In particular, the filesystem tends to undo a lot of the benefits you get from log-structuring unless you are using a filesystem designed to keep your files log-structured. Using huge writes definitely still helps, though. A paper that I really like goes deeper into this: http://pages.cs.wisc.edu/~jhe/eurosys17-he.pdf http://pages.cs.wisc.edu/~jhe/eurosys17-he.pdf Edit: I had originally said "designed for flash" instead of "designed to keep your files log-structured." F2FS is designed for flash, but in my testing does relatively poorly with log-structured files because of how it works internally. Edit 2: de-googled the link. Thank you for pointing that out.
- trulyme 5y agoDegoogled link: http://pages.cs.wisc.edu/~jhe/eurosys17-he.pdf http://pages.cs.wisc.edu/~jhe/eurosys17-he.pdf
- 10000truths 5y agoAchieving cutting-edge storage performance tends to require bypassing the filesystem anyways. Traditionally, that meant using SPDK. Nowadays, opening /dev/nvme* with O_DIRECT and operating on it with io_uring will get you most of the way there. In either case, the advice given in the article and by the OP is filesystem agnostic.
- DrNuke 5y agoA number of high-level techniques help rationalize data management and transfer, but the mileage of practical implementations may vary a lot. Generally speaking, only a small number of applications really need to take care and add a further layer of abstraction, that because the best practices already codified into any widespread language do an acceptable job already.
- rabuse 5y agoA little off topic, but I bought a new Macbook Pro with the M1 chip with 8GB of RAM, and I'm worried about the swap usage of this machine wearing out the SSD too quickly. Is this an actual concern, as my swap has been in the multiple GB range with my use?
- cbsmith 5y agoIt's an actual concern for you. For Apple it's a variant on planned obsolescence. ;-) Note though that memory use metrics on MacOS can been a misleading. Make sure that you're seeing what's actually there.
- 1-6 5y agoFrom what I’ve been able to gather, the excessive paging may actually have to do with non-native apps running on the M1. Avoid those.
- rabuse 5y agoMost of my programs are JetBrains IDE's and browsers. Don't know if they're optimized for M1.
- Liquid_Fire 5y agoAFAIK most of JetBrains' IDEs are now native (other than Android Studio, which is still WIP). The mainstream browsers are also all native. Remaining non-native apps include Dropbox, Spotify, LibreOffice and a few others. And basically all games with very few exceptions. This website has a decently up-to-date list of what has been ported and what hasn't: https://isapplesiliconready.com/ https://isapplesiliconready.com/
- Grazester 5y agoWhy did you get the 8 gig version? If you are using all this swap then your purchased the wrong MacBook.
- 5y ago
- rectang 5y agoI wince at the amount of wear the `git clean -dxf; npm ci` cycle must be putting on my SSD.
- githubalphapapa 5y agoIf you're on Linux, libeatmydata might help reduce the number of writes hitting the SSD.
- dan-robertson 5y agoSee this paper from 2017, The unwritten contract of solid state drives: https://dl.acm.org/doi/10.1145/3064176.3064187 https://dl.acm.org/doi/10.1145/3064176.3064187
- klodolph 5y agoIf you care about SSDs, one paper you should read is “Don’t Stack Your Log on My Log” by Yang et al. 2014 https://www.usenix.org/system/files/conference/inflow14/inflow14-yang.pdf https://www.usenix.org/system/files/conference/inflow14/infl... > Log-structured applications and file systems have been used to achieve high write throughput by sequentializing writes. Flash-based storage systems, due to flash memory’s out-of-place update characteristic, have also relied on log-structured approaches. Our work investigates the impacts to performance and endurance in flash when multiple layers of log-structured applications and file systems are layered on top of a log-structured flash device. We show that multiple log layers affects sequentiality and increases write pressure to flash devices through randomization of workloads, unaligned segment sizes, and uncoordinated multi-log garbage collection. All of these effects can combine to negate the intended positive affects of using a log. In this paper we characterize the interactions between multiple levels of independent logs, identify issues that must be considered, and describe design choices to mitigate negative behaviors in multi-log configurations.
- rossdavidh 5y agoInteresting, and fun to read and think about! And, as a professional programmer for 17 years now, not once have I done anything where this would have been important for me to know (even if I had been running my code on a system with SSD's). So, I'm not convinced the title is at all accurate. But, fun to read and think about.
- cottsak 5y agoI think the key is hidden in > which can help creating software that is capable of exploiting them Unless you're writing desktop software or your application behaves in a way where you have actually selected the particular hardware components (most of us in cloud hosting don't do this), you probably don't [need to] care.
- deleted 5y ago[deleted]
- wly_cdgr 5y agoThere's nothing whatsoever I should need to know about SSDs as a Javascript programmer and if there is then the programmers on the lower levels haven't done their jobs right and are wasting my time
- BrissyCoder 5y agoWhy on earth do 99.5% of programmers even need to know what SSD stands for?
- ropeladder 5y agoIf sequential and random reads are mostly the same on SSDs, does that make the distinction between columnar and row-based databases/data storage less important?
- wtallis 5y agoNope, unless your columns are all several kB wide. If you force the hardware to perform a multi-kB read for each 64-bit value you need, you're still going to waste a lot of potential performance.
- riobard 5y agoOne thing I'm still puzzled about SSD over-provisioning, which is also mentioned by the tutorial (https://codecapsule.com/2014/02/12/coding-for-ssds-part-4-advanced-functionalities-and-internal-parallelism/ https://codecapsule.com/2014/02/12/coding-for-ssds-part-4-ad...) recommended by the article: > A drive can be over-provisioned simply by formatting it to a logical partition capacity smaller than the maximum physical capacity. The remaining space, invisible to the user, will still be visible and used by the SSD controller. Does the controller read the partition table to decide that the space beyond logic partition is safe to use as scrap?
- ars 5y agoAny sector with nothing written on it can be used as scrap. So if you partition the entire thing, but just never write to the full disk (you never use all the space), that also works as overprovisioning. Partitioning just forces that to happen.
- riobard 5y agoIf I partition the entire drive, eventually all blocks will be used, depending on how the filesystem allocates, right? So to guarantee some free space it's better to over-provision by under-partitioning. Now how do I make sure that on a used drive?
- ars 5y agoThat's what the trim command does. It runs periodically and lets the SSD know about unused areas. So as long as you don't fill up the drive and let trim do its thing the unused areas effectively do the same thing as over provisioning.
- rdc12 5y agoYou could use some sort of disk quota system, to make the filesystem artificially smaller than it actually is (trim after applying this change). Or simply insure that you don't exceed 80% - 90% used space. It it is also worth noting that many SSD's are over-provisioned by the manufacture anyway, in those drives manual over-provisioning might achieve very little anyway.
- mikewarot 5y agoIf you leave un-partitioned space on the SSD, how the heck does the SSD know it is ok to erase it? Wouldn't it be safer to partition it as an extra drive letter, format it, and then leave that drive alone? That would allow the OS to trim all the "empty" blocks.
- qiqitori 5y agoNot 100% sure what you are replying to, and not sure what you meant by "safer", but this may help: The actual physical address on the storage chip and the physical address from the operating system's perspective don't have much to do with another. For harddrives, "un-partitioned space" means that there is a physical "chunk of metal" that is unused. However, that's not the case for SSDs. SSDs dynamically remap "OS-physical" block numbers to whatever they want. (Preferably addresses that have never been used before or that have been discarded/trimmed. If there aren't any available, perhaps to the address that was previously used for the same block number.)
- mikewarot 5y ago>Not 100% sure what you are replying to, and not sure what you meant by "safer", but this may help: I'm replying to the whole of comments on this article. The write amplification problem goes up as the number of "free" sectors/blocks goes down. Many solutions have been presented that don't allocate X% of the hard drive... but I'm not sure than any of them let the hard drive's SSD controller know they aren't allocated. For that to happen, the OS has to have TRIM support, AND the block in question has to be on a volume that the OS is managing. My worry is that if you have a blank partition, it's not being actively managed by anything, and thus isn't going to be TRIMed, and thus the SSD doesn't know the blocks are free for use. Thus, leaving an unpartitioned area isn't going to help.
- rdc12 5y agoThe drive can infer that the LBA hasn't been mapped, since it won't be present in the FTL, there is no need for the OS to inform the drive of this.
- Agentlien 5y agoThis reminds me of a recent interview[0] by Digital Foundry with the Core Technology Director of Ratchet and Clank: Rift Apart. Near the beginning they talk about how targeting the PlayStation 5, which has an SSD, drastically changed how they went about making the game. In short, the quick data transfer meant they were CPU bound rather than disk bound and could afford to have a lot of uncompressed data streamed directly into memory with no extra processing before use. [0] https://youtu.be/-YpCQrPRpE0 https://youtu.be/-YpCQrPRpE0
- BatteryMountain 5y agoSo.. interesting topic. Last year I experimented with some C# + Samsung 970 Evo Plus Nvme + MessagePack (with compression) + Zfs .. to benchmark how fast I could dump objects from .net memory to disk. The numbers involved was insane and I played with various scenarios, with/without compression (MessagePack feature), with/without typeless serializer (MessagePack feature), with/without async and then the difference between using sync vs async and forcing disk flushes. I also weighed the difference between writing 1 fat file (append only) or millions of small files. I also checked the difference between using .net streams versus using File.WriteAllBytes (C# feature, an all-in-memory operation, good for small writes, bad for bigger files or async serialization + writing). I also played with the amount of objects involved (100K, 1M, 10M, 50M). I cannot remember all the numbers involved, but I still have the code for all of it somewhere, so maybe I can write a blogpost about it. But I do remember being utttterly stunned about how fast it actually was to freeze my application state to disk and to thaw it again (the class name was Freezer :p). The whole reason was, I started using Zfs and read up a bit about how it works. I also have some idea about how ssd's work. I also have some idea how serialization works and writing to disk works (streams etc).. I also have a rough idea how mysql, postgres, sql server save their datafiles to disk and what kind of compromises they make. So one day I was just sitting being frustrated with my data access layers and it dawned on me to try and build my own storage engine for fun, so I started by generating millions of objects that sits in memory, which I then serialized with MessagePack using a Parallel.Foreach (C# feature) to a samsung 970 evo plus to see how fast it would be. It blew my mind and I still don't trust that code enough to use it in production but it does work. Another reason why I tried it out, was because at work we have some postgres tables with 60m+ rows that are getting slow and I'm convinced we have a bad data model + too many indexes and that 60m rows are not too much (since then we've partitioned the hell out of it in multiple ways but that is a nightmare on its own since I still think we sliced the data the wrong way, according to my intuition and where the data has natural boundaries, time will tell who was right). So I do believe there is a space in the industry where SSD's, paired with certain file systems, using certain file sizes and chunking, will completely leave sql databases in the dust, purely by the mechanism on how each of those things work together. I haven't put my code out in public yet and only told one other dev about it, mostly because it is basically sacrilege to go against the grain in our community and to say "I'm going to write my own database engine" sounds nuts even to me.
- 2OEH8eoCRo0 5y ago>Drives not Disks And where did the word "drive" come from? I thought it referred to motors that spin the media, which SSDs also do not have.