6 ms·
Yeah - the article was talking about mmap... but what i wanted was to not have to define persistence boundary. I wanted the entire in-memory state of my program
by zinodaur 3y ago
Yeah - the article was talking about mmap... but what i wanted was to not have to define persistence boundary. I wanted the entire in-memory state of my program to be persisted - perhaps even duplicated and moved elsewhere
- ddalex 3y agoBut then you can't just reset the state when crashes happen and data corrupts
- gumby 3y agoThis shouldn’t be voted down. This problem, in the more general case, was inherent at the system level with the persistent and PARC’s immersive Smalltalk and Interlisp environments on the D-Machines. It was much better to have full source you could reload into a fresh environment.
- tanelpoder 3y agommap on Optane direct-access-aware (DAX) filesystems like EXT4/XFS now, is not like mmap on block devices where the OS gets in your way and pages stuff in from disk and (maybe) later syncs it back to persistent storage. Optane is the persistent storage, it's just usable/addressable as regular RAM as it's plugged in to DIMM slots. And in the later Xeon architecture (Xeon Scalable 3rd gen, I think), intel expanded the persistence domain to CPU caches too. So, you didn't even have to bother with CLFLUSH and CLWB instructions to manually ensure that some cache lines (not 512B blocks, but 64B cache lines) get persisted. You could operate in the CPU cache and in the event of power loss, the CPU/mem controllers/and the capacitors on Optane DCPMMs ensured that the dirty cache lines got persisted to Optane before the CPU lights went off. But all this coolness a bit too late... Another note: Intel's marketing had terrible naming for Optane stuff. Optane DCPMMs are the ones that go into DIMM slots and have all the cool features. Optane Memory SSDs (like Optane H10) are just NAND SSDs with some Optane cache in front of them. These are flash disks, installed in PCIe slots but Intel decided to call these disks "Optane Memory" ...
- buildbot 3y agoThere are also fully optane based SSDs such as the 905p or the 5800x
- tanelpoder 3y agoYes, which made it even more confusing, why call the Optane-cached consumer NAND disks "memory" ... but perhaps they thought that it's easier to fool the consumer segment (?)
- zozbot234 3y ago> mmap on Optane direct-access-aware (DAX) filesystems like EXT4/XFS now, is not like mmap on block devices where the OS gets in your way and pages stuff in from disk Yes, the real case for Optane memory is that, supposedly, you don't have to fsync(). And insisting on proper fsync() tends to tank the performance of even the fastest NVMe SSD's. So the argument for a real, transformative performance improvement is there.
- adastra22 3y agoWhy would you not have to fsync? The fsync is a memory barrier that is just as useful with octane to ensure integrity. Do you mean the latency of ensuring fsync safety is lower?
- RyanHamilton 3y agoNo you don't have to fsync. Think of it like RAM. You don't fsync RAM.
- adastra22 3y agoYou do, in fact. It’s called a memory write barrier. Ensures consistency of data structures as needed. And it call stall the cpu pipeline, so there’s a nontrivial cost involved.
- jared_hulbert 3y agomemverge.com does some cool work around making that happen.
- gwd 3y agoI'll admit this sounded cool when I first heard about it; but it's actually a lot harder to program if you want to be able to recover from sudden power outages (which would be the main reason for having persistence in the first place).
- barrkel 3y agoThat's an easy way to accumulate data corruption. It's better to design for unexpected restarts than design for a golden in-memory image which needs to be carefully ported around, have all its connections wired back up, and so on. You're going to get unexpected restarts anyway. The faster and more reliable you can make recovery from that, it benefits you in the moving use case. The kinds of things you might want to do to enable reliable restart - like retry mechanisms for incoming requests - make migration work too.
- tanelpoder 3y agoYou can design for that. For example, when building a persistent-memory native database engine, you probably need some sort of data versioning anyway - either Postgres style multi-version rows (or some other memory structures) that later need to be vacuumed or Oracle/InnoDB style rollback segments that hold previous values of some modified objects. Then you probably want WAL for efficient replication to other machines and point in time recovery (in case things go wrong or just DB snapshots for dev/test). Transient & disposable memory structures like keeping track who's logged in or compiled SQL execution plans that facilitate access to the persistent "business data", much of that stuff will need to be in RAM/HBM/CPU cache anyway, for performance reasons and as these things do not necessarily need to persist across a crash/reboot. The data (and likely indexes, etc) need to. But you won't need a buffer cache manager that copies entire blocks around from storage to different places in memory and vice versa. Your giant index or graph could rely just on direct memory pointers instead of physical disk block addresses that need to get read to somewhere in memory and then are accessed via various hashtable lookups & indirect pointers. And you don't have to ship entire 512B-8kB blocks around just to access the next index/graph pointer, just access only the relevant cache line, etc. With proper design, you'd still have layers of code that take care of coherency, consistency and recovery...
- addaon 3y agoLook at developing for MSP430s with FRAM -- these microcontrollers have a decent amount of FRAM with full persistence, full XIP etc, up to 256 kB; but only 8 kB or less of traditional SRAM. Even in this world, where you /could/ have everything persisted, you still end up aware of the persistence boundary and using SRAM both for the absolutely highest-performance code (e.g. interrupt handlers; FRAM has more wait states than SRAM in this implementation), but more interestingly for things that specifically /should not/ be persisted (e.g. the bytes storing whether your POST has completed, hardware initialized, etc). You can come close to persistence-oblivious, especially at a conceptual "application layer", but the overall implementation still ends up persistence-aware.