4 ms·
https://x.com/garybernhardt/status/600783770925420546 https://x.com/garybernhardt/status/600783770925420546 (Gary Bernhardt of WAT fame): > Consulting service:
by _ugfj 2y ago
https://x.com/garybernhardt/status/600783770925420546 https://x.com/garybernhardt/status/600783770925420546 (Gary Bernhardt of WAT fame):
> Consulting service: you bring your big data problems to me, I say "your data set fits in RAM", you pay me $10,000 for saving you $500,000.
This is from 2015...
- RandomCitizen12 2y agohttps://yourdatafitsinram.net/ https://yourdatafitsinram.net/
- crowcroft 2y agoI wonder if it's fair to revise this to 'your data set fits on NVME drives' these days. Astonishing how fast and how much storage you can get these days.
- fbdab103 2y agoYou can always check available ram: https://yourdatafitsinram.net/ https://yourdatafitsinram.net/
- xethos 2y agoBased on a very brief search: Samsung's fastest NVME drives [0] could maybe keep up with the slowest DDR2 [1]. DDR5 is several orders of magnitude faster than both [2]. Maybe in a decade you can hit 2008 speeds, but I wouldn't consider updating the phrase before then (and probably not after, either). [0] https://www.tomshardware.com/reviews/samsung-980-m2-nvme-ssd-review https://www.tomshardware.com/reviews/samsung-980-m2-nvme-ssd... [1] https://www.tomshardware.com/reviews/ram-speed-tests,1807-3.html https://www.tomshardware.com/reviews/ram-speed-tests,1807-3.... [2] https://en.wikipedia.org/wiki/DDR5_SDRAM https://en.wikipedia.org/wiki/DDR5_SDRAM
- dralley 2y agoThe statement was "fits on", not "matches the speed of".
- Dylan16807 2y agoSeveral gigabytes per second, plus RAM caching, is probably enough though. Latency can be very important, but there exist some very low latency enterprise flash drives.
- int_19h 2y agoI think the point is that if it fits on a single drive, you can still get away with a much simpler solution (like a traditional SQL database) than any kind of "big data" stack.
- deleted 2y ago[deleted]
- weebull 2y agoI always heard it as "if the database index fits in the RAM of a single machine, it's not big data". The reason being that this makes random access fast. You always know where a piece of data is. Once the index is too big to have in one place, thing get more complicated.
- justsomehnguy 2y ago980 is an M.2 drive, PCIe 3.0 x4, 3 years old, up to 3500MB/s sequential read. You want something like PM1735: PCIe 4.0 x8, up to 8000 MB/s sequential read. And while DDR5 is surely faster the question is what the data access patterns are there. In almost all cases (ie mix of random access, occasional sequential reads) just reading from the NVMe drive would be faster than loading to RAM and reading from there. In some cases you would spend more time processing the data than reading it. PS all these RAM bandwidth rates are good for the sequential access, as you go random access the bandwidth drops. https://semiconductor.samsung.com/ssd/enterprise-ssd/pm1733-pm1735/ https://semiconductor.samsung.com/ssd/enterprise-ssd/pm1733-...
- datadrivenangel 2y agoYour data access patterns are fast enough in NVME. You own me $10,000 for saving you $250,000 (in ram). The value of our data skills are getting eroded!