5 ms·
In my experience, the amount of data many applications are expected to access is growing geometrically, while RAM costs are declining only arithmetically. There
by biggestdummy 9y ago
In my experience, the amount of data many applications are expected to access is growing geometrically, while RAM costs are declining only arithmetically. There certainly are in-memory use cases, but for most use cases it is not now (and probably won't become) economical.
- amelius 9y agoThat's odd, because I've seen the number of cases where in-memory databases make sense increase over the last three decades.
- biggestdummy 9y agoAgreed. For certain use cases (session store, user profile store), the lower cost of memory and the relatively slow growth of data will mean that they will usually fit in memory. But, in my experience, there's a larger, and growing, class of use cases (ML, IOT, Fraud detection, Gaming data) where both the amount of raw data captured and the expectations for data retention are causing the total dataset to spiral into the dozens or hundreds of TB, with expectations for real-time SLAs on query. That's what I was talking about.
- amelius 9y agoYes, but data on the order of TBs is usually stored on platters (as opposed to SSD). In that case, a fast data access library will not buy you much, because the disk will be slow anyway. (By the way it is common that an in-memory database is combined with on-disk storage for large blobs; those blobs are just flat files, and the OS already offers the fastest way to access them without any library)
- seastarer 9y agoMulti-Terabyte SSDs are very common. An i3.16xlarge instance on AWS, for example, has 15TB of SSD storage.
- amelius 9y agoOk, but it doesn't take away the point that very often the database can still be stored in memory, except for large blobs which are stored in flat files. Access of those files is usually sequential, so you will not need a special library.
- seastarer 9y agoIt happens sometimes. Very often? Not in my experience
- biggestdummy 9y agoYes, video streaming and management is a use case as you describe. Metadata is kept in memory, and the video files, themselves, are stored on an inexpensive media - usually as file system. So queries about "what can i watch" are quick, but loading the video can take seconds. But any use case where you want to keep state over time - sums, averages, maxes, histories - requires you to keep all the individual points of data in the database. Those are the (increasingly common) use cases that I run into where very fast disk storage like Scylla are important.
- eloff 9y agoYou're both right and wrong at the same time. The disconnect is the based on the kind of data we're talking about. Human-generated metadata (i.e. OLTP data) grows roughly with respect to the population of internet connected humans - humans only can produce so much content in their finite time. Advances in computers far exceed this growth rate, so more of these kinds of problems fit into RAM each year. The volume of media data, which depends more on the increasing fidelity of the media (videos going from 320K to 4K, photos going from 1MP to 30MP) have been increasing rapidly but that trend won't continue forever because there are diminishing returns. Do you care about the difference between 4K and 8K video? TV companies hope you do. At some point (likely soon) you will stop caring. OLAP data, which includes data produced by IoT devices, sensors, etc is increasing geometrically. But this kind of data is not best stored in RAM anyway, and is served best by scale-out column-oriented and time-series databases.