3 ms·
Can you share who your vendor/host is? In another comment you mentioned that you're mmap'ing files on SSDs... on vendors like GCP that sort of hardware (especia
by counters 5y ago
Can you share who your vendor/host is? In another comment you mentioned that you're mmap'ing files on SSDs... on vendors like GCP that sort of hardware (especially the storage volume you'd need to have access to a handful of the local/global models, as you reference on your website) isn't particularly cheap if you want to have 100% uptime.
A cheaper/scalable approach instead would be to re-process your data into an appropriately chunked, cloud-optimized storage format like Zarr and save it in object storage. Then your scaling bottleneck would just be the VMs or compute you use to query from object storage, as a function of traffic/load.
- meteo-jeff 5y agoThe host is netcup in Germany. I use regular KVMs. Yes, the amount of storage can be an issue, but I want to stay below 500GB of hot-data. One bottleneck is network traffic to copy updated files after each weather model update. My binary files are just plain Float16 files without an meta data. Logically they similar to 3D Zarr files. I know exactly which bytes I have to read and the kernel cache helps a look to keep it fast. In theory this data does not have to be a file on SSD. I could also use a block storage directly or request data from S3 via range requests. One hopelessly over-engineered approach would be to use an `AWS EBS io2 Block Express Volume` use `Multi-Attach` and spawn up to 16 API instances to serve data from it.
- tehlike 5y agoCurious how much this would cost you to do from Google cloud storage or aws. Historical data access can be quite valuable (and something that might be a premium paid feature even).
- meteo-jeff 5y agoFor the VMs, the absolut minimum memory requirement is 16GB. The more, the better. Otherwise it takes around 40 GB of disk space for every day of data. For an history of 100 days, 4000 GB are required. With compression I could save 50%, but have to invest a couple of days development time to make is work. You could calculate the AWS bill now ;-) Data on cold storage is an option, but it also super slow.... Currently, I did not yet integrate all the high-resolution models that I want to. Coverage for Europe is great. In North America I will add high resolution NOAA models next. Most likely I will keep only a limited subset of data as history, but on fast storage to make is accessible quickly
- tehlike 5y agoCurious if you could link directly to GCS standard storage and link to it directly - I don't know how slow, but could be reasonable compromise for historical access... Interesting problem to solve.