3 ms·
Congrats on getting your project to the front page of HN. With that said I think you are going to need to change your approach if you want this project to be us
by thefreeman 7y ago
Congrats on getting your project to the front page of HN. With that said I think you are going to need to change your approach if you want this project to be usable as more than a toy project in the long run.
From what I can tell it essentially saves a map of url -> response in memory as you browse. Every 10 seconds this file is serialized to json and dumped to a cache.json file. This is going to be very inefficient as the number of web pages indexed grows since you are rewriting the entire cache every 10 seconds even if only a few pages have been added to it. It also will eventually exceed the memory of the computer running the app if the content of every page ever visited needs to be loaded into memory. I highly recommend looking into some of the other suggestions mentioned here, either sqlite or mapping a local directory structure to your caching strategy so that you can easily query a given url without keeping the entire cache in memory, and also add / update urls without rewriting the entire cache.
- archivist1 7y agoMy future plan was to cache responses on disk and just keep cached keys in memory: https://github.com/dosyago/22120#future https://github.com/dosyago/22120#future
- dunham 7y agoI wrote something similar years ago in Go, and settled on writing the data to a WARC file on disk (you can gzip the individual requests and concatenate to get random access), and also concatenating to a warc index file. The working index was kept in memory, while the warc index was read at startup. My version acted as a proxy and would serve the latest entry from cache if a copy was cached. I had a special X-Skip-Cache header for when I wanted to go around the cache. (I can't remember if it handled https or if sites just didn't use https back then.) My use-case was web scraping, particularly recipe and blog sites. I wanted to be able to develop my scraping code without re-hitting the sites all the time. Structuring it as a proxy allowed me to just write my python scraping code as if I was talking to the server. Previously I'd written a layer on top of the python requests library to consult a cache stored in a directory (raw dumps of content / headers, with v2 involving git). But I found that required extra care when more than one script was running at once, and I liked the idea of storing it in a standardized format (WARC) that could be manipulated by other tools.
- breatheoften 7y agoI tried to build something like this for jest tests in an app I worked on. I wanted my jest tests to serve as both unit tests and service diagnostics - so I instrumented axios and setup a hidden cache layer within it when running inside the test suite. I was trying to figure out how to best organize the cache so I could run tests really quickly by having all results pulled from cache — or run it slow and as a service diagnostic mechanism by deleting the cache before execution ... I had to extend axios to accept a bit of additional logic from the application ... it was hard for me to get it to work properly inside of jest though ...