12 ms·
Show HN: File-based cache for slow Python functions
- andrewgazelka 3y agoThis looks cool :). A while ago, I wrote something similar that analyzes bytecode and invalidates the cache if the bytecode changes. https://github.com/andrewgazelka/smart-cache https://github.com/andrewgazelka/smart-cache
- kapilsinha 3y agoI like the simplicity. I definitely get the payoff for standalone Python scripts, where once the script errors out the memory is cleared. But do you see a similar payoff for Jupyter notebooks (or similar)?
- williamzeng0 3y agoI think the marginal gain would be a lot less for Jupyter notebooks, but I've definitely rerun individual cells and wasted time there before. I think it could help if you forget to save the output of a function within a single cell like this: 1. print(f(x)) # -> check what happened 2. out = f(x) # -> turns out we want to save this, so we have to wait again
- p10_user 3y agoFWIW There is a built-in cache system for r markdown documents. I'm not up to speed on their exact implementation but I have found it useful. https://bookdown.org/yihui/rmarkdown-cookbook/cache.html https://bookdown.org/yihui/rmarkdown-cookbook/cache.html
- williamzeng0 3y agofile.mtime. (file last modified) is an awesome way to key the cache.
- skp1995 3y agoThis is a pretty good implementation. I like the simplicity of it, reminds me of SQLite backed storage decorators we used to have, where the data was persisted to a DB instead of the file system (altho thats just a different storage engine) Does this also take care of the thundering heard problem? That was one of the cases where lru_cache really blows
- williamzeng0 3y agoUnfortunately it doesn't, we typically don't expect to handle high load with this cache and actually disable it in production with another envvar. Sometimes caching can actually be slower for certain functions, because just performing that operation is faster than pickle.load/pickle.dump.
- rmholt 3y agoI have extensively used https://pypi.org/project/diskcache/ https://pypi.org/project/diskcache/. Is there a reason you decided to make an in house solution?
- quickslowdown 3y agoI found DiskCache sometime last year, it's amazing. Very simple to set up and works great as a cache for so many different things.
- dotancohen 3y agoWhat are you using it for? A disk based cache seems almost contradictory for my use cases, I would love to hear yours. Anything that I would store on disk, even as a cache, I can generally put in SQLite.
- akx 3y agoYou'll be happy to know that Diskcache is backed by SQLite (and/or spill files for large enough (size configurable) blobs).
- rmholt 3y agoIt pretty much does that, if I was to be a little reductive, DiskCache is just a wrapper around sqlite and pickle
- p10_user 3y agoThe degree of reduction is nice considering the countless times in the past where I wrote my own file cache logic using if/else statements, temporary files, pickle, and bespoke sqlite databases. Why give myself the headache of maintaining so much extra code when someone already wrote it.
- quickslowdown 3y agoI write some toy Python scripts/apps for a couple APIs (the Pokeapi, Spacetraders, etc). I use the HTTPX library as a request client, and wasn't aware of the Hishel library for caching requests. I used DiskCache to cache responses for ~15 minutes so I wasn't sending live requests every time I tested the app. I'm not building anything "cloud scale," most things I run are off my local machine. Having a convenient, fast local cache that's simple to use (DiskCache) has so many uses, it's hard to think of them all! I might use a cache with no expiration to store some configs, or to store a serialized object for later retrieval. I might use it as an in-memory object cache while the program loads, so I don't have to spin up a Redis server.
- wildermuthn 3y agoIf you aren’t caching LLM functions during development, then you’re an even greater glutton for punishment than the normal engineer. My local file cache Python decorator also allows the decorator to define the hash manually, either by the decorator’s parameter function call that plucks a value from the cached function params, or by calling a global function from anywhere with any arbitrary value. What’s cool about caching results locally to files during development is the ease of invalidating caches — just delete the file named after the function and key you want.
- mpeg 3y agoThis is also why in my custom cache I back it with sqlite – much easier to delete one db file than thousands of pickle files.
- AlecSchueler 3y agoGlobs are a thing?
- mpeg 3y agoWeird comment... yes, they are, but what is faster for me to type `rm .cache_dir/function_name*.pickle` or just to delete the one sqlite file in my file manager / vscode file tree. Regardless, there are other reasons why sqlite is nice for this, you gain control over locking and thread safety without having to implement it all from scratch
- canadiantim 3y agoI'm sure this is a stupid question, but why is it much better to be caching LLM functions during development?
- p10_user 3y agoBecause they are generally incredibly computationally expensive operations that can take hours/days to complete (?more)
- mpeg 3y agoI recently wrote a version of this that I use in my projects, some things I do differently that you may or may not care about: - from your code it seems you're not sorting kwargs, I would strongly recommend sorting them so that whether you call f(a=1, b=2) or f(b=2, a=1) the cache key is the same - I use inspect.signature to convert all args to kwargs, this way it doesn't matter how a function gets called, the cache logic is always consistent. I know this is relatively slow but it only gets called once per function (I call it outside the wrapper) and the DX benefits are nice (in this same note, you could probably move the inspect.getsource call outside your wrapper fn for a speed boost) I also took the opposite approach to ignore_params, and made the __dict__ params that get hashed opt-in, which works well when caching instance methods
- AlecSchueler 3y agoVery insightful comment, but can I ask what DX stands for? Maybe I'm missing something obvious.
- oulipo 3y ago"Developper Experience", eg good developper tools / libs
- by_the_bay 3y agoDeveloper experience
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- williamzeng0 3y agoMaking the __dict__ opt-in makes it a lot more user-friendly at the expense of a little verbosity. That makes sense. These tips make sense, we often use named args in our function calls (not using them has caused so many bugs), but we don't really enforce the order. Copilot doesn't always get it right either. By moving inspect.getsource out of the wrapper, do you mean initializing it when the module is imported? I'm curious how that improves performance.
- epr 3y agodef hash_code(code): return hashlib.md5(code.encode()).hexdigest() Be warned. The above function is used as part of the hash. The ostensible purpose is to prevent using cached values of functions who's code has changed, but it does not handle dependencies of that function.
- martinky24 3y agoHow do you suggest one might fix that issue? Also pin the cache to a hash of all dependency versions? And then if one minor update And let's say the dependency did change, but it's generally inert (more error handling around edge cases, for example), how do you factor that in? Blow up the whole cache? Your example isn't really a problem with OPs utility, but a specific example of a broader dependency management problem that affects just about everything. The answers usually boil down to 1) invest heavily in a kick ass test suite, 2) never upgrade or 3) upgrade and pray nothing breaks.
- williamzeng0 3y ago+1, we considered traversing the function's dependencies to key the cache on (not just the initial function source code), but decided to leave this in a as a constraint. Otherwise we also blowing up the cache when we didn't want it to happen.
- epr 3y ago> How do you suggest one might fix that issue? Also pin the cache to a hash of all dependency versions? Pretty much. Recursively collect dependencies by analyzing the AST of the code. > And then if one minor update And let's say the dependency did change, but it's generally inert (more error handling around edge cases, for example), how do you factor that in? Blow up the whole cache? You're saying that like it's some kind of ridiculous ask, but yes. The current implementation is already "Blow[ing] up the whole cache" whenever the code for the decorated function is changed anyways. I'd guess that additionally handling dependencies recursively would only modestly increase the rate of "Blow[ing] up the whole cache". > Your example isn't really a problem with OPs utility... Whether or not this is a problem in practice obviously depends on your use case. Maybe you don't generally care if functions return the correct result, but many do. > [This is] a specific example of a broader dependency management problem that affects just about everything. Dependency resolution is not trivial per se, but it's a pretty common problem. Every single package manager, build system (make), etc. have all solved this.
- rassibassi 3y agoWhat's the difference to using joblibs Memory class similar to this implementation: https://github.com/stanfordnlp/dspy/blob/main/dsp/modules/cache_utils.py https://github.com/stanfordnlp/dspy/blob/main/dsp/modules/ca...
- khaledh 3y agoI was going to mention this as well. It's fairly similar: memory = joblib.memory.Memory(...) @memory.cache def slow_func(...): ...
- rassibassi 3y agoThe diskcache docs state: """ Caching Libraries joblib.Memory provides caching functions and works by explicitly saving the inputs and outputs to files. It is designed to work with non-hashable and potentially large input and output data types such as numpy arrays. """ From https://pypi.org/project/diskcache/ https://pypi.org/project/diskcache/
- williamzeng0 3y agoThis is great! I see it also supports an 'ignore' parameter.
- rthnbgrredf 3y agoRecently, I experimented with various techniques to cache some JSON responses from FastAPI, using Python decorators for both in-memory and disk caching on a single machine. After benchmarking the performance, I found the results somewhat disappointing (500 req/s vs 5k req/s). While caching did lead to a tenfold improvement in speed compared to no caching, I believe the primary bottleneck was Python's inherent performance limitations, which made it X times slower than a comparable program written in C. Consequently, I decided to remove the cache decorator and instead put a simple nginx caching reverse proxy in front of FastAPI. This resulted in performance gains that were an order of magnitude better (60k req/s) than those achieved with Python based caching.
- kevinlu1248 3y agoWe also found a lot of cases where caching ends up being actually slower than doing the operation. The 100% solution would probably be to use a SQL db the way diskcache does it, but this is easier to use for us.
- emilehere 3y agoReminds me of a little prototype I wrote a while ago that tried to do something similar with Javascript's Proxy class. https://github.com/emileindik/cashola https://github.com/emileindik/cashola The main difference is that it stores the state of an object, not a function. If your data is JSON serializable then it could be a cool way to save and resume application state.