3 ms·
This only works well if you are creating sub agents with clean contexts for each task. If you constantly are switching models part way through some work then th
by iamflimflam1 3mo ago
This only works well if you are creating sub agents with clean contexts for each task. If you constantly are switching models part way through some work then the whole session needs to be replayed each switch. You lose all the benefits of the context cache.
- xyzsparetimexyz 3mo agoWell yeah, it'd require the models to be loaded on the same system and the cache to be shared between them somehow
- inigyou 3mo agoCache sharing is not possible. The numbers in the cache are completely specific to the model.
- xyzsparetimexyz 3mo agoCurrently. I'm sure that you could make a system where the cache values are a superset C of e.g. models A and B where C is probably bigger than max(A,B) but smaller than A+B
- inigyou 3mo agoIt's equal to A+B. There is literally no sharing possible.
- xyzsparetimexyz 3mo agoEven when training both models together in a novel way?
- wonnage 3mo agoReddit post so take with a grain of salt but this does seem possible. But unclear whether this is actually a viable architecture https://www.reddit.com/r/LocalLLaMA/comments/1t8s83r/nvidia_ai_releases_star_elastic_one_checkpoint/ https://www.reddit.com/r/LocalLLaMA/comments/1t8s83r/nvidia_...