3 ms·
How is inference latency for coding use cases on a local 3090 or 4090 compared to say, hitting the GPT-4o API?
by shostack 2y ago
How is inference latency for coding use cases on a local 3090 or 4090 compared to say, hitting the GPT-4o API?
- whereismyacc 2y agoI assume the characteristics would be pretty different, since your local hardware can keep the context loaded in memory, unlike APIs which I'm guessing have to re-load it for each query/generation?
- christina97 2y agoIf you integrate with existing tooling, it won’t do this optimization. Unless of course you really go crazy with your setup.
- moffkalast 2y agoSetting one launch flag on llama.cpp server hardly qualifies as going crazy with one's setup.