3 ms·
Show HN: LLM Inference Calculator – Estimate VRAM, Latency, and Throughput
- popopanda 1mo ago[flagged]
- maestroquirk 1mo agoNeed one for vision models too tbh. Token/s doesn't really map easily
- popopanda 1mo agoYeah, tokens/s varies a lot based on workload. However, I’ve calibrated the estimator against public benchmarks, so it stays within a 30% error margin! I'll definitely explore how to estimate vision models next!
- zuuna 1mo agoHow about the other way around? I'll tell you my hardware and you tell me the possible stats? I've been having some problems with bigger models, using MoE to make it work for longer/tedious work like nightly sweeps that don't really on speed.
- popopanda 1mo ago[flagged]
- brownbear7 1mo agoCould you add cost estimates too? I usually care about the tradeoff between VRAM/throughput and roughly what the same workload would cost on different providers.