4 ms·
What we need is a platform for benchmarking hardware for AI models. With X hardware you can get X amount of tokens with X amount of latency For context token pr
by FloatArtifact 2y ago
What we need is a platform for benchmarking hardware for AI models. With X hardware you can get X amount of tokens with X amount of latency For context token pre-filled. So, standard testing methodology per model with user-supplied benchmarks. Yes, I recognize there's going to be some variability based on different versions of the software stack and encoders.
End user experience should start by selecting the models of interest to run and output hardware builds with price tracking for components.
- pmontra 2y agoAgreed. I opened the comments to write that nearly all of those articles spend very few words on the hardware, its cost and the performances compared to using a web service. The result is that I'm left with the feeling that I have to spend about $1,000 plus setup time (HW, SW) and power to get something that could be slower and less accurate than the current free plan of ChatGPT.