3 ms·
The demo looks too slow for practical usage. How much it will cost if i host it in cloud to get instant response similar to the speed of openAI?
by android521 3y ago
The demo looks too slow for practical usage. How much it will cost if i host it in cloud to get instant response similar to the speed of openAI?
- 3Sophons 3y agohttps://github.com/second-state/WasmEdge-WASINN-examples/tree/master/wasmedge-ggml-llama-interactive#performance https://github.com/second-state/WasmEdge-WASINN-examples/tre... These are the speed of different hardwares. You can rent GPU, which will be much faster
- 3Sophons 3y agoThe demos are as fast as the ChatGPT, they just look slower. -First. It might seem slow at first because it's loading the model. But once that's done, it gets much quicker for any follow-up requests. Second, Turning on streaming output helps a lot, as it shows responses as they're processed. And yes, the hardware matters too. A good GPU setup in the cloud can work wonders, though it might bump up the cost a bit. So actually the demo's not slower than ChatGPT overall; it just takes a moment to warm up at the start.