2 ms·
> generative inference this is just what the doctor ordered. Naive question, wouldn't you need a descent tool chain for inference as well ?
by soulbadguy 2y ago
> generative inference this is just what the doctor ordered.
Naive question, wouldn't you need a descent tool chain for inference as well ?
- ein0p 2y agoAssuming you mean “decent” toolchain, it’s actually pretty decent. Could it use some polish? Yes. But any decent ML engineer would be able to get a high performance server (or a batch job) running in a relatively short time. Or in almost no time at all if using a lot of the FOSS models. You just basically create a model in PyTorch and then hand it over to Gaudi stuff which patches it with optimized, Gaudi specific ops and converts things into an inference graph. “Closeness to CUDA” is less important for inference because all the experimentation is already done by then, and if need be you could just implement the model using Gaudi ops to begin with, in a span of a few days including tuning and debugging.