4 ms·
Would this also be possible with other LLM engines / GPUs? E.g. Llama / Apple Silicon or Radeon?
by terhechte 1y ago
Would this also be possible with other LLM engines / GPUs? E.g. Llama / Apple Silicon or Radeon?
- saagarjha 1y agoYeah, none of this is specific to CUDA (though the relative latencies might be different).