4 ms·
GPU utilization should be down when using this technique. I’m hoping this could allow for more efficient batch inference on GPUs. If you can predict 10 tokens f
by valine 3y ago
GPU utilization should be down when using this technique. I’m hoping this could allow for more efficient batch inference on GPUs. If you can predict 10 tokens for the price of 1 it should allow you to do tree of thought much more efficiently.
https://github.com/princeton-nlp/tree-of-thought-llm https://github.com/princeton-nlp/tree-of-thought-llm