4 ms·
The author here. If you have any questions, let me know. If there is anybody here who wants to help parallelize this, let me know!
by certik 3y ago
The author here. If you have any questions, let me know.
If there is anybody here who wants to help parallelize this, let me know!
- collaborative 3y agoThe benchmark says the fastest model takes 0.3s for 20 tokens. Does this mean it would take 30 seconds for 2000 tokens?
- jdkee 3y agoSo does the model simply extract the most likely answer to the prompt based on model weights?