4 ms·
This ignores batching - token generation is much more efficient in batch - and I strongly suspect is itself written by AI, given the heavy use of bullets
by apsec112 1y ago
This ignores batching - token generation is much more efficient in batch - and I strongly suspect is itself written by AI, given the heavy use of bullets
- biophysboy 1y agois it common for adjacent tokens to use the same weights in a memory cache?
- twoodfin 1y agoThe “X—not Y” pattern is also a dead giveaway.