3 ms·
Yes we’re running our own inference server because we do nonstandard things in the decoder loop (think logit masking, other types of sampling, etc).
by zackangelo 2y ago
Yes we’re running our own inference server because we do nonstandard things in the decoder loop (think logit masking, other types of sampling, etc).