3 ms·
I’ve been playing around with something similar for factual nouns, where the next token prediction is a token like [PNOUN] and then a downstream predictor choos
by tempusalaria 3y ago
I’ve been playing around with something similar for factual nouns, where the next token prediction is a token like [PNOUN] and then a downstream predictor chooses the actual result.
One major Problem is that you need to autoregressively access the correct embedding when doing inference. And in particular in batch inference. So waiting for the downstream prediction to complete is gonna slow inference. One way I was considering getting around it was running all auxiliary classifiers in parallel with the main LLM, and just accessing those results when needed.
Which then runs into the problem that those auxiliary classifiers have to be smallish relative to the main model, and then there are performance questions. A complicated thread to unravel.
Their approach to num doesn’t have the same performance implications but there are some clear issues - large numbers will have blowup issues, 0 and small floats have big problems here.