3 ms·
interesting paper. essentially a peek under the hood of why transformers generalize so well across multiple tasks, through the lens of hidden markov models.
by badmonster 1y ago
interesting paper. essentially a peek under the hood of why transformers generalize so well across multiple tasks, through the lens of hidden markov models.