3 ms·
With causal masking and autoregressive token generation, it’s not clear to me that it is inherently different. My original expectation was the same as the way
by deepsquirrelnet 3y ago
With causal masking and autoregressive token generation, it’s not clear to me that it is inherently different.
My original expectation was the same as the way instructor software implemented it. But I found the prompt in the article confusing toward that perspective. I’m sure it can work either way, but it should be a lot more performant (and less expensive) as a single pass.
- Terretta 3y agoI can't get it to respect this instruction in single pass mode: - Never drop entities from the previous summary. If space cannot be made, add fewer new entities. Specifically, for me it randomly drops entities. > a lot more performant Faster? Absolutely. But I'm not having luck getting it smarter.