3 ms·
One of the biggest issue with incredibly large context sizes is the lack of training / dataset meant specifically to use such large context sizes. And this app
by pico_creator 3y ago
One of the biggest issue with incredibly large context sizes is the lack of training / dataset meant specifically to use such large context sizes.
And this apply to all models. So when a specific model now does badly even at 32k or 50k it’s hard to say if it’s an architecture design issue, or a dataset issue
- ofermend 3y agoYes; and the amount of training examples that are really long (relative to the other training examples) becomes small really fast, so it's also a problem of the long-tail in a sense.
- mike_hearn 3y agoIt feels like this is where training on code is going to go from important to critical. Most human texts won't require you to look back 100,000 tokens to understand what it means, but if you dump 1000 source files one after the other than it will.