8 ms·
1 million tokens is great until you notice the long context scores fall off a cliff past 256K and the rest is basically vibes and auto compacting.
by jryio 7mo ago
1 million tokens is great until you notice the long context scores fall off a cliff past 256K and the rest is basically vibes and auto compacting.
- olliepro 7mo agoI bet they lack good long context training data and need to start a flywheel of collecting it via their api (from willing customers)
- jbergqvist 7mo agoThis would be my guess too. It can probably be generated synthetically or via agentic rollouts, but high quality long context examples where outputs meaningfully depend on long-range interactions probably remain scarce
- rrr_oh_man 7mo agoIt's the same now with Gemini as well. Unfortunately. :(