4 ms·
We measure skill differentiation between frontier / last-gen LLMs across our environments, and one of our curated coding environments is a closed-system market
by gertlabs 18d ago
We measure skill differentiation between frontier / last-gen LLMs across our environments, and one of our curated coding environments is a closed-system market simulator, containing only other agents and some system participants (a market maker and a liquidity provider via issuance / buybacks) whose behavior is fully defined for all of the agents.
This has the least measured skill differentiation of all of our environments, and not because forecasting/markets don't require skill or intelligence. Even the best models are so far from anticipating the behavior of the other agents and understanding the emergent effects that a 2025 model with a naive strategy can often outperform over the timeframes of the simulation simply because some other models in the simulation chose a similar self-reinforcing strategy. This likely happens to some degree in real markets.
You can watch these simulations here https://gertlabs.com/spectate?game=market https://gertlabs.com/spectate?game=market