3 ms·
Launch HN: EdotEnv (YC S26) – Quant Trading RL Envs to Teach LLMs Research
We are Rui and Michael and we’re building EdotEnv (https://edotenv.com https://edotenv.com): self-improving RL environments from Quant Trading workflows.
With all the benchmaxxing around, evals saturate and become meaningless for model comparison. Useful benchmarks should increase in difficulty as models advance. Back in our Quant jobs, Michael and I saw that the market has exactly this property: markets became more efficient as people profited from trading inefficiencies, making new profitable strategies harder to find and old ones decay over time.
This makes markets an ideal, continuously evolving benchmark for LLM training. The hard part is to turn professional quant workflows into reliable training envs, as this is a very niche expertise.
In our environments, we give LLMs a quant trading workflow and evaluate their performance on out-of-sample data: build predictive features/ models, design a portfolio, backtest strategies, adapt continuously to market regimes. Each step is a task with different self-built tools. For example, a predictive feature building task gives the agent cleaned market data of time period [0,T] to research ideas, a backtesting tool to test created features at time t on [0, t], an execution tool to trade strategies with the new features on [t+1, T] and a final evaluation. Our reward isolates the agent's feature building skills and yet benefits from market properties.
From running SOTA models in our environments, we see that i) they seem to struggle with iterating deeply on research ideas, preferring broad shallow searches; ii) higher reasoning does not seem to increase performance and iii) agents do not understand trading, e.g. when losing money they stop trading instead of trading smarter. Check out our blogs for more details! https://edotenv.com/?tab=blog https://edotenv.com/?tab=blog
Quant workflows are essentially applied ML research, long-horizon planning and continual learning. Through our envs, we teach these transferable research skills, rather than task specific answers. Our environments are closer to a realistic research workflow: we use real-world data instead of synthetic ones; our envs naturally contain noise and real trade-offs; our rewards are verifiable and immediate, with no need for an additional LLM judge or human expert.
We open sourced a sample task repository: https://github.com/MMcollab-dotcom/feature-engineering https://github.com/MMcollab-dotcom/feature-engineering. We plan to sell continuously improving envs to AI labs/researchers/enterprises training their own agents, who are interested in ML modelling capabilities, continual learning, long horizon planning or Quant Research in general.
We'd love feedback from anyone trying out their own agents in our envs, for either eval or post training. And of course, we are always happy to discuss the future of trading with LLMs (and no, it should not be asking the LLM to read tea leaves and give you the stock to buy tomorrow). Looking forward to your comments!
- deleted 2mo ago[deleted]
- Mzzzzz 2mo agoHere is an example rollout trace with gpt 5.6 luna. Checkout if you are interested what kind alphas the agent found XD https://hub.harborframework.com/jobs/af0299f9-a3bb-44ea-8ced-7b56682bf739 https://hub.harborframework.com/jobs/af0299f9-a3bb-44ea-8ced...
- jjallen 2mo agoI looked through the transcript/output of the model/run linked but didn't find anything that showed much, if any, alpha. Maybe I missed it?
- RuiWang0811 2mo agonot sure about your background, the trace shows the feature engineering the LLMs did
- jjallen 2mo agoThe comment in the parent said "kind alphas the agent found". It linked to a job then a trial. All of the comments I read there except for one said things like "The final alpha check is still fitting; the prior lagged backtest’s exact zero metrics make clear that model was not acceptable economically". One said "successful backtest’s" but did not expound on what that means or anything. I was looking for the alpha it found and not the feature engineering. My backgroud is in finance but do not look at things so quantitatively. Here is another comment/output: "correlation −0.007 and directional accuracy 0.498". Wouldn't directional accuracy have to be greater than .50 to be profitable? Forgive my ignorance I am interested in this though. I have asked Claude/ChatGPT to show me where the alpha is that the model found as well so I can learn.
- jjallen 2mo agoSpent more time looking at this because I was still interested in the alpha it found and found this: "cumulative_after_cost_return": -0.004003033519454746" https://hub.harborframework.com/jobs/af0299f9-a3bb-44ea-8ced-7b56682bf739/trials/5b38696a-7175-4c68-a673-24c6642c039a?tab=verifier https://hub.harborframework.com/jobs/af0299f9-a3bb-44ea-8ced... ¯\(ツ)/¯