2 ms·
Yeah, probably pretty simple compared to the methods we've publicly discussed for months before this publication. Here's the last time we showed our demo on HN
by backflippinbozo 1y ago
Yeah, probably pretty simple compared to the methods we've publicly discussed for months before this publication.
Here's the last time we showed our demo on HN: https://news.ycombinator.com/item?id=45132898 https://news.ycombinator.com/item?id=45132898
We'll actually be presenting on this tomorrow at 9am PST https://calendar.app.google/3soCpuHupRr96UaF8 https://calendar.app.google/3soCpuHupRr96UaF8
Besides ReAct, we use AG2's 2-agent pattern with Code Writer and Code Executor in the DockerCommandLineCodeExecutor
Also, using hardware monitors and LLM-as-a-Judge to assess task completion.
It's how we've built nearly 1K Docker images for arXiv papers over the last couple months: https://hub.docker.com/u/remyxai https://hub.docker.com/u/remyxai
And how we'll support computational reproducibility by linking Docker images to the arXiv paper publications: https://github.com/arXiv/arxiv-browse/pull/908 https://github.com/arXiv/arxiv-browse/pull/908