3 ms·
Hi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility
by yohji1984 3mo ago
Hi HN, I've been working on an agentic runtime framework. The earlier benchmark eval is promising. But I can see the limits of the test design and the fragility of the runtime itself. I would like to ask for your reviews of the framework and the eval process itself.
https://github.com/Tura-AI/tura https://github.com/Tura-AI/tura
https://turaai.net/docs#benchmark-current-test-set-record https://turaai.net/docs#benchmark-current-test-set-record
Tura-AI/tura https://turaai.net/blog#why-i-am-building-tura https://turaai.net/blog#why-i-am-building-tura