3 ms·
thanks for posting the question - this space certainly keeps evolving I've been trying to tackle various aspects of traceability and validation at scale... emb
by FacelessAICoder 2mo ago
thanks for posting the question - this space certainly keeps evolving
I've been trying to tackle various aspects of traceability and validation at scale... embedding the self improvement & validation in various iterations of Ralph Loops.... here's a deeper write up: https://dataspheres.ai/pages/dataspheres-ai/spec-driven-development https://dataspheres.ai/pages/dataspheres-ai/spec-driven-deve...
The tools & methodologies keep changing everyday - so I mostly just iterate on my own tooling https://github.com/geekdreamzz/ari-dai-skills https://github.com/geekdreamzz/ari-dai-skills so I can keep learning and iterating.... it's on my todo list to create some videos demoing it...
At a high-level - I find state management is really hard at scale in addition to managing all the context for specific situations. What I do is create a graph that tracks your original.... prompts <> specs <> tasks <> ai-generated code/content.... so it's never a question on why/what since it's all instrumented into a chain. The tasks have clear validation criteria and gates between statuses to adversarially challenge the status updates and validation. Where appropriate I have playwright take screenshots and review it as part of the chain. All of this gets tracked into a dashboard so it can just keep running while you're on the go..
I'm talking high-level and a bit all over the place because there is a lot of moving parts in this space. Would love to connect more on how you and others are tackling this. I'm using Claude Code to run the ari-dai-skills repo I linked above for all my initiatives now. I'm spending a lot of $$$ on Claude and more & more I'm itching to just invest in local compute and run more local LLM tasks autonomously because we're all iterating so much. I've been able to build a lot though but I think the open source models and the local inference is getting better.... so a new wave of tooling is on the horizon to make traceability and validation better so keep hacking ... best of luck... these harnesses are like a casino & sometimes we do win big ...I think we're a long way out from real consistency since there's just so many perspectives a harness facilitates.