5 ms·
This mirrors exactly what I have been doing. - Give Claude/Codex a way to verify its own work (browser, smoke tests, e2e tests, high-fidelity local environment
by shepherdjerred 4mo ago
This mirrors exactly what I have been doing.
- Give Claude/Codex a way to verify its own work (browser, smoke tests, e2e tests, high-fidelity local environment)
- Keep all context (issue tracking, docs, ideas, plans, worklogs) in-repo (https://github.com/shepherdjerred/monorepo/tree/main/packages/docs https://github.com/shepherdjerred/monorepo/tree/main/package...)
- Give Claude/Codex access to observability (Grafana, Prometheus, Tempo, PagerDuty)
- Have Claude/Codex follow good engineering guidelines like fail-fast, type safety, parse at boundaries
I haven't yet been able to achieve full autonomy due to cost and CI load on my homelab.
- para_parolu 4mo agoDoes it yield good results? I found that instead of docs it’s easier just to ask ai to read code. I feel like this is same as comments in code. Become outdated fast
- shepherdjerred 4mo agoI don't really use "docs" for documentation. I've prompted Claude/Codex to always write a "log" and save it in-repo to track what it did and why. I've found this to be really helpful, e.g. "you did this last week, and now some other thing is happening" or "you tried this approach before to solve alert X but it didn't work" -- except it can discover this itself. https://github.com/shepherdjerred/monorepo/tree/main/packages/docs/logs https://github.com/shepherdjerred/monorepo/tree/main/package... I've also used it to store TODOs and plans. For example I might want to explore some idea and defer it for later, or some weekend have it execute on some tech debt I've put off. One last use case is asking "what did I work on in the last 2-3 weeks, is it healthy, and what additional quality checks can/should I do; is there any follow-up work?"
- DenisM 4mo agoI find that preserving logs that contain errors will confuse future sessions even if the errors were corrected at the time. Do you have that problem? Essentially preserving logs extends the context window with all related problems.
- shepherdjerred 4mo agoI haven’t actually noticed that, but I’m not sure why. Maybe because I specifically describe it to the agent as a work log rather than documentation? I’m not sure
- c0rruptbytes 4mo agoit does not result in great results left unattended, it’ll start creating slop or hardcoding solutions but overtime if you adjust your verification rubric, it’s not too bad, gets pretty good, if you do make it do TDD, it gets kinda crazy and you’ll have 2000-3000 tests after awhile, or on my common case, 6000-7000 lines of code in single files (i usually have a cron to audit files for decomposition and create tickets) i wouldn’t use it at my job yet, but it’s been fun to use for personal projects - it’s like modded minecraft automation or factorio
- shepherdjerred 4mo agoStatic analysis can help here! Add CI checks for duplicated code or file length. For test growth, maybe use a coverage tracker and remove redundant tests?
- vividfrier 4mo ago[dead]
- geoffbp 4mo agoI like the idea of saving the work done into files - helps to prevent the llm from redoing the same work. Maybe one day instead of code in a repo it will just be a list of prompts.
- shepherdjerred 4mo agoYes, this was a huge help for me. For example I would have a difficult bug that requires a few sessions/deployments to truly close out. With the worklog, it can easily see "oh I've worked on something similar before"
- ndegwaduncan 4mo ago[dead]