3 ms·
Hi folks, author here and also author of the seminal OpenAI blog on this topic. Let me know how I can help you all let it rip.
by lopopolo 3mo ago
Hi folks, author here and also author of the seminal OpenAI blog on this topic. Let me know how I can help you all let it rip.
- jpitz 3mo agoThis is interesting. I'd already had a conversation with my harness ( pi ) about incorporating continuous improvement. This is a great deal better than what I came up with.
- lopopolo 3mo agoGlad to hear it! Good luck, have fun. The agents tend to do a pretty good job incorporating these ideas. This was an unexpected thing we learned when publishing the initial harness engineering post.
- mips_avatar 3mo agoOne challenge/opportunity I've had is harnessing really wide running cheap agents. Any thoughts on how to move really cheap agents beyond basic summarization so we can go broader than the pricing of frontier llms allows?
- lopopolo 3mo agoYour “really cheap” agents can’t be so cheap that they do not have good tool calling skills. But! Using bigger models to put guardrails in place as static verifiers allows lower complexity changes to “self steer” as tests fail, which means coming down on the cost curve is more effective.
- mcapodici 3mo agoI was experimenting and found deepseek-v4-flash and found it cheap (way cheap compared to sota models) and perfectly good at tool calling. I did a post on it https://martincapodici.com/2026/07/18/weekly-ai-learnings-3/ https://martincapodici.com/2026/07/18/weekly-ai-learnings-3/ So $1.5 for 40m tokens I guess would cost much more with sota (but would need less tokens perhaps).
- hankbond 3mo agoHow do you view harness engineering as an organic development that emerges from its use within a specific domain? Basically the meta-loop that allows an agent to tailor its harness to improve outcomes based on performance feedback. I use Pi a lot and I'm very interested in "self-assembling software". One concrete example might be maintaining a conventions document per-project that covers how to name things semantically from a list of nouns and verbs. The idea is that LLMs are often not very globally aware, but it's important to maintain coherence across a code base in order for it to scale (in size and over time). Sometimes an LLM might call the same concept a Materialization, sometimes a Projection, and its not useful if its using two terms interchangeably without purpose. Basically, how are you maintaining coherence when there isn't a human steering the code beyond providing requirements and validation directives? I see you have relevant context in the repo like https://github.com/lopopolo/harness-engineering/tree/trunk/docs/durable-systems/ https://github.com/lopopolo/harness-engineering/tree/trunk/d... but I'm curious what exists beyond context. Do you use any tooling to steer this type of thing more consistently?
- lopopolo 3mo agoYour example is super amenable to vibing some tests. As an example, I’ve been able to ban `number` from representing a duration by walking the AST in a linter to fail if var or param names that look like the end in millis or ms or sec appear. This is largely good enough. If you see that “drifting” behavior appear more than once, you have enough to stop and force the agent to write some static verifiers that reject all but the option you want. For a closer example, we did this with zod schemes and their corresponding inferred types to be universally ZPascalCase and PascalCase instead of camelCaseSchema and CamelCase
- lopopolo 3mo agoAnd to address your broader question, yes this is a form of RSI and to me a vastly superior approach to fine tuning since it allows adopting new model releases without throwing anything away while still having the same effect on improving adherence to local acceptance criteria.
- 3mo ago
- slopinthebag 3mo agoWhat motivated you to quote your own quote (??) in your readme claiming to boost productivity by 100x, and where did you derive that number from? "When a quote sounds profound enough, reality usually nods out of politeness, without echoes is just a sentence wearing pajamas." - slopinthebag
- hahahaa 3mo agoThanks. I am guessing you have to try stuff and build tacit experience. No other way, just get stuck in and try stuff, then try and learn bits from others?
- lopopolo 3mo agoBasically yea. It is the only way to learn how to outrun your priors on what “high ambition” looks like. The labor that goes into implementation is an uncapped resource now.
- lopopolo 3mo agoThe models are very good now so the feedback cycle on these meta adjustments is much tighter. Yesterday I was able to one shot a Liquid Glass, HIG-compliant and localized DICOM image viewer (frame by frame and looping video) with Apple Intelligence for de-jargoning the series details. Took 30 minutes. But the app had 60% CPU because it was not caching the decoded JPEGs. I can do a point in time fix for that of course, but the more interesting thing is why that misaligned code was permitted to be generated in a “done” artifact in the first place. What other misaligned code from a perf perspective might there be? And how do I intervene into the system that produced this software to make these misalignments statically not meet acceptance criteria?
- ydoc5212 3mo ago> the more interesting thing is why that misaligned code was permitted to be generated in a “done” artifact in the first place It's refreshing to me to hear slop being challenged. From first principles, why ought smart models put out slop, as opposed to self-consistent content?
- zwaps 3mo agoYou call your own blog post seminal and quote yourself in the repo. Are you … alright?
- DenisM 3mo agoIn my little ai bubble the original blog post was very influential. It does sound immodest, but could also be true.
- otekengineering 3mo ago> Let me know how I can help you all let it rip. i'd love to see a way (forum, competition, ?) for people to compare harnesses in different domains. folks like accountants, retail store owners, electrical engineers, and all sorts of niches are building personal harnesses/toolkits around claude/codex. those toolkits fit neither their domain communities nor SWE-heavy harness spaces, and a generic home for small niches could help them flourish. a current problem is nomenclature. i suspect that many people have organically grown toolkits substantial enough to call harnesses, but are not close enough to the SWE bubble to be familiar with the word harness. ai has made it easier than ever to build bridges between domains, but it's also made it easier than ever to get domain tunnel vision and reinvent a super great wheel.