4 ms·
The evaluation framework might be the real moat here. Coding agents work because "tests pass" is an instant signal, but what about domains where success is subj
by sitestable 11mo ago
The evaluation framework might be the real moat here. Coding agents work because "tests pass" is an instant signal, but what about domains where success is subjective or delayed by months?
Medical diagnosis? Legal research? Customer support quality? I wonder if agent labs in these types of domains are struggling to improve as fast, even with great workflow data.