2 ms·
the 8 hours vs 2 weeks framing is the part i'd want more on. generating candidates got cheap, checking them didn't. what does the funnel actually look like ,of
by dhchun1203 2mo ago
the 8 hours vs 2 weeks framing is the part i'd want more on. generating
candidates got cheap, checking them didn't. what does the funnel actually look
like ,of the candidates from an 8 hour run, how many make it to synthesis?
asking because i hit the same shape in a much dumber domain and what got me was that the failures were quiet. nothing errored, output looked normal, it was just
wrong in a way only someone who knew the domain would catch.
- akashramdas 2mo agoYou can see from our the benchmark that only one of the candidates proposed was determined to be worth synthesizing. Each individual candidate generation is quick, 8 hours is required for the model to iterate with various tools to find ones worth submitting. We found that speaking to domain experts was critical in desigining a rubric that could catch these silent synthesis recipe failures, before we attempt the longer 2 week synthesis effort.