Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
languid-photic
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
languid-photic
5mo ago
also feels like a good posttraining task
2.
▲
by
languid-photic
5mo ago
would be fun to do zig -> rust -> zig and to measure the delta (in a VAE-ish way, kl div on the embeddings?)
3.
▲
by
languid-photic
5mo ago
It’s mostly a bandwidth thing. We’ve seen the pattern consistently, but haven’t had time yet to write up the analysis carefully. We are not the only ones to see the reasoning inversion.: https://arxiv.org/abs/2510.11977
4.
▲
by
languid-photic
5mo ago
My point is more reasoning often leads to worse "scope creep/churn, codebase fit, maintainability".
5.
▲
by
languid-photic
5mo ago
We also use a secondary signal from blinded multi-verifier reviews. Each verifier ranks the candidates, and those verification outcomes serves as an additional quality signal. It's somewhat similar to consensus labeling. Btw, this also
6.
▲
by
languid-photic
5mo ago
Agreed. Harness is really important. Especially since many labs are now post-training agents directly in their native harness. (Which is why my prior is that third party harnesses would not perform as well. But I haven't actually measu
7.
▲
by
languid-photic
5mo ago
Yes, the signal we are measuring is quite different from most evals. We are measuring sth much closer to: when multiple agents compete on the same spec, which one produces the patch that holds up best in code review? Most evals are static &
8.
▲
by
languid-photic
5mo ago
I agree! So far we have been native harnessmaxxing, which simplifies things a lot. The configuration space around open models is much larger. Eg which models, capability heterogeneity, which harness, networking, data egress / privacy,
9.
▲
by
languid-photic
5mo ago
Yes! It depends on the extent of changes needed. If the changes needed are small, I'll apply the best implementation as a foundation and then just iterate directly. If the changes needed are drastic, it usually signals that there was s
10.
▲
by
languid-photic
5mo ago
We track performance vs. the all-in cost of completing real engineering tasks, rather than cost per token. [1] Cost per token is a bit misleading because, as others have noted, different models use tokens in different ways. (Aside - This is
11.
▲
by
languid-photic
5mo ago
Agreed. As alignment improves, I'm becoming increasingly bearish on sandboxing. Version control and isolation will probably stay useful, though, more for distributed development and workflow reasons than for safety.
12.
▲
by
languid-photic
5mo ago
It’s very hard to encode the properties that matter most in code in tests. [1] [1] https://voratiq.com/blog/your-workflow-is-the-eval
13.
▲
by
languid-photic
5mo ago
I think that's Pro. Regular 5.5 is 2x regular 5.4.
14.
▲
by
languid-photic
5mo ago
It's 2x/token, but for default reasoning we've found GPT-5.5 uses fewer tokens overall, so net cheaper on median. [1] (Note, that stops being true at higher reasoning levels, where our observed total cost goes up ~2-3x.) [1]
15.
▲
by
languid-photic
6mo ago
Appreciate the clarification. But, it's still not great. To the PM behind this - developers are sensitive to this kind of thing. Just make it opt-in instead?
16.
▲
by
languid-photic
7mo ago
it's reasonable to note that w/o sharing the data these findings can't be audited or built upon but i think the prior on 'this team fabricated these findings' is v low
17.
▲
by
languid-photic
7mo ago
makes sense! we wrote something yesterday about the weaknesses of test-based evals like swe-bench [1] they are definitely useful but they miss the things that are hard to encode in tests, like spec/intent alignment, scope creep, adhere
18.
▲
by
languid-photic
7mo ago
Yes! The spec is the source of truth. Voratiq converts it into a canonical prompt so every agent gets identical instructions.
19.
▲
Test Evals Are Not Enough
(voratiq.com)
3 points
by
languid-photic
7mo ago
|
2 comments
20.
▲
by
languid-photic
8mo ago
Good point. This post measures `1x top-N` (one attempt each from N models), not `Nx top-1` (N attempts from the best-scoring model). We should make that more clear. Part of why we chose `1x top-N` is that we expect lower error correlation c
21.
▲
by
languid-photic
8mo ago
It was not, the agent id is not overt but can be found via the workspace filepath. But that is a good point. Perhaps it should be mapped to something unidentifiable.
22.
▲
by
languid-photic
8mo ago
https://github.com/voratiq/voratiq For comparison, there's a `review` command that launches a sandboxed agent to review a given run and rank the various implementations. We usually run 1–3 review agents, pull the
23.
▲
by
languid-photic
8mo ago
Yes, understandable. The question is which multi-agent architecture, hierarchical or competitive, yields the best results under some task/time/cost constraints. In general, our sense is that competitive is better when you want bre
24.
▲
by
languid-photic
8mo ago
Yes indeed, you get a big lift out of running just the few top agents. We run big ensembles because we are doing a lot of analysis over the system etc
25.
▲
by
languid-photic
8mo ago
This was exactly the kernel of the idea :)
26.
▲
by
languid-photic
8mo ago
This is a good point! We still code via interactive sessions with single agents when the stakes are lower (simple things, one off scripts, etc). But for more important stuff, we generally want the highest quality solution possible. We also
27.
▲
Selection rather than prediction
(voratiq.com)
43 points
by
languid-photic
8mo ago
|
20 comments
28.
▲
by
languid-photic
8mo ago
They build Claude Code fully with Claude Code.
29.
▲
by
languid-photic
9mo ago
Fun! Any plan to open source it?
30.
▲
by
languid-photic
9mo ago
Sure! https://github.com/voratiq/voratiq
More ›