6 ms·
The trick they announce for Grok Heavy is running multiple agents in parallel and then having them compare results at the end, with impressive benchmarks across
by tibbar 1y ago
The trick they announce for Grok Heavy is running multiple agents in parallel and then having them compare results at the end, with impressive benchmarks across the board. This is a neat idea! Expensive and slow, but it tracks as a logical step. Should work for general agent design, too. I'm genuinely looking forward to trying this out.
EDIT: They're announcing big jumps in a lot of benchmarks. TIL they have an API one could use to check this out, but it seems like xAI really has something here.
- deleted 1y ago[deleted]
- sidibe 1y agoYou are making the mistake of taking one of Elon's presentations at face value.
- simianwords 1y agothat's how o3 pro also works IMO
- tibbar 1y agoInteresting. I'd guess this technique should probably work with any SOTA model in an agentic tool loop. Fun!
- zone411 1y agoThis is the speculation, but then it wouldn't have to take much longer to answer than o3.
- bobjordan 1y agoI can’t help but call out that o1-pro was great, it rarely took more than five minutes and I was almost never dissatisfied with the results per the wait. I happily paid for o1-pro the entire time it was available. Now, o3-pro is a relative disaster, often taking over 20 minutes just to refuse to follow directions and gaslight people about files being available for download that don’t exist, or provide simplified answers after waiting 20 minutes. It’s worse than useless when it actively wastes users time. I don’t see myself ever trusting OpenAI again after this “pro” subscription fiasco. To go from a great model to then just take it away and force an objectively terrible replacement, is definitely going the wrong way, when everyone else is improving (Gemini 2.5, Claude code with opus, etc). I can’t believe meta would pay a premium to poach the OpenAI people responsible for this severe regression.
- sothatsit 1y agoI have never had o3-pro take longer than 6-8 minutes. How are you getting it to think for 20 minutes?! My results using it have also been great, but I never used o1-pro so I don't have that as a reference point.
- irthomasthomas 1y agoLike llm-consortium? But without the model diversity. https://x.com/karpathy/status/1870692546969735361 https://x.com/karpathy/status/1870692546969735361 https://github.com/irthomasthomas/llm-consortium https://github.com/irthomasthomas/llm-consortium
- sailingparrot 1y ago> Expensive and slow Yes, but... in order to train your next SotA model you have to do this anyway and do rejection sampling to generate good synthetic data. So if you can do it in prod for users paying 300$/month, it's a pretty good deal.
- daniel_iversen 1y agoVery clever, thanks for mentioning this!
- icoder 1y agoI can understand how/that this works, but it still feels like a 'hack' to me. It still feels like the LLM's themselves are plateauing but the applications get better by running the LLM's deeper, longer, wider (and by adding 'non ai' tooling/logic at the edges). But maybe that's simply the solution, like the solution to original neural nets was (perhaps too simply put) to wait for exponentially better/faster hardware.
- cfn 1y agoMaybe this is the dawn of the multicore era for LLMs.
- the8472 1y agogrug think man-think also plateau, but get better with tool and more tribework Pointy sticks and ASML's EUV machines were designed by roughly the same lumps of compute-fat :)
- deleted 1y ago[deleted]
- SauciestGNU 1y agoThis is an interesting point. If this ends up working well after being optimized for scale it could become the dominant architecture. If not it could become another dead leaf node in the evolutionary tree of AI.
- deleted 1y ago[deleted]
- simondotau 1y agoYou could argue that many aspects of human cognition are "hacks" too.
- emp17344 1y ago…like what? I thought the consensus was that humans exhibit truly general intelligence. If LLMs require access to very specific tools to solve certain classes of problems, then it’s not clear that they can evolve into a form of general intelligence.
- JKCalhoun 1y ago> I'm genuinely looking forward to trying this out. Myself, I'm looking forward to trying it out when companies with less, um, baggage implement the same. (I have principles I try to maintain.)
- einrealist 1y agoSo the progress is basically to brute force even more? We got from "single prompt, single output", to reasoning (simple brute-forcing) and now to multiple parallel instances of reasoning (distributed brute-forcing)? No wonder the prices are increasing and capacity is more limited. Impressive. /s
- deleted 1y ago[deleted]
- nisegami 1y agoI've suspected that technique could work on mitigating hallucinations, where other agents could call bullshit on a made up source.