5 ms·
I suspect this was released by Anthropic as a DDOS attack on other AI companies. I prompted 'how do we solve this challenge?' into gemini cli in a cloned repo a
by pvalue005 9mo ago
I suspect this was released by Anthropic as a DDOS attack on other AI companies. I prompted 'how do we solve this challenge?' into gemini cli in a cloned repo and it's been running non-stop for 20 minutes :)
- bird0861 9mo agoWhich Gemini model did you use? My experience since launch of G3Pro has been that it absolutely sucks dog crap through a coffee straw.
- Mashimo 9mo ago> sucks dog crap through a coffee straw. That would be impressive.
- anematode 9mo agoNew LLM benchmark incoming? I bet once it's done, people will still say it's not AGI.
- dotancohen 9mo agoWhen they get the hardware capable of that, a different industry will be threatened by AI. The oldest industry.
- cess11 9mo agoTextile?
- nineteen999 9mo agoThe emperor's (empresses?) new textile.
- darepublic 9mo agoSong of Solomon I guess
- stronglikedan 9mo agoOnly if the dog didn't get too much human food the night before.
- pvalue005 9mo ago/model: Auto (Gemini 3) Let Gemini CLI decide the best model for the task: gemini-3-pro, gemini-3-flash After ~40 minutes, it got to: The final result is 2799 cycles, a 52x speedup over the baseline. I successfully implemented Register Residency, Loop Unrolling, and optimized Index Updates to achieve this, passing all correctness and baseline speedup tests. While I didn't beat the Opus benchmarks due to the complexity of Broadcast Optimization hazards, the performance gain is substantial. It's impressive as I definitely won't be able to do what it did. I don't know most of the optimization techniques it listed there. I think it's over. I can't compete with coding agents now. Fortunately I've saved enough to buy some 10 acre farm in Oregon and start learning to grow some veggies and raise chickens.
- apsurd 9mo agowe've lost the plot. you can't compete with an AI on doing an AI performance benchmark?
- kqr 9mo agoThis is not an AI performance benchmark, this is an actual exercise given to potential human employees during a recruitment process.
- IsTom 9mo agoDid you check that it did the things it claims it did?
- light_hue_1 9mo agoKeep in mind that the boat on competing with machines to generate assembly sailed for 99% of programmers half a century ago. It is not surprising that this is an area where AI is strong.
- triyambakam 9mo ago> grow some veggies and raise chickens. Maybe Claude will be able to do that soon, too.
- ece 9mo ago
- bird0861 9mo agoHilarious that this got a downvote, hello Satya!
- bjackman 9mo agoLately with Gemini CLI / Jules it doesn't seem like time spent is a good proxy for difficulty. It has a big problem with getting into loops of "I am preparing the response for the user. I am done. I will output the answer. I am confident. Etc etc". I see this directly in Gemini CLI as the harness detects loops and bails the reasoning. But I've also just occasionally seen it take 15m+ to do trivial stuff and I suspect that's a symptom of a similar issue.
- sva_ 9mo agoI feel like sometimes it just loops those messages when it doesn't actually generate new tokens. But I might be wrong
- bjackman 9mo agoThere are some other failure modes that all feel kinda vaguely related that probably help with building a hypothesis about what's going wrong: Sometimes Gemini tools will just randomly stop and pass the buck back to you. The last thing will be like "I will read the <blah> code to understand <blah>" and then it waits for another prompt. So I just type "continue" and it starts work again. And, sometimes it will spit out the internal CoT directly instead of the text that's actually supposed to be user-visible. So sometimes I'll see a bunch of paragraphs starting with "Wait, " as it works stuff out and then at the end it says "I understand the issue" or whatever, then it waits for a prompt. I type "summarise" and it gives me the bit I actually wanted. It feels like all these things are related and probably have to do with the higher-level orchestration of the product. Like I assume there are a whole bunch of models feeding data back and forth to each other to form the user-visible behaviour, and something is wrong at that level.
- hackpelican 9mo agoAt one point it started spitting out its CoT in the comments of the code it’s supposed to be changing.
- 9mo ago