2 ms·
I'm disappointed. After all the buzz and benchmarks, I've tested with my personal benchmark that simulates real-world day-to-day specs for agentic coding, follo
by lrsaturnino 3mo ago
I'm disappointed. After all the buzz and benchmarks, I've tested with my personal benchmark that simulates real-world day-to-day specs for agentic coding, following instructions across long time walls, changing several files and code requirements with separation of concerns to build a complete Saas e2e - it reaches a similar rating as DeepSeek V4 Flash.
- bel8 3mo agoWhat harness? These can make or break benchmarks because of tool call failures/limitations. And is the benchmark open source?
- lrsaturnino 3mo agoPi. The benchmark is local, mostly stuff from my work, I run it everytime a new model comes up. The top model rn is gpt 5.6 sol, followed by fugu ultra, fable, opus 4.8, gpt 5.5 and glm 5.2 (which is the REAL IMPRESSIVE one still). Kimi-k3 is 14th in the list.
- bel8 3mo agoCool, and how is it ranked?