2 ms·
Here is a system I developed for my own projects: 1) hand-write a simple 2d TUI-based rogue-like in Rust using pretty much just the std; 2) grab opus 5.0 (it u
by drpython 1mo ago
Here is a system I developed for my own projects:
1) hand-write a simple 2d TUI-based rogue-like in Rust using pretty much just the std;
2) grab opus 5.0 (it used to be opus 4.6, 4.7) and give it some vague "requests", and ask it to make this game "production-ready" and "blockbuster", but keep the 2d and TUI aspects so I can actually run it.
3) now the fun part, take a test subject, say GLM 5.3, and ask it to find code smell, architecture issues, duplication and all sort, and *simplify the code*
compare the result to my original version.
It's not a simple thing, but the concept is simple: can an LLM remove all the mud?
The winners so far are (ranked by the quality of the final result, not by token cost)
GPT 5.6 sol (extra high thinking);
GLM 5.3;
Grok 4.6;
Qwan 3.8;
(fable could not make it to the list because it simply cannot follow the instructions)