4 ms·
Has anyone tried to rewrite some popular open source project with IA? I imagine modern LLMs can be very effective at license-washing/plagiarizing dependencies,
by pera 9mo ago
Has anyone tried to rewrite some popular open source project with IA? I imagine modern LLMs can be very effective at license-washing/plagiarizing dependencies, it could be an interesting new benchmark too
- benhoyt 9mo agoNot me personally, but a GitHub user wrote a replacement for Go's regexp library that was "up to 3-3000x+ faster than stdlib": https://github.com/coregx/coregex https://github.com/coregx/coregex ... at first I was impressed, so started testing it and reporting bugs, but as soon as I ran my own benchmarks, it all fell apart (https://github.com/coregx/coregex/issues/29 https://github.com/coregx/coregex/issues/29). After some mostly-bot updates, that issue was closed. But someone else opened a very similar one recently (https://github.com/coregx/coregex/issues/79 https://github.com/coregx/coregex/issues/79) -- same deal, "actually, it's slower than the stdlib in my tests". Basically AI slop with poor tests, poor benchmarks, and way oversold. How he's positioning these projects is the problematic bit, I reckon, not the use of AI. Same user did a similar thing by creating an AWK interpreter written in Go using LLMs: https://github.com/kolkov/uawk https://github.com/kolkov/uawk -- as the creator of (I think?) the only AWK interpreter written in Go (https://github.com/benhoyt/goawk https://github.com/benhoyt/goawk), I was curious. It turns out that if there's only one item in the training data (GoAWK), AI likes to copy and paste freely from the original. But again, it's poorly tested and poorly benchmarked. I just don't see how one can get quality like this, without being realistic about code review, testing, and benchmarking.
- CuriouslyC 9mo agoTo be fair, good benchmarking is hard, most people get it wrong. Scientific training helps.
- dragonwriter 9mo ago> up to 3-3000x+ faster than stdlib Note that this is semantically exactly equivalent to "up to 3000x faster than stdlib" and doesn't actually claim any particular actual speedup since "up to" denotes an upper bound, not a lower bound or expected value. It’s standard misleading-but-not-technically-false marketing language to create a false impression because people tend to focus on the number and ignore the "up to".
- supriyo-biswas 9mo agoReminds me of https://xkcd.com/870/ https://xkcd.com/870/
- arcticbull 9mo agoWith the “up to 3-3000x+” language the plus leaves us with the entire number line.
- Dylan16807 9mo agoWhen you say "up to" about a list of data points, it's not just a bound. At least one has to reach that amount or it's a lie.
- nkrisc 9mo agoSaying “up to” means that bound is the maximum value of the data set. It may be far from the median value, but it is included (or you’re lying). With any other interpretation the phrase has no meaning whatsoever.
- nkrisc 9mo agoI will concede, proactively, that "up to" could refer to some maximum possible bound, even if the current set doesn't include a value at that bound, though I would argue that's likely deceptive wording. For example, you could say that each carton of of eggs on a pallet contains up to 12 eggs, because that's the maximum capacity of the carton, even if none of the actual cartons on this pallet actually have 12 eggs in them.
- DonHopkins 9mo ago3000x Faster Optimized Random Number Generator: https://xkcd.com/221/ https://xkcd.com/221/
- AlexeyBelov 9mo agoOh yeah, I recognize this guy. The author of most commits in coregex posted his vibecoded projects to Reddit. I've looked at his other repos and it's the same shit. Responses are also quite funny, does he not realize this reads like the worst of AI?
- gorkaerana 9mo agoI think it's fair enough to consider porting a subset of rewriting, in which case there are several successful experiments out there: - JustHTML [1], which in practice [2] is a port of html5ever [3] to Python. - justjshtml, which is a port of JustHTML to JavaScript :D [4]. - MiniJinja [5] was recently ported to Go [6]. All three projects have one thing in common: comprehensive test suites which were used to guardrail and guide AI. References: 1. https://github.com/EmilStenstrom/justhtml https://github.com/EmilStenstrom/justhtml 2. https://friendlybit.com/python/writing-justhtml-with-coding-agents/ https://friendlybit.com/python/writing-justhtml-with-coding-... 3. https://github.com/servo/html5ever https://github.com/servo/html5ever 4. https://simonwillison.net/2025/Dec/15/porting-justhtml/ https://simonwillison.net/2025/Dec/15/porting-justhtml/ 5. https://github.com/mitsuhiko/minijinja https://github.com/mitsuhiko/minijinja 6. https://lucumr.pocoo.org/2026/1/14/minijinja-go-port/ https://lucumr.pocoo.org/2026/1/14/minijinja-go-port/
- daxfohl 9mo agoInteresting, IIUC the transformer architecture / attention mechanism were initially designed for use in the language translation domain. Maybe after peeling back a few layers, that's still all they're really doing.
- nathan_compton 9mo agoThis has long been how I have explained LLMs to non-technical people: text transformation engines. To some extent, many common, tedious, activities basically constitute a transformation of text into one well known form from another (even some kinds of reasoning are this) and so LLMs are very useful. But they just transform text between well known forms.
- daxfohl 9mo agoAnd while it appears that lots of problems can be contorted into translation, "if all you have is a hammer, everything looks like a nail". Maybe we do hit a brick wall unless we can come up with a model that more closely aligns with actual human reasoning.
- deleted 9mo ago[deleted]
- hedgehog 9mo agoI used one of the assistants to reverse and rewrite a browser-hosted JS game-like app to desktop Rust. It required a lot of steering but it was pretty useful.