4 ms·
Even then, "resolving merge conflicts along the way" doesn't mean anything, as there are two trivial merge strategies that are always guaranteed to work ('ours'
by deng 9mo ago
Even then, "resolving merge conflicts along the way" doesn't mean anything, as there are two trivial merge strategies that are always guaranteed to work ('ours' and 'theirs').
- fzzzy 9mo agothat’s not guaranteed to work. Other parts of the CodeBase that didn’t conflict could depend on the discarded code.
- dingnuts 9mo ago[dead]
- formerly_proven 9mo agoWell they did mention the code doesn't work.
- nyeah 9mo agoWhere did Cursor say that?
- logicallee 9mo agoIt's implied by the fact that early in the post they say: >"To test this system, we pointed it at an ambitious goal: building a web browser from scratch." and then near the end, they say: >"Hundreds of agents can work together on a single codebase for weeks, making real progress on ambitious projects." This means they only make progress toward it, but do not "build a web browser from scratch". If you're curious, the State of Utopia (will be available at https://stateofutopia.com https://stateofutopia.com ) did build a web browser from scratch, though it used several packages for the networking portion of it. See my other comments and posts for links.
- deleted 9mo ago[deleted]
- deleted 9mo ago[deleted]
- madeofpalk 9mo agoThe point is that the merge conflict was resolved, regardless of whether there was a working product at the end. Which there apparently isn’t.
- paulus_magnus2 9mo agoHaha. True, CI success was not part of PR accept criteria at any point. If you view the PRs, they bundle multiple fixes together, at least according to the commit messages. The next hurdle will be to guardrail agents so that they only implement one task and don't cheat by modifying the CI piepeline
- formerly_proven 9mo agoIf I had a nickel for every time I've seen a human dev disable/xfail/remove a failing test "because it's wrong" and then proceeding to break production I would have several nickels, which is not much, but does suggest that deleting failing tests, like many behaviors, is not LLM-specific.
- vizzier 9mo ago> but does suggest that deleting failing tests, like many behaviors, is not LLM-specific. True, but it is shocking how often claude suggests just disabling or removing tests.
- ciaranmca 9mo ago100%, trying a bit of an experiment like this(similar in that I mostly just care about playing around with different agents, techniques etc.) it has built out literally hundreds of tests. Dozens of which were almost pointless as it decided to mock apis. When the number of failed tests exceeded 40 it just started disabling tests.
- icedchai 9mo agoTo be fair, many human developers are fond of pointless tests that mock everything to the extent that no real code is actually exercised. At least the tests are fast though.
- falkensmaize 9mo ago
- PunchyHamster 9mo agoSo, AI agent battle royale
- anonzzzies 9mo agoWe use claude code a lot for updating systems to a newer minor/major version. We have our own 'base' framework for clients which is a, by now, very large codebase that does 'everything you can possibly need'; so not only auth, but payments, billing, support tickets, email workflows, email wysiwyg editing, landing page editor, blogging, cms, AI /agent workflows etc etc (across our client base, we collect features that are 'generic' enough and create those in the base). It has many updates for the product lead working on it (a senior using Claude code) but we cannot just update our clients (whose versions are sometimes extremely customised/diverging) at the same pace; some do not want updates outside security, some want them once a year etc. In this case AI has been really a productivity booster; our framework always was quite fast moving before AI too when we had 3.5 FTE (client teams are generally much larger, especially the first years) on it but then merging, that to mean; including the new features and improvements in the client version that are in the new framework version without breaking/removing changes on the client side, was a very painful process taking a lot of time and at at least 2 people for an extended period of time; one from the client team, one from the framework team. With CC it is much less painful: it will merge them (it is not allowed, by hooks, to touch the tests), it will run the client tests and the new framework tests and report the difference. That difference is evaluated usually by someone from the client team who will then merge and fix the tests (mostly manually) to reflect the new reality and test the system manually. Claude misses things (especially if functionalities are very similar but not exactly the same, it cannot really pick which to take so it does nothing usually) but the biggest bulk/work is done quickly and usually without causing issues.
- efreak 9mo ago`git add .; git merge continue` also "solves" the conflict