8 ms·
You can't unit test for taste
- dionian 3mo agoGreat article, easy to read, and not ai slop! thanks for sharing
- throw93949444 3mo ago> For example, my native Iceland had a nice mix of nature, historical sites and populated places. You absolutely can unit test for taste, just put an agent into loop, and write into prompt what you like. Then do scoring... Iceland is really bad example, it basically has one populated site (capital) and circular road that goes around the island.
- voidUpdate 3mo agoI'm pretty sure there's more points of interest in the entirety of Iceland than just Reykjavík and Route Number One
- chantepierre 3mo agoIt makes me smile when runners use "X is a marathon, not a sprint" to hint at an effort that accumulates over time and an optimal use of energy. I do it too because it's a common expression, and a marathon is of course longer than a sprint, but both have in common that properly raced, they are absolutely brutal efforts that leave you without a single additional drop at the end. The effort length and instantaneous power output changes, of course. Maybe "it's a marathon build, not the race" would be more precise at the loss of nearly all its expressive power (but with a lot more pedanticism points) :-p . Nice project !
- another-dave 3mo ago"The effort length and instantaneous power output changes, of course." but that's what the phrase is meant to convey, right? Don't run through consumable X (energy/money/etc) like there's no tomorrow - even though there's <some big important milestone> now, we've got dozens more of those that we need to meet, so you're better off getting this one done at 75% than committing 100% to it and failing on all the others.
- boredumb 3mo agoDon't work 12 hour days to get milestone X out, because there are dozens more milestones so don't get burnt on trying to get this one out yesterday. It would probably be more like, don't use 200% to get this out and then quit or burn yourself to 0% or a few % in a year when we want you to extend and maintain this stuff.
- chantepierre 3mo agoYeah you're right, I hear it more like "this is a week long hike, not a sprint" as if a marathon included rest. In any length of racing there's no tomorrow. But I'm doing tongue-in-cheek pedanticness here and will stop that right now !
- dasil003 3mo agoI'd wager that if a manager says that they want you to take it more like a real marathon and less like long hike.
- jayd16 3mo agoIn a marathon, not sprinting is the rest.
- Jensson 3mo agoNo its not a rest, marathon runners are still exhausted at the end, they can't go and run another marathon right after.
- chantepierre 3mo agoThere is no rest. There is just (properly done) a continuous output from start to finish, or a very slight increase of output (negative splitting), but effort to maintain it feels exponential. In terms of feeling, it’s a 32km « dynamic run » where you should feel good, then the hardest 10k you can pull off just after that. If paced properly there should not be a « wall » but at all levels you pass people walking who disintegrated around 30km. Even people with sub-elite/elite bibs sometimes explode. A half is more intense but way easier, you’re just sub threshold but for a time short enough that you cannot really not make it.
- trjordan 3mo agoYou can't unit test for taste if you haven't written down what you mean by taste. If you can externalize it, then you can. Follow this line of thinking, and the AI-friendly answer is easy: we just have to externalize everything we know, so Claude can implement what I want. Except that I can't fully externalize myself. Debugging a system takes more resources than running the system. If I could write down everything I know and hand it to a machine, I'd do that, but it impossible. People aren't books or hashmaps. If you want to build something, you need to use the tools, not teach the tools to use you. [edit: I'm trying to figure out if there's something to be done about this. Email me if you want to chat -- tr at tern dot sh]
- bonzini 3mo agoIt can't be written down as code, that's the point. I am more familiar with taste in coding and it can at best be described—that the resulting code is too subtly different from something else in the codebase, that you're masking a different bug, that you're not following what the code tells you. The good part is that while this cannot be unit tested, you can write documentation and code comments about it that tell people what they need to know. But for taste of the kind described in the article there's not even a definition. The logic ended up being "trust a bunch of opaque weights the most"
- Chris2048 3mo agoTechnically, AI is code, just very complex code. I'd say there are "simple" simple things you can do though, like take automated screenshots and detect colours for jarring colourschemes.
- Chris2048 3mo agoMust have hit some nerves.
- fragmede 3mo agoApple's human interface guidelines says that some things can be written down though. It's a very thurough look at UX and while they don't adhere to them perfectly themselves, it's very much a north star to a some ideals. You can't unit test for taste, but you can integration test that bad tastes haven't happened.
- deleted 3mo ago[deleted]
- draw_down 3mo ago[dead]
- a_c 3mo agoI like to think of testing as making sure things not wrong, but not making it right. Working, useful, delightful, in that order. Testing can make things more likely to work, that's it.
- vshulcz 3mo ago[flagged]
- TimXare 3mo agoTaste is mostly the part of the spec you forgot to write down, plus the part you couldn't write down even if you tried.
- esafak 3mo agoWe can encode taste -- generative AI depends on it. Ask people to compare two examples and pick the one with better taste. You can even ask them to rate multiple subjective criteria at once. Use that to learn a scoring function based on the rating labels, and raw features. Now you can write tests.
- Gosper 3mo agoLanguage count is a decent notoriety signal though pretty coarse. The OP/author should take a look at QRank: https://qrank.toolforge.org/ https://qrank.toolforge.org/ > QRank is a ranking signal for Wikidata entities. It gets computed by aggregating page view statistics for Wikipedia, Wikitravel, Wikibooks, Wikispecies and other Wikimedia projects from https://github.com/brawer/wikidata-qrank/blob/main/doc/design.md https://github.com/brawer/wikidata-qrank/blob/main/doc/desig...
- TestINGNG 3mo ago[dead]
- carra 3mo agoSo now we need a framework for unit tastes
- timroman 3mo agohttps://pureinference.com/insights/taste-is-the-new-skill https://pureinference.com/insights/taste-is-the-new-skill I wrote about this a few months back. Rick Rubin is famous for this. I do think it is something that can be trained though, it just needs a lot more context. Taste builds over time through lots of unit tests, through lots of content writing, through an accumulation of product decisions. It’s hard to put it in the individual spec, but it can be teased out of 100 project specs. And when you get to that scale the AI starts to do it pretty well.
- sesm 3mo ago> Rick Rubin told Anderson Cooper he has no technical ability. Doesn't play instruments. Can't work a mixing board. If you watch his interview on Rick Beato's channel, this myth will fall apart. He plays guitar, had his own punk rock band and his guitar playing is featured on some high-profile records he produced. Also, he has a lot of practical experience with all kinds of studio equipment.
- timroman 3mo agoThat’s exactly it. His taste isn’t in any one thing. It’s the esoteric and accumulated from a variety of things. You can’t package it up. That’s the point on the project specs. I can never get it right in one, but the arc over 100 becomes visible. Especially to an LLM that has the capacity to intake and understand that.
- pixl97 3mo ago>You can’t package it up. Well, you can package it up, otherwise Rick wouldn't exist.
- themgt 3mo agoThis is exactly it - the ultimate skill now is to be Rick Rubin with an LLM. Not a comfortable transition as a coder.
- pjmlp 3mo agoExactly one of the reasons I never went down with all the TDD dogma of only writing code to fix broken tests. There is a reason conference talks are always about plain algorithms and data structures.
- bob1029 3mo agoThe biggest flaw I've seen with TDD is the fact that correctness does not compose upward. Every time two units come into contact, you've got an entirely new kind of unit. The tests from constituents do not cover emergent properties of the new things. You will repeat this same exercise the entire way up to the top, and the moment you come into contact with the customer (they want to change everything), the house of cards comes crumbling down and you have to start your agonizingly-slow process all over from the bottom again. The only thing that the business seems to care about is top-down UI testing. This is also convenient because you can leave it until the very end after the customer has already seen several prototypes. I do think TDD makes sense in isolated scopes (prove this specific custom parser works at the edges), but as the general policy for the entire product it's definitely not a viable practice. Much of the time if comes off as an ego trip to see just how cleverly we can mock something so that we can say we technically tested it.
- pjmlp 3mo agoExactly, the whole system thinking and large scale architecture also fails apart, when writing everything from little working tests.
- bluGill 3mo agoI tell people you should be testing at the level where a change would be so hard you wouldn't do it anyway. Internal helper functions - they are tested only because the code that calls them passes. Interfaces that are used thousands of places - you better test them well because you wouldn't dare change that anyway: it would break too many others. Or to put it differently: a test is an assertion that no matter what, for all time this should never change again. Even if customer requirements change in the future they won't change in such a way as to break this test (this isn't always true, but you should believe it is true). A test is most valuable when it alerts you to a real problem when it fails. If the test fails but there isn't a real problem (either because customer requirements have changed, or it is flaky) it was needless cost to investigate it. If the test passes that gives some hope of correctness, but you can never be sure it is really correct vs a bug in the test (even if you use TDD and so the test failed when you wrote it that doesn't mean a refactoring since didn't make this an always pass test). Part of the problem is if I tell you to write sort() or your new toy language's list type you have an intuitive idea of what it should look like and probably will get them right the first time (other than bugs you want the tests so you catch). These should have tiny micro tests. These things also are really easy to use as examples of how to do TDD - which they are, but they are not representative: this type of code is generally in your standard library already and you are not writing it. Instead you are writing code that isn't well defined with lots of industry experience. It is not clear what the exact interface should be (or more likely it is clear customer requirements will change but you don't know how yet). You have no idea what the best implementation is. You don't know if this will be used in this one place, or if it will become a useful key part that many future projects depend on. You have to make guesses.
- jpadkins 3mo agoI think another important question is can you distill taste? (another comment uses the phrase "externalize", which might mean something similar). I think people have been trying for the written word, with some degree of success (anti-slop skills). I have been trying for visuals, and it's pretty meh. It's easy to get a multimodal LLM to follow a style guide, but a style guide doesn't capture everything that accounts for taste. And anything that is dynamic (not a screenshot test) seems really hard or really expensive.
- ChrisMarshallNY 3mo ago> but it ended up merely in a supporting role This has been my experience, as well, but it’s a really big support. It just needs adult supervision. I can’t understand how vibe-coded apps, actually work. As far as “taste,” goes, I test my stuff constantly, checking for even minor “friction points,” sometimes, refactoring back to design, in order to resolve issues that many folks would ship. I’m pretty anal, and want my work to be the best experience possible. I can’t see any LLM coming close to being able to evaluate the user experience, like I can.
- paytonjjones 3mo agoTools like Playwright and Maestro can already give you a small taste of what that would look like. But overall I agree, LLMs are currently awful at being beta testers. They miss the most basic stuff that any human would immediately catch as being poor UX, and for all their visual prowess they are terrible at auditing UI.
- hombre_fatal 3mo ago> I can’t understand how vibe-coded apps, actually work. With a better process. e.g. plan->revision cycles, better instructions/docs like an ADR system. I don't think vibe-coding is relegated to "build me reddit but with blockchain" and then it's done. I think it instead describes the workflow where the software impl stays opaque but you evaluate the end product as an end user to step the product forward. It basically centers you as the tastemaker. I'd say I vibe-code all of my personal projects now since December where AI had a breakthrough where it required less babysitting and developed good "taste" like smart sum types without being prompted to do so. I've accumulated my own best practices like a heavy plan->revise cycle where plans ultimately promote into ./plans/impl/YYYY-MM-DD-{slug}.md, and an ADR system in ./docs/design/*.md that encodes arch/design invariants that accumulate over time, and new decisions/principles are folding back into it as they are discovered (by the AI). During the plan revision cycles, the LLMs may ask me a multiple choice question about which decision branch to take, and lately I've just been responding with "take the ideal option" with good results -- either way it will take a well-reasoned position that I can't really argue with. Meanwhile, my role is mainly to evaluate the end product and steer it directionally. How much I decide to prescribe and inject myself into technical decisions is a function of how serious the project is, but it's easy to notice that LLMs are simply better and better at arriving at well-reasoned decisions, and my interjections are more and more limited to technical/directional taste rather than necessity.
- tuo-lei 3mo agothe taste part for me is cutting what the agent generated. 200 lines come back, i keep 80, no test for which 80.
- thomasfl 3mo agoThat's what linters are for. Linters can prevent SQL code from spilling out to code outside the model layer. Even more important when vibecoding.
- fotoblur 3mo agoNo but you can add selection as part of your workflow. Governance is something AI agents have allowed me to focus on more and more and this IMHO is where taste lands for me: https://github.com/lramoth/infoPipeline/blob/main/governance.md https://github.com/lramoth/infoPipeline/blob/main/governance...
- deleted 3mo ago[deleted]
- zamalek 3mo agoUnrelated to code, but along the same lines. I've been keeping track of the Reckless Ben case to fuel my unhealthy indignation, and we just had a like-for-like comparison between a human and an LLM. Human: well-scoped argument that does just enough to get the job done with minimal risk. AI: Extremely clever and correct legal argument that almost any lawyer would have said not to file (at least as written). It tries to burn the world and seriously risks pissing off the judge. https://www.youtube.com/watch?v=YRXJnKP6Tu0 https://www.youtube.com/watch?v=YRXJnKP6Tu0
- jdlshore 3mo agoInteresting video, thanks for sharing it.
- layer8 3mo agoYou can’t even unit-test for correct program logic, unless you’re able to enumerate all possible inputs and states within a short time frame.
- bluGill 3mo agoYou can get close enough by testing only the known edge cases. If you need more mathematical proofs can give it but they are much harder.
- ddemian 3mo ago[dead]
- jt2190 3mo ago> Overall the evaluation of success was one of the most challenging parts of the project. As a developer, I’m used to building features that either work or don’t and there is often an objective way to measure how well a feature performs. For messy real world data it was hard to evaluate how good or bad the pipeline was. Furthermore, it was easy to start optimising for a specific parameter or route and find later that this work led to severe degradations in other areas. > Verification becomes hard to reason about because there is no ground truth for points of interest, there are no red/green unit tests for taste. I’m sure these are familiar challenges to data scientists and that there are frameworks and evals for working on them. This will require more iteration and manual overrides. Hopefully with feedback and collaboration from the community. But for now I’ve shipped V1… I suspect LLMs may be able to help us quantify our taste because they can keep track of so many data points all at once, where we have to lossily abstract these details away.
- HoldOnAMinute 3mo agoI am quite confident I could take a series of photos of various designs and classify them as "tacky" or not, and train a neural network to recognize tackiness.
- m0nacle 3mo agothe whole “taste” thing is so trite lmao
- yiyingzhang 3mo agoIsn't this true since the beginning of software development? AI hasn't changed that yet
- brap 3mo agoThe thing I struggle with wrt to taste is that LLMs just don’t get it. Even if I write down every single thing it did wrong and how I’d do it, and even if I turn those into rules, it will know how to follow these specific rules, but for some reason it can’t seem to generalize beyond that. And the real list of rules seem truly infinite.
- deleted 3mo ago[deleted]
- dirkc 3mo ago> So with my friend Claude I set about building After this line all the references becomes *we*. I can't help but be a little disturbed by that > To begin with we downloaded ... For instance we excluded ... We also selected ... We used this as a notoriety ... <and many more> I am increasingly concerned about how LLMs are anthropomorphizing and how that affects our judgement?
- kalli 3mo agoThis was a conscious choice, addressed in a footnote in the blog post: > This is my first time writing up a project that I worked on using an AI agent. I kept writing “we” because the project felt like a collaboration.[...] On reading it back, saying we feels like an accountability dodge, because of course I’m fully and solely responsible for any errors in this write-up or code. But just using I/me also feels dishonest, because so much of the implementation here isn’t fully mine so I feel like I’m taking too much credit for my collaboration with the machines. I figure this is a new kind of pronouns debate we’ll be having for the foreseeable future. I think it is an interesting topic.
- dirkc 3mo agoThanks for pointing out the footnote, I did not get that far. And like you say, I agree it's interesting. The footnote however does re-enforce my concern - in what other ways do we alter our behavior when it feels like we're interacting with another human?
- kalli 3mo agoThat's fair, we can disagree. I don't think I'm personally anthropomorphising llms (I think my mental model of how they work is rough but fairly accurate), but at a population level it might be something to be concerned about (see all the ai-psychosis talk) What I was getting at with the "we" in the post is more how we talk and think about work like this. I think it is different in kind to previous projects I've done where a relied on google, stack overflow and elbow grease. Programming has always been "standing on the shoulders of giants" kind of work, but doing it with agents feels different from that. Maybe it was a poor stylistic choice, but I think we need a way to talk about it in an honest way.
- ahmedehab_01 3mo agoI think, when using LLMs, learning to accept some mediocrity sometimes is a necessity. It will never even have "acceptable" taste, let alone yours.
- gafferongames 3mo agoNonsense. I unit taste all the time, it's called taste driven develoment (TDD) for a reason.
- sgarland 3mo agoI’m confused about the choice for Parquet and DuckDB here. PostGIS is arguably a better match for what this project is doing, and would let you skip most of if not all uses of Shapely and Pyproj.
- qsoomro 3mo ago[flagged]
- cadamsdotcom 3mo agoYou can’t unit test for all the aspects that make up taste, it’s true. But if you break off parts of that - eg. by looking at what is codified out there as “good” design, what’s considered best practice etc - you can create tools the agent can call on that let it get critiques of its own work. What’s really cool about this is those tools can be code, written by agents and committed to your repo. Put together a script that for example makes sure your brand colors are enforced (eg. https://github.com/cadamsdotcom/CodeLeash/blob/main/scripts/check_brand_colors.py https://github.com/cadamsdotcom/CodeLeash/blob/main/scripts/...) and then put it in your pre-commit checks (https://github.com/cadamsdotcom/CodeLeash/blob/main/.pre-commit-config.yaml#L40 https://github.com/cadamsdotcom/CodeLeash/blob/main/.pre-com...), and the agent will get feedback on its use of tasteless defaults and adjust accordingly (partly because you blocked commits that contain said tasteless defaults!)
- grepzero 3mo ago[flagged]
- kimjune01 3mo agothis is known as the oracle problem
- zdmgg 3mo agoYou might want to consider looking at Wikipedia's internal article quality assessments (these include Featured Articles, Good Articles, B-class, C-class, Start, and Stub). I use these as a rule of thumb for how popular the topic is, it's a solid proxy for both significance and the richness of the available content.
- GreenJacketBoy 3mo agoI may be the only one feeling this way, but the repetitive mention of Claude – worded as if it was a coworker ("we", "me and my friend") to the point that somebody reading it just 3 years ago would reasonably assume this "Claude" was in fact a human – made it hard to read. How much am I reading a behind the scenes of the "making of" of the application VS an essay on what somebody else (Claude) did ? I don't know. The reason I browse this website is to see what other humans are saying, inventing, using. But in some cases like this one, I see the line between tool and co-author being blurred for LLMs. And unless what they did is a specifically impressive thing on its own, I do not want to know what an LLM did. (Don't get me wrong, I would much rather have this than people lying, but I would also much rather people treat LLMs as tools.)
- farid_mth 3mo ago[flagged]
- amelialucas 3mo ago[flagged]