6 ms·
Simulacrum of Knowledge Work
- balamatom 5mo ago>We've automated ourselves into Goodhart's law. Yes. This does not however mean that progress is not being made. It just means the progress is happening along such dimensions that are completely illegible in terms of the culture of the early XXI century Internet, which is to say in terms of the values of the society which produced it.
- firefoxd 5mo agoEverybody's output is someone else's input. When you generate quantity by using an LLM, the other person uses an LLM to parse it and generate their own output from their input. When the very last consumer of the product complains, no one can figure out which part went wrong.
- balamatom 5mo agoWell the last consumer is holding it wrong of course. Why? The last consumer is present, and everyone else is behind 7 proxies.
- mrtesthah 5mo ago>"is the RLHF judge happy with the answer." Reinforcement Learning with Verifiable Rewards (RLVR) to improve math and coding success rates seems like an exception.
- rowanG077 5mo agoI don't really agree with the premise of the article. Sure proxy measures are everywhere. But for knowledge work specifically you can usually check real quality. Of course it's not as extremely easy as "oh this report contains a few spelling errors", but it is doable. If you accepted work purely based on superficial proxy measures you were not fairly evaluating work at all.
- zingar 5mo agoI think there’s a weaker claim that holds true: we were able to ignore lots of content based on the superficial (and pay proper attention to work that passed this test) and now we are overwhelmed because everything meets the superficial criteria and we can’t pay proper attention to all of it.
- thehappyfellow 5mo agoThat's what I had in mind! The whole post is a claim that evaluating knowledge work got more expensive because cheaper measures stopped correlating well with quality. If someone was already evaluating the work output using a metric closer to the underlying quality then it might not have been a big shift for them (other than having much more work to evaluate).
- Izkata 5mo agoYou may have benefited from using the term we already had for the cheaper measures of negative code quality: code smells.
- zingar 5mo agoI find misapplication of anti-smell techniques a pretty cheap indicator that I’m looking at LLM garbage. I think they’re not really usefully engaging with that stuff yet.
- rowanG077 5mo agoYes, I agree that this is true! You could however only do that if you were fine with unfairly judging the quality of work, as you now readily discarded quality work based on superficial proxies. Which admittedly is done in a lot of cases.
- bensyverson 5mo agoThe article asserts that the quality of human knowledge work was easier to judge based on proxy measures such as typos and errors, and that the lack of such "tells" in AI poses a problem. I don't know if I agree with either assertion… I've seen plenty of human-generated knowledge work that was factually correct, well-formatted, and extremely low quality on a conceptual level. And AI signatures are now easy for people to recognize. In fact, these turns of phrase aren't just recognizable—they're unmistakable. <-- See what I did there? Having worked with corporate clients for 10 years, I don't view the pre-LLM era as a golden age of high-quality knowledge work. There was a lot of junk that I would also classify as a "working simulacrum of knowledge work."
- downboots 5mo agoYes. I think the main warning here is that it is an added risk. A little glitch here and there until something breaks.
- bambax 5mo agoIt's not that pre-LLM era was a "golden age of quality", far form it. It's that LLMs have removed yet another tell-tale of rushed bullshit jobs.
- bensyverson 5mo agoHave they though?
- happytoexplain 5mo agoAbsolutely. Our heuristics for judging human output are useless with LLMs. We can either trust it blindly, or tediously pick over every word (guess which one people do). I've watched this cause havoc over and over at my job (I work with many different teams, one at a time). AI signatures don't mean low quality, they just mean AI. And humans do use them (I have always used the common AI signatures). And yes, humans produce good-looking garbage, but much more commonly they produce bad-looking garbage. This is all tangential to the point.
- zby 5mo agoIf you have a test that fails 50% times - is that test valuable or not? A 50% failure rate alone looks like a coin toss, but by itself that does not tell us whether the test is noise or whether it is separating bad states from good ones. For a test to be useful it needs to have positive Youden’s statistic (https://en.wikipedia.org/wiki/Youden%27s_J_statistic https://en.wikipedia.org/wiki/Youden%27s_J_statistic): sensitivity + specificity - 1. A 50% failure rate alone does not let us calculate sensitivity and specificity. I can see a similar problem with this article - the author notices that LLMs produce a lot of errors - then concludes that they are useless and produce only simulacrum of work. The author has an interesting observation about how llms disrupt the way we judge knowledge work. But when he concludes that llms do only simulacrum of work - this is where his arguments fail.
- card_zero 5mo agoGee, a thing by a guy, with a name. What are you saying exactly? So the test in question is a test the LLM is asked to carry out, right? Then your point is that if it's a load of vacuous flannel 49% of the time, but meaningful 51% of the time, on average this is genuine work so we can't complain about the 49%? Wait, you're probably talking about the test of discarding a report based on something superficial like spelling errors. Which fails with LLMs due to their basic conman personalities and smooth talking. And therefore ..?
- jszymborski 5mo ago> For a test to be useful it needs to have positive Youden’s statistic This is not true as stated. I'd try to gloss over the absolutes relative to the context, but if I'm totally honest, I'm not sure I understand what idea you're trying to communicate.
- jdw64 5mo ago[dead]
- simianwords 5mo agoThe FUD about LLM's will never get old. The way I know and trust LLM's is the same way a manager would trust their reportees to do good work. For most tasks, the complexity/time required to verify a task is << the time required to do the task itself. Sure there can be hallucinations on the graph that the LLM made. But LLMs are hallucinating much less than before. And the time to verify is much lower than the time required for a human to do the task. I wrote a post detailing this argument https://simianwords.bearblog.dev/the-generation-vs-verification-delta-explains-why-llms-are-useful/ https://simianwords.bearblog.dev/the-generation-vs-verificat...
- JackSlateur 5mo agoFUD ? You are missing the point entierly, and so does your blog post Are LLM a good dictionary of synonyms ? Perhaps, but is it relevant ? Not at all Are you biased when a solution is presented to you ? Yes, like all humans. Is it damageful when said solution is brain-dead ? Obsiously. Are you failing to understand that most (if not all) manager's work is human centric and, as such, cannot be applied to a non-human ? Obviously .. You trust a machine's intent. Joke's on you, it has no intent at all, it will breaking that "trust" your pour in it without even realizing-it You say that LLM does better job than you. Perhaps this says it all ?
- doggers246 5mo ago[dead]
- simianwords 5mo agoAre you asking yourself questions and answering them without seeing my point? Yes
- deleted 5mo ago[deleted]
- wxw 5mo agoUltimately to understand a thing is to do the thing. And to not understand (which is ok!) is to trust others to, proxy measures or not. Agreed that the future of work is in a precarious place: doing less and trusting more only works up to a point. `simulacrum` is a great word, gotta add that to my vocabulary.
- slickytail 5mo agoThe idea of Simulacrum comes from Baudrillard. His essay "Simulation and Simulacra" is highly recommended for understanding what is so strange about the modern economy.
- mplanchard 5mo agoFun fact: the Wachowskis made this required reading for most of the actors in The Matrix
- wxw 5mo agoThank you, I'll take a look.
- NickNaraghi 5mo agoIt's a funny thing to write, like an article in an old newspaper that aged quickly. I suspect that this will be wildly out of date within 2-3 years.
- krackers 5mo agoI think it's already out of date with verifiable reward based RL, e.g. on maths domain. When "correctness" arguments fall, the argument will probably just shift to whether it's just "intelligent brute force".
- TheOtherHobbes 5mo ago"stochastic genius"
- gipp 5mo agoThe set of tasks for which "correctness" is formally verifiable (in a way that doesn't put Goodharts Law in hyperdrive) is vanishingly small.
- kj4211cash 5mo agoThere is this belief that in 2-3 years AI will be much better and all the gripes people have with AI use today will be solved. Honestly, personally, I think that optimism will age poorly. But to say it out loud at work or post publicly probably hurts my career prospects.
- dcre 5mo agoIt's already out of date because it makes no sense. If it's true that the superficial signals of quality were once somehow good enough to keep the entire economy on the rails (it's not true), surely you can have an LLM look at given piece of work and extract comparably useful signals of quality or effort.
- deleted 5mo ago[deleted]
- 5mo ago
- deleted 5mo ago[deleted]
- sendes 5mo agoThis is an already apparent problem in academia, though not for the reasons the article suggests. It is not so much that the "tells" of a poor quality work are vanishing, but that even careful scrutiny of a work done with AI is going to become too costly to be done only by humans. One only has so much time to read while, say, in economics journals, the appendices extend to hundreds of pages. Would love to hear if other fields' journals are experiencing a similar pressure in not only at the extensive margin (no of new submission) but the intensive margin (effort needed to check each work).
- Daishiman 5mo agoTo be fair, a lot of academic fields are such that anything at a Master's level or above requires serious competence to judge and for anyone below there's no distinction between what's right and what looks right.
- tkiolp4 5mo agoI think this is pretty obvious for many of us in the industry. Unfortunately, there is so much money on the table that the big players will shove whatever they want down our throats
- happytoexplain 5mo ago"They sound very confident," was a warning a gave a lot on a project a year ago, before I gave up trying to get developers to stop blindly trusting the output and submitting things that were just wrong. The documentation of that team went to absolute shit because the developers thought LLMs magically knew everything.
- larrytheworm 5mo ago[dead]
- throwaway_sydn 5mo ago“/reliable-resources-skill Claude, using the list of approved resources, evaluate the report I’m attaching”
- vivid242 5mo agoWith AI, we‘re cargo-culting understanding. We‘re reproducing the surface of having understood something, but we‘re robbing ourselves the time and effort to truly do it.
- hellohello2 5mo agoAI can do things on its own, without you understanding them yes. But if you are trying to understand something well, there is no better tool for helping you than AI.
- bluefirebrand 5mo ago> But if you are trying to understand something well, there is no better tool for helping you than AI Could not disagree more. The best way to understand something deeply is to practice it. AI is anti-practice. It's like trying to learn something by following a YouTube video step by step. It has an outcome and it feels productive but it's not going to stick in your head at all. It's not practice
- matrix87 5mo agoyou can use AI to get a faster explanation for what's happening in a big codebase, it makes the timelines on developing features much lower from my experience am I losing out on something by not having to spend hours clicking through redundant parts of a large codebase to get a concrete answer on something? doesn't feel like it
- aroman 5mo agoI would say a better analogy is using Google… you can use it as a tool to seek information and deepen your understanding. But it requires your brain to be engaged and to be putting that stream of knowledge into practice.
- hellohello2 5mo agoAI is only anti-practice if you are uncreative with it. It is completely different from youtube; Youtube is passive, but AI is interactive. As an example, you can ask AI to create problems for you to practice with, tailored for whichever specific difficulty you are having.
- hellohello2 5mo ago"How do you know the output is good without redoing the work yourself?" Verifying the correctness of solutions is often much easier than finding correct solutions yourself. Examples: Sudoku and most practical problems in just about any field. - "The training doesn't evaluate 'is the answer true' or "is the answer useful.'" Lets pretend RLVF does not exist to give this argument a chance. Then, while the training loop does not validate accuracy directly I guess, the meta-training loop still does. When someone prompts a model, the resulting execution trace shows if the generated answer is correct or not, and this trace is kept for subsequent training runs. The way coding agents are used productively is not: a) generate code with AI and b) run it yourself; its a) ask the AI to do something, including generating the code and running it too, no step b. This naturally creates large training sets of correct and incorrect solutions. - "We spent billions to create systems used to perform a simulacrum of work." Have you even tried using these systems to produce valuable work? How could this possibly be your conclusion after having tried them?
- nlawalker 5mo ago>"We spent billions to create systems used to perform a simulacrum of work." >Have you even tried using these systems to produce valuable work? How could this possibly be your conclusion after having tried them? The operative words there are used to, as opposed only able to. The conclusion isn't derived from using the tools, it's from observing how other people tend to use them.
- bluefirebrand 5mo ago> Verifying the correctness of solutions is often much easier than finding correct solutions yourself In order to verify correctness you need to understand what correctness is in context, which is actually pretty hard to do if you can't actually find correct solutions yourself, or even if you can but haven't bothered to do so
- kj4211cash 5mo ago"Verifying the correctness of solutions is often much easier than finding correct solutions yourself." Honestly, this has not been my experience at all. Defining what a good solution looks like is most of the battle in Operational Research. But, trying to be constructive, maybe we have identified a sort of diving line between areas where AI is more vs. less helpful.
- adampunk 5mo agoWhy is it not more of a scandal that all these anti-AI articles are written, using large language models? Why is that not an embarrassment for everyone who moans and carps and complains about the craft?
- coppsilgold 5mo ago> The training doesn't evaluate "is the answer true" or "is the answer useful." It's either "is the answer likely to appear in the training corpus" or "is the RLHF judge happy with the answer." We are optimising LLMs to produce output which looks like high quality output. It's not quite as dire as this. One of the main reasons why LLM's are getting better over time is that they are used themselves to bootstrap the next generation by sifting through the training set to do 'various things' to it. People often forget that the training corpus contains everything humanity ever produced and anything new humanity will produce will likely come from it as well. Torturing it with current generation models is among the most productive things you can do to improve the next generation systems.
- cyber_kinetist 5mo ago"The simulacrum is never what hides the truth - it is truth that hides the fact that there is none. The simulacrum is true." - Jean Baudrillard Aligned with the theory of Bullshit Jobs - LLMs expose the fact that the white collar work most of us have been doing at this point were actually bullshit. When LLMs "fake" work, it actually hides the reality that there was no meaningful work here in the first place.
- glaslong 5mo agoLayers of reading internal docs to synthesize new docs to turn into slides to aggregate into docs, where a different set of people only partially read or understand what they're seeing at any given mutation cycle...... it's all a farce of earnest but ultimately useless Productivity. The LLMs are just making it more obvious.
- monocasa 5mo agoI think this is why middle managers seemed to be the first acolytes to the church of llm supremacy. It's a weird space in middle management where all of the incentives other than true competency in the role push you to abstract the knowledge work that you're managing, and that abstraction seems to well describable in embedding space.
- rushabh 5mo agoA corollary of this could be that people interested in Serious Work will never use LLMs. Could be the new "tell".
- loa_in_ 5mo agoWhat if subatomic particles are actually whole universes, and their properties are a reflection of... what kind of peoples dominated, conquered their universe, and what kind of automation was left running after them themselves were gone. Some kinds of entropy harvesting automata that perpetually self build and become everything in their spacetime. We're creating forces bigger than ourselves, and we may reach a point of no return.
- lioeters 5mo agoI don't totally understand, but I like where you're going with this. I picture a cosmological history, the rise and fall of billions of subatomic universes and civilizations, many of them consumed by their own autonomous pseudo-intelligent technologies for better or worse, which on a macro scale are behaviors of particles. We're currently working on our particle, making collective decisions that will affect the super-universe we're a part of, in a tiny but significant way.
- somesortofthing 5mo agoI find AI code usually looks worse than it actually is. It's overly verbose, confusing, and littered with fallbacks that mean that if something goes wrong it falls through a million layers of try/catch and moves the stack trace somewhere completely unrelated to where the error actually happened, but in terms of the actual functionality it works much better than any similar-looking code written by a human would.
- layer8 5mo agoSuch code as you describe is still bad code, though, because it’s difficult to reason about it, both for humans and LLMs.
- somesortofthing 5mo agoIt's bad code, but it's not pretending to be better than it is.
- openclawclub 5mo ago[flagged]
- openclawclub 5mo ago[flagged]
- docheinestages 5mo agoI truly wish more people would write blogs with this style. Adqeuate length, gets the message across, while still telling a story. What do I find often nowadays? LLM-written AI slop that qualifies as a novel in terms of length. Great work!
- learningstud 5mo agoToo true. This is why Rudin's little book, Principles of Mathematical Analysis, normally takes a whole year to cover: one has to work through the proof line by line in order gain enough understanding to do the exercises. Programming gives you the false sense of ease with leaky "interfaces" and quantum-entangled "decoupling". LOL, one microservice in one separate git repo, and tested with mocks and Gherkin syntax BDD. Just LOL, I call this hyperreal programming. With "spec-driven" AI generated code that is reviewed and tested by AI, it is certain that "the programming did not take place". In this "brave new world", "the map has overcome the territory", and "the twelfth camel" was never returned. Programmers certainly don't need "the mythical man-ager" to go delulu and full ouroboros, i.e. "there is no big Other".