5 ms·
I like how critique of LLMs evolved on this site over the last few years. We are currently at nonsensical pacing while writing novels.
by comboy 2y ago
I like how critique of LLMs evolved on this site over the last few years.
We are currently at nonsensical pacing while writing novels.
- ruraljuror 2y agoWe are, if this comment is the standard for all criticism on this site. Your comment seems harsh. Perhaps novel writing is too low-brow of a standard for LLM critique?
- jorl17 2y agoI didn't quite read parent's comment like that. I think it's more about how we keep moving the goalposts or, less cynically, how the models keep getting better and better. I am amazed at the progress that we are _still_ making on an almost monthly basis. It is unbelievable. Mind-boggling, to be honest. I am certain that the issue of pacing will be solved soon enough. I'd give 99% probability of it being solved in 3 years and 50% probability in 1.
- jiggawatts 2y agoIn my consulting career I sometimes get to tune database servers for performance. I have a bag of tricks that yield about +10-20% performance each. I get arguments about this from customers, typically along the lines of "that doesn't seem worth it." Yeah, but 10% plus 20% plus 20%... next thing you know you're at +100% and your server is literally double the speed! AI progress feels the same. Each little incremental improvement alone doesn't blow my skirt up, but we've had years of nearly monthly advances that have added up to something quite substantial.
- eru 2y agoYes, if you are Mary Poppins, each individual trick in your bag doesn't have to be large. (For those too young or unfamiliar: Mary Poppins famously had a bag that she could keep pulling things out of.)
- rafaelmn 2y agoExcept at some point the low hanging fruit is gone and it becomes +1%, +3% in some benchmarked use case and -1% in the general case, etc. and then come the benchmarking lies that we are seeing right now, where everyone picks a benchmark that makes them look good and its correlation to real world performance is questionable.
- dalmo3 2y agoWhat exactly is the problem with moving the goalposts? Who is trying to win arguments over this stuff? Yes, Z is indeed a big advance over Y was a big advance over X. Also yes, Z is just as underwhelming. Are customers hurting the AI companies' feelings?
- TeMPOraL 2y ago> Are customers hurting the AI companies' feelings? No. It's the critics' feelings that are being hurt by continued advances, so they keep moving goalposts so they can keep believing they're right.
- HelloMcFly 2y agoThe goalposts should keep moving. That's called progress. Like you, I'm not sure why it seems to irritate or even amuse people.
- ripped_britches 2y agolol wouldn’t that be great to read this comment in 2022
- skyechurch 2y agoThe most straightforward way to measure the pace of AI progress is by attaching a speedometer to the goalposts.
- kaliqt 2y agoOh, that's a good one. And it's true. There seems to be a massive inability for most people to admit the building impact of modern AI development on society.
- benterix 2y agoOh, we do admit impact and even have a name for it: AI slop. (Speaking on LLMs now since AI is a broad term and it has many extremely useful applications in various areas)
- Workaccount2 2y agoAI slop is soon to be "AI output that no one wanted to take credit for".
- munksbeer 2y agoI love this comment.
- josefx 2y agoThey certainly seem to have moved from "it is literally skynet" and "FSD is just around the corner" in 2016 to "look how well it paces my first lady Trump/Musk slashfic" in 2025. Truly world changing.
- Nition 2y agoHaha, so that's the first derivative of goalpost position. You could take the derivative of that to see if the rate of change is speeding up or slowing.
- orena 2y agoI've asked claude to explain what you meant... https://claude.ai/share/391160c5-d74d-47e9-a963-0c19a9c7489a https://claude.ai/share/391160c5-d74d-47e9-a963-0c19a9c7489a
- solardev 2y agoIt's not really passing the Turing Test until it outsells Harry Potter.
- eru 2y agoWell, strictly speaking outselling the Harry Potter would fail the Turing test: the Turing test is about passing for human (in an adversarial setting), not to surpass humans. Of course, this is just some pedantry. I for one love that AI is progressing so quickly, that we _can_ move the goalposts like this.
- jychang 2y agoTo be fair, pacing as a big flaw of LLMs has been a constant complaint from writers for a long time. There were popular writeups about this from the Deepseek-R1 era: https://www.tumblr.com/nostalgebraist/778041178124926976/hydrogen-jukeboxes https://www.tumblr.com/nostalgebraist/778041178124926976/hyd...
- newswasboring 2y agoThis was written on march 15. Deepseek came out in January. "Era" is not a language I would use for something that happened few days ago
- dragonwriter 2y ago> It's not really passing the Turing Test until it outsells Harry Potter. Most human-written books don't do that, so that seems to be a ceiteria for a very different test that a Turing test.
- mirekrusin 2y agoThe joke is that the goalpost is constantly moving.
- TeMPOraL 2y agoThis subgoal post can't move much further after it passes "outsells the Bible" mark.
- krzat 2y agoThis either ends at "better than 50% of human novels" garbage or at unimaginably compelling works of art that completely obsoletes fiction writing. Not sure what is better for humanity in long term.
- WindyMiller 2y agoThat could only obsolete fiction-writing if you take a very narrow, essentially commercial view of what fiction-writing is for. I could build a machine that phones my mother and tells her I love her, but it wouldn't obsolete me doing it.
- bergundytomato 2y agoAhh, now this would be a great premise for a short story (from the mom's POV).
- leokennis 2y agoNot really new is it? First cars just had to be approaching horse and cart levels of speed. Comfort, ease of use etc. were non-factors as this was "cool new technology". In that light, even a 20 year old almost broken down crappy dinger is amazing: it has a radio, heating, shock absorbers, it can go over 500km on a tank of fuel! But are we fawning over it? No, because the goalposts have moved. Now we are disappointed that it takes 5 seconds for the Bluetooth to connect and the seats to auto-adjust to our preferred seating and heating setting in our new car.
- rafaelmn 2y agoPeople are trying to use gen AI in more and more use-cases, it used to fall flat on its face at trivial stuff, now it got past trivial stuff but still scratching the boundaries of being useful. And that is not an attempt to make the gen AI tech look bad, it is really amazing what it can do - but it is far from delivering on hype - and that is why people are providing critical evaluations. Lets not forget the OpenAI benchmarks saying 4.0 can do better at college exams and such than most students. Yet real world performance was laughable on real tasks.
- parineum 2y ago> Lets not forget the OpenAI benchmarks saying 4.0 can do better at college exams and such than most students. Yet real world performance was laughable on real tasks. That's a better criticism of college exams than the benchmarks and/or those exams likely have either the exact questions or very similar ones in the training data. The list of things that LLMs do better than the average human tends to rest squarely in the "problems already solved by above average humans" realm.
- ksec 2y agoDo we have any simple benchmarks ( and I know benchmarks are not everything ) that tests all the LLMs? The pace is moving so fast I simply cant keep up. Or a ELI5 page which gives a 5 min explanation of LLM from 2020 to this moment?
- basch 2y agoIt’s more a bellwether or symptom of a flaw where the context becomes poisoned and continually regurgitates the same thought over and over.
- stickfu 2y agoI don’t know why I keep submitting myself to hacker news but every few months I get the itch, and it only takes a few minutes to be turned off by the cynicism. I get that it’s from potentialy wizened tech heads who have been in the trenches and are being realistic. It’s great for that, but any new bright eyed and bushy tailed dev/techy, whatever, should stay far away until much later in their journey