13 ms·
Every time I see an article like this, it's always missing --- but is it any good, is it correct? They always show you the part that is impressive - "it walked
by mynameisjody 11mo ago
Every time I see an article like this, it's always missing --- but is it any good, is it correct? They always show you the part that is impressive - "it walked the tricky tightrope of figuring out what might be an interesting topic and how to execute it with the data it had - one of the hardest things to teach."
Then it goes on, "After a couple of vague commands (“build it out more, make it better”) I got a 14 page paper." I hear..."I got 14 pages of words". But is it a good paper, that another PhD would think is good? Is it even coherent?
When I see the code these systems generate within a complex system, I think okay, well that's kinda close, but this is wrong and this is a security problem, etc etc. But because I'm not a PhD in these subjects, am I supposed to think, "Well of course the 14 pages on a topic I'm not an expert in are good"?
It just doesn't add up... Things I understand, it looks good at first, but isn't shippable. Things I don't understand must be great?
- apendleton 11mo agoI think they get to that a couple of paragraphs later: > The idea was good, as were many elements of the execution, but there were also problems: some of its statistical methods needed more work, some of its approaches were not optimal, some of its theorizing went too far given the evidence, and so on. Again, we have moved past hallucinations and errors to more subtle, and often human-like, concerns.
- adamors 11mo ago> Things I don't understand must be great? Couple it with the tendency to please the user by all means and it ends up lieing to you but you won’t ever realise, unless you double check.
- JumpCrisscross 11mo ago> Couple it with the tendency to please the user by all means Why aren't foundational model companies training separate enterprise and consumer models from the get go?
- Herring 11mo agoI think the point is we’re getting there. These models are growing up real fast. Remember 54% of US adults read at or below the equivalent of a sixth-grade level.
- lm28469 11mo ago> Remember 54% of US adults read at or below the equivalent of a sixth-grade level. The sane conclusion would be to invest in education, not to dump hundreds of billions of llms, but ok
- Herring 11mo agoIn theory yeah, but in practice 54% will also vote against funding education. Catch-22.
- tehjoker 11mo agoNot true, most people are not upper-middle class anti-tax wackos. They benefit from those people being taxed.
- wmeredith 11mo agoThe electorate in the U.S. commonly votes against its own interests.
- krainboltgreene 11mo agoPithy, but not true.
- mavhc 11mo agoThat's why you phrase it as "woke liberals turning your children gay!" In USA K-12 education costs about $300k 350 million people, want to get 175 million of them better educated, but we've already spent $52 trillion dollars on educating them so far
- stavros 11mo agoIt's gotten more and more shippable, especially with the latest generation (Codex 5.1, Sonnet 4.5, now Opus 4.5). My metric is "wtfs per line", and it's been decreasing rapidly. My current preference is Codex 5.1 (Sonnet 4.5 as a close second, though it got really dumb today for "some reason"). It's been good to the point where I shipped multiple projects with it without a problem (with eg https://pine.town https://pine.town being one I made without me writing any code).
- Madmallard 11mo agoIt's not really any different in my experience
- mirekrusin 11mo agoStochastic parrot? Autocomplete on steroids? Fancy autocorrect? Bullshit generator? AI snake oil? Statistical mimicry? You don't hear that anymore. Feels like whole generation of skeptics evaporated.
- bigstrat2003 11mo agoI certainly hold those opinions still, because the models still have yet to prove they are anything worth a person's time. I don't bother posting that because there's no way an AI hype person and I are ever going to convince each other, so what's the point? The skeptics haven't evaporated, they just aren't bothering to try to talk to you any more because they don't think there's value in it.
- mirekrusin 10mo ago[flagged]
- csomar 10mo agoThe earth is flat until you have evidence of the contrary. It's you who should provide that evidence. We had physics, navigation and then space shuttles that clearly showed the earth is not flat. We are yet to have a fully vibe-coded piece of software that actually works. The blog post is actually great because LLMs are very good are regurgitating pieces of code that already exist on a single prompt. Now ask them to make a few changes and suddenly the genie is back in the bottle. Something doesn't math out. You can't be both a genius and extremely dumb (retarded) at the same time. You can be, however, good at information retrieval and presenting it in a better way. That's what LLMs are and am not discounting the usefulness of that.
- pojzon 11mo agoTruth is you still need human to review all of it, fix it where needed, guide it when it hallucinate and write correct instructions and prompts. Without knowledge how to use this “PROBALISTIC” slot machine to have better results ypu are only wasting energy those GPUs need to run and answer questions. Majority of ppl use LLMs incorrectly. Majority of ppl selling LLMs as a panacea for everyting are lying. But we need hype or the bubble will burst taking whole market with it, so shuushh me.
- brightball 11mo agoI keep trying out different models. Gemini 3 is pretty good. It’s not quite as good at one shotting answers as Grok but overall it’s very solid. Definitely planning to use it more at work. The integrations across Google Workspace are excellent.
- cgh 11mo agoThis is a variation of the Gell-Mann amnesia effect: https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect
- monooso 11mo agoThe author goes into the strengths and weaknesses of the paper later in the article.
- tsss 11mo agoFor what it's worth I have been using Gemini 2.5/3 extensively for my masters thesis and it has been a tremendous help. It's done a lot of math for me that I couldn't have done on my own (without days of research), suggested many good approaches to problems that weren't on my mind and helped me explore ideas quickly. When I ask it to generate entire chapters they're never up to my standard but that's mostly an issue of style. It seems to me that LLMs are good when you don't know exactly what you want or you don't care too much about the details. Asking it to generate a presentation is an utter crap shoot, even if you merely ask for bullet points without formatting.
- ammbauer 10mo ago> It's done a lot of math for me that I couldn't have done on my own (without days of research), Isn't the point of doing the master's thesis that you do the math and research, so that you learn and understand the math and research?
- ragequittah 10mo agoI bet they were talking about how people didn't do long division when the calculator first came out too. Is using matlab and excel ok but AI not? Where do we draw the line with tools?
- bleepblap 10mo agoOP said they "generated entire chapters"
- dickersnoodle 10mo agoApparently not. This is the most perfect example I've seen of "I can recite it, but I don't understand it so I don't know if it's really right or not" that I've seen in a while.
- tsss 10mo agoI do understand it. I just don't have the overview of all the algorithms that LLMs have.
- Lerc 11mo agoI guess you have a couple of options. You could trust the expert analysis of people in that field. You can hit personal ideologies or outliers, but asking several people seems to find a degree of consensus. You could try varying tasks that perform complex things that result in easy to test things. When I started trying chatbots for coding, one of my test prompts was Create a JavaScript function edgeDetect(image) that takes an ImageData object and returns a new ImageData object with all direction Sobel edge detection. That was about the level where some models would succeed and some will fail. Recently I found Can you create a webgl glow blur shader that takes a 2d canvas as a texture and renders it onscreen with webgl boosting the brightness so that #ffffff is extremely bright white and glowing, Produced a nice demo with slider for parameters, a few refinements (hierarchical scaling version) and I got it to produce the same interface as a module that I had written myself and it worked as a drop in replacement. These things are fairly easy to check because if it is performant and visually correct then it's about good enough to go. It's also worth noting that as they attempt more and more ambitious tasks, they are quite probably testing around the limit of capability. There is both marketing and science in this area. When they say they can do X, it might not mean it can do it every time, but it has done it at least once.
- taurath 11mo ago> You could trust the expert analysis of people in that field That’s the problem - the experts all promise stuff that can’t be easily replicated. The promises the experts send doesn’t match the model. The same request might succeed and might fail, and might fail in such a way that subsequent prompts might recover or might not.
- Lerc 11mo agoThe experts I am talking about trusting here are the ones doing the replication, not the ones making the claims.
- timschmidt 11mo agoThat's how working with junior team members or open source project contributors goes too. Perhaps that's the big disconnect. Reviewing and integrating LLM contributions slotted right into my existing workflow on my open source projects. Not all of them work. They often need fixing, stylistic adjustments, or tweaking to fit a larger architectural goal. That is the norm for all contributions in my experience. So the LLM is just a very fast, very responsive contributor to me. I don't expect it to get things right the first time. But it seems lots of folks do. Nevertheless, style, tweaks, and adjustments are a lot less work than banging out a thousand lines of code by hand. And whether an LLM or a person on the other side of the world did it, I'd still have to review it. So I'm happy to take increasingly common and increasingly sophisticated wins.
- secondbreakfast 11mo agoLoads of AI chatter is the Murray Gell-Mann Amnesia Effect on steroids
- jrumbut 11mo agoWell, that's why people still have jobs but I appreciate the idea of the post that the neat demo was a coherent paragraph or silly poem. The silly poems were all kind of similar, not very funny, and the paragraphs were a good start but I wouldn't use them for anything important. Now the tightrope is a whole application or a 14 page paper and the short pieces of code and prose are now professional quality more often than not. That's some serious progress.
- visarga 10mo agoYou don't use it that way. You use it to help you build and run experiments, and help you discuss your findings, and in the end helps you write your discoveries. You provide the content, and actual experiments provide the signal.
- ManlyBread 10mo agoLike clockwork. Each time someone criticizes any aspect of any LLM there's always someone to tell that person they're using the LLM wrong. Perhaps it's time to stop blaming the user?
- becquerel 10mo agoYou can recognise that the technology has a poor user interface and is wrought with subtleties without denying its underlying capabilities. People misuse good technology all the time. It's kind of what users do. I would not expect a radically new form of computing which is under five years old to be intuitive to most people.
- sandspar 10mo agoIf someone says that they can't get a camera to work, you tell them how to fix it, right? I can't think of what other response is appropriate.
- ManlyBread 10mo agoWhy would their response be appropriate when even the creators of the LLM doesn't clearly state the purpose of their software, yet alone instruct users how to use it? The person I replied to said that this software should be used yo "help you build and run experiments, and help you discuss your findings, and in the end helps you write your discoveries" - I dare anyone to find any mention of this workflow being the "correct" way of using any LLM in the LLM's official documentation.
- evilduck 10mo agoValidation that cameras will never work and photographs aren't real.
- Glemkloksdjf 10mo ago[dead]
- eckesicle 10mo ago> It just doesn't add up... Things I understand, it looks good at first, but isn't shippable. Things I don't understand must be great? It’s like the Gell-Mann amnesia effect applied to AI. :) https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect
- seidleroni 10mo agoThe author actually discusses the results of the paper. He's not some rando but a Wharton Professor and when he is comparing the results to a grad student, it is with some authority. "So is this a PhD-level intelligence? In some ways, yes, if you define a PhD level intelligence as doing the work of a competent grad student at a research university. But it also had some of the weaknesses of a grad student. The idea was good, as were many elements of the execution, but there were also problems..."
- leeoniya 10mo ago> But because I'm not a PhD in these subjects, am I supposed to think, "Well of course the 14 pages on a topic I'm not an expert in are good"? https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect https://en.wikipedia.org/wiki/Gell-Mann_amnesia_effect