4 ms·
I've worked with these systems for four years now and they have not meaningfully improved in that time frame. We still have: - statistical correlation between
by rf15 3mo ago
I've worked with these systems for four years now and they have not meaningfully improved in that time frame.
We still have:
- statistical correlation between two things will always cause one thing to lead to the other, no matter how much you prompt it to not have that connection (to be expected with a stochastic system)
- Math completely fails in longer contexts
- "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion
- smearing of properties between logically distinct objects (a red ball and a green cube can quickly become a red cube and a green ball)
- Barbing 3mo agoLike how toddlers’ skills don’t meaningfully improve on infants’, because either could wake up in a wet bed.
- mdp2021 3mo agoLet us be more clear: there is no structural jump, no architectural overcoming of the original fault. (Edit: and on a similar point, structural properties such as having static ntetworks, as opposed to continuously learning and improving architectures (such as us), will reveal that there is still road ahead.)
- Barbing 3mo agoMaybe more fair then would be: “I've worked with these systems for four years now and while they _have_ meaningfully improved in that time frame, they’re not perfect and remain fundamentally flawed in various ways.” You prompt less. You need not inject search results into the context window yourself, a window much larger than years ago. You get code that’s already been run successfully once instead of finding an obvious show stopping bug yourself. The technology is not a brand new one that fixed everything wrong with the old one, no, but not sure I would’ve noticed your comment if it had been such a bland observation. I genuinely assume good faith here… will say am tempted to assume the standards of someone posting such a thing might be impossibly high. Glad to be having a fun conversation instead of getting your grades on my work product or something :)
- orwin 3mo agoIf you go to a LLM without harness, GP original point in completely right. LLMS by themselves are still shit at math, they still confuse weird correlation to causation every time (and sometimes in ways even a 9 year old would say "no, that's dumb"), and confuse original parameters very often. I disagree with " "thinking" token generation being on the correct track just to 'no, wait' on an already correct conclusion", because i think that is an effect of the harness, not the LLMs. 80% off all the improvements since ChatGPT4 are in the harnesses, and the LLMs by themselves, while they improved in areas they already were good at (translation especially) did not fix any of they original issues (object permanence, calculusm correlation). Just run old models in the playground and get them to play chess (maybe make a small custom harness if you feel like it), then replace it with a frontier model (i don't know if you still have API access without harness on US models, but if you don't try K3), you will see LLMs weaknesses were not at all fixed, even marginally. They're way better and not inducing bugs in the code, so that make them usable since Opus4.5 (anyone using them prior to that either had a greenfield project or like spending hours debugging).
- rf15 3mo agoI agree, harnesses is where everyone improved the most. Our internal experimental tool can now semi-reliably formulate small programs to assert the correctness of their theories, for example. Context length is still a weird factor that we haven't sensibly solved - if anything, the lesson learned was to keep the context as small as possible and do most of the true "thinking" in the harness and temporary generated code.
- Barbing 3mo agoSounds very fair, thanks :)
- user43928 3mo ago> 80% off all the improvements since ChatGPT4 are in the harnesses That seems easily falsifiable by putting an old model into the current harness and comparing it to 5.6 Sol or Fable.
- fakwandi_priv 3mo ago> - Math completely fails in longer contexts Not sure what longer contexts we're talking about but didn't we have an old math problem optimized, which even the LLM itself was surprised about, just a week ago? Something which wasn't possible 6 months ago.
- rf15 3mo agoI mean calculations, not mathematical proofs
- Barbing 3mo agoIf they use Python to fill the gap, and the end user doesn’t have to know or care, is it unfair to assess this as progress and attribute the progress to the _system_? OK, the core technology that is the language model still can’t math as well as you’d hope, but how about the end result users see from the system when they interface with it? “Did you know humans are better at flying today than they were a thousand years ago?” ‘No they’re not, they need planes.’ Technically correct in a way but isn’t it kind of annoying to be so stubbornly pedantic when the context is speed of reaching Point B from Point A?
- rf15 3mo agoYou are correct, the frameworks around it have improved. In that regard, my assessment is unfair: I only judge the underlying technology and what is sold by the sota providers, with the premise of what it's like when you start fresh. You can achieve a lot by coding around the issues, but that's kinda against the point of 'AI', is it?
- wahnfrieden 3mo agoNo, its capabilities with a harness are what we are interested in. Your assessment is only relevant to benchmarking, not practical value.
- Barbing 3mo ago
- IanCal 3mo ago> I've worked with these systems for four years now and they have not meaningfully improved in that time frame. Not meaningfully improved?! Four years ago was gpt *3.5*! ChatGPT hadn’t been released!
- rf15 3mo agoYes! Impressive, isn't it? I see how it has improved for some minor points, that the big models can cover more finetuning ground, but my big gripes are still the same - you could do the same back then with multiple models and more targeted finetuning.
- anonzzzies 3mo ago> you could do the same back then with multiple models and more targeted finetuning Are you one of those anonymous billionaires as if you did this a few years ago, you would've been famous and rich.
- danielbln 3mo agoOP is delusional or deliberately optuse. I work in the space and stare down these systems 12h/day, and saying the systems haven't meaningfully improved is ludicrous.
- rf15 3mo agoOP is largely pissed with what OAI/Antrophic are trying to sell as meaningful improvements and the market-bending money they ask for it. I work in the space and we trained LLM models on conceptual tokens, not language tokens, for example. See Symbolic AI and all the attempts at hybrid models. Also, uh, fame and riches are not really my thing. Middle income is fine. My mistake was speaking up here because I got carelessly annoyed because I have skin in the game, research-wise. I'm sorry for that.
- mdp2021 3mo ago
- glimshe 3mo agoMessages like this in the training data are how LLMs learn to say absurd things with total confidence.
- azan_ 3mo ago> I've worked with these systems for four years now and they have not meaningfully improved in that time frame. That's absolutely insane. Is it some case of anti-AI psychosis?