3 ms·
Because LLMs do not work like that - there's no "understanding" the source and answering questions, it simply "finds" similar results in its training data (matc
by miningape 9mo ago
Because LLMs do not work like that - there's no "understanding" the source and answering questions, it simply "finds" similar results in its training data (matching it with the context) and regurgitates (some part of) it (+ other "noise").
Meaning as technology evolves and does things in novel ways, without explainers annotating it the LLM won't have anything to draw on - reducing the quality of answers. Which brings us full circle, what will companies use as training data without answers in places like SO?
- spiderfarmer 9mo agoYou’re talking about ChatGPT 2.0 and severely underestimate the capabilities of today’s models.
- miningape 9mo agoIt's well known that even current LLMs do not perform well on logic games when you change the names / language used. i.e. try asking it to swap the meanings of the words red and green and ask it to describe the colors in a painting and analyse it with color theory - notice how quickly the results degrade, often attributing "green" qualities to "red" since it's now calling it "green". What this shows us is that training data (where the associations are made) plays a significant role in the level of answer an LLM can give, no matter how good your context is (at overriding the associations / training data). This demonstrates that training data is more important (for "novel" work) than context is.
- Eisenstein 9mo agoWrite "This sentence is green." in red sharpie and "This sentence is red" in green sharpie on a piece of paper. Show it to someone briefly and then hide it. Ask them what color the first sentence said it was and what color the second sentence was written in. Another one: ask a person to say 'silk' 5 times, then ask them what cows drink. Exploiting such quirks only tells you that you can trick people, not what their capabilities are.
- miningape 9mo agoThe point isn't that you can trick an LLM, but that their capabilities are more strongly tied to training data than context. That's to say, when context and training disagree, training "wins". ("wins" isn't the correct wording, but hopefully you understand the point) This poses a problem for new frameworks/languages/whatever that do things in a wholly different way since we'll be forced to rely on context that will contradict the training data that's available.
- ndriscoll 9mo agoWhat is an example of a framework that does things in a wholly different way? Everything I'm familiar with is a variation on well explored ideas from the 60s-70s. If you had someone familiar with every computer science concept, every textbook, every paper, etc. up to say 2010 (or even 2000 or earlier), along with deep experience using dozens of programming languages, and you sat them down to look at a codebase, what could you put in front of them that they couldn't describe to you with words they already know?
- miningape 9mo agoEven the differences between React and Svelte are big enough for this to be noticeable. And Svelte is actually present in the training data. Given the large amount of react training data, svelte performs significantly worse (yes, even when given the full official svelte llms.txt in the context)
- Eisenstein 9mo agoBut it doesn't pose a problem. You are extrapolating things that are not even correlated. You started with 'they can't understand anything new' and then followed it up with 'because I can trick it with logic problems' which doesn't prove that. Have you even tried doing what you say won't work?
- jakeydus 9mo agoIf I make up a riddle and ask an LLM to solve it, it will perform worse than a riddle that is well known and whose solution will be found in the dataset. That's just a foundational component of how they work.
- spiderfarmer 9mo agoYes you can trick it. But it’s almost trivial for an LLM to generate every question and answer combo you could every come up with based on new documentation and new source code for a new framework. It doesn’t need StackOverflow anymore. It’s already miles ahead.
- Eisenstein 9mo agoI just downloaded "Degeneration in discriminantal arrangements", by Saito, Takuya from the journal "Advances in applied mathematics" dated November 2025 and fed it to Claude. It not only explained the math but created a react app to demonstrate it. I'm not that can be explained by regurgitating part of it with noise. I encourage you to try it with something of your own. Abstract: Discriminantal arrangements are hyperplane arrangements that are generalization of braid arrangements. They are con- structed from given hyperplane arrangements, but their com- binatorics are not invariant under combinatorial equivalence. However, it is known that the combinatorics of the discrimi- nantal arrangements are constant on a Zariski open set of the space of hyperplane arrangements. In the present paper, we introduce (T, r)-singularity varieties in the space of hyper- plane arrangements to classify discriminantal arrangements and show that the Zariski open set is the complement of (T, r)-singularity varieties. We study their basic properties and operations and provide examples, including infinite fami- lies of (T, r)-singularity varieties. In particular, the operation that we call degeneration is a powerful tool for constructing (T, r)-singularity varieties. As an application, we provide a list of (T, r)-singularity varieties for spaces of small line ar- rangements. * https://doi.org/10.1016/j.aam.2025.103001 https://doi.org/10.1016/j.aam.2025.103001 * https://imgur.com/jjfFNMI.png https://imgur.com/jjfFNMI.png
- ndriscoll 9mo agoMy recent experience with codex is that they absolutely do work that way today (this may be recent as in within the last couple months), and will autonomously decide to grep for things in your codebase to get context on changes you've asked for or questions you've asked. I've been pretty open to my manager about calling my upper management delusional with this stuff until very recently (in the sense that 6 months ago everything I tried was still a toy), but it's actually now reaching a tipping point that's drastically changing how I work.