4 ms·
> the kind of analysis the program is able to do is past the point where technology looks like magic. I don’t know how you get here from “predict the next word.
by wavemode 7mo ago
> the kind of analysis the program is able to do is past the point where technology looks like magic. I don’t know how you get here from “predict the next word.”
You're implicitly assuming that what you asked the LLM to do is unrepresented in the training data. That assumption is usually faulty - very few of the ideas and concepts we come up with in our everyday lives are truly new.
All that being said, the refine.ink tool certainly has an interesting approach, which I'm not sure I've seen before. They review a single piece of writing, and it takes up to an hour, and it costs $50. They are probably running the LLM very painstakingly and repeatedly over combinations of sections of your text, allowing it to reason about the things you've written in a lot more detail than you get with a plain run of a long-context model (due to the limitations of sparse attention).
It's neat. I wonder about what other kinds of tasks we could improve AI performance at by scaling time and money (which, in the grand scheme, is usually still a bargain compared to a human worker).
- selridge 7mo ago>You're implicitly assuming that what you asked the LLM to do is unrepresented in the training data. This is just as stuck in a moment in time as "they only do next word prediction" What does this even mean anymore? Are we supposed to believe that a review of this paper that wasn't written when that model (It's putatively not an "LLM", but IDK enough about it to be pushy there) was trained? Does that even make sense? We're not in the regime of regurgitating training data (if we really ever were). We need to let go of these frames which were barely true when they took hold. Some new shit is afoot.
- wavemode 7mo agoStatistical models generalize. If you train a model that f(x) = 5 and f(x+1) = 6, the number 7 doesn't have to exist in the training data for the model to give you a correct answer for f(x+2) Similarly, if there are millions of academic papers and thousands of peer reviews in the training data, a review of this exact paper doesn't need to be in there for the LLM to write something convincing. (I say "convincing" rather than "correct" since, the author himself admits that he doesn't agree with all the LLM's comments.) I tend to recommend people learn these things from first principles (e.g. build a small neural network, explore deep learning, build a language model) to gain a better intuition. There's really no "magic" at work here.
- selridge 7mo agoOk cool cool. Instead of pretending you need to teach me, you could engage with what I'm saying or even the OP! "I don't know how you get here from "predict the next word"" is not really so much a statement of ignorance where someone needs you to step in but a reflection that perhaps the tech is not so easily explained as that. No magic needs to be present for that to be the case.
- c22 7mo ago> If you train a model that f(x) = 5 and f(x+1) = 6, the number 7 doesn't have to exist in the training data for the model to give you a correct answer for f(x+2) This is an interesting claim to me. Are there any models that exist that have been trained with a (single digit) number omitted from the training data? If such a model does exist, how does it represent the answer? (What symbol does it use for the '7'?)
- wavemode 7mo agoWhen I say "model" here I'm referring to any statistical model (in this example, probably linear regression). Not specifically large language models / neural networks.
- c22 7mo agoGotcha, I don't think I know enough about it. What constitutes training data for a for a (non neural network) statistical model? Is this something I could play around with myself with pen and paper?
- anon7725 7mo ago“Represented in the training data” does not mean “represented as a whole in the training data”. If A and B are separately in the training data, the model can provide a result when A and B occur in the input because the model has made a connection between A and B in the latent space.
- selridge 7mo agoYes. I’m saying that “it’s just in the training data” is a cognitive containment of these models which is incomplete. You can insist that’s what’s happening, but you’ll be left unable to explain what’s going on beyond truisms.
- WithinReason 7mo agoIt's called "generalization": https://en.wikipedia.org/wiki/Generalization_(learning) https://en.wikipedia.org/wiki/Generalization_(learning)
- selridge 7mo ago>"If A and B are separately in the training data, the model can provide a result when A and B occur in the input because the model has made a connection between A and B in the latent space." This statement (The one I was replying to) is fundamentally unbounded. There's nothing that can't be explained as a combination of "A" and "B" in "training data" because practically speaking we can express anything as such where the combination only needs to be convex along some high-dimensional semantic surface. Add on to that my scare quotes around "training data" because very few people have any practical idea of what is or isn't in there, so we can just make claims strategically. Do we need to explain a success? It was in the training data. A failure, probably not in the training data. Will anyone call us on this transparent farce? Not usually, no. If a statement can--at will--explain everything and nothing, what's it worth?
- jjmarr 7mo agoI created a code review pipeline at work with a similar tradeoff and we found the cost is worth it. Time is a non-issue. We could run Claude on our code and call it a day, but we have hundreds of style, safety, etc rules on a very large C++ codebase with intricate behaviour (cooperative multitasking be fun). So we run dozens of parallel CLI agents that can review the code in excruciating detail. This has completely replaced human code review for anything that isn't functional correctness but is near the same order of magnitude of price. Much better than humans and beats every commercial tool. "scaling time" on the other hand is useless. You can just divide the problem with subagents until it's time within a few minutes because that also increases quality due to less context/more focus.
- smallpipe 7mo ago> This has completely replaced human code review for anything that isn't functional correctness Isn’t functional correctness pretty much the only thing that matters though?
- grey-area 7mo agoWell no, style is important too for humans when they read a codebase, so the LLMs the parent is running clearly have some value for them. They're not claiming LLMs solved every problem, just that they made life easier by taking care of busywork that humans would otherwise be doing. I think personally this is quite a good use for them - offering suggestions on PRs say, as long as humans still review them as well.
- 1718627440 7mo agoBut isn't style already achievable by running e.g. GNU indent?
- jjmarr 7mo agoSome examples of complex transformations linters can't catch: * Function names must start with a verb. * Use standard algorithms instead of for loops. * Refactor your code to use IIFEs to make variables constexpr. The verb one is the best example. Since we work adjacent to hardware, people like creating functions on structs representing register state called "REGISTER_XYZ_FIELD_BIT_1()" and you can't tell if this gets the value of the first field bit or sets something called field bit to 1. If you rename it to `getRegisterXyzFieldBit1()` or `setRegisterXyzFieldBitTo1()` at least it becomes clear what they're doing.
- Kim_Bruning 7mo ago> You're implicitly assuming that what you asked the LLM to do is unrepresented in the training data. That assumption is usually faulty - very few of the ideas and concepts we come up with in our everyday lives are truly new. I made a cursed CPU in the game 'Turing Complete'; and had an older version of claude build me an assembler for it? Good luck finding THAT in the training data. :-P (just to be sure, I then had it write actual programs in that new assembly language)
- deleted 7mo ago[deleted]
- withinboredom 7mo agoBut the ideas are not 'new'. A benchmark that I use to tell me if an AI is overfitted is to present the AI with a recent paper (especially one like a paxos variant) and have it build that. If it writes general paxos instead of what the paper specified, its overfitted. Claude 4.5: not overfitted too much -- does the right thing 6/10 times. Claude 4.6: overfitted -- does the right thing 2/10 times. OpenAI 5.3: overfitted -- does the right thing 3/10 times. These aren't perfect benchmarks, but it lets me know how much babysitting I need to do. My point being that older Claude models weren't overfitted nearly as much, so I'm confirming what you're saying.
- Kim_Bruning 7mo agoCould also be that the model has stronger priors wrt Paxos (and thus has Opinions on what good Paxos should look like) At any rate, with an assembler, you end up with a lot of random letter-salad mnemonics with odd use cases, so that is very likely to tokenize in interesting ways at the very least.
- withinboredom 7mo agoI was just using paxos as an example. Any paper will do.