8 ms·
I don't know how you get here from “predict the next word”
- leptons 7mo ago[flagged]
- margalabargala 7mo agoThe article author claims AI was not used to write the article. Personally I believe them, considering the content of the article.
- Doc-Bok 7mo agoFor what it's worth, my AI content blocker blocked the article.
- qsi 7mo agoMaybe because it quotes (at length) AI-generated output?
- st-keller 7mo agoHappens a lot to highly reasonable, well written texts without orthographical errors.
- alex_smart 7mo agoIt’s worth nothing, to the point it makes people wonder why you even volunteer that information.
- qsi 7mo agoThe article as written is entirely consistent with John Cochrane's style. I have been reading him off and on over the years so I think I have a decent baseline for comparison. It doesn't smell of AI to me. If anything, even the included quotes from Refine don't smell much of typical AI, but maybe I am less discerning there. I did notice the em-dashes though!
- throw310822 7mo agoWe flew past the Turing test so brilliantly that now people have to try to convince interviewers that they're real.
- pushedx 7mo agoYes, most people (including myself) do not understand how modern LLMs work (especially if we consider the most recent architectural and training improvements). There's the 3b1b video series which does a pretty good job, but now we are interfacing with models that probably have parameter counts in each layer larger than the first models that we interacted with. The novel insights that these models can produce is truly shocking, I would guess even for someone who does understand the latest techniques.
- measurablefunc 7mo agoWhat's the latest novel insight you have encountered?
- brookst 7mo agoNot the person you asked, and “novel” is a minefield. What’s the last novel anything, in the sense you can’t trace a precursor or reference? But.. I recently had a LLM suggest an approach to negative mold-making that was novel to me. Long story, but basically isolating the gross geometry and using NURBS booleans for that, plus mesh addition/subtraction for details. I’m sure there’s prior art out there, but that’s true for pretty much everything.
- measurablefunc 7mo agoI don't know, that's why I asked b/c I always see a lot of empty platitudes when it comes to LLM praise so I'm curious to see if people can actually back up their claims. I haven't done any 3D modeling so I'll take your word for it but I can tell you that I am working on a very simple interpreter & bytecode compiler for a subset of Erlang & I have yet to see anything novel or even useful from any of the coding assistants. One might naively think that there is enough literature on interpreters & compilers for coding agents to pretty much accomplish the task in one go but that's not what happens in practice.
- pushedx 7mo ago
- belZaah 7mo agoIt’s called emergent behavior. We understand how an llm works, but do not have even a theory about how the behavior emerges from among the math. We understand ants pretty well, but how exactly does anthill behavior come from ant behavior? It’s a tricky problem in system engineering where predicting emergent behavior (such as emergencies) would be lovely.
- devmor 7mo agoThe good news is that despite being incredibly complex, it’s still a lot simpler than ants because it is at least all statistical linguistics (as far as LLMs are concerned anyways).
- themafia 7mo ago> but do not have even a theory about how the behavior emerges We fully do. There is a significant quality difference between English language output and other languages which lends a huge hint as to what is actually happening behind the scenes. > but how exactly does anthill behavior come from ant behavior? You can't smell what ants can. If you did I'm sure it would be evident.
- kristiandupont 7mo agoI am very curious about this significant hint, could you point me to some material?
- spiralcoaster 7mo agoTwo very big revelations here that I would love to know more about: 1. Can you reveal "what's actually happening behind the scenes" beyond the hint you gave? I can't figure it out. 2. Can you explain how an ants sense of smell leads to anthills?
- jen729w 7mo ago> 2. Can you explain how an ants sense of smell leads to anthills? Ant 0: doesn’t seem to be dangerous here. I’ll drop a scent. Ant 1: oh cool, a safe place. And I didn’t die either. I’ll reinforce that. Ant 142,857,098,277: cool anthill.
- WD-42 7mo agoThis is really hard to judge because by the looks of it, finance papers mostly consist of gobbledygook and extensive filler to begin with.
- sp4cemoneky 7mo agoThis. Verbalism lands really well to verbalism.
- cyanydeez 7mo agoEconomics is the attempt to take sociology and add numbers to make it look like a hard science. The fintechbros then seem to think because they can make numbers go up that this proof it's a hard science.
- Tarq0n 7mo agoThat's entirely missing the point. "All models are wrong, but some are useful". You can test hypotheses and learn things even about chaotic or emergent systems.
- friendzis 7mo ago> You can test hypotheses and learn things even about chaotic or emergent systems. Ah yes, the famous "Cut GDP in half, abolish public schooling and use that as a control" experiment. Majority of economic "models" are entirely correlational without any mechanistic explanation whatsoever or an explanation so superficial that it contradicts either itself or observed reality. If you look deeper and read explanatory notes of economic laws, the model may refer some publications, but then the actual figures plugged in the model are explained as "these values have been observed to lead to the desired outcomes, therefore are set without any modeling or validation, hope for the best, lesssgoooo".
- tolerance 7mo agoIt’s interesting to read about the use and leverage of LLMs outside of programming. I’m not too familiar with the history, but the import of this article is brushing up on my nose hairs in a way that makes me think a sort of neo-Sophistry is on the horizon.
- themafia 7mo ago> The comments it offered were on the par of the best comments I’ve received on a paper in my entire academic career. Sort of the lowest hanging fruit imaginable. Just because it became "fundamental" to the process doesn't mean it gained any quality.
- libraryofbabel 7mo agoI have come to think “predict the next token” is not a useful way to explain how LLMs work to people unfamiliar with LLM training and internals. It’s technically correct, but at this point saying that and not talking about things like RLVR training and mechanistic interpretability is about as useful as framing talking with a person as “engaging with a human brain generating tokens” and ignoring psychology. At least AI-haters don’t seem to be talking about “stochastic parrots” quite so much now. Maybe they finally got the memo.
- qsera 7mo ago>“predict the next token” is not a useful way That is the exact thing to say because that is exactly what it does, despite how it does so. It is not useful to say it if you are an AI-shill though. You bought up AI-hater, so I think I am entitled to bring up AI-shills.
- vasco 7mo agoMy neurons are also just passing electric signals back and forward and exchanging water and salts with the rest of my body.
- wavemode 7mo ago> the kind of analysis the program is able to do is past the point where technology looks like magic. I don’t know how you get here from “predict the next word.” You're implicitly assuming that what you asked the LLM to do is unrepresented in the training data. That assumption is usually faulty - very few of the ideas and concepts we come up with in our everyday lives are truly new. All that being said, the refine.ink tool certainly has an interesting approach, which I'm not sure I've seen before. They review a single piece of writing, and it takes up to an hour, and it costs $50. They are probably running the LLM very painstakingly and repeatedly over combinations of sections of your text, allowing it to reason about the things you've written in a lot more detail than you get with a plain run of a long-context model (due to the limitations of sparse attention). It's neat. I wonder about what other kinds of tasks we could improve AI performance at by scaling time and money (which, in the grand scheme, is usually still a bargain compared to a human worker).
- selridge 7mo ago>You're implicitly assuming that what you asked the LLM to do is unrepresented in the training data. This is just as stuck in a moment in time as "they only do next word prediction" What does this even mean anymore? Are we supposed to believe that a review of this paper that wasn't written when that model (It's putatively not an "LLM", but IDK enough about it to be pushy there) was trained? Does that even make sense? We're not in the regime of regurgitating training data (if we really ever were). We need to let go of these frames which were barely true when they took hold. Some new shit is afoot.
- wavemode 7mo agoStatistical models generalize. If you train a model that f(x) = 5 and f(x+1) = 6, the number 7 doesn't have to exist in the training data for the model to give you a correct answer for f(x+2) Similarly, if there are millions of academic papers and thousands of peer reviews in the training data, a review of this exact paper doesn't need to be in there for the LLM to write something convincing. (I say "convincing" rather than "correct" since, the author himself admits that he doesn't agree with all the LLM's comments.) I tend to recommend people learn these things from first principles (e.g. build a small neural network, explore deep learning, build a language model) to gain a better intuition. There's really no "magic" at work here.
- mnewme 7mo agoIs this an ad? Seems like it. The text is not really what the headline suggests.
- pianom4n 7mo agoDo you think the submitter intended this as an ad? His post history doesn't seem suspicious. Or do you think article's author wrote this an an ad? He's a reputable academic who seems impressed with an AI tool he used and is honestly sharing his thoughts. For reference he published the 80 page inflation mini-book 2 weeks ago asking for feedback: https://www.grumpy-economist.com/p/inflation https://www.grumpy-economist.com/p/inflation
- deaux 7mo ago> Or do you think article's author wrote this an an ad? He's a reputable academic who seems impressed with an AI tool he used and is honestly sharing his thoughts. Ghuntley used to be reputable on here, then the crypto money looked too juicy.
- pianom4n 7mo agoAre you seriously comparing a random hacker to a lifelong academic for their odds of becoming a crypto shill?
- callmeal 7mo agoThe "predict the next word" to a current llm is at the same level as a "transistor" (or gate) is to a modern cpu. I don't understand llms enough to expand on that comparison, but I can see how having layers above that feed the layers below to "predict the next word" and use the output to modify the input leading to what we see today. It is turtles all the way down.
- brookst 7mo agoIt’s a good comparison. It’s about abstraction and layers. Modern LLMs aren’t just models, they’re all the infrastructure around promoting and context management and mixtures of experts. The next-word bit may be slightly higher than an individual transistor, possibly functional units.
- echelon 7mo agoHumans are future predictors. Our vision systems, our mental models of our careers. People that predict the future tend to do well financially. Now the machines are getting better than we are. It's exciting and a little bit terrifying. We were polymers that evolved intelligence. Now the sand is becoming smart.
- qsera 7mo ago>Now the machines are getting better than we are Then AI companies should stop looking for investors and instead play stock markets with all that predictive powers!
- echelon 7mo agoThe real money is in using the models to build utility and money-making companies. You're removed from orders of magnitude in upside potential if you have to wait for the public markets.
- qsera 7mo ago> money-making companies You mean, money sucking companies, right? >You're removed from orders of magnitude in upside potential if you have to wait for the public markets. because that won't work. That is why!
- visarga 7mo ago> Nothing you write will matter if it is not quickly adopted to the training dataset. That is my take too, I was surprised to see how many people object to their works being trained on. It's how you can leave your mark, opening access for AI, and in the last 25 years opening to people (no restrictions on access, being indexed in Google).
- mbgerring 7mo agoPeople who produced the works LLMs are trained on are not compensated for the value they are now producing, and their skills are increasingly less valued in a world with LLMs. The value the LLMs are producing is being captured by employees of AI companies who are driving up rent in the Bay Area, and driving up the cost of electricity and water everywhere else. Your surprise to people’s objections makes sense if you can’t count.
- chii 7mo ago> People who produced the works LLMs are trained on are not compensated for the value they are now producing the value being extracted via LLM techniques is new value, which did not previously exist. The producer(s) of the old data had an asking price, which was taken by the LLM trainers. They cannot make the argument that since the LLM is producing new value, they should retroactively update their old asking price for their works. They could update their asking price for any new works they produce. They also have the right to ask their works not be used for training, etc. But they cannot ask their old works to be paid for by the new uses in LLM in a retroactive way.
- GolfPopper 7mo ago>The producer(s) of the old data had an asking price, which was taken by the LLM trainers. This is... blatantly untrue? https://arstechnica.com/tech-policy/2026/02/microsoft-removes-guide-on-how-to-train-llms-on-pirated-harry-potter-books/ https://arstechnica.com/tech-policy/2026/02/microsoft-remove... https://www.theatlantic.com/technology/archive/2025/03/libgen-meta-openai/682093/ https://www.theatlantic.com/technology/archive/2025/03/libge...
- retrac 7mo agoI know this sounds insane but I've been dwelling on it. Language models are digital Ouija boards. I like the metaphor because it offers multiple conflicting interpretations. How does a Ouija board work? The words appear. Where do they come from? It can be explained in physical terms. Or in metaphysical terms. Collective summing of psychomotor activity. Conduits to a non-corporeal facet of existence. Many caution against the Ouija board as a path to self-inflicted madness, others caution against the Ouija board as a vehicle to bring poorly understood inhuman forces into the world.
- brookst 7mo agoOuija boards are just collective negotiation among people.
- nekusar 7mo agoThere's 2 completely different ways to understand how a Ouija board works. Occult, and Scientific. Scientific: It's a combined response from everyone's collective unconscious blend of everyone participating. In other words, its a probabilistic result of an "answer" to the question everyone hears. Occult: If an entity is present, it's basically the unshielded response of that entity by collectively moving everyone's body the same way, as a form of a mild channel. Since Ouija doesn't specific to make a circle and request presence of a specific entity, there's a good chance of some being hostile. Or, you all get nothing at all, and basically garbage as part of the divination/communication. But comparing Ouija to LLMs? The LLM, with the same weights, with the same hyperparameters, and same questions will give the same answers. That is deterministic, at least in that narrow sense. An Ouija board is not deterministic, and cannot be tested in any meaningful scientific sense.
- ChaitanyaSai 7mo agoThe whole next word thing is interesting isn't it. I like to see it with Dennett's "Competence and comprehension" lens. You can predict the next word competently with shallow understanding. But you could also do it well with understanding or comprehension of the full picture. A mental model that allows you to predict better. Are the AIs stumbling into these mental models? Seems like it. However, because these are such black boxes, we do not know how they are stringing these mental models together. Is it a random pick from 10 models built up inside the weights? Is there any system-wide cohesive understanding, whatever that means? Exploring what a model can articualate using self-reflection would be interesting. Can it point to internal cognitive dissonance because it has been fed both evolution and intelligent design, for example? Or these exist as separate models to invoke depending on the prompt context, because all that matters is being rewarded by the current user?
- halyconWays 7mo agoSearle's Chinese Room experiment but without knowing what's in the room, and when you try to peek in you just see a cloud of fog and are left to wonder if it's just a guy with that really big dictionary or something more intelligent.
- selridge 7mo agoIt's an octopus, perhaps: https://aclanthology.org/2020.acl-main.463.pdf https://aclanthology.org/2020.acl-main.463.pdf There's also this blog post: https://julianmichael.org/blog/2020/07/23/to-dissect-an-octopus.html https://julianmichael.org/blog/2020/07/23/to-dissect-an-octo... (which IMO is better to read than the paper)
- grey-area 7mo agoGiven their failure on novel logic problems, generation of meaningless text, tendency to do things like delete tests and incompetence at simple mathematics, it seems very unlikely they have built any sort of world model. It’s remarkable how competent they are given the way they work. Predict the next word is a terrible summary of what these machines do though, they certainly do more than that, but there are significant limitations. ‘Reasoning’ etc are marketing terms and we should not trust the claims made by companies who make these models. The Turing test had too much confidence in humans it seems.
- bdhcuidbebe 7mo ago[flagged]
- tomhow 7mo agoWe've banned this account. You can't post vile comments like this, no matter who or what it's about, and we obviously have to ban accounts that do this if we're to have any standards at all. You've been warned before about your style of commenting, so it's not that you don't know the expectations here.
- intended 7mo agoThe article talks about LLMs reviewing Econ papers. I’m hesitant to call this an outright win, though. Perhaps the review service the author is using is really good. Almost certainly the taste, expertise and experience of the author is doing unseen heavy lifting. I found that using prompts to do submission reviews for conferences tended to make my output worse, not better. Letting the LLM analyze submissions resulted in me disconnecting from the content. To the point I would forget submissions after I closed the tab. I ended up going back to doing things manually, using them as a sanity check. On the flip side, weaker submissions using generative tools became a nightmare, because you had to wade through paragraphs of fluff to realize there was no substantive point. It’s to the point that I dread reviewing. I am going to guess that this is relatively useful for experts, who will submit stronger submissions, than novices and journeymen, who will still make foundational errors.
- ruhith 7mo ago[dead]
- pharrington 7mo agoWhy do the deliverables always take about 1 hour? Is this fully automated?
- gammalost 7mo agoIt is really interesting how great and also how terrible LLMs can be at the same time. For example, I had a really annoying bug yesterday, I missed one character, "_". Asking ChatGPT for help led to a lot of feedback that was arguably okay but not currently relevant (because there was a fatal flaw in the code). Remade the conversation with personal information stripped here https://chatgpt.com/share/699fef77-b530-8007-a4ed-c3dda9461d03 https://chatgpt.com/share/699fef77-b530-8007-a4ed-c3dda9461d...
- GodelNumbering 7mo agoIt is probably the first-time aha moment the author is talking about. But under the hood, it is probably not as magical as it appears to be. Suppose you prompted the underlying LLM with "You are an expert reviewer in..." and a bunch of instructions followed by the paper. LLM knows from the training that 'expert reviewer' is an important term (skipping over and oversimplifying here) and my response should be framed as what I know an expert reviewer would write. LLMs are good at picking up (or copying) the patterns of response, but the underlying layer that evaluates things against a structural and logical understanding is missing. So, in corner cases, you get responses that are framed impressively but do not contain any meaningful inputs. This trait makes LLMs great at demos but weak at consistently finding novel interesting things. If the above is true, the author will find after several reviews that the agent they use keeps picking up on the same/similar things (collapsed behavior that makes it good at coding type tasks) and is blind to some other obvious things it should have picked up on. This is not a criticism, many humans are often just as collapsed in their 'reasoning'. LLMs are good at 8 out of 10 tasks, but you don't know which 8.
- Kim_Bruning 7mo agoIn your model, explain the old trick "think step by step"
- GodelNumbering 7mo agoIt simply forces the model to adopt an output style known to conduce systematic thinking without actually thinking. At no point has it through through the thing (unless there are separate thinking tokens)
- modeless 7mo agoIt's clear that in the general case "predict the next word" requires arbitrarily good understanding of everything that can be described with language. That shouldn't be mysterious. What's mysterious is how a simple training procedure with that objective can in practice achieve that understanding. But then again, does it? The base model you get after that simple training procedure is not capable of doing the things described in the article. It is only useful as a starting point for a much more complex reinforcement learning procedure that teaches the skills an agent needs to achieve goals. RL is where the magic comes from, and RL is more than just "predict the next word". It has agents and environments and actions and rewards.
- sasjaws 7mo agoA while ago i did the nanogpt tutorial, i went through some math with pen and paper and noticed the loss function for 'predict the next token' and 'predict the next 2 tokens' (or n tokens) is identical. That was a bit of a shock to me so wanted to share this thought. Basically i think its not unreasonable to say llms are trained to predict the next book instead of single token. Hope this is usefull to someone.
- sputknick 7mo agoI'd like to explore this idea, did you make a blog post about it? is it simple enough to post in the reply?
- WithinReason 7mo agoLook up attention masks
- krackers 7mo agoUnless I've misunderstood the math myself, I don't think GPs comment is quite right if taken literally since "predict the next 2 tokens" would literally mean predict index t+1, t+2 off of the same hidden state at index t, which is the much newer field of multi-token prediction and not classic LLM autoregressive training. Instead what GP likely means is the observation that the joint probability of a token sequence can be broken down autogressively: P(a,b,c) = P(a) * P(b|a) * P(c|a,b) and then with cross-entropy loss which optimizes for log likelihood this becomes a summation. So training with teacher forcing to minimize "next token" loss simultaneously across every prefix of the ground-truth is equivalent to maximizing the joint probability of that entire ground-truth sequence. Practically, even though inference is done one token at a time, you don't do training "one position ahead" at a time. You can optimize the loss function for the entire sequence of predictions at once. This is due the autoregressive nature of the attention computation: if you start with a chunk of text, as it passes through the layers you don't just end up with the prediction for the next word in the last token's final layer, but _all_ of the final-layer residuals for previous tokens will encode predictions for their following index. So attention on a block of text doesn't give you just the "next token prediction" but the simultaneous predictions for each prefix which makes training quite nice. You can just dump in a bunch of text and it's like you trained for the "next token" objective on all its prefixes. (This is convenient for training, but wasted work for inference which is what leads to KV caching). Many people also know by now that attention is "quadratic" in nature (hidden state of token i attends to states of tokens 1...i-1), but they don't fully grasp the implication that even though this means for forward inference you only predict the "next token", for backward training this means that error for token i can backpropagate to tokens 1...i-1. This is despite the causal masking, since token 1 doesn't attend to token i directly but the hidden state of token 1 is involved in the computation of the residual stream for token i. When it comes to the statement >its not unreasonable to say llms are trained to predict the next book instead of single token. You have to be careful, since during training there is no actual sampling happening. We've optimized to maximize the joint probability of ground truth sequence, but this is not the same as maximizing the probability the the ground truth is generated during sampling. Consider that there could be many sampling strategies: greedy, beam search, etc. While the most likely next token is the "greedy" argmax of the logits, the most likely next N tokens is not always found by greedily sampling N times. It's thought that this is one reason why RL is so helpful, since rollouts do in fact involve sampling so you provide rewards at the "sampled sequence" level which mirrors how you do inference. It would be right to say that they're trained to ensure the most likely next book is assigned the highest joint probability (not just the most likely next token is assigned highest probability).
- Alex_L_Wood 7mo ago[flagged]
- rossant 7mo agoDon't assume bad intent. I use LLMs for different research-related tasks and I surely can relate. In the past few months, the latest models have become better than me at many tasks. And I am not an ad.
- tsunamifury 7mo agoI think it’s funny that at Google I invented and productized next word (and next action) predictor in Gmail and hangouts chat and I’ve never had a single person come to me and ask how this all works. To me LLMs are incredibly simple. Next word next sentence next paragraph and next answer are stacked attention layers which identify manifolds and run in reverse to then keep the attention head on track for next token. It’s pretty straight forward math and you can sit down and make a tiny LLM pretty easily on your home computer with a good sized bag of words and context To me it’s baffling everyone goes around saying constantly that not even Nobel prize winners know how this works it’s a huge mystery. Has anyone thought to ask the actual people like me and others who invented this?
- booleandilemma 7mo agoA lot of people in tech thrive on the mystery and don't like explaining things in simple terms. It makes what they do seem more valuable if no one can understand what they're talking about. At the same time, being vague and mysterious can help hide someone's own misunderstandings. When you speak clearly you need to be accurate, because it's more obvious when you're wrong.
- tsunamifury 7mo agoI agree -- or the math is just way over peoples heads -- even word points to word N times.
- kosh2 7mo agoThis is like saying quantum mechanics is really simple to understand, all you have to do is find the right formula and plug in the numbers. When people talk about understanding, they mean as knowing how the underlying mechanism works often by finding an analog in real life.
- tsunamifury 7mo agoIt is a sophisticated way of putting your foot in front of you and taking a step while keeping your head up and looking at your destination.
- teekert 7mo agoI think this is a thing not often discussed here, but I too have this experience. An LLM can be fantastic if you write a 25-pager then later need to incorporate a lot of comments with sometimes conflicting arguments/viewpoints. LLMs can be really good at "get all arguments against this", "Incorporated this view point in this text while making it more concise.", "Are these views actually contradicting or can I write it such that they align. Consider incentives". If you know what you're doing and understand the matter deeply (and that is very important) you'll find that the LLM is sometimes better at wording what you actually mean, especially when not writing in your native language. Of course, you study the generated text, make small changes, make it yours, make sure you feel comfortable with it etc. But man can it get you over that "how am I going to write this down"-hump. Also: "Make an executive summary" "Make more concise", are great. Often you need to de-linkedIn the text, or tell it to "not sound like an American waiter", and "be business-casual", "adopt style of rest of doc", etc. But it works wonders.
- mrorigo 7mo agoAttention is all you need.
- deleted 7mo ago[deleted]
- deleted 7mo ago[deleted]
- throawayonthe 7mo ago[flagged]
- trhway 7mo ago>I don’t know how you get here from “predict the next word.” The question puts horse behind the buggy. The main point isn't "from", it is how you get to “predict the next word.” During the training the LLM builds inside itself compressed aggregated representation - a model - of what is fed into it. Giving the model you can "predict the next word" as well as you can do a lot of other things. For simple starting point for understanding i'd suggest to look back at the key foundational stone that started it all - "sentiment neuron" https://openai.com/index/unsupervised-sentiment-neuron/ https://openai.com/index/unsupervised-sentiment-neuron/ "simply predicting the next character in Amazon reviews resulted in discovering the concept of sentiment. ... Digging in, we realized there actually existed a single “sentiment neuron” that’s highly predictive of the sentiment value."
- asymmetric 7mo agoThe ideas in the update were previously explored by Gwern 2 years ago: https://www.lesswrong.com/posts/PQaZiATafCh7n5Luf/gwern-s-shortform?commentId=KAtgQZZyadwMitWtb https://www.lesswrong.com/posts/PQaZiATafCh7n5Luf/gwern-s-sh...
- gwern 7mo agoSpecifically, Cochrane wrote: > On reflection I have started to worry again. In 10 to 20 years nobody will read anything any more, they just will read LLM digests. So, the single most important task of a writer starting right now is to get your efforts wired in to the LLMs. Nothing you write will matter if it is not quickly adopted to the training dataset. As the art of pushing your results to the top of the google search was the 1990s game, getting your ideas into the LLMs is today’s. Refine is no different. It’s so good, everyone will use it. So whether refine and its cousins take a FTPL or new Keynesian view in evaluating papers is now all determining for where the consensus of the profession goes. For more recent comments, see https://dwarkesh.com/p/gwern-branwen https://dwarkesh.com/p/gwern-branwen https://gwern.net/llm-writing https://gwern.net/llm-writing https://www.lesswrong.com/posts/34J5qzxjyWr3Tu47L/is-building-good-note-taking-software-an-agi-complete?commentId=WW2uRJdonqEw9krqm#WW2uRJdonqEw9krqm https://www.lesswrong.com/posts/34J5qzxjyWr3Tu47L/is-buildin... https://gwern.net/blog/2025/ai-cannibalism https://gwern.net/blog/2025/ai-cannibalism https://gwern.net/blog/2025/good-ai-samples https://gwern.net/blog/2025/good-ai-samples https://gwern.net/style-guide https://gwern.net/style-guide The scaling will continue until morale improves. I advise people to skate to where the puck will be, and to ask themselves: "if I knew for a fact that LLMs could do something I am doing in 1-2 years, would I still want to do it? If not, what should I be doing now instead?"
- keybored 7mo ago> On reflection I have started to worry again. In 10 to 20 years nobody will read anything any more, they just will read LLM digests. So, the single most important task of a writer starting right now is to get your efforts wired in to the LLMs. Nothing you write will matter if it is not quickly adopted to the training dataset. As the art of pushing your results to the top of the google search was the 1990s game, getting your ideas into the LLMs is today’s. Refine is no different. It’s so good, everyone will use it. So whether refine and its cousins take a FTPL or new Keynesian view in evaluating papers is now all determining for where the consensus of the profession goes. I expected a bit more cynicism and less merrily going with the downward spiral flow from a “grumpy” blogger. Oh, but it’s a grumpy economist.
- singularity2001 7mo agoSuperhuman chess engines are now trained just from one bit reward signal: win / lose. This says absolutely nothing about the complexity that the model develops inside. They even learned the rule of the games just from that reward.
- vb7132 7mo agoIMO, the writer is overzealous with their comments on LLMs. As a coder, it feels like an outsider trying out a product that was amazed me over and over so many times. > They aren’t perfect, but the kind of analysis the program is able to do is past the point where technology looks like magic. But as you use this product over a long period of time, there are many obvious gaps - hallucinations / repeated tool calls / out of context outputs / etc. To me, refine.ink sounds like a company that has built heavy tooling around some super high context window LLMs and then some very good prompts. Their claim is to compare it against any good off-the-shelf LLM with any prompt. But when you are spending bunch of money to build a whole ecosystem around LLMs, it's obvious that it's not going to beat their output. I won't be surprised if the next version of an LLM within the next few months completely outperforms their output -- that's usually the case with all the coding tools and scaffoldings. They are rendered useless by a superior LLM.
- TYPE_FASTER 7mo agoThe best overview I've found so far of how LLMs work: https://www.youtube.com/watch?v=7xTGNNLPyMI https://www.youtube.com/watch?v=7xTGNNLPyMI
- joquarky 7mo agoThe links on this site are low contrast.