20 ms·
LLM Daydreaming
- zwaps 1y agoWasn't this already implemented in some agents? I want to remember I heard about it in several podcasts
- johnfn 1y agoIt's an interesting premise, but how many people - are capable of evaluating the LLM's output to the degree that they can identify truly unique insights - are prompting the LLM in such a way that it could produce truly unique insights I've prompted an LLM upwards of 1,000 times in the last month, but I doubt more than 10 of my prompts were sophisticated enough to even allow for a unique insight. (I spend a lot of time prompting it to improve React code.) And of those 10 prompts, even if all of the outputs were unique, I don't think I could have identified a single one. I very much do like the idea of the day-dreaming loop, though! I actually feel like I've had the exact same idea at some point (ironic) - that a lot of great insight is really just combining two ideas that no one has ever thought to combine before.
- cantor_S_drug 1y ago> are capable of evaluating the LLM's output to the degree that they can identify truly unique insights I noticed one behaviour in myself. I heard about a particular topic, because it was a dominant opinion in the infosphere. Then LLMs confirmed that dominant opinion (because it was heavily represented in the training) and I stopped my search for alternative viewpoints. So in a sense, LLMs are turning out to be another reflective mirror which reinforces existing opinion.
- MrScruff 1y agoYes, it seems like LLMs are system one thinking taken to the extreme. Reasoning was supposed to introduce some actual logic but you only have to play with these models for a short while to see that the reasoning tokens are a very soft constraint on the models eventual output. Infact, they're trained to please us and so in general aren't very good at pushing back. It's incredibly easy to 'beat' an LLM in an argument since they often just follow your line of reasoning (it's in the models context after all).
- cantor_S_drug 1y agoThis is also true in a sense nuance will dropped in the compression mechanism and overrepresentation in the training data will get more weightage to be retained.
- zyklonix 1y agoTotally agree, most prompts (especially for code) aren’t designed to surface novel insights, and even when they are, it’s hard to recognize them. That’s why the daydreaming loop is so compelling: it offloads both the prompting and the novelty detection to the system itself. Projects like https://github.com/DivergentAI/dreamGPT https://github.com/DivergentAI/dreamGPT are early steps in that direction, generating weird idea combos autonomously and scoring them for divergence, without user prompting at all.
- apples_oranges 1y agoIf the breakthrough comes, most if not all links on HN will be to machine generated content. But so far it seems that the I in current AI is https://www.youtube.com/watch?v=uY4cVhXxW64 https://www.youtube.com/watch?v=uY4cVhXxW64 ..
- NitpickLawyer 1y agoSomething I haven't seen explored, but I think could perhaps help is to somehow introduce feedback regarding the generation into the context, based on things that are easily computed w/ other tools (like perplexity). In "thinking" models we see a lot of emerging behaviour like "perhaps I should, but wait, this seems wrong", etc. Perhaps adding some signals at regular? intervals could help in surfacing the correct patterns when they are needed. There's a podcast I listened to ~1.5 years ago, where a team used GPT2, further trained on a bunch of related papers, and used snippets + perplexity to highlight potential errors. I remember them having some good accuracy when analysed by humans. Perhaps this could work at a larger scale? (a sort of "surprise" factor)
- aredox 1y agoOh, in the middle of "AI is PhD-level" propaganda (just check Google News to see this is not a strawman argument), some people finally admit in passing "no LLM has ever made a breakthrough". (See original argument: https://nitter.net/dwarkesh_sp/status/1727004083113128327 https://nitter.net/dwarkesh_sp/status/1727004083113128327 )
- bonoboTP 1y agoI agree there's an equivocation going on for "PhD level" between "so smart, it could get a PhD" (as in come up with and publish new research and defend its own thesis) and "it can solve quizzes at the level that PhDs can".
- washadjeffmad 1y agoServices that make this claim are paying people with PhDs to ask their models questions and then provide feedback on the responses with detailed reasoning.
- NetRunnerSu 1y ago[dead]
- ashdksnndck 1y agoI’m not sure we can accept the premise that LLMs haven’t made any breakthroughs. What if people aren’t giving the LLM credit when they get a breakthrough from it? First time I got good code out of a model, I told my friends and coworkers about it. Not anymore. The way I see it, the model is a service I (or my employer) pays for. Everyone knows it’s a tool that I can use, and nobody expects me to apportion credit for whether specific ideas came from the model or me. I tell people I code with LLMs, but I don’t commit a comment saying “wow, this clever bit came from the model!” If people are getting actual bombshell breakthroughs from LLMs, maybe they are rationally deciding to use those ideas without mentioning the LLM came up with it first. Anyway, I still think Gwern’s suggestion of a generic idea-lab trying to churn out insights is neat. Given the resources needed to fund such an effort, I could imagine that a trading shop would be a possible place to develop such a system. Instead of looking for insights generally, you’d be looking for profitable trades. Also, I think you’d do a lot better if you have relevant experts to evaluate the promising ideas, which means that more focused efforts would be more manageable. Not comparing everything to everything, but comparing everything to stuff in the expert’s domain. If a system like that already exists at Jane Street or something, I doubt they are going to tell us about it.
- Yizahi 1y agoThis is bordering conspiracy theory. Thousands of people are getting novel breakthroughs generated purely by LLM an not a single person discloses such result? Not even one of the countless LLM corporation engineers who depend on the billion dollar IV injections from deluded bankers just to continue surviving, and not one has bragged about LLM doing that revolution? Hard to believe.
- esafak 1y agoCountless people are increasing their productivity and talking about it here ad nauseam. Even researchers are leaning on language models; e.g., https://mathstodon.xyz/@tao/114139125505827565 https://mathstodon.xyz/@tao/114139125505827565 We haven't successfully resolved famous unsolved research problems through language models yet but one can imagine that they will solve increasingly challenging problems over time. And if it happens in the hands of a researcher rather than model's lab, one can also imagine that the researcher will take credit, so you will still have the same question.
- blueflow 1y agoI have not yet seen AI doing a critical evaluation of data sources. AI willcontradict primary sources if the contradiction is more prevalent in the training data. Something about the whole approach is bugged. My pet peeve: "Unix System Resources" as explanation for the /usr directory is a term that did not exist until the turn of the millenium (rumor is that a c't journalist made it up in 1999), but AI will retcon it into the FHS (5 years earlier) or into Ritchie/Thompson/Kernigham (27 years earlier).
- _heimdall 1y ago> Something about the whole approach is bugged. The bug is that LLMs are fundamentally designed for natural language processing and prediction, not logic or reasoning. We may get to actual AI eventually, but an LLM architecture either won't be involved at all or it will act as a part of the system mimicking the language center of a brain.
- zhangjunphy 1y agoI also hope we have something like this. But sadly, this is not going to work. The reason is this line from the article, which is so much harder that it looks: > and a critic model filters the results for genuinely valuable ideas. In fact, people have tryied this idea. And if you use a LLM or anything similar as the critic, the performance of the model actually degrades in this process. As the LLM tries too hard to satisfy the critic, and the critic itself is far from a good reasoner. So the reason that we don't hear too much about this idea is not that nobody tried it. But that they tried, and it didn't work, and people are reluctant to publish about something which does not work.
- imiric 1y agoExactly. This not only affects a potential critic model, but the entire concept of a "reasoning" model is based on the same flawed idea—that the model can generate intermediate context to improve its final output. If that self-generated context contains hallucinations, baseless assumptions or doubt, the final output can only be an amalgamation of that. I've seen the "thinking" output arrive at a correct solution in the first few steps, but then talk itself out of it later. Or go into logical loops, without actually arriving at anything. The reason why "reasoning" models tend to perform better is simply due to larger scale and better training data. There's nothing inherently better about them. There's nothing intelligent either, but that's a separate discussion.
- yorwba 1y agoReasoning models are trained from non-reasoning models of the same scale, and the training data is the output of the same model, filtered through a verifier. Generating intermediate context to improve the final output is not an idea that reasoning models are based on, but an outcome of the training process. Because empirically it does produce answers that pass the verifier more often if it generates the intermediate steps first. That the model still makes mistakes doesn't mean it's not an improvement: the non-reasoning base model makes even more mistakes when it tries to skip straight to the answer.
- 1y ago
- jumploops 1y agoHow do you critique novelty? The models are currently trained on a static set of human “knowledge” — even if they “know” what novelty is, they aren’t necessarily incentivized to identify it. In my experience, LLMs currently struggle with new ideas, doubly true for the reasoning models with search. What makes novelty difficult, is that the ideas should be nonobvious (see: the patent system). For example, hallucinating a simpler API spec may be “novel” for a single convoluted codebase, but it isn’t novel in the scope of humanity’s information bubble. I’m curious if we’ll have to train future models on novelty deltas from our own history, essentially creating synthetic time capsules, or if we’ll just have enough human novelty between training runs over the next few years for the model to develop an internal fitness function for future novelty identification. My best guess? This may just come for free in a yet-to-be-discovered continually evolving model architecture. In either case, a single discovery by a single model still needs consensus. Peer review?
- n4r9 1y agoIt's a good question. A related question is: "what's an example of something undeniably novel?". Like if you ask an agent out of the blue to prove the Collatz conjecture, and it writes out a proof or counterexample. If that happens with LLMs then I'll be a lot more optimistic about the importance to AGI. Unfortunately, I suspect it will be a lot murkier than that - many of these big open questions will get chipped away at by a combination of computational and human efforts, and it will be impossible to pinpoint where the "novelty" lies.
- jacobr1 1y agoGood point. Look at patents. Few are truly novel in some exotic sense of "the whole idea is something never seen before." Most likely it is a combination of known factors applied in a new way, or incremental development improving on known techniques. In a banal sense, most LLM content generated is novel, in that the specific paragraphs might be unique combinations of words, even if the ideas are just slightly rearranged regurgitations. So I strongly agree that, especially when are talking about the bulk of human discovery and invention, the incrementalism will be increasingly in striking distance of human/AI collaboration. Attribution of the novelty in these cases is going to be unclear, when the task is, simplified something like, "search for combinations of things, in this problem domain, that do the task better than some benchmark" be that drug discovery, maths, ai itself or whatever.
- OtherShrezzing 1y agoGoogle's effort with AlphaEvolve shows that the Daydream Factory approach might not be the big unlock we're expecting. They spent an obscene amount of compute to discover a marginal improvement over the state of the art in a very narrow field. Hours after Google published the paper, mathematicians pointed out that their SOTA algorithms underperformed compared to techniques published in the 50 years ago. Intuitively, it doesn't feel like scaling up to "all things in all fields" is going to produce substantial breakthroughs, if the current best-in-class implementation of the technique by the worlds leading experts returned modest results.
- khalic 1y agoUgh, again with the anthropomorphizing. LLMs didn't come up with anything new because _they don't have agency_ and _do not reason_... We're looking at our reflection and asking ourselves why it isn't moving when we don't
- yorwba 1y agoIf you look at your reflection in water, it may very well move even though you don't. Similarly, you don't need agency or reasoning to create something new, random selection from a large number of combinations is enough, correct horse battery staple. Of course random new things are typically bad. The article is essentially proposing to generate lots of them anyway and try to filter for only the best ones.
- RALaBarge 1y agoI agree that brute forcing is a method and how nature does it. The problem would still be the same, how would it or other LLMs know if the idea is novel and interesting? Given access to unlimited data, LLMs likely could spot novel trends that we cant but still cant judge the value of creating something unique that it has never encountered before.
- RALaBarge 1y agoYet.
- amelius 1y ago> anthropomorphizing Gwern isn't doing that here. They say: "[LLMs] lack some fundamental aspects of human thought", and then investigates that.
- cranium 1y agoI'd be happy to spend my Claude Max tokens during the night so it can "ultrathink" some Pareto improvements to my projects. So far, I've mostly seen lateral moves that rewrites code rather than rearchitecture/design the project.
- precompute 1y agoVariations on increasing compute and filtering results aside, the only way out of this rut is another breakthrough as big, or bigger than transformers. A lot of money is being spent on rebranding practical use-cases as innovation because there's severe lack of innovation in this sphere.
- pilooch 1y agoAlphaEvolve and similar systems based on map-elites + DL/LLM + RL appears to be one of the promising paths. Setting up the map-elites dimensions may still be problem-specific but this could be learnt unsupervisedly, at least partially. The way I see LLMs is as a search-spqce within tokens that manipulate broad concepts within a complex and not so smooth manifold. These concepts can be refined within other spaces (pixel -space, physical spaces, ...)
- guelo 1y agoIn a recent talk [0] Francois Chollet made it sound like all the frontier models are doing Test-Time Adaptation, which I think is a similar concept to Dynamic evaluation that Gwern says is not being done. Apparently Test-Time Adaptation encompasses several techniques some of which modify model weights and some that don't, but they are all about on-the-fly learning. [0] https://www.youtube.com/watch?v=5QcCeSsNRks&t=1542s https://www.youtube.com/watch?v=5QcCeSsNRks&t=1542s
- anupj 1y ago[dead]
- LourensT 1y agoRegardless of accusations of anthropomorphizing, continual thinking seems to be a precursor to any sense of agency, simply because agency requires something to be running. Eventually LLM output degrades when most of the context is its own output. So should there also be an input stream of experience? The proverbial "staring out the window", fed into the model to keep it grounded and give hooks to go off?
- daxfohl 1y agoThough it kind of reminds me of The Shining (too much time just thinking drives one to insanity). It seems like we need to evolve intelligence, perception, and agency very closely in tandem, or an imbalance in any will send it off the rails.
- amelius 1y agoHumans daydream about problems when they think a problem is interesting. Can an LLM know when a problem is interesting and thereby prune the daydream graph?
- zild3d 1y ago> The puzzle is why The feedback loop on novel/genuine breakthroughs is too long and the training data is too small. Another reason is that there's plenty of incentive to go after the majority of the economy which relies on routine knowledge and maybe judgement, a narrow slice actually requires novel/genuine breakthroughs.
- cs702 1y agoThe question is: How do we get LLMs to have "Eureka!" moments, on their own, when their minds are "at rest," so to speak? The OP's proposed solution is a constant "daydreaming loop" in which an LLM is does the following on its own, "unconsciously," as a background task, without human intervention: 1) The LLM retrieves random facts. 2) The LLM "thinks" (runs a chain-of-thought) on those retrieved facts to see if they are any interesting connections between them. 3) If the LLM finds interesting connections, it promotes them to "consciousness" (a permanent store) and possibly adds them to a dataset used for ongoing incremental training. It could work.
- NetRunnerSu 1y ago[dead]
- epcoa 1y agoThe step 3 has been shown to not work over and over again, the “find interesting connections” is the hand wavy magic at this time. LLMs alone don’t seem to be particularly adept at it either.
- cs702 1y agoHas this been tried with reinforcement learning (RL)? As the OP notes, it is plausible from a RL perspective that such a bootstrap can work, because it would be (quoting the OP) "exploiting the generator-verifier gap, where it is easier to discriminate than to generate (eg laughing at a pun is easier than making it)." The hit ratio may be tiny, so doing this well would be very expensive.
- epcoa 1y agoRun ML of any combination and form in a for loop for higher order is one of the most obvious avenues. If it worked you would have heard about it a long time ago.
- varelse 1y ago[dead]
- kookamamie 1y ago> The puzzle is why The breakthrough isn't in their datasets.
- velcrovan 1y agoI’m once again begging people to read David Gelernter’s 1994 book “The Muse in the Machine”. I’m surprised to see no mention of it in Gwern’s post, it’s the exact book he should be reaching for on this topic. In examining the possibility of genuinely creative computing, Gelernter discovers and defends a model of cognition that explains so much about the human experience of creativity, including daydreaming, dreaming, everyday “aha” moments, and the evolution of human approaches to spirituality. https://uranos.ch/research/references/Gelernter_1994/Muse%20In%20The%20Machine.pdf https://uranos.ch/research/references/Gelernter_1994/Muse%20...
- sneak 1y agoSeems like an easy hypothesis to quickly smoke test with a couple hundred lines of script, a wikipedia index, and a few grand thrown at an API.
- dr_dshiv 1y agoYes! I’ve been prototyping dreaming LLMs based on my downloaded history—and motivated by biomimetic design approaches. Just to surface ideas to myself again.
- A_D_E_P_T 1y ago> You are a creative synthesizer. Your task is to find deep, non-obvious, and potentially groundbreaking connections between the two following concepts. Do not state the obvious. Generate a hypothesis, a novel analogy, a potential research question, or a creative synthesis. Be speculative but ground your reasoning. > Concept 1: {Chunk A} > Concept 2: {Chunk B} In addition to the other criticisms mentioned by posters ITT, a problem I see is: What concepts do you feed it? Obviously there's a problem with GIGO. If you don't pick the right concepts to begin with, you're not going to get a meaningful result. But, beyond that, human discovery (in mechanical engineering, at least,) tends to be massively interdisciplinary and serendipitous, so that many concepts are often involved, and many of those are necessarily non-obvious. I guess you could come up with a biomimetics bot, but, besides that, I'm not so sure how well this concept would work as laid out above. There's another issue in that LLMs tend to be extremely gullible, and swallow the scientific literature and University press releases verbatim and uncritically.
- sartak 1y agoFrom The Metamorphosis of Prime Intellect (1994): > Among Prime Intellect's four thousand six hundred and twelve interlocking programs was one Lawrence called the RANDOM_IMAGINATION_ENGINE. Its sole purpose was to prowl for new associations that might fit somewhere in an empty area of the GAT. Most of these were rejected because they were useless, unworkable, had a low priority, or just didn't make sense. But now the RANDOM_IMAGINATION_ENGINE made a critical connection, one which Lawrence had been expecting it to make [...] > Deep within one of the billions of copies of Prime Intellect, one copy of the Random_Imagination_Engine connected two thoughts and found the result good. That thought found its way to conscious awareness, and because the thought was so good it was passed through a network of Prime Intellects, copy after copy, until it reached the copy which had arbitrarily been assigned the duty of making major decisions -- the copy which reported directly to Lawrence. [...] > "I've had an idea for rearranging my software, and I'd like to know what you think." > At that Lawrence felt his blood run cold. He hardly understood how things were working as it was; the last thing he needed was more changes. "Yes?"
- js8 1y agoI am not sure why tie this to any concrete AI technology such as LLMs. IMHO the biggest issue we have with AI right now is that we don't know how to philosophicaly formalize what we want. What is reasoning? I am trying to answer that for myself. Since every logic is expressible in untyped lambda calculus (as any computation is), you could have a system that just somehow generates terms and beta-reduces them. In even so much simpler logic, what are the "interesting" terms? I have several answers, but my point is, you should simplify the problem and this question has not been answered even under such simple scenario.
- HarHarVeryFunny 1y agoReasoning is chained what-if prediction, together with exploration of alternatives (cf backtracking), and leans upon general curiosity/learning for impasse resolution (i.e. if you can't predict what-if, then have the curiosity to explore and find out). What the LLM companies are currently selling as "reasoning" is mostly RL-based pre-training whereby the model is encouraged to predict tokens (generate reasoning steps) according to similar "goals" seen in the RL training data. This isn't general case reasoning, but rather just "long horizon" prediction based on the training data. It helps exploit the training data, but isn't going to generate novelty outside of the deductive closure of the training data.
- js8 1y agoI am talking about reasoning in philosophical not logical sense. In your definition, you're assuming a logic in which reasoning happens, but when I am asking the question, I am not presuming any specific logic. So how do you pick the logic in which to do reasoning? There are "good reasons" to use one logic over another. LLMs probably learn some combination of logic rules (deduction rules in commonly used logics), but cannot guarantee they will be used consistently (i.e. choose a logic for the problem and stick to it). How do you accomplish that? And even then reasoning is more than search. If you can reason, you should also be able to reason about more effective reasoning (for example better heuristics to cutting the search tree).
- 1y ago
- deleted 1y ago[deleted]
- HarHarVeryFunny 1y ago> Despite impressive capabilities, large language models have yet to produce a genuine breakthrough. The puzzle is why. I don't see why this is remotely surprising. Despite all the hoopla, LLMs are not AGI or artifical brains - they are predict-next-word language models. By design they are not built for creativity, but rather quite the opposite, they are designed to continue the input in the way best suggested by the training data - they are essentially built for recall, not creativity. For an AI to be creative it needs to have innate human/brain-like features such as novelty (prediction failure) driven curiosity, boredom, as well as ability to learn continuously. IOW if you want the AI to be creative it needs to be able to learn for itself, not just regurgitate the output of others, and have these innate mechanisms that will cause it to pursue discovery.
- NetSurferSu 1y ago[dead]
- karmakaze 1y agoYes LLMs choose probable sequences because they recognize similarity. Because of that, it can diverge from similarity to be creative: increase the temperature. What LLMs don't have is (good) taste—we need to build an artificial tongue and feed it as a prerequisite.
- grey-area 1y agoWell they also don't have understanding, a model of the world, and the ability to reason (no chain-of-thought created by AI companies is not reasoning), as well as having no taste. So there is quite a lot missing.
- HarHarVeryFunny 1y agoIt depends on what you mean by "creative" - they can recombine fragments of training data (i.e. apply generative rules) in any order - generate the deductive closure of the training set, but that is it. Without moving beyond LLMs to a more brain-like cognitive architecture, all you can do is squeeze the juice out of the training data, by using RL/etc to bias the generative process (according to reasoning data, good taste or whatever), but you can't move beyond the training data to be truly creative.
- _acco 1y agoThis is a good way of framing that we don't understand human creativity. And that we can't hope to build it until we do. i.e. AGI is a philosophical problem, not a scaling problem. Though we understand them little, we know the default mode network and sleep play key roles. That is likely because they aid some universal property of AGI. Concepts we don't understand like motivation, curiosity, and qualia are likely part of the picture too. Evolution is far too efficient for these to be mere side effects. (And of course LLMs have none of these properties.) When a human solves a problem, their search space is not random - just like a chess grandmaster's search space of moves is not random. How our brains are so efficient when problem solving while also able to generate novelty is a mystery.
- vintagedave 1y ago> Hypothesis: Day-Dreaming Loop This mirrors something I have thought of too. I have read multiple theories of emerging consciousness, which touch on things from proprioception to the inner monologue (which not everyone has.) My own theory is that -- avoiding the need for an awareness of a monologue -- a LLM loop that constantly takes input and lets it run, saving key summarised parts to memory that are then pulled back in when relevant, would be a very interesting system to speak to. It would need two loops: the constant ongoing one, and then for interaction, one accessing memories from the first. The ongoing one would be aware of the conversation. I think it would be interesting to see what, via the memory system, would happen in terms of the conversation emitting elements from the loop. My theory is that if we're likely to see emergent consciousness, it will come through ongoing awareness and memory.
- yahoozoo 1y agoI once asked ChatGPT to come up with a novel word that would return 0 Google search results. It came up with “vexlithic” which does indeed return 0 results, at least for me. I thought that was neat.
- nsedlet 1y agoI believe an important reason for why there are no LLM breakthroughs is that humans make progress in their thinking through experimentation, i.e. collecting targeted data, which requires exerting agency on the real world. This isn't just observation, it's the creation of data not already in the training set.
- haolez 1y agoMaybe also the fact that they can't learn small pieces of new information without "formatting" its whole brain again, from scratch. And fine tuning is like having a stroke, where you get specialization by losing cognitive capabilities.
- ramoz 1y agoIve walked 10k steps everyday the past week and produced more code in that period than most would over months. Using Claude Code (and vibetunnel over tailscale to my phone- that I speak instructions into). There is a breakthrough happening. in real time.
- CaptainFever 1y agoCan we see an example please?
- zyklonix 1y agoThis idea of a “daydreaming loop” hits on a key LLM gap, the lack of background, self-driven insight. A pragmatic step in this direction is https://github.com/DivergentAI/dreamGPT https://github.com/DivergentAI/dreamGPT , which explores divergent thinking by generating and scoring hallucinations. It shows how we might start pushing LLMs beyond prompt-response into continuous, creative cognition.
- zby 1y agoThe novelty part is a hard one - but maybe in many cases we could substitute something else for it? If an idea promises to beat state of the art in some field - and it is not yet actively researched - then it is novel. But most promising would be to use the Dessalles theories. Here is 4.1o expanding this: https://chatgpt.com/s/t_6877de9faa40819194f95184979b5b44 https://chatgpt.com/s/t_6877de9faa40819194f95184979b5b44 By the way - this could be a classic example of this day dreaming - you take two texts: one by Gwern and some article by Dessalles (I read "Why we talk" - a great book! - but maybe there is some more concise article?) and ask LLM to generate ideas connecting these two. In this particular case it was my intuition that connected them - but I imagine that there could be an algorithm that could find this connection in a reasonable time - some kind of semantic search maybe.
- throwaway328 1y agoThe fact that LLMs haven't come up with anything "novel" would be a serious puzzle - as the article claims - only if they were thinking, reasoning, being creative, etc. If they aren't doing anything of the sort, it'd be the only thing you'd expect. So it's a bit of an anti-climactic solution to the puzzle but: maybe the naysayers were right and they're not thinking at all, or doing any of the other anthropomorphic words being marketed to users, and we've simply all been dragged along by a narrative that's very seductive to tech types (the computer gods will rise!). It'd be a boring outcome, after the countless gallons of digital ink spilled on the topic the last years, but maybe they'll come to be accepted as "normal software", and not god-like, in the end. A medium to large improvement in some areas, and anywhere from minimal to pointless to harmful in others. And all for the very high cost of all the funding and training and data-hoovering that goes in to them, not to mention the opportunity cost of all the things we humans could have been putting money into and didn't.
- aaron695 1y ago[dead]
- krunck 1y agoI've said before that until these "AI" systems become always-on, always-thinking, always-processing, progress is stuck. The current push-button AI - meaning it only processes when we prompt it - is not how the kind of AI that everyone is dreaming of needs to function. ( https://news.ycombinator.com/item?id=44423983#44426438 https://news.ycombinator.com/item?id=44423983#44426438 )
- bwfan123 1y agoA genuine new insight is a "theory" which is a deterministic symbolic representation of the chaotic worlds of meaning. A theory is a formalism which can explain and predict phenomena. A theory consists of symbols to represent concepts, and rules to associate symbols. for example, euclidean geometry organizes the perceptions of "space" around us, or newtonian mechanics organizes the perceptions of "motion" or turing machine organizes computation. LLMs cannot create theories yet since they are missing the capability to create symbols to represent things and link them via rules.
- rotexo 1y agoI feel like this approach might be relevant to achieve the equivalent of “go sit in the corner and think about what you’ve done.” Pull out two random poorly-rated chatbot conversations and have the model look for commonalities in its shortcomings.