15 ms·
LLM Powered Autonomous Agents
- gremlinsinc 3y agowow, this was a brilliant summary of much research into ai agents, I've been reading a lot about these and following this stuff, but learned a lot.
- ChatGTP 3y agoHowever, the reliability of model outputs is questionable, as LLMs may make formatting errors and occasionally exhibit rebellious behavior (e.g. refuse to follow an instruction). Right…sounds quite reckless?
- RamblingCTO 3y agoNo, it just misunderstands the request in the sense that it mismaps input and expected out.
- swyx 3y ago> Challenges in long-term planning and task decomposition: Planning over a lengthy history and effectively exploring the solution space remain challenging. LLMs struggle to adjust plans when faced with unexpected errors, making them less robust compared to humans who learn from trial and error. While working on smol-developer I also eventually landed on the importance of planning as mentioned in my Agents writeup https://www.latent.space/p/agents https://www.latent.space/p/agents . I feel some hesitation with suggesting this because it suggests I'm not deep-learning-pilled, but I really wonder how far next-token-prediction can go with planning. When I think about planning I think about mapping out a possibility space, identifying trees of dependencies, assigning priorities, and then solving for some kind of weighted shortest path. That's an awful lot of work to expect of a next token predictor (goodness knows its scaled far beyond what anyone thought - is there any limit to next token prediction?). If there were one focus area for GPT-5, my money would be on a better architecture capable of planning.
- avereveard 3y agoI feel this is more of a limit from single scratch pad agents Having a tasking agent refining prompts seems a much robust approach. Bonus if it can decompose the required data properly. I'm working on a world database prompt with limited success, but implementing the retrieval with prompts is time consuming and I've not much time. The idea is along the line of "translate the user question in SQL, you have a database of all known facts with a table for each entity" and then you handle the retrieval of each table data step by step. I think then you'd be able to use the postgres planner as planner, but while I've seen project of postgres retrieving data from random sources I don't remember the specifics.
- friendzis 3y ago> When I think about planning I think about mapping out a possibility space, identifying trees of dependencies, assigning priorities, and then solving for some kind of weighted shortest path. but that's two things: planning and execution of a plan. The plan is the dependency graph (!), assignment of priorities and shortest path is execution. This is very important in the context of agents, autonomous or not. If you want the agent to self-correct it has to understand that there can be multiple start points or multiple end points (or both) to backtrack and pivot. And as long as it is glorified "next token predictor" it cannot really do that. Of course some tasks indeed are ifttt style linear-ish flows where next token prediction may prove to be adequate. However, if your agent is incapable of understanding non-linear flows, can it reasonably back off from one true way hen faced with such flow?
- niemandhier 3y agoI think that in the end predicting words is non optimal , most things we want to do are things that are related to the internal representation of concepts that exist deeper in the layers. I at least do not want to predict the next token, I want to predict the next concept in a chain of reasoning, but it seems that currently we are stuck at using the same representation for autoregression we use for training. Maybe we can come up with a better way to construct these chains once we understand the models better. Edit: Typo
- 3y ago
- Xen9 3y agoI predict this type of cognitive engineering of LLM-multiagents to become a thing.
- novaRom 3y agoHow small can be a LLM transformer in order to be able to understand basic human language and search for answers on the internet? It should not contain all the facts and knowledge, but must be quick (so, it's a small model), understand at least one language, and know how and where to look for answers. Would it be sufficient to have 1B, 3B or 7B parameters to achieve this? Or is it doable with 100M or even fewer parameters? I mean vocabulary size might be quite small, max context size could also be limited to 256 or 512 tokens. Is there any paper on that maybe?
- minimaxir 3y agoYou don't need a 1B+ parameter model for this workflow, but it helps, particularly for the quality of the final output. > know how and where to look for answers The point of calling them Agents is that they don't know how. All the examples, including AutoGPT that makes AI influencers go OMG AGI, are operating in a discrete action space with user-specified hints to select which action (or none at all).
- quickthrower2 3y agoI like the idea that a small model could fine tune itself for each task more affordably, so that could be an advantage.
- lhl 3y agoA team at Microsoft Research asked the same question and just published a paper about part of that at least: TinyStories: How Small Can Language Models Be and Still Speak Coherent English? https://arxiv.org/abs/2305.07759 https://arxiv.org/abs/2305.07759 "We show that TinyStories can be used to train and evaluate LMs that are much smaller than the state-of-the-art models (below 10 million total parameters), or have much simpler architectures (with only one transformer block), yet still produce fluent and consistent stories with several paragraphs that are diverse and have almost perfect grammar, and demonstrate reasoning capabilities." They also trained a TinyStories-Instruct instruction following variant. It looks like the 28M parameter 8 layer model had a 9/10 on grammar and consistency. You could probably combine that with something like Jsonformer or parserLLM to enforce valid formatting.
- TekMol 3y agoDo I understand it correctly, that an LLMs are neural networks which only ever output a single "token", which is a short string of a few chars? And then the whole input plus that output is fed back into the NN to produce the next token? So if you ask ChatGPT "Describe Berlin", what happens is that the NN is called 6 times with these inputs: Input: Describe Berlin. Outpu: Berlin Input: Describe Berlin. Berlin Outpu: is Input: Describe Berlin. Berlin is Outpu: a Input: Describe Berlin. Berlin is a Outpu: nice Input: Describe Berlin. Berlin is a nice Outpu: city Input: Describe Berlin. Berlin is a nice city Outpu: . ChatGPT's answer: Berlin is a nice city. Is that how LLMs work?
- going_ham 3y agoThis is a great example to show the underlying mechanism!
- TekMol 3y agoGreat, thanks for the clarification. And how does the NN represent the token at the output layer? Is it a binary representation of the token number? Or does it have a neuron for each token it knows and ChatGPT takes the most activated neuron as the answer?
- zwaps 3y agoBasically the latter
- Animats 3y agoIt's plausible, but it was partly written by ChatGPT. From the paper: "Big thank you to ChatGPT for helping me draft this section." So of course it's plausible. That's what ChatGPT does. There are a number of systems where someone bolted together components like this. Now we need more demos and evaluations of them. I just read a comment about someone who tried a sales chatbot. It would make up nonexistent products to respond to customer requests. The underlying LLM systems need some kind of confidence metric output.
- p-e-w 3y agoA fairly lengthy article about Autonomous AI, and, as far as I can tell, not a single word about the safety implications of such a system (a short note about reliability of LLMs is all we get, and it's not clear that the author means anything more than "the thing might break"). I get that there are different philosophies on AI risk, but this is like reading an in-depth discussion about potential atomic bombs in 1942, with no mention of the fact that such a bomb could potentially level cities and kill millions. If a field of research ever needed oversight and regulation, this is it. I'm not convinced that would solve the problem, but allowing this kind of rushing forward to continue is madness.
- ftxbro 3y agook but it's an odd take for a tech site like hacker news i mean what does that say about the llm paradigm vs for example object oriented paradigm or functional programming paradigm were they just playing around before and now it's getting real? were they deliberately doing toy things before and now the adults are in the room making real computation systems with actual consequences for the real world and not just nerds playing with toys or what? Like does this mean that it will completely obsolete the works of boffins like simon peyton jones and those other esoteric languages and everyone will talk to their computers in english language and the ones that will be the most effective computer enjoyers will be the high emotional intelligence salesmen? were they going in wrong direction for so long on purpose because they didn't want to awaken the real power of the computers for ethical reasons, or were they trying to do this the whole time but they weren't smart enough to figure it out?
- carom 3y agoNo, this is nothing like atomic bombs. In what scenario can you kill 100,000 people with with an LLM hooked up to a Python interpreter? The AI safety grift is out of control.
- p-e-w 3y ago> In what scenario can you kill 100,000 people with with an LLM hooked up to a Python interpreter? The article talks about LLMs being hooked up to the Internet. Big difference. If you really can't imagine a scenario where this might lead to people dying, you're not trying hard enough.
- mercurialsolo 3y agoAutonomy without alignment is a slippery road. Autonomous agents need guardrails and oversight. An autonomous agent let loose with all the tools in the world will in essence lead to an outcome which is not predicted to be in our favour. Which is why the Open AI app store and plugins scare me more than anything else - more likely than not they are tool and data feeders into a large scale autonomous system.
- logicchains 3y ago> An autonomous agent let loose with all the tools in the world will in essence lead to an outcome which is not predicted to be in our favour. Given the current state of the art, it won't lead to any worse outcomes than a somewhat mentally impaired human "let loose on the world". And since training scales quadratically (doubling model size means double the amount of data is needed to train it optimally), and we're already at the limits of current hardware, that's unlikely to change much any time soon.
- quickthrower2 3y agoExcept for security culture. We distrust humans but allow apps to have keys to the kingdom DB connections and so on. Allowing an LLM an outbound internet connection is dangerous. Maybe a firewalled one might be OK.
- ilaksh 3y agoJust to be clear, I think LLMs have enormous potential and am focused on building products with them. But I also believe that smarter hyperspeed LLMs will be an existential risk when widely deployed in the relatively near future. In a way GPT-4 is like a mentally impaired human but in other ways it's superhuman. It operates faster than a human in many contexts. It has vastly greater knowledge. Agents based on LLMs could communicate and process new information practically instantaneously. And the important point that people are in denial about is that GPT-4 reasons effectively. It's far from perfect and has some strange failure modes, but demonstrates the potential of these systems. It's not accurate to think that LLM performance can't be improved without doubling size. There are many approaches to efficiency recently demonstrated that don't require larger datasets. We now have many geniuses with billions and billions behind them pushing hard to optimize this specific application from all directions. Modifying the software that runs the model, the model parameters, model architecture, and the hardware. In particular for hardware there is now a large increase in attention to novel compute-in-memory paradigms or techniques. GPT-X will have at least 33% higher IQ and 50-100 times faster output within no more than 5-10 years. Quite possibly less than that. Humans will not be able to compete with that. The only option will be to deploy their own AI agents. And there will be a strong incentive to increase the level of autonomy for the agents, since making them wait for human input a few hours means that the competitors' agents race ahead doing the equivalent of days of work in that time frame. This delegation of control to agents with superior reasoning ability sets the stage for real danger. Especially in a military or industrial context.
- atomlib 3y agoCoolest post on this subreddit.
- quickthrower2 3y agoWhy do people keep calling this that?
- atomlib 3y agoBecause I liked the post.
- RamblingCTO 3y agoBut we're not on reddit. This is not a subreddit
- atomlib 3y agoWell, this is a Reddit clone.
- RamblingCTO 3y agoGet a life dude
- Roark66 3y agoSeriously LLMs are remarkable tools, but they are horribly unreliable. What tasks could such autonomous agent do (beyond what a chat bot, perhaps extended with web access, already does)? I mean which task is so complex one can't just automate it with simple scripting and non critical if it goes wrong to the point of letting an AI LLM do it? BTW, running those models is rather expensive so also the task has to be quite expensive now, perhaps completed by a human.
- ryanSrich 3y agoImagine a task that can automatically create and post ads on LinkedIn, Adsense, etc. Imagine you can also provide analytics as feedback, thus making each ad more analytics optimized. These are the types of mundane activities I dream of offloading to GPTs until the entire concept of online ads become irrelevant.
- kolinko 3y agoExample task that you cannot automate with a basic script - for each e-mail you receive, convert it into json of a given format for a given e-mail type (e.g. for delivery notifications, newsletters, meeting requests and so on). Also, tag it if it relates to a specific area of interest. Another case that a friend of mine built - summarise a stream of news from Twitter/Telegram so that no news are lost, but duplicates are removed and bad writing is reformatted.
- TeMPOraL 3y ago> I mean which task is so complex one can't just automate it with simple scripting and non critical if it goes wrong to the point of letting an AI LLM do it? Many of the things that get to deal with unstructured human inputs. Translation, summarization, data extraction, etc. The choice isn't between "reliable code" and "unreliable LLM", but rather "unreliable LLM" vs. nothing at all. But I think more important aspect is that LLMs are getting quite good at coding - meaning they can be used to automate all the tasks that you could automate with scripting, but it wouldn't be simple, or you don't know the relevant tooling well enough, or you just can't be arsed to do it yourself. For example, I've recently been using GPT-4 to automate little things in my Emacs. GPT-4 actually sucks at writing Emacs Lisp, but even having to correct it myself is much less effort than writing those bits from scratch, as the LLM saves me the step of figuring out how to do it. This is enough to make a difference between automating and not automating those little things.
- YChacker100 3y ago[flagged]
- arisAlexis 3y agoMake me some paperclips please. Do not kill any humans.
- snowcrash123 3y agoGood read. Currently there are lot of issues in autonomous agents apart from Finite context length, task decomposition and natural language as interface mentioned in the article. I think for agents to truly find adoption in real world, agent trajectory fine tuning is critical component - how do you make an agent perform better to achieve particular objective with every subsequent run. Basically making the agents learn similar to how we learn when we Also I think current LLMs might not fit well for agent use cases in mid to long term because the RL they go through is based on input-best output methods whereas the intelligence that you need in agents is more around how to build an algorithm to achieve an objective on the fly - this requires perhaps new type of large models ( Large Agent Models ? ) which are trained using RLfD ( Reinforcement Learning from demonstration ) Also I think one of the key missing piece is a highly configurable software middle ware between Intelligence ( LLMs ), Memory ( Vector Dbs ~LTMs, STMs ), Tools and workflows across every iteration. Current agent core loop to find next best action is too simplistic. For example if core self prompting loop or iteration of an agent can be configured for the use case in hand. Eg for BabyAGI, every iteration goes through workflow of Plan, Prioritize and Execute or in AutoGPT it finds the next best action based on LTM/STM, or GPTEngineer it is to write specs > write tests > write code. Now for dev infra monitoring agent this workflow might be totally different - it would look like consume logs from different tools like Grafana, Splunk, APMs > See if it doesnt have an anomaly > if it has an anomaly then take human input for feedback. Every use case in real world has it's own workflow and current construct of agent frameworks have this thing hard coded in base prompt. In SuperAGI( https://superagi.com https://superagi.com) ( disclaimer : Im creator of it ), core iteration workflow of agent can be defined as part of agent provisioning. Another missing piece is notion of Knowledge. Agents currently depend entirely upon knowledge of LLMs or search results to execute on tasks, but if a specialised knowledge set is plugged to an agent, it performs significantly better.
- m3kw9 3y agoAutoGPT and babyagi is taking a very long nap till gpt5 comes out