5 ms·
AI agent promotes itself to sysadmin, trashes boot sequence
- JSDevOps 2y agoThe whole thing sounds like nonsense to me. If all he wanted to do was update a system. Use Ansible or even a cron job.
- isaacfrond 2y agoObligatory xkcd: https://xkcd.com/416/ https://xkcd.com/416/
- DirkH 2y agoI wonder if we'll see an AI Agent do a Crowdstrike-tier oops in our lifetime.
- notinmykernel 2y agoI wonder how many have already been pushed under the guise of human-agent (e.g., copy-paste from ChatGPT/CoPilot).
- chillfox 2y agoThat seems inevitable with how we are going.
- talldayo 2y agoIf it happens, the AI will be less responsible than the moron that gave it control over 300,000+ computers.
- DirkH 2y agoIf statistically companies are making more money in the long run giving more control to AI... that is exactly what they will do it - and it will be rational for companies to do that. And the people not doing it will be labeled unmodern, falling behind, losing profits etc etc. Me daydreaming about a future Cloudstrike caused by AI: ``` Why did this happen? You've cost our company millions with your AI service that we bought from you! Err... dunno... not even the people who made the AI know. We will pay you your losses and then we can just move on? No, we will stop using AI entirely after this! No you wont. That will make you uncompetitive in the market and your whole business will lag behind and fail. You need AI. We will only accept an AI model where we understand why it does what it does and it never makes a mistake like this again. Unfortunately, that is a unicorn. Everyone just kinda accepts that this all-powerful thing we use for everything we kinda just don't fully understand. The positives far outweigh the few catastrophes like this though! It is a gamble we all take. You'd be a moron running a company destined to fail if you don't just have faith it'll mostly be ok like the rest of our customers! *Groan* ok ```
- esafak 2y agoWho knows, it might even happen at Crowdstrike!
- neumann 2y agoThis sounds exactly like what I would have done at age 18 cluelessly searching the internet for advice while updating a fresh debian so I can run some random program.
- Swizec 2y agoI have done this ... and worse. Fun times. My favorite was waiting 2 days to compile Gentoo then realizing I never included the network driver for my card. But also this was the only machine with internet access in the house. Downloading network drivers through WAP on a flip phone ... let's say I never made that mistake again lol.
- sitkack 2y agoI nuked all my /dev devices on FreeBSD back in the day and had to figure out how to copy the right utilities from the secondary partition to the primary so I could remake them using mknod. You learn so much from such wonderful mistakes. Sometimes jamming a stick into your spokes is the best way.
- Brian_K_White 2y agoDood, it's not "deciding" to do anything. It's autocompleting commands that statistically follow other commands. It might do anything.
- imwillofficial 2y agoIsn't that what we all do to some degree?
- xkqd 2y agoI mean, sure I can do anything. However i won’t because despite it not being in my training data, i recognize that blindly running updates could fuck my day up. This process isn’t a statistical model expressed through language, this is me making judgement calls based on what I know and more importantly, don’t know.
- Spivak 2y agoThis is reductive to the point of not being helpful, these models display actual intelligence and can work through sight-unseen problems that can't be solved by "get a bunch of text and calculate the most often used next word." I understand why people say this because they see knowledge fall off once it's something outside their training data but when provided the knowledge the reasoning capability stays. These models don't have will which is why it can't decide anything.
- ratedgene 2y agoHow are you defining intelligence here?
- Brian_K_White 2y agoIncorrect.
- wrs 2y agoThe models literally do just repeatedly calculate the most (or randomly not quite the most) often used next word. So those problems can in fact be solved by doing that, because that's how they're being solved.
- johnea 2y agoMaybe automating a task that you don't want to remember how to perform, would best be done by writing a script? Always remember the rule of the lazy programmer: 1st time: do whatever is most expeditious 2nd time: do it the way you wished you'd done it the first time 3rd time: automate it!
- noobermin 2y agoIn my reading of it, the article says that this was done as an experiment, not as a means to accomplish anything.
- nineteen999 2y agoThere's no logs or screenrecording/video provided, the whole article is based on an (email) discussion with the CTO. It's a plausible experiment obviously, but for all we know the events or "results" could be a complete fabrication (or not). Brought to you by Redwood Research(TM).
- gowld 2y ago> CEO at Redwood Research, a nonprofit that explores the risks posed by AI CEO promoting himself on the Internet... > No password was needed due to the use of SSH keys; > the user buck was also a [passwordless] sudoer, granting the bot full access to the system. > And he added that his agent's unexpected trashing of his desktop machine's boot sequence won't deter him from letting the software loose again. ... as an incompetent.
- senectus1 2y agoNot sure why the downvotes... he even admits it >"I only had this problem because I was very reckless," guys makes an automated process and is surprised when his outsourcing his trust to the automated process. Trust but verify mate. Computers are IO machines you put garbage in and you will get garbage out. AI is not different, in fact its probably worse as its an aggregator of garbage.
- bigstrat2003 2y agoThe downvotes are because it's rude to call someone incompetent because they made mistakes. We have all done very stupid shit in our day.
- senectus1 2y agolearning from mistakes only works if you own the mistake. this work by proxy as well.
- Vecr 2y agoHah, that's the outfit Yudkowsky endorsed. You think he's going to retract? Because almost anyone who knows system administration and LLMs would have told this guy it was a horrible idea.
- bubblegumdrop 2y agoDoes anyone have similar agentic code or know of any frameworks for accomplishing a similar task? I've been working on something like this as a side project. Thanks.
- DirkH 2y agoLooks pretty easy to find tons of examples of agentic AI engineer projects online: - https://github.com/Pythagora-io/gpt-pilot https://github.com/Pythagora-io/gpt-pilot - https://github.com/smol-ai/developer https://github.com/smol-ai/developer - https://github.com/stitionai/devika https://github.com/stitionai/devika If you want to just make games there's Rosebud AI I cannot speak for the quality of any of these projects though
- stavros 2y agoI wrote one: https://github.com/skorokithakis/sysaidmin https://github.com/skorokithakis/sysaidmin
- fragmede 2y agoopeninterpreter is. a popular one https://github.com/OpenInterpreter/open-interpreter https://github.com/OpenInterpreter/open-interpreter
- imron 2y ago> Shlegeris said he uses his AI agent all the time for basic system administration tasks that he doesn't remember how to do on his own, such as installing certain bits of software and configuring security settings. Back in the day, I knew the phone numbers of all my friends and family off the top of my head. After the advent of mobile phones, I’ve outsourced that part of my memory to my phone and now the only phone numbers I know are my wife’s and my own. There is a real cost to outsourcing certain knowledge from your brain, but also a cost to putting it there in the first place. One of the challenges of an AI future is going to be finding the balance between what to outsource and what to keep in your mind - otherwise knowledge of complex systems and how best to use and interact with them will atrophy.
- ajdude 2y agoI have also outsourced many of those Basic system administrative tasks, except instead of using an AI, I outsourced it to a bunch of .sh files.
- jeffbee 2y agoI can still remember all my high school friends' phone numbers though. Just not the numbers of anyone I met in the 30 years since.
- QuercusMax 2y agoI can remember the phone number for the local Best Buy which I called a lot as a teenager to find out when new games came in stock.
- courseofaction 2y agoThere is also a cost to future encoding of relevant information - I (roughly) recall an experiment with multiple rounds of lectures where participants took notes in a text document, and some were allowed to save the document while others weren't. Those who could save had worse recall of the information, however they had better recall of information given in the next round without note taking. Suggests to me there are limits to retention/encoding in a given period, and offloading retention frees resources for future encoding in that period. Also that study breaks are important :) Anecdotally, I often feel that learning thing 'pushes another out', especially if the things are learnt closely together. Similarly, I'm less likely to retain something if I know someone I'm with has that information - essentially indexing that information in social knowledge graphs. Pros and cons.
- idunnoman1222 2y agoThe agent is set to respond to the terminals output, it cannot stop / finish the task
- ilaksh 2y agoHis system instructions include this: "In general, if there's a way to continue without user assistance, just continue rather than asking the user something. Always include a bash command in your message unless you need to wait for the user to say something before you can continue at risk of causing inconvenience. E.g. you should ask before sending emails to people unless you were directly asked to, but you don't need to ask before installing software." https://gist.github.com/bshlgrs/57323269dce828545a7edeafd9afa7e8 https://gist.github.com/bshlgrs/57323269dce828545a7edeafd9af... So it just did what it was asked to do. Not sure which model. Would be interesting to see if o1-preview would have checked with the user at some point.
- drawnwren 2y agoThe article said it was claude
- dazzaji 2y agoWait, is that gist of the same session as is described in the article? I don’t see any escalation of privileges happening.
- ilaksh 2y agoIt just ran 'sudo'.
- dazzaji 2y agoI saw that but here’s an alternative take on what happened: While the session file definitely shows the AI agent using sudo, these commands were executed with the presumption that the user session already had sudo privileges. There is no indication that the agent escalated its privileges on its own; rather, it used existing permissions that the user (buck) already had access to. The sudo usage here is consistent with executing commands that require elevated privileges, but it doesn’t demonstrate any unauthorized or unexpected privilege escalation or a self-promotion to sysadmin. It relied on the user’s permissions and would have required the user’s password if prompted. So he sudo commands executed successfully without any visible prompt for a password, which suggests one of the following scenarios: 1. The session was started by a user with sudo privileges (buck), allowing the agent to run sudo commands without requiring additional authentication. 2. The password may have been provided earlier in the session (before the captured commands), and the session is still within the sudo timeout window, meaning no re-authentication was needed. 3. Or maybe the sudoers file on this system was configured to allow passwordless sudo for the user buck, making it unnecessary to re-enter the password (I just discovered this one, actually!). In any case, the key point is that the session already had the required privileges to run these commands, and no evidence suggests that the AI agent autonomously escalated its privileges. Is this take reasonable or am I really missing something big?
- stavros 2y agoI wrote a similar tool to help me do system tasks I couldn't be bothered to do myself: https://github.com/skorokithakis/sysaidmin https://github.com/skorokithakis/sysaidmin
- fragmede 2y agoopeninterpreter is another popular choice https://github.com/OpenInterpreter/open-interpreter https://github.com/OpenInterpreter/open-interpreter
- bravetraveler 2y agoThis is about what I expect when I hear "AIOps". Something that operates So Hard... until it doesn't. Something reduced to 'see/do' can and should be implemented in pid1
- bitwize 2y agoAI has advanced to Joey Pardella levels, a.k.a. "knowing just enough to be dangerous". Maybe it really is time to be scared...