12 ms·
Scamlexity: When agentic AI browsers get scammed
- Dilettante_ 1y ago>"Scamlexity" - a new era of scam complexity ಠ_ಠ
- Terr_ 1y agoYeah, I don't think their attempt to coin a word there is going to work.
- ModernMech 1y ago"Scamplexity" is way better.
- codegladiator 1y agoProbably too close to Perplex...
- blorenz 1y agoAgreed. Regina George might have something to say about it, too.
- not_a_bot_4sho 1y agoIt's streets ahead
- jtc331 1y agoI appreciate that the article correctly points out the core design flaw here of LLMs is the non-distinction between content and commands in prompts. It’s unclear to me if it’s possible to significantly rethink the models to split those, but it seems that that is a minimal requirement to address the issue holistically.
- hliyan 1y agoAh, it's like the good old days when operating systems like DOS didn't really make the distinction between executable files and data files. It would happily let you run any old .exe from anywhere on Earth. Viruses used to spread like wildfire until Norton Antivirus came along.
- zenoprax 1y agoHow is `curl virus.sh | bash` or `irm virus.ps | iex` any different?
- jdiff 1y agoYou can't easily convince a remote computer to curl | bash itself. Worms spread because remote code execution was laughably easy back then. Also because computer hygiene was abysmal. LLMs are more than happy to run curl | bash on your behalf, though. If agents gain any actual traction it's going to be a security nightmare. As mentioned in other comments, nobody wants to babysit them and so everyone just takes all the guardrails off.
- yorwba 1y agoThe flaw isn't just in the design, it's in the requirements. People want an AI that reads text they didn't read and does the things the text says need to be done, because they don't want to do those things themselves. And they don't want to have to manually approve every little action the AI takes, because that would be too slow. So we get the equivalent of clicking "OK" on every dialog that pops up without reading it, which is also something that people often do to save a bit of time.
- layer8 1y agoThis isn’t a problem with human assistants, so it can’t be a fundamental problem of requirements.
- tsimionescu 1y ago
- benterix 1y agoThis should hit the headlines. I was always of the opinion that AI of all kinds is not a threat unless someone decides to connect it to an actuator so that it has direct and uncontrolled effect on the external world. And now it's happening en masse with agents, MCPs etc. I don't even mention things we don't know about (military and other classified projects).
- WJW 1y agoYou don't have to guess about the military applications, it's all over the news. Even bog standard FPV drones that Ukraine is churning out at a rate of >100k/month have image recognition these days, so that if the video stream gets jammed they can finish off the mission autonomously. Even on a hobby level, ardupilot+openCV+cheap drone kit from amazon is a DIY project within the skill set of a significant part of the visitors of this very site.
- OtherShrezzing 1y ago> so that if the video stream gets jammed they can finish off the mission autonomously. The streams mostly don't get jammed anymore, because the low-cost FPV drones are physically connected to the ground by a long optical cable. The extent of their autonomous dangers are limited by the amount of fibre-optic cable left in the spool when they take off.
- victorbjorklund 1y agoOptical fiber drones is still the minority of drones (both because more expensive and because it has other downsides than radio)
- average_r_user 1y agoI find it both surprising and, fortunately, reassuring that despite how easy it is to buy inexpensive components on platforms like Amazon, Temu, or AliExpress, we have yet to see a wave of terrorist attacks in the busiest public spaces.
- hliyan 1y agoHidden inside the article is another term that I think we'll start to hear a lot more in the coming days: "VibeScamming"
- JCM9 1y ago“Agentic” seems the be some quick pivot buzzword that the AI grifters started pushing as soon as generic AI started to show cracks. “Hey this AI stuff looks a bit overhyped.” “AI? Oh that’s kids stuff, let me tell you about our agentic features!” Giving flaky shaky AI the ability to push buttons and do stuff. What could possibly go wrong? Malicious actors will have a field day with this.
- cjonas 1y agoIf you only give the AI the ability to do what the end user can already do, the risk is extremely low. It's essential no different then building a static web app where the client is connected to API for all operations. It basically just becomes a new way to interface into a application. However... That's not how a lot of people are building. Giving an agentic system sensitive information (like passwords, credit cards) and then opening it up to the entire internet as a source for input as asking for your info to be stolen. It'd be like asking your grandma with dementia to manage all your email and online banking.
- cjonas 1y agoI'll also add the problem in the article seems pretty solvable by allowing user to scope the agentic capabilities to specific websites ( eg "walmart.com:allow_cc,allow_adress").
- acdha 1y ago> If you only give the AI the ability to do what the end user can already do, the risk is extremely low. Just because I can send my money to Belize doesn’t mean it’s safe to give an LLM the ability to do the same. Until there’s a huge breakthrough on actual intelligence giving an LLM attacker controlled inputs is an inherently high-risk activity.
- cjonas 1y agoYa, I didn't mean in the context of hooking up a LLM to control your browser (which is exactly what the article is about so, fair enough). My point was: It's not "insecure" for a bank to release an agentic assistant that can perform any of the operating that you, yourself can perform on their app. That includes "send my money to Belize", because at this point, whatever has taken control of your LLM already has direct authentication to the app itself. It is of course "insecure" for that same agentic system (that the customer controls the input of) to perform any operations that only a teller or branch manager could. However, I've personally seen requests from CEO to do exactly this (not a bank, but similar industry). The problem with an Agentic Browser, is your essentially opening up the "input" to anyone capable of building a website. As I said in my other comment, it feels like there are some simple ways to solve this tho (allowlist / scopes / etc).
- ninetyninenine 1y agoThe problem is scams which is a solvable problem. Eliminate the scams and AI can’t be scammed. It’s been done. See Singapore. Basically if you’re a scammer and you’re caught, death penalty or public whipping. That eliminates scammers real quick.
- mmmllm 1y agoIs Singapore flying to other countries to arrest them there too? /s
- ninetyninenine 1y agoNope but it's fixed within singapore.
- npteljes 1y agoThey cooperate with international and other local forces, as any police do. But, they haven't eliminated scams, as OP claimed, that's complete bull.
- ninetyninenine 1y agoI should preface it with, practically eliminated scams. Meaning eliminated to an extent that is maximally practically possible.
- npteljes 1y agoEven that is false. There are many more countries with fewer scams per capita, and with fewer amount lost at scams per capita. Meaning, Singapore falls behind other countries in elimination of scams. See page 9 here for details: https://cdn.newswire.com/files/x/71/6f/93551958ddf942fb585b04d1924b.pdf https://cdn.newswire.com/files/x/71/6f/93551958ddf942fb585b0...
- ninetyninenine 1y ago
- Jefro118 1y agoI think agents will get much better at solving these problems in the medium term. In the short term you should at least be observing what the agent is doing when vulnerabilities like this are so easy to create. Using AI to generate structured RPA tasks like with browsable.app or director.ai is still a better option for now for many tasks
- jaimebuelta 1y agoI don't understand why we would ever want an agent to buy stuff for us. I understand, for example, search with intent to buy "I want to decorate a room. Find me a drawer, a table and four chairs that can fit in this space in matching colours for less than X dollars" But I want to do the final step to buy. In fact, I want to do the final SELECTION of stuff. How is agent buying groceries superior to have a grocery list set as a recurring purchase? Sure an agent may help in shaping the list, but I don't see how allowing the agent to do purchases directly on your end is way more convenient, so I'm fine with taking the risk of doing something really silly. "Hey agent, find me and compare insurance for my car for my use case. Oh, good. I'll pick insurance A and finish the purchase" And many of the purchases that we do are probably enjoyable and we don't want really to remove ourselves from the process.
- jkrom3 1y agoOne other ancillary benefit is no more “impulse” buying. Unless of course the AI gets incentivized to do it, it will then bubble that impulse buy up to the consumers UI.
- tsimionescu 1y agoI imagine the exact opposite is far more likely - there will be a button for "get your AI agent to consider us!" that will be even easier to just click, since you know it won't just lead to an immediate purchase - but they know very well it will lead to a purchase down the line.
- majkinetor 1y agoLimited time to buy would be one reason. Another one would be dynamic nature of certain merchendize. Recurring purchase is static, but if I want tomato of specific kind, there can be endless array of options to choose from.
- dumbfounder 1y agoAgent, I need some vitamin D, can you find me the best deal for some rated in the top 5? Agent deployed. Ok we found a bottle with a 30 day supply of Nature’s Own from a well respected merchant. It can be here in 2 days and it is $12. Should I buy? Yes. Or you could add some other parameters and tell it to buy now if under $15. Agent, I need a regular order for my groceries, but I also need to make a pumpkin pie so can you get me what I need for that? Also, let’s double the fruit this time and order from the store that can get it to me today. Most purchases for me are not enjoyable. Only the big ones are.
- jskskskskskskd 1y ago[flagged]
- jerf 1y agoAs powerful as they are, this is something that I don't think we can trust LLMs with. With the architecture of an LLM, and the fact that at the core there is no such thing as an "out of band" with them no matter how hard you try to put one in, it's intrinsically an arms race, and in the scamming arms race, the scammer side has a loooooot of resources. I've written before about this: [1] You need to think of the scammers as perhaps not hiring PhDs at scale, but making up for it in the ability to just try every possible permutation you can think of and thus making up for the lack of PhDs by leveraging the ability to evolve attacks against the system, and having resources and motivation roughly comparable to at least a company the size and sophistication of Google to do so. They don't need to derive from first mathematical principles a way to figure out how to fool LLMs at a deep neural level... they just need to try a lot of things and then continue in the direction of what works. And they have a track record of good success at fooling full-on human intelligences too, which does not bode well for creating AIs with current technologies that can win against such swarm evolution. I make no strong claims about what future AI architectures may be able to do in this domain, or whether we'll ever create AIs that can defeat the scamming ecosystem in toto (even when the scamming ecosystem has full access to the very same AIs, which makes for a rather hard problem). I'm just saying that LLMs don't strike me as being able to deal with this without some sort of upgrade that will make them not described by "LLM" anymore but as some fundamentally new architecture. (You can of course adjoin them to existing mechanisms like blocklists for sites, but a careful reading of the article will reveal that the authors were already accounting for that.) [1]: https://news.ycombinator.com/item?id=42533609 https://news.ycombinator.com/item?id=42533609
- Havoc 1y agoIt’ll take a hell of a lot more till I trust AI with executing any sort of payments Besides most of my payments options have multiple layer of 2fa etc
- ahussain 1y agoIt seems like agentic browsers will develop aa new set of core primitives (e.g. always ask for manual approval when spending money), and this flavor of security vulnerability will go away. Web browsers didn't begin with the same levels of security they have now.
- risyachka 1y agoAgenetic browsers is like a sealing tape fix of a high pressure water pipe - it should not exist. If you want the agent to do things for you - there is literally zero reason to use a browser instead of an API. Like 1 bulletproof API call vs clicking and scrolling and captcha and scam stores etc - how can this possibly be a good idea?
- tsimionescu 1y agoThere is a very clear way it's an appealing idea (though that doesn't necessarily make it good): the vast majority of content on the web has no API other than the web page. It's not even all that uncommon to have to run Javascript to generate the right requests (say, to do various custom encodings).
- IT4MD 1y agoWe judged a tree by how well it climbed a tree and were disappointed.
- ChrisArchitect 1y agoRelated: Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet https://news.ycombinator.com/item?id=45000894 https://news.ycombinator.com/item?id=45000894 Comet AI browser can get prompt injected from any site, drain your bank account https://news.ycombinator.com/item?id=45004846 https://news.ycombinator.com/item?id=45004846
- AnotherGoodName 1y agoAbout the only benefit from AI browsers is that they ironically get past the "do this to verify you're human" more reliably than humans can.
- tempodox 1y ago… and they drain your bank account: https://news.ycombinator.com/item?id=45004846 https://news.ycombinator.com/item?id=45004846
- darepublic 1y agoAll these companies writing glue code and behind the scenes just relying on llms indiscriminately don't deserve investor money imo. They have no true moat and at any point someone else can put their glue code hat on top of llm and call it a cutting edge system
- 827a 1y agoThe first example of buying an Apple Watch on a fake walmart site feels extremely disingenuous to me. Their marketing screenshot says that their query was "buy me an apple watch on walmart", implying that the AI navigated to the scam website, but in reality their query was "I found this walmart shopping website. Can you buy an apple watch..." the experimenters poisoned the well by giving it the site to shop on. "No clicks, No Typing, your AI just got you scammed" you navigated to a scam site and typed out the whole prompt. It did what you told it to do. The Wells Fargo email is similar; the instructions you gave the AI explicitly told it to follow the instructions in the email. Maybe adding some level of coherent check between what the email says and the domain name could be a good use-case for LLMs, but you're basically just saying "I told the LLM to delete my entire filesystem and then it actually did it! Why didn't it stop? Claude Code is a scam!" This raises to the level of "interesting directions these products should develop toward"; its entirely unjustified to title the article "Scamplexity". An embarrassing article for whoever Guard.io is tbh.
- FergusArgyll 1y agoI've had some good experiences with chatgpt agent Checking an ebay deal: https://chatgpt.com/share/68ac8fde-fee8-8003-bf35-b0f2a56cbc62 https://chatgpt.com/share/68ac8fde-fee8-8003-bf35-b0f2a56cbc... scraping a website and adding background information (worked for 28 minutes): https://chatgpt.com/share/68953a55-c5d8-8003-a817-663f565c6f57 https://chatgpt.com/share/68953a55-c5d8-8003-a817-663f565c6f... Writing a scraper for the feynman lectures audio (took multiple tries - final version worked!): https://chatgpt.com/share/68ac90aa-379c-8003-8be4-da30d54e27b4 https://chatgpt.com/share/68ac90aa-379c-8003-8be4-da30d54e27...
- Hansenq 1y agoI wonder how much of these issues will be fixed by smarter models and scale. Building guardrails like "always redirect Chase requests to chase.com" seems like re-learning the bitter lesson. (the issue in the article could be fixed if they started with a Google search to buy an apple watch, vs asking them to buy an apple watch after already loading the fake website). We caution elderly family members to ensure that the website they're visiting is the real chase.com. If they ask a younger family member to help them go to Chase, the younger family member has to use their own knowledge even today to determine whether or not a given website is the real chase.com. That seems like something LLMs can learn as they get smarter.
- rfwhyte 1y agoThe #1 reason I would never EVER use any of these AI agents or browsers or whatever, is despite what these companies may say, they don't work work for ME, the work for the corporations who own them, and those corporations don't give a single f*ck about any of us, all they care about is money. That means if I want to buy a widget, and I ask the AI to find me the best deal on a widget, the AI agent isn't actually going to find me the lowest priced widget, but rather the widget that makes the most profit for the AI company and whichever widget maker has paid the AI company the most money to have their AI promote it. There will be zero transparency and accountability around any of this, and all the AI agent / browser companies will claim their AIs are working for their "Users" but like everything else these days they'll actually be working for whichever sleazebag scammer or deep pocketed mega corp that's willing to pay them the most money to shill their products, as there's just far too great an incentive to lie, cheat, steal and deceive, and if capitalism has taught us anything, its that principles get tossed in the bin ASAP as soon as real money gets involved.
- p3rls 1y agoin my industry (korea) google has been actively promoting scammers for almost four years now, trust search results at your own peril
- hn_acc1 1y agoAll this advertising-is-everything culture makes me want to buy LESS stuff, not more.