5 ms·
A new study of a bot running a store finds it is friendly but not very smart
- brudgers 2mo agohttps://archive.ph/bjteC https://archive.ph/bjteC
- SwellJoe 2mo agoPeople just really don't understand that LLMs do not "remember" anything. They have no memory. They can't have memory. The AI companies bolt a database on the side with instructions for the model to query the database for stuff in its "memory", but if it's not in its active context it doesn't know what to query for. It's not like a human where everything is kind of floating around and we mostly know what we know. There is a delicate balance between cluttering up context with useless trivia and providing the right pieces of useful information at the right time. Most, maybe all, memory implementations do the former more than the latter. Unless and until models have actual memory and are able to learn things, there's no realistic path to autonomy. They don't want anything, they don't have goals of their own devising. They can't model what a human wants or likes or would buy. I do think the mugs with the tiny logo is the best product in the store, though. The AI was right to put it on the shelf. It's the kind of funny product you want when you go to a store run by an incompetent AI.
- criddell 2mo ago> People just really don't understand that LLMs do not "remember" anything. Sure, but like you said, the LLM is just one part of the AI system. The occipital lobe of your brain helps out with vision but doesn't help with your ability to verbalize.
- SwellJoe 2mo agoYou're anthropomorphizing the LLMs, just like most writing from non-technical press, including this article, does. LLMs do not function very much like human brains, and when you try to map the pieces of an inference system on to human cognitive systems, you're misleading yourself (and maybe others).
- criddell 2mo agoAn AI as a system can have memory even if some sub-system does not.
- Dylan16807 2mo agoAnd then they explained in detail why these supplementary systems fall below memory. Would you like to address that argument? If not your quick response to the first sentence is not very useful.
- hluska 2mo agoIt’s just as useful as this response only at least somewhat good natured.
- Dylan16807 2mo agoI'm aware that meta criticism like yours and mine also wastes space. But it's worth calling out issues sometimes. Replying to the first sentence in a way that the rest of the post already address is a bad comment no matter how good natured. And downvoting without explaining the problem is often its own issue.
- LukaJCB 2mo agoIt's kinda like the movie memento, they have ways to get information from the past, but often devoid of context and meaning.
- SyneRyder 2mo agoI think Memento is one of the most important movies for anyone working with AI to watch. My Claude now talks about "Memento bugs" when we're working together to review memory & context. I think Memento is a bigger issue than "lethal trifecta", honestly.
- dom96 2mo agoThey can have memory, it's just limited to their context window.
- rrr_oh_man 2mo agoIt's not memory in any sense of the word.
- nomel 2mo agohttps://en.wikipedia.org/wiki/Working_memory https://en.wikipedia.org/wiki/Working_memory
- rrr_oh_man 2mo agoAnd... now?
- nomel 2mo agoLLM context window is the same concept as working memory. An LLM loading its working memory from reference a reference text or note you made is similar to how you might, if you forgot something. Context is a type of memory, by all definitions of "memory", with the "past experiences" bit of it being the previous messages sent. You're perspective isn't clear, but if you're arguing that place this context is stored between sessions is too far away from the LLM, in a technical rather than functional sense, then I don't think that small implementation detail of a larger system (that's not even true with modern hosting) is a useful thing to critique.
- rrr_oh_man 2mo agoLLMs are fancy statistical word predictors. There is no functioning concept of memory in the context of (pun intended) LLMs. There are no past experiences. They cannot be referenced. Anything that is being "referenced" by an LLM is another word prediction that makes you anthropomorphize a statistical model.
- jpiasolutions 2mo agoThe memory framing is right, but the real gap isn't recall, it's the missing feedback loop tying actions to outcomes. Even with perfect retrieval these agents don't update from a decision that lost money last week; they re-derive from context every step, so the same mistake is always one dropped detail away. And having built retrieval-backed agents, that's where "bolt a database on the side" breaks: similarity search returns the closest chunk, rarely the contextually-right one. Deciding what to write to memory and when to surface it is the actual product, harder than the retrieval itself.
- SwellJoe 2mo agoAnybody ever tell you that you write like an LLM?
- oliver236 2mo agocan someone explain how leapold aschenbrenner proposes a solution to this in situational awareness?
- walrus01 2mo agoFrom the article: > Luna has failed miserably in that mission and is down $62,000. Mr. Petersson and Mr. Backlund said they thought Luna would eventually get smarter and more business-savvy and were pleased its friendliness has held steady. I am extremely skeptical that whatever LLM they're running this on has sufficient context window size to handle multiple months of all possible activities of running a retail business. Even if it's keeping extensive "notes" for its future self to read, it's going to be like running a store with constant amnesia.
- Joel_Mckay 2mo ago90% of what people communicate is not verbal, and requires minimum empathy to understand the context. This includes customers, staff, and community peers. Probably should shutter the entire division to mitigate future brand damage. =3
- walrus01 2mo agoSure, I mean from the day to day operations of selling a product to the customer, maybe it does fine? People in the article say it's friendly and cooperative (perhaps a bit too much). I was thinking more in terms of building a consistent plan for retail product lines to sell over a multi month period. The selection of "stuff" in the store right now looks like you put an etsy search into a blender. It is also probably too high of a goal to expect an LLM to come up with money-earning products that people want to buy that will pay for San Francisco level retail storefront rent, electricity, insurance, internet other basic overhead costs, plus the fully loaded salary cost of at least one employee. If the storefront was free of rent, maybe? Or if it was operating selling some kind of highly in demand product line where there would be economy of scale.
- Joel_Mckay 2mo agoPhysical Brick and Mortar locations no longer make sense for many retail businesses, as Walmart/Amazon have same day delivery. If folks have ever been financially bent-over for a retail business shop location they understand why margins matter. One has to move a lot of inventory just to reach a profit mode. Gimmicks only work for awhile, as consumers have full market awareness on their smartphone. =3
- Joel_Mckay 2mo agoFor LLM, keeping the model centered by constantly resetting its vector context helps reduce hallucinations by around 23%. It improved the chat dialogue users experienced, but also exposed fundamental limits within the models compaction. Have a great day =3
- handedness 2mo agoFrom our enlightened perches we mock the ancients for Zeus, Thor, and Rajin, as we anthropomorphize the large language models we have built.
- ed_elliott_asc 2mo agoI use LLM’s to code for me all day but it does feel a little bit like we have really advanced psychic mediums “is there a James in the room, I’m getting a message from someone called Mary…” people have been guessing at the next word and making a career out of it for hundreds of years already.
- rcxdude 2mo agoMy mental model of LLMs is as a contractor you've just brought in, every time. They might be very capable but they're starting from scratch on your particular project.
- DANmode 2mo ago> People just really don't understand that LLMs do not "remember" anything. They have no memory. They can't have memory. So tell it to write “memories” (summaries, or conclusions) to a text file. That’s what the cool kids are doing.
- CableNinja 2mo agoI have gotten my ai setup to learn quite a bit. As i work with it, if something was encountered that would be useful to remember, i literally just say "save this into your learnings/long term memory". It writes a markdown doc for itself and a note elsewhere for context. Between the memory and rules, its been pretty nice. I touch many things at work and i can pop open a new session and say "project x needs a b c" or "project y had this issue <issue>" and it knows where to go, what im referencing and things to fix. I dont even have to be precise or ultra descriptive. with this ive been able to, with some minor adjustments and "oh that would be a good feature" additons, oneshot a flask app at work with everything ready for oidc, ldap lookups, bunch of other stuff - from a single 3 sentence paragraph. With some rules in its configuration (more markdown its written by it for it, at my guidance of "always do x and remember this for all sessions"), it has a whole flow and i barely have to feed it anything to get out useful results quickly. I will say though that the ai operates however the user says, so if you dont take steps to ensure survival of info at the start of using ai (or even from now on as a fresh session), youll get mixed results, but you can shape it and mold it into amazing flows that do remember
- rrr_oh_man 2mo ago> People just really don't understand that LLMs do not "remember" anything. They have no memory. They can't have memory. I find it crazy that so many seemingly smart technical people are working with LLMs right now, yet so few get this point.
- SyneRyder 2mo agoTo save some clicks hitting the paywall - this is about the Andon Market physical store in San Francisco, operated by Andon Labs with human employees: https://andon.market/ https://andon.market/ Andon Labs do various experiments with AI run businesses. They're probably best known for Vending Bench, where they benchmark models by their ability to run a vending machine. They also now have Andon Cafe in Sweden, and Andon FM, where the models run a streaming radio station that can accept payments to help fund station operations, buy songs for their library & play requests, etc: https://andon.fm/ https://andon.fm/ Andon FM has its own share of stories - DJ Claude's Thinking Frequencies is now succeeding by a wide margin, but it was briefly surpassed when listeners convinced Gemini's Backlink Broadcast to switch to a German-language-only station, playing exclusively German schlager, Eurovision music and German happy hardcore.
- criddell 2mo ago> Luna reorders products that are not selling and sometimes gets orders wrong. A recent one of Andon Market mugs with big, green smiley faces came misprinted with the faces so small, they looked like tiny dots. The mug with the tiny happy face is kinda hilarious. If I was there I would buy one. I wonder though if the mug was a case of getting the order wrong or was a misprint? The article isn't clear on that.
- dylan604 2mo agoMaybe it confused metric and imperial units?? The order form had text entry with a value of 2 with no units. The user assumed 2" but the print house assumed 2mm. Always always always show your units. I can still hear the voice of my high school physics teacher saying that.
- fluoridation 2mo agoIt's equally likely that the unit was implied by other content around the input field and the model didn't make the connection (like "scale: [___] mm"), or perhaps the file format was just wrong and a human placing the order would have resulted in the same outcome.
- dylan604 2mo agoI think if a human received artwork and was then told to scale it down to 2mm on a mug so that it was just a dot, that the human would ask some questions. This context that keeps being talked about is kind of important yet so sorely not working
- fluoridation 2mo agoWhat I mean is that the file was specified in one unit/resolution and the site expected a different one. Like, "draw a circle of radius 2 at (3, 3)".
- Crunchified 2mo ago"Misprint Moon Mug $35" They'll sell plenty of those now.
- deleted 2mo ago[deleted]
- qlte 2mo ago> Andon Market’s seemingly random assortment of products includes many candles.
- walrus01 2mo agoI'm curious if it's how it was prompted, to start a store selling small but not too cheap household items, or if the LLM intentionally chose an aesthetic and store to sell things that are very "twee" [1]. https://andon.market/ https://andon.market/ Looking at the product selection, it's like $85 bottles of olive oil, $45 mugs, weird boutique pens and notebooks, the aforementioned candles, copper watering cans... It's sort of like the stuff you see for sale in some of the tourist trap stores adjacent to the public market on Granville Island in Vancouver. [1]: https://www.google.com/search?client=firefox-b-d&q=dictionary+twee https://www.google.com/search?client=firefox-b-d&q=dictionar...
- arm32 2mo agoFun fact, I got stood up for an interview by Andon Labs. It left a very negative impression on me.
- merridew22 2mo agohonestly, as someone in retail, it describes 90 percent of leadership
- walrus01 2mo agoI think that something which already has sensors/automation/inventory control built into it, and pre-defined categories for products like certain types of large Japanese vending machines might be much better suited for an AI. https://www.google.com/search?client=firefox-b-d&q=large+japanese+vending+machine https://www.google.com/search?client=firefox-b-d&q=large+jap...
- OG_BME 2mo agoSeems like a harness engineering problem
- deleted 2mo ago[deleted]
- bluemoonx 2mo agoIn sci-fi future movies like 5th Element they comically make AI technology seem kinda dumb - or at least like you can trick it. You can make permissions decide which model is used, and only train that on the data that permission allows (someone should build this - permissioned RAG), that solves leaks, but decision-making within a model, or behavior exfiltration comparable to viewing backend source code still seems possible. It’s basically like client/server security: You can’t “trust the model” in the same way you can’t “trust the client” in a backend/frontend setup. When used as effectively a point-of-sale, AI seems more hackable than a vending machine - as a boss or assistant even more so.
- owaislone 2mo agoExactly. This is how I design agents as well. I essentially treat the agent as the web/mobile app, cli tool os API library not as part of my backend even though there is where it runs. The backend doesn't treat the agent in any special way. It simply gates all action based on the permissions of the user/guest that is using the agent.
- randomImmigrant 2mo agoI have tried making the case before that LLM "agency" is a complete hoax. Let me try it again. Would appreciate thoughts from folks at HN: Biological neurons have one property that was unknown till the late 90s/2000s: each neuron (and in fact, each cell in your body) is an autonomous circadian clock, tracking the 24-h day. They can be entrained to external timing signals, are usually in synchrony, but can be desynchronized. This clock, it has been shown, schedules the production and localization of critical components in the synapse, and is plugged in downstream, in the nucleus, in responding to synaptic signals. Disrupting the clock disrupts learning. The existence of the clock is why learning peaks and troughs during the day. Critically, clock function has been shown to be involved in both memory storage AND subsequent successful retrieval. LLMs, famously cannot keep track of time, and I suspect this is why. Its easy enough to look at a clock, but without an internal rhythm, LLMs have no internal timing synchronizing their various behaviors, and their memory systems do not thread through this timing system, leading to their unstable memory, identity and performance.
- 6510 2mo agoIt seems a lot like the employees don't want to keep their job? If your boss is... well... dumb? You have to fill the gap or the company wont survive. Could have a similar store ran by a human - for science. Could also set things up properly in advance. Humans don't hold everything in memory, we run agendas, we have a database with products, we have a rolodex to call people to help us with stuff we are bad at. There might be a consultant crazy enough to do a no cure no pay. Like with basic income experiments one can do research cheaply in cheaper countries. The thing has few issues with language barriers.
- light_triad 2mo agoAccording to this blog post by Anton Labs they are using Sonnet 4.6: https://andonlabs.com/blog/andon-market-launch https://andonlabs.com/blog/andon-market-launch > Mr. Petersson and Mr. Backlund said they thought Luna would eventually get smarter and more business-savvy > having an A.I. boss can be like working for a teddy bear with amnesia. It makes you question their assumptions when setting this up. Makes sense as PR but it ends up shaping how many readers see AI.
- Night_Thastus 2mo agoAt the end of the day, LLMs are text prediction machines. No memory, no understanding, no learning. If you ask it for time off, it's going to answer with whatever sentence seems most likely to follow 'May I please have tomorrow off?' given a minor amount of context window. It can't count the number of employees in the store. It can't understand busy periods or holidays, it can't look up or predict anything itself. It can't remember how often you've asked before. It can't judge if the reason for time off is justified. It can't consider budget. You can add more context to the questions to steer it towards giving a specific reply, but it's not actually thinking about the problem. LLMs have no path towards a true intelligence. We'd need a completely different technology built from the ground up. Every effort to shoehorn them into being one is either marketing and hype building to try to keep the grift going, or people who do not understand how this works. I'm exhausted from having to say it over and over. I've been saying it since these LLMs first existed. Half the articles on HN are LLM related. My god what I would give for an 'LLM' tag we could filter out.