17 ms·
I see this sentiment pretty regularly, and I don’t get it. Variable rewards is not sufficient to establish that it is “ basically gambling”. Everything in lif
by dpark 27d ago
I see this sentiment pretty regularly, and I don’t get it. Variable rewards is not sufficient to establish that it is “ basically gambling”.
Everything in life is variable reward. You invite a friend over, they might accept or they might not. Drive to work, traffic might be good or might be bad. You ask a colleague to finish a task, they might do it or might not or might do a good job or might not.
Everything is variable reward. Is everything gambling?
- howunfortunate 27d agoSame vibe as people saying "addicted to sugar" or "sugar hijacks your reward system" Sugar is the original point of the reward system!
- Barbing 27d ago& you can cheat the reward system. Do hard work (takes time), get dopamine for successful completion. Find berries, taste sweet (hopefully safe), eat all, get calories. Doordash Krispy Kreme instead = few too many calories. (I’m no Luddite in the sense popularly thought of them pre-‘22 [1], though we have to watch skill atrophy) [1] regressionist? Decelerationist, too loaded perhaps. Someone remembers or knows the word…
- utopiah 27d agoNo, if I use a ruler or a pocket calculator they will reliabley give me the correct result. There is no gambling.
- smugglerFlynn 27d agoYou invite a friend over, but raccoon appears. Then pigeon appears. Then friend appears but at the last second suddenly becomes a banana. You remember you are out of bananas so you order more and also some cola zero cans on your local grocery delivery app. You are back to the party, but now you have 5 friends in the room, and you run de-duplication query. Now half of your friend is sitting at the sofa, and another half becomes a quarter of banana. Suddenly bananas arrive so you need to open the door. Once you are back there are no friends, pigeons or raccoons but also no bananas and no cola - all the delivery results are gone. This seems to be urgent and important, gotta fix this first before going back to that friend invitation...
- rnjesus 27d agoi’m not sure “variable rewards” is the right term, but i do agree with the op that it is very similar to gambling. regarding your examples, i think the difference is that with ai, you’re literally sitting in front of a machine, pressing a button, and (almost instantly) getting a result that, if not desired, can immediately be tried for again. you even spend “tokens” to do this, and at least in my native language, “token” brings to mind the coins you’d stick in a slot machine
- dpark 27d agoI don’t see much similarity beyond the most superficial. If you sit at a slot machine and pump quarters into it, each “turn” is independent. You spin and you win or lose. It’s pure chance and there is no destination. You execute the exact same action over and over and hope random chance brings you more money. If you sit down in front of a coding harness, the progress is incremental and directed. You ask for a thing, the LLM produces something that is hopefully close to what you wanted. You give it more direction to prod it closer to the end state you want. You are not executing the same action, but incrementally nudging it in the right direction. I’ve literally never restarted from the same initial state with the same prompt and hoped for a different result and I don’t know why anyone would. Rarely I’ve thrown away the progress made and started over but always with a very different prompt that includes learnings from the failed attempt.
- rnjesus 27d agoi agree with you that it’s principally different from a slot machine, and that it’s possible to use it in a way (like you describe) that is much more focused, for lack of a better term, to great effect most people don’t use ai this way though, and i still feel like the end-psychological reward mechanism is very, very similar to gambling regardless of how well one utilizes it (and this is even more obvious with image generation as you chase that perfect output) perhaps it’s better to compare it to gacha than slots?
- dpark 27d ago> most people don’t use ai this way though How do they use it? Surely no one is just repeating the same prompt over and over (except as a Ralph loop perhaps, which is automated). I’m really struggling with the notion that most people just throw the same prompt repeatedly hoping it eventually works. Because that doesn’t sound like gambling. It sounds crazy (and frustrating). > and i still feel like the end-psychological reward mechanism is very, very similar to gambling regardless of how well one utilizes it In the sense that you get a dopamine reward when you succeed, sure, but I get the same reward when I code by hand and achieve a successful result. > and this is even more obvious with image generation as you chase that perfect output This is fair, because sometimes with image generation the same exact prompt will produce very different output. This is becoming less true as the models get better and it becomes more effective to direct image generation iteratively than to keep starting from scratch with a barely tweaked prompt.
- Tadpole9181 27d agoI would agree that 1-2 years ago models were more "slot machine"-esque - sometimes the output was good, sometimes the output was bad. And as a result, I primarily used them for auto-complete functionality and bouncing ideas around. In those workflows, you can easily ignore it if the spin is wrong. Not everyone has the desire to work around the system, and many are diametrically opposed to the concept of AI. They get this perception that it's a slot machine because of that inconsistency, and then do the human thing of assuming that other people must just be flawed if they're different from them. They're "addicted to gambling". Obviously, things have changed. Open models can still be like that, but are often so fast and cheap at iterating it doesn't matter. SOTA models aren't perfect, but are to the point that they're generally much better than the average developer. But once that perception set in and the meme spreads, it's really hard for some to break out of it. Especially at the pace AI development has been moving. It's just that simple.
- jodrellblank 27d ago> "Everything is variable reward. Is everything gambling?" well, no. If you work overtime and get paid overtime, you are not gambling and that is not a variable reward. Humans engage more with rewards that are intermittent and variable. Like Futurama's scene from 'The Scary Door' where the character says "A casino where I'm winning, I must be in heaven! A casino where I always win, that's boring, I must really be IN HELL!". A constant predictable reward is boring, less engaging. So if you know you get no overtime, but sometimes your boss rewards you with $5 coffee voucher, sometimes a free pizza dinner, sometimes double-time pay for the time worked or a half-day off, now you might be gambling 1hr overtime for an intermittent variable reward. > "Drive to work, traffic might be good or might be bad." Good traffic is not a "reward" for driving to work(!) and you have to drive to work regardless so you are not risking anything [you might be risking your life, but you are not making a choice which can reward you with good traffic]. You might say that going a different route is a choice and a gamble which could reward you with good traffic, but traffic engineering does not work that way because if there was a consistently low-traffic route, everyone else would take that route until it was no faster than any other route. Traffic will generally be the predictable and similar every day, plus 'arriving at work early' is not much of a reward.
- dpark 27d ago> A constant predictable reward is boring, less engaging. Perhaps but predictable outcome is a very desirable quality. No one wants a hammer that sometimes drives nails and sometimes doesn’t. All of the current harness engineering work is about squeezing predictability out of the LLM. > Good traffic is not a "reward" for driving to work(!) Like hell it’s not. I drove into work last Friday and there was no traffic because of the holiday weekend. It was amazing. Had me considering whether Friday should be one of my standard RTO days.
- jodrellblank 27d agoI think you're missing the point. Predictable outcomes are desirable, but they aren't addictive or gambling. People quickly get used to opening the faucet and seeing water come out and stop doing it, whereas people scroll TikTok or channel surf for hours at a time. In what way was amazing no-traffic "a reward"? What system was rewarding you for what change in behaviour?