4 ms·
> It is very much like playing an instrument. Or it is more like playing a slot machine and you imagine the rest.
by h05sz487b 4mo ago
> It is very much like playing an instrument.
Or it is more like playing a slot machine and you imagine the rest.
- psychoslave 4mo agoI play slot machines as instrument! ;)
- dotancohen 4mo agoRoger Waters and Nick Mason were playing the cash register in 1973!
- glerk 4mo agoIt is a bit of both. A non-deterministic instrument and a predictable slot machine.
- cube00 4mo agoThis is how I feel whenever I see bold all caps instructions in a system prompt or someone claims they conducted "research" and found the magic prompt template that makes the model pay out. Maybe it works some of the time but it isn't a solution that works everytime. It reminds me of people hovering to play a slot machine when someone gets up and it hasn't paid out as if they've solved slot machines. While I don't mind putting something in a loop until the tests pass, I'm less comfortable doing that when providers are silently rerouting to lower quality models, or in Google's case burning quota faster to ease their own server load without being transparent about what the "standard limits" are to begin with. [1] I'm hopeful I'll be more comfortable with these "slot machines" when frontier models get to the point where they can be run locally on hardware I can actually afford so I know exactly what I'm getting and not jumping at shadows with providers playing tricks behind the scenes to ease their own load without admitting the customer is getting less for their money as they get more popular. [1]: https://support.google.com/gemini/answer/16275805?hl=en&sjid=16852286370759456931-NC#zippy=%2Cusage-limit-changes:~:text=Limits%20may%20change%20without%20notice%2C%20including%20due%20to%20capacity%20constraints.%20When%20there%E2%80%99s%20a%20large%20increase%20in%20activity%20in%20Gemini%20Apps%2C%20we%20may%20change%20limits%20to%20maintain%20a%20high%20standard%20of%20quality.%C2%A0 https://support.google.com/gemini/answer/16275805?hl=en&sjid...
- coldtea 4mo ago>This is how I feel whenever I see bold all caps instructions in a system prompt or someone claims they conducted "research" and found the magic prompt template that makes the model pay out. Maybe it works some of the time but it isn't a solution that works everytime. For such thing to be useful, it's enough that they works substantially more times that not having those instructions in.
- Planktonne 4mo agoEvery gambler thinks their system works, given enough chances.
- user43928 4mo agoHas there been any evidence of a well known provider rerouting to lower quality models? Last I saw, engineers working at OpenAI denied this on HN. I saw that someone set up a tracker that aims to record the performance of the models, and so far it has not shown any statistically significant deviation in performance for Codex, and not yet enough data for Claude: https://marginlab.ai/trackers/codex/ https://marginlab.ai/trackers/codex/
- cube00 4mo ago> Has there been any evidence of a well known provider rerouting to lower quality models? The firm [Anthropic] would deliberately degrade the model’s performance in ways that were invisible to the user. https://news.ycombinator.com/item?id=48485958 https://news.ycombinator.com/item?id=48485958
- dannyw 4mo agoYes, OpenAI admits they silently reroute sensitive requests to different models for user welfare at least: https://openai.com/index/building-more-helpful-chatgpt-experiences-for-everyone/ https://openai.com/index/building-more-helpful-chatgpt-exper... The implementation was so borked, SamA went back on Reddit and apologised: https://old.reddit.com/r/ChatGPT/comments/1o6jins/updates_for_chatgpt/ https://old.reddit.com/r/ChatGPT/comments/1o6jins/updates_fo... Model re-routing happens for coding tasks too. For example, in OpenAI support pages used to (at least 1 month ago when I checked) mention that if they automatically use a cheaper -mini to accomplish the task behind the scenes, you’ll be charged -mini prices even if you selected a more expensive model. I just checked again and they’ve removed it, but there’s probably archives. Finally, even if they’re the same weights, you don’t know what quantisation you’re running at. Adaptive quantisation based on load (given workday peaks), or similar techniques, have been happening since the ChatGPT 3.5 days; the techniques are probably more advanced now.
- ramon156 4mo agoInstruments are pseudo-random until you know what you're doing. Slot machines are just slot machines
- deleted 4mo ago[deleted]
- Forgeties79 4mo agoMusical instruments are not random. You’re just doing random inputs. Instruments are consistent, even if the “flavor” and quality varies with different builds. Playing a B on a saxophone always plays a B.
- dotancohen 4mo agoSaxophone, being a wind instrument was a bad choice. I can definitely tell which student was blowing when hearing a note. But your analogy remains solid if you substitute e.g. a piano and a reasonably proficient player. A single note would be nearly indistinguishable between players... But a full piece most certainly will sound different.
- palata 4mo agoWhile I agree with you, I think it's diverging from the initial point. The original take was "LLMs are very much like playing an instrument". I think they are very much NOT like playing an instrument. While different musicians will produce different results, one musician won't get drastically different results on different days or when trying a different "copy" of the same instrument. If you can play the violin on your violin and I lend you my violin, you will still be able to play very consistently. You may argue that the sound will differ and you will have to adapt slightly, but that's not remotely similar to the randomness coming from LLMs.
- tekne 4mo agoWill you? That's only if both violins are tuned the same way, and one must continually tune them lest they get out of sync. Similarly, an LLM can be extremely consistent if tuned properly -- indeed, if you fix the weights and settings, they can be made "essentially deterministic" for many prompts!
- hodgehog11 4mo agoA poor analogy depending on the setting because you can't adjust the odds with a slot machine, and the ROI is negative by design. If that's your experience, yeah, I wouldn't use an LLM either.
- victorbjorklund 4mo agoPretty sure most modern slot machines are digital and you could adjust the odds (even to a positive EV) if you change the code.
- hodgehog11 4mo agoYou're being unfaithful to the original statement. The whole point of saying something is like a slot machine is that there are significant odds that you lose. If you ever have access to a casino slot machine that has a positive EV, there are no tangible negative aspects anymore; you would use it over and over again and accumulate significant wealth from the house. That's my point.