6 ms·
I don't get your point. Web tools have been doing A/B feature testing all the time, way before we had LLMs.
by danielbln 7mo ago
I don't get your point. Web tools have been doing A/B feature testing all the time, way before we had LLMs.
- reconnecting 7mo agoThis is very different from the A/B interface testing you're referring to, what LLMs enable is A/B testing the tool's own output — same input, different result. Your compiler doesn't do that. Your keyboard doesn't do that. The randomness is inside the tool itself, not around it. That's a fundamental reliability problem for any professional context where you need to know that input X produces output X, every time.
- huflungdung 7mo ago[dead]
- orf 7mo agoIt’s exactly the same as A/B testing an interface. This is just testing 4 variants of a “page” (the plan), measuring how many people pressed “continue”.
- stavros 7mo agoYou've groupped LLMs into the wrong set. LLMs are closer to people than to machines. This argument is like saying "I want my tools to be reliable, like my light switch, and my personal assistant wasn't, so I fired him". Not to mention that of course everyone A/B tests their output the whole time. You've never seen (or implemented) an A/B test where the test was whether to improve the way e.g. the invoicing software generates PDFs?
- applfanboysbgon 7mo ago> LLMs are closer to people than to machines. jfc. I don't have anything to say to this other than that it deserves calling out. > You've never seen (or implemented) an A/B test where the test was whether to improve the way e.g. the invoicing software generates PDFs? I have never in my life seen or implemented an a/b test on a tool used by professionals. I see consumer-facing tests on websites all the time, but nothing silently changing the software on your computer. I mean, there are mandatory updates, which I do already consider to be malware, but those are, at least, not silent.
- johnisgood 7mo agoWhy are you calling it out? You are interpreting the statement too literally. The point is probably about behavior, not nature. LLMs do not always produce identical outputs for identical prompts, which already makes them less like deterministic machines and superficially closer to humans in interaction. That is it. The comparison can end here.
- applfanboysbgon 7mo agoThey actually can, though. The frontier model providers don't expose seeds, but for inferencing LLMs on your own hardware, you can set a specific seed for deterministic output and evaluate how small changes to the context change the output on that seed. This is like suggesting that Photoshop would be "more like a person than a machine" if they added a random factor every time you picked a color that changed the value you selected by +-20%, and didn't expose a way to lock it. "It uses a random number generator, therefore it's people" is a bit of a stretch.
- deleted 7mo ago[deleted]
- johnisgood 7mo agoYou are right, I was wrong. I think anthropomorphizing LLMs to begin with is kind of silly. The whole "LLMs are closer to people than to machines" comparison is misleading, especially when the argument comes down to output variability. Their outputs can vary in ways that superficially resemble human variability, but variability alone is a poor analogy for humanness. A more meaningful way to compare is to look at functional behaviors such as "pattern recognition", "contextual adaptation", "generalization to new prompts", and "multi-step reasoning". These behaviors resemble aspects of human capabilities. In particular, generalization allows LLMs to produce coherent outputs for tasks they were not explicitly trained on, rather than just repeating training data, making it a more meaningful measure than randomness alone. That said, none of this means LLMs are conscious, intentional, or actually understanding anything. I am glad you brought up the seed and determinism point. People should know that you can make outputs fully predictable, so the "human-like" label mostly only shows up under stochastic sampling. It is far more informative to look at real functional capabilities instead of just variability, and I think more people should be aware of this.
- doc_ick 7mo agoAs far as I can tell, llms never give the exact same output every time.
- johnisgood 7mo ago> same input, different result. What is your point? You get this from LLMs. It does not mean that it is not useful.
- freeone3000 7mo agoYes! And it was bad then too!! I want software that does a specific list of things, doesn’t change, and preferentially costs a known amount.