6 ms·
To some extent it’s not that they don’t live up to the hype - rather that the gains are hard to measure. Llms have spared me hours of research on exotic topics
by Agingcoder 1y ago
To some extent it’s not that they don’t live up to the hype - rather that the gains are hard to measure.
Llms have spared me hours of research on exotic topics actually useful for my day job However, that’s the whole problem - I don’t know how much.
If they had a real price ( accounting for OpenAI losses for example) with ChatGPT at 50 usd/month for everyone, OpenAI being profitable, and people actually paying for this, I think things might self adjust and we’d have some idea.
Right now, we live in some kind of parallel world.
- alganet 1y ago> I don’t know how much. If you're not willing to measure how it helps you, then it's probably not worth it. I would go even further: if the effort of measuring is not feasible, then it's probably not worth it. That is more targeted at companies than you specifically, but it also works as an individual reflection. In the individual reflection, it works like this: you should think "how can I prove to myself that I'm not being bamboozled?". Once you acquire that proof, it should be easy to share it with others. If it's not, it's probably not a good proof (like an anecdote). I already said this, and I'll say it again: record yourself using LLMs. Then watch the recording. Is it that good? Notice that I am removing myself from the equation here, I will not judge how good is it, you're going to do it yourself.
- aeon_ai 1y agoI just did it. You were right. It is, in fact, that good.
- alganet 1y agoYou could have recorded, found it to be good, and didn't shared the news. Only used for your self. But you decided to share only the news, not the recording. That tells me something. To be more clear, I can move this argument further. I promise you that if you share the recording that led you to believe that, I will not judge it. In fact, I will do the opposite and focus on people who judge it, trying my best to make the recording look good and point out whoever is nitpicking.
- wongarsu 1y agoThere is a difference between confirming that something is worth it and quantifying the benefit though. One only requires satisfying a lower bound, the other requires an exact number. For example I use a $30/month chatbot subscription for various utility tasks. If I value my time at above $60/hour I need to save half an hour each month (a minute a day) to make the investment worth it. That is absolutely true, just with simple googleable questions and light research tasks I save much more than 7 minutes a week. But how much do I actually save? What exactly is my time actually worth? Those are much more difficult questions to answer
- brailsafe 1y agoFor me it's less important how much time I think I save on any discrete task, and how much time I net over that time, accounting for how much time the set of tasks I'd be working on would have otherwise taken had I just manually done them. Right now, that means debits and credits in the ledger of time. Sometimes I gain a lot on tasks I probably otherwise wouldn't have done, but I also don't gain much overall by doing, and sometimes I lose a ton of time simply by leaning on a loop of re-doing inaccurate agent work in a way that's actually more time intensive than had I internalized the system in working memory and produced functionality more slowly. If I save an hour, but lose 6, when I'd otherwise have spent 2, then I net -4, but sometimes overall it's positive, so the value is more ambiguous. If my employer didn't pay for the tools, I really don't know whether I would. A good price and conservative usage pattern might net more.
- alganet 1y agoThe user @brailsafe gave an answer that embodies some things I was going to say. You're accounting for the time wins, not accounting for the time losses. For a human chat user, that's when the LLM fails an answer or answers wrong. For an LLM coder, that's when context rot creeps in and you have to restart your work, and so on. There are people who don't care much if they are being bamboozled for $30/mo, they have nothing to prove nor grand expectations for the thing. To them, cargo culting might be fun and that's what they extract from the bargain. I am directing my answers mostly to people, companies or individuals, who have something to prove (evangelists, AI companies, etc). To those, a series of imperceptible small losses that results in debt in the long run is a big problem. My suggestion (the recording session) also works as a metaphor. That could be, instead of video, metrics about how contexts are discarded. It is, in that sense, also something they can decide to share or not, and the extent to what they share should be a sign of confidence in their product. Makes sense?
- mattlutze 1y ago> exotic topics [...] I don't know how much We also don't know, in situations like this, whether all of or how much of the research is true. As has been regularly and publicly demonstrated [0][1][2], the most capable of these systems still make very fundamental mistakes, misaligned to their goals. The LLMs really, really want to be our friend, and production models do exhibit tendencies to intentionally mislead when it's advantageous [3], even if it's against their alignment goals. 0: https://www.afr.com/companies/professional-services/oversight-was-not-followed-deloitte-apologises-for-ai-report-20251012-p5n1ty https://www.afr.com/companies/professional-services/oversigh... 1: https://www.nbcnews.com/world/australia/australian-lawyer-sorry-ai-errors-murder-case-fake-quotes-made-cases-rcna225220 https://www.nbcnews.com/world/australia/australian-lawyer-so... 2: https://calmatters.org/economy/technology/2025/09/chatgpt-lawyer-fine-ai-regulation/ https://calmatters.org/economy/technology/2025/09/chatgpt-la... 3: https://arxiv.org/pdf/2509.18058 https://arxiv.org/pdf/2509.18058?
- autoexec 1y ago> The LLMs really, really want to be our friend They want you to think they are your friend but they actually want to be your master and steal your personal data. It's what the companies who want to be masters over you and the AI have programed them to do. LLMs want to gain your confidence, and then your dependence, and then they can control you.
- dangus 1y agoThis seems hyperbolic to me. Sometimes companies just want to make money. Similarly, a SaaS company that would very much prefer you renew your subscription isn’t trying to make you into an Orwellian slave. They’re trying to make a product that makes me want to pay for it. 100% of paid AI tools include the option to not train on your data, and most free ones do as well. Also, AI doesn’t magically invalidate GDPR.
- ToucanLoucan 1y ago> This seems hyperbolic to me. Sometimes companies just want to make money. It's not hyperbolic at all. The entire moat is brand lock-in. OpenAI owns the public impression of what AI is- for now- with a strong second place going to Claude for coders in specific. But that doesn't change that ChatGPT can generate code too, and Claude can also write poems. If you can't lock users into good experiences with your LLM product, you have no future in the market, so data retention and flattery are the names of the game. All the transformer-based LLMs out there can all do what all the other ones can do. Some are gated off about it, but it's simulated at best. Sometimes even circumvent-able with raw input. Twitter bots regularly get tricked into answering silly prompts by people simply requesting they forget current instructions. And, between DeepSeek's incredibly resource-light implementations of solid if limited models, which do largely the same sort of work without massive datacenters full of GPUs, plus Apple Intelligence rolling out experiences that largely run on ML-specific hardware in their local devices which immediately, full stop, wins the privacy argument, OpenAI and co are either getting nervous, or they're in denial. The capex for this stuff, the valuations, and the actual user experiences are simply not cohering. If this was indeed the revolution the valley said it was, and the people were lining up to pay prices that reflected the cost of running this tech, then there wouldn't be a debate at all. But that's simply not true: most LLM products are heavily subsidized, a lot of the big players in the space are downsizing what they had planned to build out to power this "future," and a whole lot of people cite their experiences as "fine." That's not a revolution.