9 ms·
If I'm tired of one thing related to AI/llm/chatbots it's the claims that it's not useful. It 100% is. We have to separate the massive financial machinations fr
by tra3 1y ago
If I'm tired of one thing related to AI/llm/chatbots it's the claims that it's not useful. It 100% is. We have to separate the massive financial machinations from the actual tech.
Reading this article though, I'm questioning my decision to avoid hosting open source LLMs. Supposedly the performance of Owen-coder is comparable to the likes of Sonnet4. If I invest in a homelab that can host something like Qwen3 I'll recoup my costs in about 20 months without having to rely on Anthropic.
- mynameisash 1y agoI don't think I've ever seen anyone say they're not useful. Rather, they don't appear to live up to the hype, and they're sure as hell not a panacea. I'm pretty bearish on LLMs. I also think they're over-hyped and that the current frenzy will end badly (global economically speaking). Than said, sure, they're useful. Doesn't mean they're worth it.
- tra3 1y agoFair enough, I may have conflated "there's an AI bubble" with "AIs aren't useful". My employer pays for Claude pro access, and if they stopped paying tomorrow I'd consider paying for it myself. Although, it's much more likely for me to start self hosting them. So that's what it's worth to me, say $2500 USD in hardware over the next 3 years. I'd love to hear what your take on this is.
- brailsafe 1y ago$2500 is a relatively small investment for any sort of useful tool over 3 years, but that seems very low to me for this specific self-hosting endeavor
- Agingcoder 1y agoTo some extent it’s not that they don’t live up to the hype - rather that the gains are hard to measure. Llms have spared me hours of research on exotic topics actually useful for my day job However, that’s the whole problem - I don’t know how much. If they had a real price ( accounting for OpenAI losses for example) with ChatGPT at 50 usd/month for everyone, OpenAI being profitable, and people actually paying for this, I think things might self adjust and we’d have some idea. Right now, we live in some kind of parallel world.
- alganet 1y ago> I don’t know how much. If you're not willing to measure how it helps you, then it's probably not worth it. I would go even further: if the effort of measuring is not feasible, then it's probably not worth it. That is more targeted at companies than you specifically, but it also works as an individual reflection. In the individual reflection, it works like this: you should think "how can I prove to myself that I'm not being bamboozled?". Once you acquire that proof, it should be easy to share it with others. If it's not, it's probably not a good proof (like an anecdote). I already said this, and I'll say it again: record yourself using LLMs. Then watch the recording. Is it that good? Notice that I am removing myself from the equation here, I will not judge how good is it, you're going to do it yourself.
- aeon_ai 1y agoI just did it. You were right. It is, in fact, that good.
- alganet 1y agoYou could have recorded, found it to be good, and didn't shared the news. Only used for your self. But you decided to share only the news, not the recording. That tells me something. To be more clear, I can move this argument further. I promise you that if you share the recording that led you to believe that, I will not judge it. In fact, I will do the opposite and focus on people who judge it, trying my best to make the recording look good and point out whoever is nitpicking.
- wongarsu 1y agoThere is a difference between confirming that something is worth it and quantifying the benefit though. One only requires satisfying a lower bound, the other requires an exact number. For example I use a $30/month chatbot subscription for various utility tasks. If I value my time at above $60/hour I need to save half an hour each month (a minute a day) to make the investment worth it. That is absolutely true, just with simple googleable questions and light research tasks I save much more than 7 minutes a week. But how much do I actually save? What exactly is my time actually worth? Those are much more difficult questions to answer
- James_K 1y agoSomething not being useful is distinct from it having no uses. It could well be the case that the use of AI creates more damage than it does good. Many people have found it a useful tool to create the appearance of work where none is happening.
- QuantumGood 1y agoThose that claim not useful usually link it to something like "never trust because hallucinations", or backtrack when called out like "yes, I should have added details", or speak of problems outweighing usefulness hence not useful, etc. But online, people do make this statement.
- thatjoeoverthr 1y agoThe thing with the hype is it's always the same hype. "If you can just 3D print another 3D printer ..." "Apps are dead, everything will be AJAX" etc. I no longer believe the hype itself warrants attention or pushback. Let the hype boys raise money. No need to protect naive VCs.
- bigfishrunning 1y agoBut if the hype boys manage to capture big portions of the market (Microsoft, Amazon, etc...) it starts affecting pensions and retirement accounts. The next few years are gonna be rough because of this hype.
- watwut 1y ago> Let the hype boys raise money. No need to protect naive VCs. I genuinely 100% believe that ability of hype boys to raise money is harming the economy and us all. Whatever structural reason for it existing is there, it would be the best to end it.
- peteforde 1y agoMy dude, there is a small but weirdly dedicated group of people on this site that are hellbent on demanding "proof" that the wins we've personally gained from using LLMs in an intelligent way are real. It's actually been kind of exhausting, leading me to not weigh in on many threads.
- hatthew 1y agoBecause there's a lot of evidence that people tend to overestimate/overstate how useful LLMs are. Everyone says "I wrote this thing using AI" but most of the time reading the prompt would be just as useful as reading the final product. Everyone says "I wrote this large codebase using AI" but most of the time the code is unmaintainable and probably could have been implemented with much less code by a real human, and also the final software isn't actually ready for prod yet. Everyone says "I find AI coding very useful" and neglects to mention that they are making small adhoc scripts, or they're in a domain that's mostly boilerplate anyways (e.g. some parts of web dev). The one killer application of LLMs seems to be text summarization. Everything else that I have seen is either a niche domain that doesn't apply to the vast majority of people, a final product that is slop and shouldn't been made in the first place, or minor gains that are worthwhile but nowhere near as groundbreaking as people claim. To be clear, I think LLMs are useful, and I personally use them regularly. But I've gained at most 5% productivity from them (likely much less). For me, it's exhausting to keep on trying to realize these gains everyone is talking about, while every time I dig into someone claiming to get massive gains I find that the actual impact is highly questionable.
- peteforde 1y agoYour position implies that we need to prove that we're not smoking our own supply. I would argue that you are the one who should prove that we're not working (conservatively) 5-8x faster. The most telling part is when you said "most of the time reading the prompt". That strongly implies that you're attempting to one-shot whatever it is that you're working on. There is no "the prompt" in my current application. It's a 275k LoC ESP-IDF app spread across ~30 components that interact via FreeRTOS mechanisms as well as an app-wide event bus. It manages non-blocking UI, IO over multiple protocols, drives an OLED using a customized version of lvgl. It is, by any estimation, a serious and non-trivial application, and it was almost entirely crafted by LLM coding models being closely driven by yours truly across several hundred distinct Cursor conversations. It's probably taken me 10% of the time it would have taken me to do by hand, and that's precisely because I lean on it so heavily for initial buildout, thoughtful troubleshooting (it is never tired, never not available, and also knows more than I do about electronics as a bonus) and the occasional large cross-component refactor. I don't suspect that you're wrong. I know that you're wrong.
- llm_nerd 1y ago>I don't think I've ever seen anyone say they're not useful. https://news.ycombinator.com/item?id=45577203 https://news.ycombinator.com/item?id=45577203 There are thousands and thousands of comments just like this on this site. I would dare say tens of thousands. They regularly appear in any AI-related discussion. I've been involved in many threads on here where devs with Very Important Work announce that none of the AI tools are useful for them or for anyone with Real Problems, and at best they work for copy/paste junior devs who don't know what they're doing and are doing trivial work. This is right after they declare that anyone that isn't building a giant monolithic PHP app just like them are trend-chasers who are "cargo culting, like some tribe or something". >I also think they're over-hyped and that the current frenzy will end badly (global economically speaking) In a world where Tesla is a trillion dollar company based upon vapourware, and the president of largest economy (for now) is launching shitcoins and taking bribes through crypto, and every Western country saw a massive real-estate ramp up by unmetered mass migration, and Bitcoin is a $2T "currency" that has literally zero real world use beyond betting on itself, and sites like Polymarket exist for insiders to scam foolish rube outsiders out of their money, and... Dude, the AI bubble doesn't even remotely measure.
- deleted 1y ago[deleted]
- 1vuio0pswjnm7 1y ago"I don't think I've ever seen anyone say they're not useful." That's because no one has said that "AI" hype is the issue, not "AI" The hype machine and its followers have no tolerance for skepticism Any perceived skepticism of "AI", no matter how reasonable, triggers absurd accusations The author, like many others, tries to avoid the kneejerk defensiveness of "AI" hype subscribers: "Don't get me wrong: I am not denying the extraordinary potential of AI to change aspects of our world, nor that savvy entrepreneurs, companies and investors will win very big. It will - and they will." But this does not work. There is zero tolerance for skepticism. All disbelief must be countered "Crypto" hype was like this, before one of its ringleaders went to prison It's unlikely that fraud will be prosecuted under current political environment Fasten your seatbelts
- mrbungie 1y agoIt's hell useful, I use Cursor several times a week (and I'm not even working as a dev full time rn), and ChatGPT is my daily driver. Yet, it's weird to me that we're 3 years into this "revolution" and I can't get a decent slideshow from an LLM without having to practically build a framework for doing so.
- jacobr1 1y agoIt is a focus, data, and benchmarking problem. If someone comes up with good benchmarks, which means having a good dataset, and gets some publicility around, they can attract the frontier labs attention to focus training and optimization effort on making the models better for that benchmark. This is how most the capabilities we have today have become useful. Maybe there is some emergent initial detection of utility, but the refinement comes from labs beating others on the benchmarks. So we need a slideshow benchmark and I think we'd see rapid improvement. LLMs are actually ok at a building html decks, not great, but ok. Enough so that if we there was some good objective criteria to tune things toward I think the last-mile kinks would get worked out (formats, object/text overlaps). the raw content is mainly a function of the core intelligence of model, so that wouldn't be impacted (if you get get it to build a good bullet-point markdown of you presentation today it would be just a good as a prezo, but maybe not as visually compelling as you like. Also this might need to be an agentic benchmark to allow for both text and image creation and other considerations like data sourcing. Which is why everyone doing this ends up building their own mini framework. A ton of the reinforcement type training work really just aligning the vague commands a user would give to the same capability a model would produce with a much more flushed out prompt.
- huevosabio 1y agoThe problem of self-hosting is that you increase the friction to swap models and use whatever is SOTA or whatever fits your purpose best. Also, I've heard from others that the Qwen models are a bit too overfit to the benchmarks and that their real-life usage is not as impressive as they would appear on the benchmarks.
- acutesoftware 1y agoSwitching models when running locally is fairly easy - as long as you have them downloaded you can switch them in and out with a just a config setting - cant quite remember, but you may need to rebuild the vectorstore when switching though. LangChain has the embeddings for major providers: def build_vectorstore(docs): """ Create vectorstore from documents using configured embedding model. """ # Choose embedding model if cfg.EMBED_MODEL.lower() == "openai": embeddings = OpenAIEmbeddings(model="text-embedding-3-small") elif cfg.EMBED_MODEL.lower() == "huggingface": from langchain_community.embeddings import HuggingFaceEmbeddings embeddings = HuggingFaceEmbeddings(model_name="sentence-transformers/all-MiniLM-L6-v2") elif cfg.EMBED_MODEL.lower() == "nomic-embed-text": from langchain_ollama import OllamaEmbeddings embeddings = OllamaEmbeddings(model=cfg.EMBED_MODEL)
- silversmith 1y agoThe issue is that the field is still moving too fast - in 20 months, you might break even on costs, but the LLMs you are able to run might be 20 months behind "state of the art". As long as providers keep selling cheap inference, I'm holding out.
- wmf 1y agoThe gap between local models and SOTA is around 6 months and it's either steady or dropping. (Obviously this depends on your benchmark and preferences.)
- criddell 1y agoSeriously? So I can run the best models from 2024 at home now? For example, what would I need to run Open AI's o1 model from 2024 at home? Are there good guides for setting this up?
- wmf 1y agoIt's not the same model, but for example GPT-OSS-120B is smarter than o1. The guide is buy 128 GB of VRAM then install LM Studio.
- criddell 1y agoAn NVIDIA 5090 with 128 GB of VRAM is $13k. It doesn’t make any sense to run that at home when you can pay OpenAI $20 / month to use it (it would take more than 50 years to spend $13k at OpenAI this way). So technically you might be able to run a six month old model at home, but it would be foolish to do so from a financial point of view. Or is there a way to get 128 GB of VRAM for a lot less than that?
- wmf 1y agoRyzen AI Max is $2,000, M4 Max is $3,500, and DGX Spark is $4,000. Still not really economically feasible but I see it as an insurance policy. And that's the most expensive local model; smaller models will run on any PC.
- didibus 1y ago> it's the claims that it's not useful I think the reason is because it depends what impact metrics you want to measure. "Usefulness" is in the eye of the beholder. You have to decide what metric you consider "useful". If it's company profit for example, maybe the data shows it's not yet useful and not having impact on profit. If it's the level of concentration needed by engineers to code, then you probably can see that metric having improved as less mental effort is needed to accomplish the same thing. If that's the impact you care about, you can consider it "useful". Etc.
- imiric 1y ago> If I'm tired of one thing related to AI/llm/chatbots it's the claims that it's not useful. It 100% is. We have to separate the massive financial machinations from the actual tech. It's indisputable that the tech is and can be very useful, but it's also surrounded by a bubble of grifters and opportunists riding the hype and money train. The sooner we start ignoring the "AI", "ASI", "AGI", anthropomorphization, and every other snake oil these people are peddling, the sooner we can focus on practical applications of the tech, which are numerous.
- reissbaker 1y agoQwen3 Coder unfortunately isn't on par with Sonnet, no matter what the benchmarks say. GLM-4.6 does feel pretty competitive though. You'll need a pretty expensive home lab to run it though... I'd be surprised if you could do it at long context with only 20 months of Sonnet usage.
- Octoth0rpe 1y ago> It 100% is [useful] It's worth disambiguating between "worth $50b of investment" useful versus "worth $1t of investment" useful
- mattlutze 1y agoEspecially when, as it is currently in vogue to observe, the difference between $50b and $1t is roughly $1t.
- marcosdumay 1y agoUp to 2 significant figures...
- pseudosavant 1y agoFor perspective, there are 10 companies with a market cap over $1T. Is the value of LLMs greater than Tesla? Absolutely. The problem of course is that plenty of that $1T in investment will go to stupid investments. The people whose investments pan out will be the next generation of Zuckerbergs. The rest will be remembered like MySpace or Webvan.
- rz2k 1y agoTo be fair, while the incremental value of each additional year that Tesla remains in existence may not be so great, it did finally change the conventional wisdom about the viability of electric vehicles which will continue to have substantial impact. Furthermore the price of the most recently sold share times the number outstanding does not represent the total R&D or spending to make Teslas.
- pseudosavant 1y agoI'll add that MSFT, AAPL, GOOGL, AMZN, and META generated >$450B in net income in the last 4 quarters. It can't be overstated how much profits they can burn on AI without losing money.
- ants_everywhere 1y agoThe other thing that's tiring is talking about how AI is a bubble as if that's an indictment of AI. Being a bubble is a statement about the value of the stock market, not about the technology. There was a dotcom bubble, but that does not mean the internet wasn't valuable. And if you bought at the top of the dotcom bubble you'd be much wealthier now than you were when you bought. But it would have taken you a significant time to break even.
- decimalenough 1y agoIf you bought ETFs, yes, but not if you bought Pets.com and Yahoo.
- ants_everywhere 1y agoRight, which is a distinction that matters if you have a sensible view of what it means to be in a bubble. But many people talking about AI being a bubble aren't trying to figure out which ticker is going to win in the long run, they're trying to convey a belief that AI is bogus altogether. There's widespread agreement that nobody knows whether the AI valuations we see are right. What I'm saying is tiring is people who confuse that idea with an indictment of the technology.
- player1234 1y agoIt say something about how useful is it when it is not subsidiced. The bubble exists be cause no AI company gets any profits and laughable revenue. Is it 2000$/month useful is the interesting question.
- arjie 1y agoI used Qwen3-480B-Coder with Cerebras and it was not very good for my use case. You can run these models online first to see if they will work for you. I recommend you try that first.
- mrdependable 1y agoThey are useful, but I find it is only slightly more convenient than a Google search. Losing something like GPS on my phone would be a much bigger disruption to my life.
- criemen 1y ago> Supposedly the performance of Owen-coder is comparable to the likes of Sonnet4. If I invest in a homelab that can host something like Qwen3 I'll recoup my costs in about 20 months without having to rely on Anthropic. You can always try it via openrouter without investing in the home setup first. That allows you to evaluate whether it hits your quality bar or not, and is much cheaper. It is less fun than self-hosting though.
- noosphr 1y agoI had a hilarious exchange on here where I used an LLM to explain to a poster at length why they fundamentally didn't understand what I said. It did a bang up job. The poster, and a lot of other people, got mad I used AI and they still didn't understand my original post, or the AI explanation. LLMs aren't terribly useful to people who fundamentally can't read. When those people can also type very fast you get the current situation.
- Jensson 1y ago> I used an LLM to explain to a poster at length why they fundamentally didn't understand what I said. It did a bang up job. The poster still didn't understand my original post. It didn't do a bang up job if the poster still didn't understand you, so sorry this example doesn't prove what you think it does. You have to measure actual results, your own take will always be biased so you can't say "I thought it was great but it didn't work" and expect people to get convinced by that. Edit: And if that doesn't convince you, why not read what this AI has to say about your post, if you like them so much you should read this right and acknowledge you were wrong just like you expected those people to: https://chatgpt.com/s/t_68f2ae740f98819183539767b921965b https://chatgpt.com/s/t_68f2ae740f98819183539767b921965b
- noosphr 1y agoClaude, explain the fallacy fallacy to someone fond of pointing fallacies in arguments: #### *The Fallacy Fallacy: A Metacognitive Error in Logical Analysis* The fallacy fallacy, also known as the argument from fallacy or argumentum ad logicam, represents a second-order logical error wherein one incorrectly infers that a conclusion must be false solely because it has been argued through fallacious reasoning. This metacognitive error constitutes a significant impediment to rigorous philosophical discourse and warrants careful examination. #### *Theoretical Framework and Definition* Within the domain of informal logic, fallacies constitute "mistakes of reasoning, as opposed to making mistakes that are of a factual nature". The fallacy fallacy emerges when interlocutors conflate the validity of argumentative structure with the truth value of propositional content. Specifically, this error manifests when one advances the following invalid inference pattern: 1. Argument X contains logical fallacy F 2. Therefore, the conclusion C of argument X is false This inference pattern itself represents a non sequitur, as the presence of fallacious reasoning does not necessarily bear upon the truth or falsity of the conclusion in question. #### *Epistemological Implications* The commission of the fallacy fallacy reveals a fundamental misunderstanding of the relationship between logical validity and factual accuracy. *Truth values of propositions exist independently of the quality of arguments marshaled in their support*. A proposition may be demonstrably true despite being defended through specious reasoning, just as a false proposition may be supported by formally valid argumentation with false premises. Consider the following syllogistic example: - Major premise: All mammals are warm-blooded - Minor premise: Dogs are mammals because they bark - Conclusion: Dogs are warm-blooded While the minor premise employs irrelevant reasoning (dogs' classification as mammals is unrelated to their vocalization), the conclusion remains factually correct. #### *Methodological Considerations for Critical Analysis* Scholars engaged in the identification of logical fallacies must exercise epistemic humility regarding the scope of their critique. As noted in the academic literature, "fallacies are common errors in reasoning that will undermine the logic of your argument", yet this undermining pertains exclusively to the argumentative structure rather than to the ontological status of the conclusion. The appropriate scholarly response to encountering fallacious reasoning involves: 1. *Methodological separation* - Distinguishing between the evaluation of argumentative form and the assessment of propositional content 2. *Constructive engagement* - Requesting alternative justification rather than dismissing conclusions outright 3. *Epistemic charity* - Acknowledging that interlocutors may possess valid intuitions despite articulating them through flawed logical frameworks #### *Conclusion* The fallacy fallacy represents a particularly insidious form of intellectual error, as it masquerades as sophisticated logical analysis while itself committing a fundamental category mistake. Academics and scholars must remain vigilant against this metacognitive trap, recognizing that the identification of fallacious reasoning, while valuable for improving argumentative rigor, does not constitute sufficient grounds for rejecting the truth claims embedded within poorly constructed arguments. The pursuit of truth demands that we evaluate propositions on their merits, independent of the quality of their initial presentation.
- electroglyph 1y agoyou need at least an RTX 6000 pro, maybe 2 to run local models on that level. probably only worth it if you plan on doing other workloads like finetuning or generating a lot of synthetic data
- dvfjsdhgfv 1y ago> If I'm tired of one thing related to AI/llm/chatbots it's the claims that it's not useful. That is the best example of straw argument I've seen this year. I enjoy reading discussions on LLMs and have seen a huge number of arguments, some reasonable and some ridiculous, but one thing I haven't seen is someone claiming that LLMs are not useful. We can discuss usefulness for a particular purpose, or the level of its fitness for it, but not the fact that millions of people find LLMs useful enough to pay for them.
- somewhereoutth 1y ago> the claims that it's not useful There are many credible claims that not only is it not useful, but that it is actually causing serious damage.
- andrepd 1y agoIt might be "useful" as in "it has a non-zero number of use cases", and still being massively overhyped (orders of magnitude in my opinion). I guess there are use cases for it, even if we discount undisputed net negatives like the proliferation of slop online, scam calls, deepfakes, etc. That doesn't mean it provides an amount of utility that justifies pivoting a significant portion of world capital and production towards that end. It will never be AGI, by the way. We are way past the inflection point of the logistic curve, so this is more or less what it is.
- zmmmmm 1y ago> If I invest in a homelab that can host something like Qwen3 I'll recoup my costs in about 20 months without having to rely on Anthropic For me it's equally that I don't trust any of these service providers to keep maintaining whatever service or model I'm relying on. Imagine if I build a whole entire process and then the bubble bursts and they either take away what I'm using or start charging outrageous amounts for it. I feel we are well into the point where the base technology is useful enough and all the work is in how you implement and adapt it in to your process / workflow. A new model coming out that is 3% better is relatively meaningless compared to me figuring out how better to integrate what I already have which might give me a 20% bump for very little effort. So at this point all I really want is stability in the tech so I can optimise everything else. Constant churn of hosted providers thrusting change at me every second day is actively harmful to my productive use of it at this point. Hence I want local models so I can just tune out the noise and focus on getting things done.
- satisfice 1y agoFew people say they are not useful. But when people like me say they aren’t reliable and worthy of trust, AI fanboys like to pretend we are saying something else.
- Ferret7446 1y agoDoes that include electricity and maintenance costs?
- KronisLV 1y ago> Supposedly the performance of Owen-coder is comparable to the likes of Sonnet4. If I invest in a homelab that can host something like Qwen3 I'll recoup my costs in about 20 months without having to rely on Anthropic. Presently, look up the Cerebra Coder subscription. It’s cut down my reliance on paying per token by about 80% due to the model being good for most development tasks and the rate limits are such that I never hit them per day, alongside being faster than anything else out there. Lots of folks also just explore new models on OpenRouter as they come on, albeit they don’t seem to have caching support so it can get expensive. Aside from that, self-hosting can be worth it but you need lots of memory and beefy compute to have good performance without quantizing things super far. There’s a really big difference between the 30B and 480B versions of Qwen Coder and while the smaller models are getting better, feels like there are diminishing returns there.