16 ms·
LLM discourse needs more nuance
- BonoboIO 4y agoI think everybody overestimates in the beginning, here it’s the same. AI will solve every problem, nope it won’t, but it will help.
- teekert 4y agoI feel like we already passed that low point, I mean, years ago that was a promise and a trend, then I got disillusioned, but now, talking to ChatGPT is actually making me more positive about AI again.
- nikanj 4y agoNot sure about ChatGPT, but Midjourney et al have definitely transformed many markets already. Why would I get any illustrations from fiverr.com, when I can get better turnaround times for $0 from Midjourney?
- Forgeties79 4y agoCorrect me if I’m wrong, but aren’t systems like Midjourney inadvertently stealing from the people who are paid to make it? Also, in some ways that’s like asking “why should I shop local when I can buy from Walmart and get the exact same product?“ Well obviously if it’s all down to simple dollar costs, you do you. But there are several motivations/considerations for choosing why we shop where we shop.
- NateEag 4y agoI suspect you're wrong that the stealing is inadvertent. I'm suspicious the makers of the art generator AI systems know exactly what they're doing and see no issues with it.
- Forgeties79 4y agoI guess I’m just trying to be a little generous. They probably know it’s a risk and either rationalize it away with something along the lines of “well, it’s such a tiny minuscule piece from so many different people that it’s not really stealing“ or just pretend they don’t know. At this point the ignorance does seem willful. Either way, I’m more just curious what other people have to say on the matter as I am not as knowledgeable of subject. I am on the creator side and it’s all bad news over here.
- pixl97 4y ago> “well, it’s such a tiny minuscule piece from so many different people that it’s not really stealing“ RMS wrote 'the right to read' because issues like this, and if you've not read it, I recommend you do. "A bunch of little bits of different things" isn't stealing, it's society and culture. If you got your wish of everything I create if fully copyrighted, you'd find two things. One, that you're last in line and don't come up with original ideas often, even in original artwork. And two, that monied corporations would quickly buy up rights to everything and make life as an artist completely impossible.
- Forgeties79 4y agoI'm not even sure that's what is happening to be honest, I'm just speculating as to how somehow could reason it away for themselves. I don't know what the line is for "too much inspiration" where it becomes stealing. I'll make sure to check that out though, appreciate it!
- NateEag 4y agoYour choice to be generous is a kind and good one. I don't have great insight into copyright issues, but as one who loves music and used to play a lot of it, I'm right there with you on the "it's all bad news over here". It seems to me that the current crop of art generators ape styles but don't come up with their own. I worry that if they put most human artists out of work, the arts will stagnate badly. I guess we'll see.
- ben_w 4y agoEven lawyers aren't sure, the arguments can go either way: https://youtu.be/G08hY8dSrUY https://youtu.be/G08hY8dSrUY Also note some new research: https://arxiv.org/abs/2301.13188 https://arxiv.org/abs/2301.13188
- KaiserPro 4y agoSo the legal thinking I have been reading is around the reproduction of copyrighted works. However to get to a point where you _can_ recreate copyrighted works implies copyright infringement in the dataset. Thats not really been tackled. Also, there is the GDPR angle here as well, collecting people's faces on the internet for something other than the person reasonably intended is also on grey legal ground I would argue.
- runako 4y agoI tried to use a couple of the AI image generators when building a landing page recently, so I can answer this. Look at the homepage for a popular site like Stripe, Figma, Digital Ocean, whatever, take your pick. Try to get Midjourney to generate high-quality art like you see on those homepages, in a format where it is ready to go directly into a site without further processing in an image tool. My experience was that I could have gotten better results at Fiver.
- plif 4y agoThere's no way you're going to get top quality like those sites at a place like Fiverr either. They've spent millions on branding and marketing. I can see Midjourney already replacing the low end. The question is how good can it get, and then as the bar is raised what can be done to differentiate. Answer to the latter ironically may be going back to old school human interfaces.
- runako 4y ago> There's no way you're going to get top quality like those sites at a place like Fiverr either. There is a grain of truth to this, but again just choose whatever lower-budget website you want. And then just try the exercise I suggested. Midjourney isn't producing work at a "bad Fiverr" quality level, it's not producing usable work at all yet (at least without a lot of hands-on prompt tuning, at which point why not just use Fiverr?). At least with Fiverr, I am likely to end up with something I can put on the site, which is not true of Midjourney yet.
- fock 4y agosupposedly the Fiverr-contractor soon just will be a prompt-tuner with a large cache of "good" images.
- ghaff 4y agoThe language models are probably more actually useful today. I've played around with writing hypothetical articles using them. For topics I know where it's a fairly straightforward, e.g. Five qualities some job role needs, it does a "decent" job. Decent in this context doesn't mean something I could just hand to an editor. But it does mean a pretty good starting point that I could amend, flesh out, add some links, add a quote or two. I could certainly see using it to give me a sort of pre-draft on some fairly evergreen topic. I could also use it to generate some boilerplate definitions or historical background to include in an article. But, sadly, I think you'll see the LLMs being used to generate a lot of blog/article content with a minimum of human effort for even less than the small amount being paid for a lot of this today.
- mnd999 4y agoYou’re definitely screwing the other person in the prisoners dilemma.
- rco8786 4y agoIs the author conflating the cost to train an AI model with the cost to use the resulting models, here? I can certainly believe that model training cost is exponential as the # of parameters goes exponential. But that is a one time(-ish) cost relative to actually using those models. Or am I completely off base here?
- habibur 4y agoI pretty much always wanted to know the cost of running those models, not training those. That should be order of magnitude less.
- jackmott42 4y agostable diffusion will generate an image in ~30 seconds on an NVidia 3060, so the cost is on the order of cents per image, power use on the order of 170 watts during the generation, or 0.0014166666667 kwh. @ around 5 cents per kwh
- rileymat2 4y agoYou also have to calculate the life expectancy of the card against 30 seconds of its life as an ongoing cost.
- MagnumOpus 4y agoWhich is also on the order of 0.01c per query (answering 5 million queries in the lifetime of a $500 card).
- smulc 4y agoJust ask ChatGPT what the cost is? ;)
- npongratz 4y agoSure, but just remember that it can't do math, so any mathy answer it gives is just bullshit.
- tmountain 4y agoIt’s an interesting situation because on one level, the technology is at least somewhat misrepresented (best case results rising to the top of everyone’s feeds), but on the other side, we haven’t even begun to see what integrated Chat GPT will bring. If I could replace Siri with Chat GPT, I would do so immediately, as it’s objectively better. The author makes some points about the eye watering costs; however, if this is something that people truly want—and I think we do want it—the market will find a way to deliver it at scale in a cost effective way.
- flerchin 4y agoDo we have an idea of what the costs to run ChatGPT might be? Are we costing OpenAI $5 everytime we ask it to generate python in the style of Donald Trump?
- scaredginger 4y agoI've read estimates of $0.02 per query
- deleted 4y ago[deleted]
- HPsquared 4y agoProbably training it on specific data sets and contexts is the expensive part.
- scaredginger 4y agoRunning a gigantic model across multiple GPUs isn't cheap either
- Const-me 4y agoI’ve asked ChatGPT about hardware. Here’s the response: > As for your specific computer, with 64 GB of RAM and a high-performance GPU like the GeForce 1080 Ti, it should have sufficient resources to run a language model like me for many common tasks. Based on the models open sourced by OpenAI, they are using PyTorch and CUDA. This means their stack requires nVidia GPUs. I think the main reason for their high costs is a single sentence in the EULA of GeForce drivers: https://www.datacenterdynamics.com/en/news/nvidia-updates-geforce-eula-to-prohibit-data-center-use/ https://www.datacenterdynamics.com/en/news/nvidia-updates-ge... It’s technically possible to port their GPGPU code from CUDA somewhere else. Here’s a vendor-agnostic DirectCompute re-implementation of their Whisper model: https://github.com/Const-me/Whisper https://github.com/Const-me/Whisper On servers, DirectCompute is not great ‘coz Windows server licenses are expensive. Still, I did that port alone, and spent couple weeks doing that. OpenAI probably has resources to port their inference to vendor-agnostic Vulkan Compute, running on Linux servers equipped with reasonably-priced AMD or Intel GPUs. For instance, Intel A770 16GB only costs $350, but delivers similar performance to nVidia A30 which costs $16000. Intel consumes more electricity but not by much, 225W versus 165W. That’s like 40x difference in cost efficiency of that chat.
- api 4y agoThe part about compute cost ruining the economics of AI companies is somewhat hilarious and eye rolling to me for the following simple reason: users have computers right in front of them. Some models are small enough to run locally. For the rest is it conceivable to ask the customer to contribute cycles? Especially for anything cheap or free? There might be major technical challenges here but that’s not my point. My point is that I get the distinct impression this has barely occurred to most of these people as a possibility to even explore. Are we so far into peak SaaS now that the industry has forgotten about local compute the same way we forgot about data centers during peak PC?
- beepbooptheory 4y ago> Are we so far into peak SaaS now that the industry has forgotten about local compute the same way we forgot about data centers during peak PC? I mean, yes? You can host your own email server at home, but Google has still made billions selling email. Regardless, paying a company to allow them to run cycles of their proprietary model on my hardware doesn't feel like a great deal to me. Why would I choose that?
- api 4y agoIf it cost less? Or if you bought the model and ran it locally instead? The reason it’s hard to DIY a mail server is spam and all the attendant black lists and white lists. It’s not a resource problem.
- HPsquared 4y agoIt's not just a question of cycles; training these large models requires a lot of RAM and I/O, far beyond what's on a personal machine. Training these kinds of large models is not conducive to being parcelled-out in small chunks.
- api 4y agoTraining maybe but most of a model’s use is evaluation.
- didntreadarticl 4y agoWhile personally, the experience lived up to some of the hype, social media claims of it implementing entire front ends turned out to be fake. Clearly, this models performs well on in-sample tasks, and is far off on anything else. That wasn't my experience. ChatGPT dealt with everything I threw at it, except when it refused to, and even then I could get round it by telling it to pretend etc. Every wild example that I saw on twitter I was able to reproduce
- f0e4c2f7 4y agoAgree. Here is a twitter thread with video examples of using ghostwriter, an LLM similar to ChatGPT to build several example sites. https://twitter.com/Replit/status/1620445121476202497?t=6Mm9ErhDSBRoNMz-sozauQ&s=19 https://twitter.com/Replit/status/1620445121476202497?t=6Mm9... I think the nuance here is that using an LLM is a bit like googling. It doesn't seem like it would be a skill but it is. You kind of have to nudge it in a way similar to what you would do in your own mind after finding the almost relevant stackoverflow post.
- timdaub 4y agoThings it performed badly for me (our of sample): - Implement quadratic voting (a concept defined in Radical Markets by Weyl). When trying to explain it, it still invented smth that looked good - Implement a First in first out calculation to file tax returns where within a 1 year holding period, tax rate is 0%. It calculated some really implausible stuff and it seemed easier to just figure it out myself. - there were many such cases for me, but also good outcomes too
- rvz 4y agoAs long as these AI models are unable to transparently explain themselves then all of this is essentially a vacuum of hype to those who don't understand that such black box models are essentially 'throw more money and data on it' for the gamble on accuracy and the risk of overfitting and wasting money. No different to the majority of proof-of-waste cryptocurrencies like Bitcoin.
- ramesh31 4y agoI don't think anyone is being "deceived" when it comes to ChatGPT. In fact I've found it to be just the opposite. People are still completely unaware of how game changing and revolutionary it is until they sit down and actually use it. It's impossible to describe to someone just how different this is than anything else before, as the AI well has been so thouroughly poisoned by the false promises of prior tools.
- distantsounds 4y agoImagine spending hundreds of thousands of dollars on infrastructure to serve out content that isn't even yours to begin with, and then wonder why there's such a blowback in its usage
- jimbokun 4y agoIs this a good summary? LLMs have solved the language processing problem. Its responses are fluent and rapidly becoming indistinguishable from human output. Or if anything, its responses are too good to be mistaken for an average human. However, that’s different from producing accurate knowledge and insights. It’s often bullshitting, in the sense that it’s convincing prose often turns out to not reflect reality. The reference problem seems to be one example of this.
- soneca 4y agoPeople were worried that if we created an AI on our image, it would be a belligerent entity. But we did create an AI on our image, and it’s a bullshitter.
- api 4y agoHook a gun up to it and you have a belligerent bullshitter. Seems fairly accurate for a good bit of humanity.
- recuter 4y agoSeems to me it is more of a babbler. You know, like a young toddler just learning to talk.
- dandellion 4y agoOr a million monkeys with typewriters, only these monkeys are not entirely random, they have been conditioned to write things similar to what people actually write. Most of the time the output looks very natural, but others you get nonsense because the noise is too loud.
- sph 4y agoOn a philosophical level, I wonder if ChatGPT gives us mostly nonsense because most human output is bullshit. To paraphrase a famous quip, what if this is proof that 99% of everything humans have created is bullshit? The problem perhaps has never been machine learning but identifying the process that enables we humans to somehow go through life and create objective societal and scientific progress out of a massive pile of nonsense. I might be completely wrong here, but the next ChatGPT will assume this comment of mine to be absolute truth and happily ingest it wholesale.
- RandomLensman 4y agoIn a sense, some people are getting ahead of themselves - which by the way is normal in technological cycles, with typically financial innovation going beyond what the underlying innovation can support. If AI fails to move into high risk applications and instead stays largely confined to low risk applications, then the revenue pools will be stay limited to more consumer and small business facing services. Those are big pools but it will also not create the trust needed for certain high risk applications.
- rvz 4y ago> If AI fails to move into high risk applications and instead stays largely confined to low risk applications, then the revenue pools will be stay limited to more consumer and small business facing services. This. The output from the majority (if not all) of black box neural net based AI is not transparent or trustworthy and cannot explain its own decisions. Usually it either gets confused on a single pixel or generates garbage output confidently; like what ChatGPT does and there is no transparent way for it to explain its decisions. It is for that reason why it is unsuitable for applications that involve high risk to human safety, financial matters and legal situations which all three are life changing which the AI hypesquad seems to be forgetting.
- estevaoam 4y agoWhy most people are blind to the most important point of those models: the progress is moving at astonishing pace. IMO it is a useless discussion to debate about the current SOTA or economics. In two years we will certainly have some breakthrough as we have been seen for the past years. And things are speeding up. edit: typo.
- fock 4y ago> we will certainly have some breakthrough as we have been seen for the past years. And things are speeding up. In the past years we have seen a tremendous scaling of a) workforce, b) hardware, c) data. For me a breakthrough would be on data/energy efficiency? What areas are promising there?
- Jensson 4y ago> And things are speeding up. Gpt-3 came 3 years after transformers were invented. Now we are close to 3 years after Gpt-3, did anything nearly as big happen during that time? Things aren't speeding up at all.
- arcturus17 4y agoThere are extremely interesting technical and economical critiques in the article. They may be right or wrong but at least they’re intellectually rigorous - much more than the average GPT commentary. Your comment, in a nutshell and if I’m reading it right, is that it’s not worth to engage with any of these ideas because progress - whatever that means, however that’s measured - is fast.
- pixl97 4y agoUnfortunately the measurement if something was worth engaging in is only really able to determine its value in hindsight. If it were 1890 we'd be talking about how our cities will soon be buried in animal dung which will lead to the collapse of mankind. The people debating that could not have reasonably foreseen in 100 years that CO2 would be the greater risk, and 100 years from now the greater risk will be something most of us have not imagined. Are the issues you point out worth talking about? Of course, but are they worth the amount of time and effort that we will debate them? Get back to me in 5 years and we'll see.
- ArjenM 4y agoI'll stick to waiting this one out, because I know "investor" personality types are mostly comparable to an individual sitting on a pink cloud wearing sunglasses that won't allow a singly light particle to enter. I'm hearing quite little about the downfall of AI route that science fiction predicted at times. Always good to waddle into that middle outcome.
- fortran77 4y agoNot the plane with the dots again!
- larve 4y agoA lot of what this takes are missing is that the unreliability of those models is not a hurdle for a tremendous amount of applications. If there is a human in the loop that can now write a couple of words in a fuzzy human language, check or select or edit the resulting response, and thus do their job twice as fast or god beware 10x as fast, that is absolutely a game changer. My main criticism for LLMs are: - the way they were rolled out was counterproductive. Unleashing a chatbot that pretends it knows everything, without any background and guardrails, is directly responsible for the hype untethered discourse that is prevalent in the mainstream. In the literature and among practitioners, every body is well aware that these things don't "think" - for the first time, it feels that a significant amount of what my value as a aprogrammer will be fully owned by a corporation and trickled out to me for $99.95 a month. It's already the case with copilot. I can't imagine going back to a world where I work without gpt3 and copilot, which gives me no choice but to fully embrace my corporate overlords. I fully feel what farmers feel wrt their tractors. The best I can do for now is figure out what the real usecases are for me and how to leverage GPT3, and start looking heavily into open models, so that I can help out with whatever unix <> bsd situation we are going to end up with. None of this has anything to do with the end of human culture, education, discourse, or the end of quality in software. If software quality could be any lower, under capitalism, it would be. It's not like I can get really shitty code that pretends to do something for $5/h on upwork.
- jjtheblunt 4y agoThe tweet from Sam Altman in this article made me laugh out loud.
- timdaub 4y agoWas he trolling?
- jjtheblunt 4y agoi think, but am not sure, that he was trying to sell OpenAI as unimaginably (near exponentially) awesome
- timdaub 4y agohttps://imgur.com/SurVVD5 https://imgur.com/SurVVD5
- jacobsenscott 4y agoInvestors don't "believe" in AI anymore than they "believed" in crypto. But they can smell a bubble, and hope to bail out at the top.
- mjr00 4y agoThey'll never admit it now, but a lot of investors genuinely bought into crypto and blockchain bullshit. That doesn't change that their plan was to bail at the top (isn't that what VCs do normally?), but many of them really thought that crypto had a chance to eat into Visa/MC's market share, or that blockchain would somehow "solve" supply chain, identity management, etc. The idea that crypto/blockchain investors knew all along it was a ponzi/pyramid/casino/scam/whatever-you-want-to-call it is revisionist history.
- foobarbecue 4y agoWhat a pompous style. I got annoyed trying to read it. By the way, https://www.theatlantic.com/technology/archive/2023/01/chatgpt-ai-language-human-computer-grammar-logic/672902/ https://www.theatlantic.com/technology/archive/2023/01/chatg... shows that at least some of the broader press gets what's going on with "AI."
- zozbot234 4y agoI'm not sure that there's anything truly new in this article. Transformer-based models are very compute- and parameter-heavy, yes - but that's because they're optimized for generality and easily parallelizable training, even at a very large scale. The ML community has always been aware of this as an unaddressed issue. Once compute cost for both training and inference becomes a relevant metric, there's lots of things you can do as far as model architecture goes to make smaller and leaner varieties more applicable.
- smoldesu 4y agoArguably, that was the defense for cryptocurrency too. "Oh, power consumption? Totally on our radar... what if we switch to Proof of Work? Personally, I don't think power consumption is the only legitimate argument against the application of GPT. There's the fact that it's proprietary, unreliable and even consistently wrong on certain topics. It's expensive to apply at-scale and most AI-generated text sounds sterile and impersonal. Even now, after years of development, GPT is too unreliable to be called anything other than a novelty. The dream of a "suitable application" for AI is like the dream of a "suitable application" for the blockchain. Both are technology-first solutions to social problems. Cool in concept, but regularly broken in execution.
- zozbot234 4y ago> Personally, I don't think power consumption is the only legitimate argument against the application of GPT. Well, it depends what you mean by "the application of GPT". As a toy and a technical proof of concept, it works quite well. The problem is that people want something that's mostly correct and won't just make stuff up - and that's just not what GPT is for.
- smoldesu 4y agoRight. Fiction-as-a-service is a fantastical API, but not a very practical one. When you boil it down, you're paying a lot of money for some carefully-arranged noise to get sent over the internet.
- 4y ago
- JacksonGariety 4y agoReading this article is like walking through deep mud. This is largely due to the author‘s excessive and baffling use of commas. But also due to myriad grammatical errors and ambiguous sentence constructions. Parsing some of his sentences is like looking at an optical illusion or an ambiguous painting. All things considered, I had a bad time attempting to read this article. I do not look forward to reading more of this author‘s writing in the future.
- timdaub 4y agoGranted. I didn‘t use Grammarly this time to fix the article‘s numerous mistakes. But frankly, it‘s also a way to prove that my article wasn‘t chatGPT generated hahahahaha
- hoechst 4y agocontent warning: crypto/web3 bro
- timdaub 4y ago:‘(
- college_physics 4y agoIf chatgpt and AI friends is the solution, what is the problem? If we push the AI madness aside the problem definition seems to be: I want to query large amounts of textual data to find an "optimal match/answer" to a given user input/"question". What is the precise nature of this optimal match? It is unclear (as it is essentially encoded in a black box algorithm and its training data set). Different algorithms and different training data sets would provide different "answers".
- sandworm101 4y ago>> But only those that users liked and retweeted ended up circulating, which contributed to a strange spread between a users' expectations and experiences they had when first talking to chatGPT. Which is exactly what makes this so dangerous. In our current culture, only those voices that are promoted and amplified by social media matter. So a tech than can produce material in industrial quantities, even if most of it is tripe, could be a powerful manipulation tool. It will certainly be cheaper than hiring flesh-and-blood people for man a troll farm.
- timdaub 4y agoJUMP! (Mike Solana)
- cma 4y agoTo show the cost of extra parameters is exponentional, they made a graph with price linear and parameter count logarithmic and then fit an exponential? It seems very misleading: https://proofinprogress.com/assets/images/cost-llms.png https://proofinprogress.com/assets/images/cost-llms.png With linear or polynomial increase per parameter it would also look exponential with that graph setup.
- timdaub 4y agomhh yeah you might have a point. I didn‘t do it on purpose though. I‘ll go through this again
- cma 4y agoMatrix multiply scales polynomially so don't be thrown off by that. You'd also expect some increases in cost where things move from single GPU to multi-GPU to networked machines. I'm still not sure if that pushes things to exponential or would just step up cost at each transition.
- timdaub 4y agoI spent lots of time to think of the title „The AI Crowd is Mad“ because I wanted to hit all triggers for virality (belief, belonging, behavior). A bit disappointed it was moderated/changed