15 ms·
The more I use ChatGPT (including 4.0 version, available with Plus subscription) the more I think these "Studies" and articles are simply made up. For my exper
by andreagrandi 3y ago
The more I use ChatGPT (including 4.0 version, available with Plus subscription) the more I think these "Studies" and articles are simply made up.
For my experience is terrible for any question I ask:
- if it doesn't know well a person it simply makes up things
- if I ask code examples or scripts, most of the time they are wrong, I need to fix them, they contain obsolete syntax etc...
- if I'm asking a question, I'm expecting being asked for more context if the subject is not clear, instead it starts spitting text without even realising I asked a completely different thing
etc...
I could go on for hours with other examples, but I'm seriously not finding it useful
- sabellito 3y agoIs there something about the premise of the study or its method that you feel are not good? After reading the article, what you wrote doesn't seem to make much sense.
- andreagrandi 3y agoI did not read this particular article, I was explaining that from my own experience, all these articles telling how great ChatGPT is seem to be made up because my experience (and from what I read I'm not alone) is completely opposite. Maybe it's not able to solve the type of questions I ask? Fine. But it's not how ChatGPT is presented most of the times.
- raincole 3y ago> I did not read this particular article It costs you zero dollar to not post an irrelavant comment then. I really wonder how you justify this "I didn't read the article, but I have a very, very strong opinion (straight up calling it a made-up) on it" behavior. The internet is rotting people's brains I guess.
- andreagrandi 3y agoAnd it costs zero to you to ignore my comment especially if you don't understand it. Other people seem to have understood what I meant and posted constructive responses, you didn't, but honestly it's not my problem.
- raincole 3y agoAh, I see. You just don't realize that your comment is not relavent to the original article... (of course not, because you didn't read it)
- andreagrandi 3y agoI've now read the article and I'm not changing my opinion. ChatGPT can do some things very well and people tend to hype those things and claim it's better than humans rather than recognising its limits. And again, you are still missing the point of my comment despite having read it, so I'm out of patience. Maybe try to ask ChatGPT to explain what I meant :)
- notahacker 3y agoTo be fair, the opinion is thoroughly justified by the article, which might have been more honestly titled Physicians' Reddit comments shorter than ChatGPT responses; relative accuracy unknown...
- Sunhold 3y agoNot really. A team of licensed health care professional rated "the quality of information provided".
- notahacker 3y agoAnd unsurprisingly, the average 52 word Reddit comments [isolated from the context of other comments] didn't provide very much information compared with a much more verbose chatbot. The relevance of the ChatGPT response to the actual patient condition remains unknown. This is relevant to the real world of primary care only if your sole access to a medical professional is Reddit...
- jprete 3y agoTheir first example of a good ChatGPT answer - about bleach in the eye - feels like copypasted SEOified liability-proof WebMD copy. Every medical site has that crap and it’s useless once you have a moderately difficult question. N.B. as well: If someone thinks they have bleach in their _eye_ and can still open their eyes enough to write a Reddit post, much less read through ChatGPT’s extremely long answer, they’re almost certainly fine.
- ceejayoz 3y ago> Is there something about the premise of the study or its method that you feel are not good? Yes. Down in the limitations section of the study (https://jamanetwork.com/journals/jamainternalmedicine/fullarticle/2804309?guestAccessKey=6d6e7fbf-54c1-49fc-8f5e-ae7ad3e02231 https://jamanetwork.com/journals/jamainternalmedicine/fullar...): > evaluators did not assess the chatbot responses for accuracy or fabricated information That is a... significant issue with the methodology of the study.
- jstanley 3y agoIf you tell it it's wrong, it often comes up with a better answer on the second try. It seems like you could maybe automate that. Let it spit out its first draft of an answer, have the framework tell it "please correct the errors" and then let it have another go and only present the second attempt to the user.
- simmerup 3y agoBut how do you know it's wrong until a human tries the output
- mirekrusin 3y agoYou don't need to, you always tell it it's wrong.
- notahacker 3y agoBut if you always tell it it's wrong, it will sometimes come up with worse answers on the second try. Which means you still need to know whether the first answer is correct or not (or neither of them) I'm reminded particularly of a screenshot of someone gaslighting ChatGPT into repeatedly apologising and providing different suggestions for Neo's favourite pizza topping, despite it answering correctly that the Matrix did not specify his favourite pizza topping first time round, but it applies equally to non-ridiculous questions
- jstanley 3y agoThe idea isn't that you tell it to simply change what it wrote on the first try. The idea is that having a first draft to work with allows it to rewrite a better version.
- notahacker 3y agoThis technique works fine if you're creating new iterations of creative work, or have a specific thing you want it to fix It's not much use if ChatGPT gives you a diagnosis of your symptoms which may or may not be accurate.
- AshamedCaptain 3y agoYour experience matches mine. Everything I have asked is usually _horribly_ wrong, and even asking things in a different order makes it completely change its responses, even for otherwise binary questions. Even the "snippet" part of a Google search with the same prompt normally contains enough information to contradict it... I'll also note that when there was hype around Stable Diffusion, one of the images shared around was that of an astronaut riding a horse. If you actually run Stable Diffusion with its default tuning and ask for that prompt, you will get 6 images, of which 5 of them are outright disasters (horses with 6 legs and going downhill from there), and then the 6th image, the only one which could possibly pass as a decent result, is the one that everyone shared and reshared and hyped. Usually other prompts give even more terrible results where there are 0 passable images without extensive tuning. Stable Diffusion now is acknowledged to actually be crap --despite the hype-- and I supposedly need to try the next best thing, whatever that is. But I find myself facing the same situation with ChatGPT 3.5, and now with ChatGPT 4, despite the fact there is no "next best thing", and I don't even know how they could even possible try to fix the problem of it being just wrong.
- sorokod 3y agoKindest thing one could say is that there is massive cherry picking going on which is borderline dishonest.
- shanebellone 3y ago"massive cherry picking" I think the coined phrase is "prompt engineering". Side note, where's the eye roll emoji hiding?
- whateveracct 3y agoreminds me of the BTC pivot from "decentralized currency" to "store of value" goalposts
- BoorishBears 3y agoIf you understand what an LLM is, chain of thought isn't something you eye roll at.
- asimpletune 3y agoI always tell people that if they want learn more about the hype to just try and use it to do actual work. Almost no one ever does, but when they do it becomes almost immediately clear how limited it is and what it is and isn’t good for.
- jstanley 3y agoIt's good at writing. It's not good at knowing facts. If you give it all the relevant facts and ask it to do a writeup in a certain style it does a better first draft than a typical human, and in a lot less time. If you ask it what the facts are, it just gives you a load of nonsense.
- asgerhb 3y agoEven this is not always the case. I gave it a very rough draft and it actually made my structure worse, while at the same time using imprecise language. It looked like the correct style, but the content was not salvageable.
- Abroszka 3y agoI use it almost every day for work. It has mostly replaced Google for me. A lot more convenient. Now whenever I use Google it's more or less just to look up the address of a specific site.
- kaba0 3y agoWell, google has become utterly bad at its job — I fail to find sites I remember verbatim quotes from, so expert google-usage is no longer a possibility. It will gladly leave out any of your important keywords, even if you add quotes around it, absolutely useless. Sure, the average person will search for “how old is X” not “x age”, but for more complex queries the first form is not a good fit. That said, I can’t really use ChatGPT as a search engine, but I did plug it into a self-hosted telegram bot and I do ask it some basic questions from time to time - telegram is a good UI for it.
- afro88 3y agoIs it everyone else that's wrong or.... > if it doesn't know well a person it simply makes up things Asking it for factual information about a subject can be a bit hit/miss depending on the subject. Better to use bing chat, because it will use info from the web to inform the response > if I ask code examples or scripts, most of the time they are wrong, I need to fix them, they contain obsolete syntax etc... How wrong? More wrong than having a junior or mid level developer contributing code? Think about it a different way: you just gained an assistant developer that writes mostly correct code in seconds. Big time saver. Also: if you want it to use a particular code style etc, give it few shot examples. > if I'm asking a question, I'm expecting being asked for more context if the subject is not clear Then you need to tell it that in you prompt: "if the subject isn't clear, ask me some clarifying questions. Don't respond with your answer until I have answered your clarifying questions first". Or: "ask me 3 clarifying questions before answering" to force it to "consider" how well it "knows" the subject first. ChatGPT isn't an AI in the sci fi sense of the word. It's a language model that needs to be prompted the right way to get the results you want. You will get a feel for that the more you use it.
- deleted 3y ago[deleted]
- raincole 3y agoAnd even if ChatGPT is always 100% wrong with code, I still failed to see how it is relevant to this particular article. The article compares verified responses on r/AskDocs (yeah, a subreddit) and those from ChatGPT. That's it. How is its coding compatibility even remotely relevant? It's like saying "Excel is bad in editing photos, so it must be a bad spreadsheet software as well."
- rafaelero 3y agoIndeed. I don't understand why these types of comments get so many interactions. It makes me think they are relevant but on second inspection they are mostly unrelated grievances.
- isaacremuant 3y ago
- 13415 3y agoTo add to this, I tried to use it professionally but the answers were too general and generic. I suppose it has been prompted to put things simple, which prohibits it from saying meaningful things about certain topics. It did give one or two useful references, though.
- geonnave 3y agoMy experience is completely different. I have successfully used GPT-4 to: - write a contract for the sale of my motorcycle: put all details, names and numbers with labels on a spreadsheet, paste on the chat and ask for a contract, then edit. - learn french: I told gpt "when I write wrong stuff in french, always let me know and teach me the correct ways". Then, after a few weeks I asked for a .csv with the stuff that he corrected so I could import into Anki, which actually worked. - coding on a daily basis: I am learning Rust on my new job, so I ask it things all the time, it helps me a lot.
- andreagrandi 3y agoGood to know it's able to do some useful stuff. In my case I mostly ask Python related questions, because it's the one I know better so I can check if the answer is right or wrong. I will try with different languages, but I will be less capable of knowing if I got a good answer or not. It may take more time, but I find the combination of Google + Stack Overflow more accurate than asking ChatGPT
- isaacremuant 3y agoYour experience is not different from OPs. You're just ok with the mistakes or are unaware of them because you don't know how to judge them. I both use ChatGPT to boost productivity but also see the amount of mistakes it makes and will keep making and am surprised at the extreme denial of anyone who tries to shut down criticism of the wrong type of hype (the one that sells something that is not there)
- geonnave 3y agoOh but it is: I find it useful. For example, I would not pay a lawyer for that contract, so having it draft me a mediocre contract is still better than having no contract.
- hammyhavoc 3y agoA mediocre contract can be worse than having no contract, or as bad as having no contract if it isn't actually legally binding. Yikes.
- vidarh 3y agoIf you expect to be able to ask it an underspecified question without context and without telling it what role it should take and how it should act, sure, that often fails entirely. It's not a productive use of ChatGPT at all. If, on the other hand you actually put together a prompt which tells it what you expect, the results are very different. E.g. I've experimented with "co-writing" specs for small projects with it, and I'll start with a prompt of the type "As a software architect you will read the following spec. If anything is unclear you will ask for clarification. You will also offer suggestions for how to improve. If I've left "TODO" notes in the text you will suggest what to put there." and a lot more steps, but the key element is to 1) tell it what role it should assume - you wouldn't hire someone without telling them what their job is, 2) tell it what you expect in return, and what format you want it in if applicable, 3) if you want it to ask for clarifications, either ask for it and/or tell it to follow a back and forth conversational model instead of dumping a large / full answer on you. The precise type of prompt you should use will depend greatly on the type of conversation you want to be able to have.
- skepticATX 3y agoThis type of usage is rapidly approaching a Clever Hans type of situation: https://en.wikipedia.org/wiki/Clever_Hans https://en.wikipedia.org/wiki/Clever_Hans. An intelligent agent shouldn't need this type of prompting, in my opinion.
- hammyhavoc 3y agoIs the agent the LLM or the user who needs an LLM?
- vidarh 3y agoIt's perfectly fine if it's approaching a Clever Hans type situation as long as it's producing sufficient quality output fast enough that it's producing it faster than I can do manually. There are many categories of usage for them, and relatively "dumb" completion and boilerplate is still hugely helpful. In fact, probably 3/4 of my use of ChatGPT are uses where I have a pretty good idea what it'll output for a given input and that is why I'm using it, because it saves me writing and adjusting boilerplate that it can produce faster. Most of the time I don't want it to be smart, I want it to reliably do almost the same as it it's done for me before, but adjusted to context in a predictable way (the reason I'll reach for it over e.g. copying something and adapting it manually). We use far dumber agents all the time and still derive benefits from it. Sure it'd be nice if it gets smarter, but it's already saving me a tremendous amount of time.
- AnIdiotOnTheNet 3y agoMy experience is the same, and yet I am not surprised that it still scores better than the average physician.
- broast 3y agoYour experience matches mine other than I still find it extremely useful regardless of errors
- allisdust 3y agoMy experience has been polar opposite with GPT4. As long as I structure my thoughts and present it with what needs to be done - not like a product manager but like a development lead, it spits out stuff that works on first try. It also writes code with a lot of best practices baked in (like better error handling, comments, descriptive names, variable initialization). Some times this presenting of problem to it means I spend anywhere from 5-10 mins actually writing the points down that describes the requirement - which would result in a working component/module (UI/backend). We have been trialing GPT4 in my company and unfortunately almost everyone's experience is more on the lines of yours than mine. I know it shouldn't, but honestly it frustrates me a lot when I see people complain that it doesn't work :). It definitely works but it depends on the problem domain and inputs. Often people forget that it has no other context about the problem than just the input you are providing. It pays to be descriptive.
- wouldbecouldbe 3y agoCode yeah if you know it's pitfalls you can get it correct pretty fast. But I just don;t believe it gives better answers doctors, it makes silly mistakes that signal it doesn't understand things deeply. I only believe if they actually trained chatgpt on those type of tests specifically. Not the actualy dynamic nature of dealing with patients & lawsuits.
- mensetmanusman 3y agoI wonder which personality profiles interact with it best. Probably some function of which abstract layers people start with when thinking.
- croes 3y agoHow do you do that for patients?
- stainablesteel 3y agothere's definitely an art to asking the questions, likely because of subtle differences in how a lot of people communicate in writing. NLP can recognize alt accounts of individuals on places like HN and reddit, but a person would probably need to study the comments pretty hard to determine the same thing, its not natural for people imo but it seems to be the foremost aspect of any kind of model that's processing human writing.
- SkyMarshal 3y agoI’ve only asked GPT4 a few niche questions on subjects of interest to me, so I can’t really judge it yet. But so far its answers can’t compete with Wikipedia. However, it seems good at drilling down and doing followup questions that build on the prior questions, which is interesting. I can see that natural language give-and-take back-and-forth being useful for things like early education, diagnosing non-emergency patients, troubleshooting home PC problems with non-computerphiles, etc.
- hammyhavoc 3y agoIt hallucinates a lot. Any time statistics or specifications are involved, don't trust it whatsoever.
- usrusr 3y agoIt quite literally writes whatever sounds about right. Which is certainly very impressive if you happen to assess by exactly the same metric... It's more artificial overconfidence than artificial intelligence
- hammyhavoc 3y agoYes. It is inappropriate for most things. It's an LLM. It predicts the next word. People are throwing it at all kinds of problems that are not only inappropriate, but their ability to assess the quality of its output is questionable. E.g., lots of HN users claim to use it for dev or learning new programming languages. Given the frequency of hallucination and their Dunning-Kruger complexes in full-effect, they don't know when it's teaching bad information or functions that don't exist. It's an LLM. Not an AGI.
- amelius 3y agoFor us to get a better understanding of how well this tech works I suggest ChatGPT becomes integrated in HN in this way: it generates 1 response per comment; the responses written by the AI are clearly marked as such (e.g. different color); the user can turn them off; these comments can be up/down voted and the votes can be seen by any user; of course users can reply to the generated comments.
- ziml77 3y agoI've seen a lot of positivity on the output of ChatGPT for coding tasks in my workplace. And it does seem to have some use in that area. But there is just no way in hell it's replacing a human in its current state. If you ask it for boilerplate or for something that's a basic combination of things its seen before, it can give you something decent, possibly even useable as-is. But as soon as you step into more novel territory, forget it. There was one case where I wanted it to add an async method to an interface as a way of seeing if it "understood" the limitations of covariant type parameters in C# with regards to Task<T>. It did not. I replied explaining the issue and it actually did come back with a solution, but it wasn't a good solution. I told it very specifically that I wanted it to instead create a second interface for holding the async method. It did that but made the original mistake despite my message about covariance still being within the context fed back in for generating this response. I corrected it again, but the output from that ended up being so stupid I stopped trying. And at no point was it actually doing something that's very important when given tasks that are not precisely specified: ask me questions back. This seems equally likely to be a problem for one of these language models replacing a doctor. It doesn't request more context to better answer questions so the only way to know it needs more is if you already know enough to be able to recognize that the output doesn't make sense. It basically ends up working like a search engine that can't actually give you sources.
- danielbln 3y agoYou should try the model not in isolation, but hooked up to search so it stuffs the context with current and verifiable data. Check out phind.com.
- SilkRoadie 3y agoI use it at work and it clearly has strengths and weaknesses. My two use cases are initial research and generating prototype code. I find it very helpful to ask a series of questions and see a number of examples to get a primer on what to expect with something. The main benefit over Google or going straight to the docs is I can start with my specific requirements. I then dig into the documentation to deepen my understanding. I can typically move forward with ChatGPT generating some code as a starting point. It can be incorrect or out of date but combined with my experience I find myself being more productive with it. A weakness I see is complex code requirements. It knows what it knows. I note that you seem a little frustrated with vague or incorrect responses. It helps to tell ChatGPT the role it should play. It helps as well to instruct it to ask questions of you to improve the response. Personally I prefer to tell it keep its answers brief, I get less walls of text and I can narrow in on the specific answer I am after more quickly.
- mock-possum 3y agoRight? Like I had a batshit insane conversation about song lyrics the other night, where the chatbot repeatedly generated patently incorrect responses - close enough to seem reasonable, but utterly incorrect, mistakes that I can’t imagine a human making, just straight up false statements that didn’t hold up to the slightest scrutiny. Incredibly frustrating. Imagine having that kind of experience with a medical professional, when you’re sick and impatient to receive care. awful.