8 ms·
The Forer Effect can't make chatGPT generate syntactically correct and functional code snippets that I use every day at work.
by fourseventy 3y ago
The Forer Effect can't make chatGPT generate syntactically correct and functional code snippets that I use every day at work.
- Filligree 3y agoThe goalposts are moving so fast they're red-shifting.
- scott_w 3y agoThat’s not really an issue when we’re talking about AI. I said in a separate comment on the topic: intelligence is such a complicated thing that we seem to only be able to define it by pointing to things and saying “that’s not it.” If we didn’t move the goalposts, we’d have declared Stockfish to be full AI, despite it only being a chess-playing program, long ago.
- lostmsu 3y agoExcept in a few years you may realize it was "it" all along.
- scott_w 3y agoPerhaps but I’m pretty confident a ChatGPT is not “it” this time.
- pixl97 3y agoSortie's Paradox.
- scott_w 3y agoMy confidence comes from watching GothamChess play it at… chess. The problem with ChatGPT isn’t that it’s bad at chess or doesn’t understand the rules of the game. It obviously has no concept of what it’s doing. It can’t even keep track of which pieces are still on the board: a trivial task for even a small child.
- pixl97 3y agoSmall children track in their 'minds eye' the entire layout of a chess board and where each piece is? Tokenization starts screwing stuff up if you're start printing out a board after each move. Push out a plugin that sets up a virtual board GPT-4 can read after each move and see if its any better.
- scott_w 3y agoNo, they look at the board, though strong chess players can do it blindfolded. My point is that the hyperbole about “we’ve cracked AI but they changed the goalposts” is self evidently not true. You just proved it right there: I need to add more plugins because ChatGPT does not understand what it’s doing. It’s a potentially useful tool but that doesn’t make it intelligent.
- pixl97 3y agoI mean, all you've managed to state is that gpt isn't a strong AGI, not that it isn't intelligent. Stop thinking of intelligence as a binary option of 'is or !is' and more of a capability gradient. Being able to properly use tools is a sign of intelligence in of itself.
- scott_w 3y agoI’m using “intelligent” in the “intelligent life” meaning, not the “how intelligent is this person” meaning. By this meaning, any reasonable person would realise it’s not crossing that threshold. Where is this threshold? I don’t think anyone really knows yet. That’s why I don’t see a problem with “moving the goalposts,” because doing so is the best way to help us truly understand what it means to be an intelligent life form (artificial or otherwise).
- pixl97 3y agoAnd this is why the word intelligent is useless. In your use of the word we're not getting anywhere then suddenly terminator kicks in the door and steps on your head and then "Oh, yea, I guess we reached the intelligent point". This is a piss poor predictor of capabilities. And again "reasonable" is a pretty useless metric. Asking a 'reasonable' person about any system that requires expert knowledge to understand is going to derive an unreasonable answer. This is because they'll conflate intelligence with human behavior.
- marcusverus 3y agoWe agreed on a definition of artificial intelligence--the Turing test--for 50 years. The goalpost was clearly established, widely agreed upon, and promptly abandoned when chatbots blew past it. I'm convinced that when the dust has settled and historians look back to decide on THE point in time at which we achieved AI or even AGI, that time will not be in the future, but in the past.
- scott_w 3y agoThe fact it was defined a long time ago isn’t good enough. If, upon passing the test, you realise the test was insufficient, you change the test. You don’t shrug your shoulders and go “well it must be right because a guy said so 50 years ago!”
- dragonwriter 3y ago> We agreed on a definition of artificial intelligence–the Turing test–for 50 years. No, we didn’t. There was plenty of disagreement over it. Heck, the Chinese Room is very popular rejection fundamentally of the premise of it. That aside, even if in a blind scenario (where the builders didn’t know the criteria used to test) being able to fool humans in linguistic interaction would be reasonably likely to be a good test of general intelligence, LLM’s are about as an obvious of a direct and deliberate attempt to Goodhart’s Law the Turing Test as one could imagine. Using a particular capacity to test a more general capacity is obviously vulnerable to systems built to specialize in the tested capacity, as opposed to those that have it as a consequence of general ability.
- JimtheCoder 3y ago"I'm convinced that when the dust has settled and historians look back to decide on THE point in time at which we achieved AI or even AGI, that time will not be in the future, but in the past." If historians are looking back to decide the point, by definition doesn't it have to be in the past? Historians looking back to the future doesn't make sense...
- epylar 3y agoHistorians in 2050 could look back at 2040, which would be their past but our future.
- chongli 3y agoThe whole history of AI debates seems to be an exercise in goalpost-moving. Whether it's AI proponents or skeptics, no one can seem to agree on a stable benchmark. Personally, I think it's due to intelligence itself being ill-defined as a concept. How are we supposed to build something when we don't even agree on what it is we're trying to build?
- lisasays 3y agoThat's ... not the point of the article.
- PaulHoule 3y agoThe odd thing is that some people report much better results than others. Most literature on the subject points out the importance of a "human in the loop" https://arxiv.org/abs/2304.13187 https://arxiv.org/abs/2304.13187 Generally, chatbots function in direct communication with humans and do not function autonomously. Thus they are dependent on not just human meaning-making but human thirst for meaning which will find it even when it isn't there.
- Filligree 3y ago> Generally, chatbots function in direct communication with humans and do not function autonomously. Thus they are dependent on not just human meaning-making but human thirst for meaning which will find it even when it isn't there. The latter doesn't follow. They may just be dependent on good prompt design. People who aren't used to the AIs, and don't know how to use them, get worse results. This shouldn't be a surprise; it's how every tool works.
- deleted 3y ago[deleted]
- Joeri 3y agoThe prompt is everything. It’s sort of like knowing how to formulate a google search query, but cubed. You need to learn a whole bag of tricks and you can’t use the intuitions of talking to people for talking to the AI. Some people are by now really adept at prompt engineering and that makes them skilled at setting the AI to work. That’s why people’s first impressions are often so wildly off. People who have the fortune of asking the right prompts for their first conversations walk away believing they’ve spoken to AGI, people who ask a less than optimal prompt (or things it is bad at like math problems) walk away not seeing what the hype is about. Same thing with copilot. The quality of the code that it suggests depend a lot on the context, so if someone isn’t writing good comments and giving their variables proper names the suggested code will be terrible. These systems are very capable, but using them is a skill that must be learned, and there’s little material out there to teach that skill because things move so quickly.
- choeger 3y agoLet me make this very clear: Generating syntactically correct code in various languages that also looks plausible is no small feat. It is, in fact, extremely impressive and will certainly have an impact on SE. But. Every single test I ran lead to functionally wrong designs from smallish memory errors in C (that hilariously ChatGPT was able to correct ND explain when pointed to) over misplaced/hallucinated methods in python APIs to completely hallucinated perl packages. It never gave me a useful answer. I don't know if that tells us something about my work vs. your work or my approach to the model vs yours. But it definitely tells me that such a model cannot replace a developer.
- lostmsu 3y agoChatGPT which version? Even Copilot, which is a smaller model, generates working code most of the time.
- vsareto 3y ago>But it definitely tells me that such a model cannot replace a developer. Certainly not an average developer (however that's decided in C's culture). Not making any statements on skill of anyone here, but I'd bet it could replace the very low ends of skill as it is (GPT-4 and maybe Bard). I doubt there's too many of those kinds of developers though. It's not a guaranteed positive for all languages/people, but it's farther along being useful as an enhancement technology than it is as replacement technology. I think it's making progress on both axes, but one is clearly ahead of the other and it's still possible it stops progress until other innovations are made.
- skyechurch 3y agoI can guarantee you that the equivalent code which I generate (taking >100 times as much time) has a comparable error/typo rate. I can say this with confidence because ChatGPT, while not perfect, "knows" lots of things I don't know and absolutely makes me 10x as efficient. Now, I am a pretty crummy software engineer so ymmv, but it's worth noting how far the goalposts have moved for assessing AI: from "simulates understanding of language in a general way" to "comparable in any way to a professional specialist in a random task (and much faster)". A few months ago everyone was losing their minds over how good ChatGPT was, and while some of that was overblown, nothing has fundamentally changed except we have become used to it and are looking for the next shiney tech thing. At least we're not talking about crypto anymore.
- deleted 3y ago[deleted]
- thefreeman 3y agoCounter point that literally just bit me this morning. ChatGPT completely lied to me about the scaling properties of Kinesis. I was confused about the 1000 write / second but only 5 read / second throughput of a kinesis shard and was asking it a number of questions, including specifically if a batch read would count as only a single read operation. It explicitly told me that each record in the batch counted as an individual read. Most of the other information it gave me was accurate so I took this to be true. Only the next day when I still couldn't wrap my head around why a real time processing system like be designed like this and did more google searching did I find it explicitly spelled out in the docs that a batch read counts as only a single read operation and can retrieve up to 10,000 records. It's the first time I've been bit by this and it will definitely make me wary of any information it provides me in the future for technologies I am not intimately familiar with (at which point it's unlikely I'd need to consult it in the first place).
- greiskul 3y agoYes, please. Do not use LLMs as a substitute for Google search. If you are looking for factual information, just use google, bing, or duckduckgo. You should only use ChatGPT for things that you are able to review it's work. Technology is supposed to make us smarter. Blindly believing an AI that we know can hallucinate makes us dumb with confidence.
- mooxie 3y ago> You should only use ChatGPT for things that you are able to review it's work. This keeps being my argument when people at work daydream about time and cost savings by offloading non-critical business functions to AI. I say, "Great, so it can produce 1000x more work than a person. But then what army of people are we planning to use to check those outputs?" I'm super-impressed with the current crop of language models for their ability to so accurately simulate correctness, but their inability to understand what they don't know - because, in fact, they don't 'know' any of it in the sense that we do - makes them like very productive but completely untrustworthy employees. A junior dev who monopolizes his mentor's time through inconsistent performance is not a good hire.
- ITB 3y agoI think the interpretation is that there is a certain statistical distribution of statements that appear “more broadly true” than they are. I think there is some meat to this, even with code that works as intended. The degree of creativity might be lower than we perceive.
- geraneum 3y agoThe compilers _usually_ generate correct machine instructions (code) better than us. This, alone, is not really a meaningful measure and doesn’t refute the claim.
- x0x0 3y agoDo you mind sharing what you're using / how you're doing it? I've been using github copilot and I'm waving back and forth between impressed and unimpressed; it regularly generates syntactically invalid code for me. Eg calling lib functions that don't exist. Thanks!
- Dr_Birdbrain 3y agoCame here to say this. The code it generates for me is usually not perfect, but for me it provides conceptual insights that have gotten me unstuck from gnarly situations. The conventional narrative is that chatGPT can code as well as a junior engineer, but I feel like it’s more like a senior engineer scribbling on a whiteboard—no, it’s not perfect code, but it’s the right idea.
- boobooby 3y ago[dead]
- gumballindie 3y agoI am sorry to break it to you but if the level of code required in your job is doable by chatgpt then your job will be automated soon.
- mcstafford 3y agoYou make a decent point. There's a clear limitation... but a clearly worded prompt often produces more valuable content than an SEO optimized guess article.