8 ms·
I thought these were some good ideas. In my experience I vacillate between "omg the singularity is here" and "this actually isn't that good for X specific task"
by jackconsidine 3y ago
I thought these were some good ideas. In my experience I vacillate between "omg the singularity is here" and "this actually isn't that good for X specific task".
I very much trust the output of LLMs to be well-designed, but I don't trust things to just work, especially if the system is complicated. I experimented a bit the past few days doing a task myself (building an interface in an existing project), using AI assist, and trying to get AI to solve completely (GPT-4). The solve completely pathway failed and I found myself in an interminable loop. AI-assist was a solid experience.
Anecdotal but consistent with Figma's observation
- tymscar 3y agoI have ran 5 “test cases” where I would jump on a video call with some friends that are not software engineers but are very technically savvy, and would give them a simple-ish task. They are allowed to use ChatGPT as well as Google. None managed to do it in 4 hours each, and that was with me giving them hints when the AI would inevitably get into a loop. The task was to install docker and using docker compose host a reverse proxy with Træfik with self signed SSL, as well as web-server in Rust. All the rust app had to do was read the kernel version of the host machine and return it as html. I have then ran the same exact test, but this time around, all you had to do was install docker, and with docker compose run a Træfik reverse proxy with self signed SSL, and two other containers, UptimeKuma and Audiobookshelf. No dice, nobody managed to do it.
- jackconsidine 3y agoWow kudos to you for running such a controlled and extensive experiment
- rolisz 3y agoAre they still friends with you? Are they still returning your calls? I've been coding for 15 years, but on my NAS I still use the Synology reverse proxy so I don't have to deal with Traefik and self signed SSL.
- tymscar 3y agoNow that you mention it, they haven’t really answered my calls in weeks. Joking aside, we actually had a lot of fun and they thought it was very eye opening, not only about AI, but also about what my work life is like.
- SomeCallMeTim 3y agoThe output of LLMs is ... rarely well-designed. Well-documented (with often incorrect documentation), well-formatted for sure, but profoundly not well-designed, unless you're asking for something so small that the design is trivial. Even with GPT-4, if you ask it for anything interesting, it often produces code that not only won't work, but that couldn't possibly work without a major rewrite. Not sure what you've been requesting if it's always been good output. Even when asking GPT-4 for docs I've had it hallucinate imaginary APIs and parameters more often than not. Maybe the questions I ask are not as common? Given my experiences, though, I wouldn't recommend it to anyone for fear it gave them profoundly bad advice.
- nimithryn 3y agoMy experience as well. Heavy GPT-4 use (for a variety of things). Great for boilerplate, great for retrieving well-known examples from documentation, saves a fair amount of time typing and googling, but often completely wrong (majorly and subtly) and anything non-trivial I have to do myself. Great tool! Saves a ton of time! Not a dev replacement (yet)
- MeImCounting 3y agoNow I am definitely doing much simpler things I would wager than you folks are but I have found that with a bit of back and forth you can get pretty good results that work with only a bit of revision. I have found reminding it of the purpose or goals of whatever it is youre working on at the moment tends to make the output a bit more consistent
- koreth1 3y ago> with a bit of back and forth you can get pretty good results that work with only a bit of revision The problem for me is that the "back and forth" and "a bit of revision" steps very often end up taking more time than writing the code myself would have.
- MeImCounting 3y ago
- teaearlgraycold 3y agoThere are posts on X about how powerful GPT4 is. And the videos are really impressive. But then in my own experiences it’s only really good if you know what you’re doing and can carefully guide it into taking a single step in a process. Anything more and the failure rate explodes upwards. I love using it as a copilot (github copilot chat in vscode is great). But it’s so far from “singularity” that I don’t fear for my job as a programmer yet.