7 ms·
there was literally a post a few days ago where someone from cloudflare or fly (can't remember) released a new feature built with AI (albeit heavily supervised)
by ookblah 1y ago
there was literally a post a few days ago where someone from cloudflare or fly (can't remember) released a new feature built with AI (albeit heavily supervised) and answered q's along the way about it.
it's funny because what i'd also like to see is people who are skeptics make a video as well since sometimes i also have the opposite suspicion. i get a lot of the criticism, but i don't get the "it produces pure garbage" type ones.
- solarwindy 1y agoMy comment was admittedly a little inflammatory, so I’ll try to elaborate on what I find so exasperating about the whole experience of using LLMs for coding (in any form, chat, agent, etc.): at root, it’s their total lack of a model of computation. This leads to mistakes that are difficult to catch, because while I know the kinds of mistakes another human might make, having been through the process of learning not to make them myself, LLMs produce whole classes of bizarre mistakes that I have no interest in learning to catch. There is no discernible flawed mental model behind these errors—which with a human could be discussed and corrected—just an opaque stochastic process which I can tediously try to set on a better course with ‘incantations’, attempting to dial in to a better part of the training data that avoids the relevant class of error. It’s honestly amusing how much of ‘prompt engineering’ (if it can be dignified with that term) amounts to a modern-day kind of mysticism. What better can we really do though when these models’ structure and operation is utterly opaque, on the one hand through deliberate, commercially-oriented obfuscation, and on the other because we still just cannot explain how a multi-billion parameter model works. It’s rewarding to work with human juniors because people actually learn and improve. My learning how to coax these models into producing better than trash just is not, especially on anything either remotely novel, or in a legacy codebase that requires genuine understanding of how an existing system functions at runtime. Once those two ends of the spectrum are ruled out, I find there’s little left that an LLM can accomplish, without necessitating a loop of prompt refinement that leaves me feeling like a worse developer for not just having thought through the problem myself, and resentful of the time wasted. Edit: this entire thread about coercing Gemini into behaving is exactly the kind of crap I have zero interest in: https://news.ycombinator.com/item?id=44194061 https://news.ycombinator.com/item?id=44194061
- ookblah 1y agoit's really interesting to me because i feel like legacy code is where it actually excels very well. i have a large mental map in my head, so i usually just feed it some vague notion of what i want it do and use existing files as a reference ("i want to build this feature, look at X, Y, Z for reference. approach it this way" if i have some notion of how i want it go). i usually don't let it go run off on its own unless it's a very defined task that i can review quickly later, i just review and approve every change and it takes big cognitive load off for me. at some point maybe this doesn't feel like "programming", but then i'll just tweak something else or modify it and then go onto the next review. i find i can't have it produce the entire thing and then review it since i have no idea how it got to where it did or takes just as much time to understand. but doing it this way it's faster + i gain understanding. the prompts aren't overly complex or take a lot of time, certainly way less than speccing something out for a junior. all i have a is a base file for style and structure and then i describe the general problem and reference files ad-hoc. where i find it actually fails a lot is in novel code because it has nothing to ground it and starts exploring random stuff. i only use it for novel exploration to see what approaches it comes up with. still trying to understand why there's this huge chasm between the two viewpoints. like a lot of the things you just said i can't resonate at all with. like maybe the 20% i feel like i'm "fighting" the LLM i just stop and go in myself. does that suck? sort of, but it's certainly way less tedious than directing some other person to do it or the time saved had i not used it at all. edit: but to your point, yeah it really is just like magic with no way to like actually direct it in a way you would where a human would learn. maybe over a year ago i tried and wrote anything AI off beyond basic co-pilot completions (same issues, "fighting" the AI, having to specify a tons of exceptions in some god awful file). the new agents changed everything for me, esp claude code. i think it will only get better, so it's best to pick it up. my only fears are 1) no juniors being trained, thus no future seniors. part of the power is that you have experienced people using it to enhance their context or understanding. for those with no experience and no drive to "improve" (honestly, think of the avg dev at big co) or straight up "vibe coding" i shudder at the output. my hypothesis we are now going to enter a period where a LOT of shitty code is going to be created. it's already happening in education with people just cheating w/o learning. i already had issues trying to hire people who were using AI to get past initial exercises but failing on complex issues because they were just probably copy and pasting everything. best time to be a nimble startup. 2) top-down mandates to use this stuff. you should only use it when you want and if it helps you. i think there's this element of companies buying into the hype 110% and that puts a bad taste in everyone's mouth. "all devs replaced in 1 year!" type stuff.