3 ms·
I'd assume the person giving the praise is at least a bit of all 3. > It’s a weird catch-22 giving praise like that to LLMs. It's a bit asymmetrical though is
by kajecounterhack 2y ago
I'd assume the person giving the praise is at least a bit of all 3.
> It’s a weird catch-22 giving praise like that to LLMs.
It's a bit asymmetrical though isn't it -- judging quality is in fact much easier than producing it.
> you might be able to intuit and fill in the gaps left my the LLM and not even know it
Just because you are able to fill gaps with it doesn't mean it's not good. With all of these tools you basically have to fill gaps. There are still differences between Cline vs Cursor vs Aider vs Codebuff.
Personally I've found Cline to be the best to date, followed by Cursor.
- dartos 2y ago> judging quality is in fact much easier than producing it There’s still a skill floor required to accurately judge something. A layman can’t accurately judge the work of a surgeon. > Just because you are able to fill gaps with it doesn't mean it's not good. If I had to fill in my sysadmin’s knowledge gaps I wouldn’t call them a good sysadmin. Not saying the tool isn’t useful, mind you, just playing semantics with calling a tool a “good sysadmin” or whatever.
- kajecounterhack 2y ago> There’s still a skill floor required to accurately judge something. Sure but it's not high at all. Your typical sysadmin is doing a lot of Googling. If perplexity can tell you exactly what to do 90% of the time without error, that's a pretty good sysadmin. Your typical programmer is doing a lot of googling and write-eval loops. If you are doing many flawless write-eval loops with the help of cline, cline is a pretty good programmer. A lot of things AI is helping with also have good, easy to observe / generate, real-time metrics you can use to judge excellence.
- dartos 2y ago> Sure but it's not high at all. It depends. For a sysadmin maybe not, but for data scientists, the bar would be pretty high just to understand the math jargon. > If perplexity can tell you exactly what to do 90% of the time without error That “if” is carrying a lot of weight. Anecdotally I haven’t seen any llm be correct 90% of the time. IIRC SOTA on swebench (which tbf isn’t a great benchmark) is around 30%. > flawless write-eval loops with the help of cline, cline is a pretty good programmer. I’m not really sure what you mean by “flawless” but having a rubber duck is always more helpful than harmful. > A lot of things AI is helping with also have good, easy to observe / generate, real-time metrics you can use to judge excellence. Like what?
- kajecounterhack 2y ago> A lot of things AI is helping with also have good, easy to observe / generate, real-time metrics you can use to judge excellence. Exactly what I illustrated earlier: your developer productivity metrics. If you're turning code around faster, setting up your network better, turning around insights faster, the AI is working. > It depends. For a sysadmin maybe not, but for data scientists, the bar would be pretty high just to understand the math jargon. Why does an AI coding agent need to understand math jargon -- it just helps you write better code. Are you even familiar with what data scientists do? Seems not because if you were, you'd see clearly where the tool would be applied and do a good/bad job. Reminder: we're talking about evaluating whether Codebuff / alternatives are "pretty good" at X. Just go play with the tools. tgtweak expressed their opinion on how good the tool rates at some tasks {sysadmin, data engineering, cloud architecture} and your response was to question how someone could have an opinion about it. The obvious answer is that they used the tools and found it useful for those tasks. It may only be _subjectively_ good at what they're using for but it's also a rando's opinion on the internet. As another rando I very much agree with what the person you responded to is saying. You're not going to get more rigor from this discourse - go form a real opinion of your own.
- dartos 2y agoWow… why’d you get so defensive and presumptuous? I have my opinion, it’s just not the same as yours.
- kajecounterhack 2y agoLabel it what you want, I'm responding directly to questions you posed. > I have my opinion, it’s just not the same as yours. This is literally the TL;DR for what I wrote.