4 ms·
I am used to seeing technical papers from ieee, but this is an opinion piece? I mean, there is some anecdata and one test case presented to a few different mode
by ronbenton 9mo ago
I am used to seeing technical papers from ieee, but this is an opinion piece? I mean, there is some anecdata and one test case presented to a few different models but nothing more.
I am not necessarily saying the conclusions are wrong, just that they are not really substantiated in any way
- troyvit 9mo agoYeah I saw the ieee.org domain and was expecting a much more rigorous post.
- ronbenton 9mo agoThis may be a situation where HackerNews' shorthand of omitting the subdomain is not good. spectrum.ieee.org appears to be more of a newsletter or editorial part of the website, but you wouldn't know that's what this was just based on the HN tag.
- preommr 9mo agoI've been on this site for over a decade now and didn't know this. That's a genuinely baffling decision given how different content across subdomains can be.
- bee_rider 9mo agoOn the other hand, “ieee spectrum” is directly at the top of the page, then “guest article.”
- ronbenton 9mo agoWell, as much as I'm sure HN is a special place ;) , it is well documented that a lot of people on the internet just read the headlines
- hxugufjfjf 9mo agoArticles? Headlines? I go right to the comments!
- badc0ffee 9mo agoMaybe an exception could be made here, like HN does for medium.com.
- causal 9mo agoAnd the example given was specific to OpenAI models, yet the title is a blanket statement. I agree with the author that GPT-5 models are much more fixated on solving exactly the problem given and not as good at taking a step back and thinking about the big picture. The author also needs to take a step back and realize other providers still do this just fine.
- wavemode 9mo agoHe tests several Claude versions as well
- causal 9mo agoAh you're right, scrolled past that - the most salient contrast in the chart is still just GPT-5 vs GPT-4, and it feels easy to contrive such results by pinning one model's response as "ideal" and making that a benchmark for everything else.
- wavemode 9mo agoTo be fair, it's very rare that articles praising the power of AI coding assistants are ever substantiated, either. In the end, everyone is kind of just sharing their own experiences. You'll only know whether they work for you by trying it yourself.
- franktankbank 9mo agoAnd you can't try it out without for the most part feeding the training machine for at best free.
- pc86 9mo agoAre there a lot of products or services you can try out without using the product or service?
- franktankbank 9mo agoWithout sending info to their servers, yes.
- Leynos 9mo agoCodex and Claude Code allows you to opt out of model training. Perhaps you don't believe OpenAI and Anthropic when they say this, but it is a requirement upon which most enterprise contracts are predicated.
- deleted 9mo ago[deleted]
- mrguyorama 9mo ago> You'll only know whether they work for you by trying it yourself. But at the same time, even this doesn't really work. The lucky gambler thinks lottery tickets are a good investment. That does not mean they are. I've found very very limited value from these things, but they work alright in those rather constrained circumstances.
- verdverm 9mo agoand they are using OpenAI models, who haven't had a successful training run since Ilya left, GPT 5x is built on GPT 4x, not from scratch aiui I'm having a blast with gemini-3-flash and a custom copilor replacement extension, it's much more capable than Copilot ever was with any model for me and a personalized dx with deep insights into my usage and what the agentic system is doing under the hood.
- RugnirViking 9mo agocan you talk a little more about your replacement extention? I get copilot from my worksplace and id love to know what I can do with it, ive been trying to build some containerized stuff with copilot cli but im worried I have to give it a little more permissions than im comfortable with around git etc
- verdverm 9mo agoI post a lot about it on Bluesky https://bsky.app/profile/verdverm.com https://bsky.app/profile/verdverm.com The container stuff that backs it is built on Dagger https://github.com/hofstadter-io/hof/tree/_next/examples/env https://github.com/hofstadter-io/hof/tree/_next/examples/env The entire extension and agent framework is in that repo too extensions/vscode and lib/agent I let my agent do whatever because I know exactly what it can and can't do. For example, it can use git, but cannot push, and any git changes are local to its containerized environment and don't get exported back to my filesystem where I do real git work. I could create an envelope where they could push git, and more likely I'll give them something where they can call GitHub ali, that's really more useful anyway
- esafak 9mo agoThis is the Spectrum magazine; the lighter fare. https://en.wikipedia.org/wiki/IEEE_Spectrum https://en.wikipedia.org/wiki/IEEE_Spectrum