4 ms·
[flagged]
by fhfncjcc 2mo ago
[flagged]
- Legend2440 2mo agoProof? In my experience modern models are better at all tasks than models from two years ago, especially complex multi-step tasks.
- alightsoul 2mo agoyou are working on coding. they are working on things like "creative writing" remember that gpt 4o was popular among those who had ai as a romantic partnet?
- Legend2440 2mo agoWell that's on purpose lol. OpenAI does not want you falling in love with their chatbot and have been deliberately training it to be less romantic.
- vablings 2mo agoThere have been several cases of suicide and self-harm related to 4o, AI psychosis is a real risk and will probably be in the DSM
- desterothx 2mo agosure they do if it makes them money, probably just not worth the controversy right now
- michaelmrose 2mo agosycophancy It wasn't "better" it was better at kissing your ass which matches what a lot of people want in a partner.
- Marha01 2mo ago> remember that gpt 4o was popular among those who had ai as a romantic partner I suspect GPT 5.6 would be even better at it, if given the same sycophantic system prompt and lack of guardrails.
- QwenGlazer9000 2mo agoDude it's not a system prompt, it's the training.
- HDBaseT 2mo agoThe API lets you adjust the temperature. Lower values introduce more deterministic outputs, which likely helps with the hallucination rates. If you want creative writings, use the API and play with the sliders.
- moyix 2mo agoThey actually removed the temperature parameter starting with GPT-5.
- Marha01 2mo agoI highly doubt that.
- whimsicalism 2mo agogpt4o & associated parasociality is considered an alignment failure and is actively trained out of the model, so that is a terrible example of regression
- criddell 2mo agoDo any of the big AI companies have a model that are good at tasks that require learning? For example, every day people teach teenagers how to drive and with only dozens of hours of practice, they are on the road.
- whimsicalism 2mo agois this not essentially what ARC-AGI-3 is? i agree that in-context/continual learning is somewhere the models are still mostly weak at
- criddell 2mo agoI don't think any of the ARC-AGI-3 tests are very interesting. At least not as interesting as driving a car. Children literally do a similar task in go karts every day. Another interesting task would be to take the AI in a robot body into a vegetable garden and teach it to pull weeds. This is another task that lots of children help out with.
- nostrebored 2mo agoFor customer support I don't think models have gotten better since gpt-4.1. The class of small models, with limited to no reasoning, that need to handle a complex issue with a touch of empathy, has not improved much. I think most are actually worth, as agentic harnesses seem to optimize for solving poorly described problems rather than following complex procedures as written. In other words, instruction following maximizing models seem to make worse free-form agents, but they're really all that some domains need.
- jstummbillig 2mo agoI understand the point (I don't agree with it; tool calling has gotten much better/reliable and that is very important for customer support) but consider: If you can get same for a lot less, that's an improvement. If we found a way to supply fresh water and electricity for -90% cost after 2 years, that would be fantastic. You can do many more things, when stuff is cheaper, even if the stuff were otherwise unchanged.
- heaney-555 2mo ago>The class of small models, with limited to no reasoning What? GPT-4.1 was not a small model! And why wouldn't you use reasoning? You're of course going to see poor results when you restrict yourself to small non-reasoning models, but why would you?
- nostrebored 2mo ago"Small" was a poor choice of words here, "low compute budget" is more what I'm getting at. In voice interactions, ttfat is actually relatively important. If you look at models with a <1s ttfat you eliminate almost every reasoning model, less some of the diffusion models and more obscure ddtree/dflash like speculative decoding implementations.
- heaney-555 2mo agohttps://openai.com/index/introducing-gpt-live/ https://openai.com/index/introducing-gpt-live/ GPT-Live, which is coming to the API soon, responds instantly while reasoning in the background. So it can say "Hold on, I'll look that up for you" and continue to respond to the user conversationally while running an asynchronous reasoning task in the background. It's not in the API yet, but it should be in the coming weeks. You'll see an enormous improvement compared to GPT-4.1.
- analoger 2mo agoData 'compression' collapse. People publish AI generated slop on the internet -> next generation of AI is trained on that data -> the lossy/fuzzy training make the output worse -> rinse and repeat.
- Chance-Device 2mo agoSo your answer is: ignore the progress, it’s not really happening, actually it’s getting worse. That’s not a credible position, but there isn’t anything that I or anyone else can say to someone who simply doesn’t want to believe something.
- cmdli 2mo agoIt sounds like they are making a clear argument: models are getting worse for certain domains even while they are getting better at others. I don't know if I agree with that but it doesn't seem like an irrational claim and does seem credible to me.
- Chance-Device 2mo agoI just don’t think it’s true, or is significant enough to matter to the direction of travel of AI even if there were something to it. It’s another cope post being lobbed at the idea of AI going somewhere and I’m sick of them.
- lioeters 2mo ago"The sooner you can be broken out of your denial about all this the better, and we can start actually taking you seriously."
- teravor 2mo agoevery lab independently discovered that getting good at bit alchemy (coding and related tasks) should come first as it will enable the formation of training pipelines that will then solve everything else. so far there is no end to this progress in sight so it's full steam ahead on this singular domain. once it plateaus you should expect to see the greatest disruptions in human endeavors ever as all the training flops will start flowing to other domains to disrupt and dominate.