16 ms·
Opus remained better than GPT for me, even after the release of GPT-4o. VERY happy to see an even further improvement beyond that, Claude is a terrific product
by netsec_burn 2y ago
Opus remained better than GPT for me, even after the release of GPT-4o. VERY happy to see an even further improvement beyond that, Claude is a terrific product and given the news that GPT-5 only began its training several weeks ago I don't see any situation where Anthropic is dethroned in the near term. There are only two parts of Anthropic's offering I'm not a fan of:
- Lack of conversation sharing: I had a conversation with Claude where I asked it to reverse engineer some assembly code and it did it perfectly on the first try. I was stunned, GPT had failed for days. I wanted to share the conversation with others but there's no way provided like GPT, and no way to even print the conversation because it cuts off on the browser (tested on Firefox).
- No Android app. They're working on this but for now, there's only an iOS app. No expected ETA shared, I've been on the waitlist.
I feel like both of these are relatively basic feature requests for a company of Anthropic's size, yet it has been months with no solution in sight. I love the models, please give me a better way of accessing them.
- gotrythis 2y agoWhat I understand is that it's GPT 6 that just went into training, and that GPT 5 is complete and being delayed until after the U.S. election.
- r2_pilot 2y ago(assuming you are correct) It says something about how a company feels about the safety of their products when they feel like they should time the releases based on political events.
- futureshock 2y agoThis is speculation because I don’t think any of the key players ever explicitly stated this is their strategy, but this year it feels like there’s some significant foot dragging on things like Sora and GPT-5. The big AI players really don’t want AI to become an election year punching bag and don’t want any major campaign promises around AI to placate a spooked electorate. And they really don’t want it to be revealed that generative AI powered bot armies outnumber real human political discourse 10-1. And they absolutely do not want an AI generated hoax video to have a measurable effect on the polls. It’s a stopgap. If we get through this election without a major public freak out, it gives the industry 4 more years to take LLMs out to the point of diminishing returns and figure out safety before we get knee jerk regulation.
- viraptor 2y agoIt there any online confirmation of this, that's more than speculation?
- icpmacdo 2y agoNo there is not
- PaulWaldman 2y agoAnd after GPT-5's release, what would be the plan for subsequent elections? This seems to be a temporary play in delaying AI regulation if public sentiment further becomes that AI can have a strong influence in the elections.
- imjonse 2y agoGPT-5 will make elections obsolete :)
- futureshock 2y agoIt’s absolutely temporary, but 4 years feels like an eternity in this field and the m sure the major players would love to have that much time to entrench themselves before they have to battle “AI ban” legislation.
- modeless 2y agoThis is pure speculation, right?
- gotrythis 2y agoI've listened to so many interviews that I couldn't tell you who said what at this point, but that is what I understood from somewhere. So, sure, take it as speculation.
- gotrythis 2y agoHere's something that talks about it. I can't speak for the legitimacy, but I'm not pulling it out of my ass. They may be pulling it out of theirs. :-) https://lifearchitect.ai/gpt-6/ https://lifearchitect.ai/gpt-6/
- sva_ 2y agoSource: trust me bro
- ilaksh 2y agoI also believe that gpt-4o was originally called gpt-5. If you look at the image generation on their website from gpt-4o which has not been released, I believe that along with the voice caused Ilya to declare mission accomplished (AGI) and that is why there was a coup. The coup failed because no one wanted to wrap up the company or change the way it operated because they would lose a lot of money. The reason the name was changed was because there was a big public scare about gpt-5 taking over and so Altman had to promise not to release gpt-5 soon. So they changed the name to gpt-4o (omni). Which is A) obviously dramatically a different architecture, B) a huge step up in capabilities (most still unreleased) C) very general purpose. Because of A) and B), this should obviously be a new major version (5). Yes, this is speculation, but it's very obvious speculation to me. It's weird for me that most people not only don't share this view but seem to absolutely hate when I say it.
- christianqchung 2y agoI don't hate this speculation, I just don't buy it at all. 4o's about the same in terms of reasoning as 4. People don't find the text abilities that much more usable over 4 (at least on the LMS leaderboard). It's faster and has audio2audio capabilities alongside new native image stuff I think, but how exactly is that AGI if 4 isn't? These models understanding and reasoning ability is still far too weak to do any serious economic shifts yet.
- ilaksh 2y agoScroll to Explorations of Capabilities: https://openai.com/index/hello-gpt-4o/ https://openai.com/index/hello-gpt-4o/ That combined with the voice was probably considered AGI by Ilya.
- christianqchung 2y agoYes, I've seen this. Read my comment.
- Rastonbury 2y agoIt's speculation with no basis at all, OAI has a track record of releasing half step models and 4o is no different just like 3 to 3.5 and the numerous subsequent 3.5 releases. If you've used 4 and 4o they are too similar for 4o to have been trained from scratch
- deleted 2y ago[deleted]
- viraptor 2y agoOn the plus side, at least ChatBoost supports both openai and claude API. But for this specific model it seems to be broken... I hope that gets noticed and fixed soon.
- coreylane 2y agoI recently released Slackrock [https://github.com/coreylane/slackrock https://github.com/coreylane/slackrock] that you may find helpful, it's a Slack chat app that can access several FMs (including Claude 3.5) via AWS Bedrock. Responses can be easily shared with others by inviting them to your channels, and Slack has an Android app. It doesn't support attachments (yet) but I'm working on it!
- natsucks 2y agocool!
- sk11001 2y agoBoth GPT-4 and 4o have been completely useless for coding in the past couple of weeks for me - constant errors, and not just your typical LLM inaccuracies but incapable of producing a few lines of self-consistent code e.g. defines variables foo on one line and refers to it as bar on the next, or it misspells it as foox.
- esafak 2y agoFor me it has been very repetitious despite my instruction to the contrary.
- ipsum2 2y agoIt's the same model though. Maybe your perception has changed.
- ndr_ 2y agoI have first noticed logprob fluctuations in GPT-4o. Perhaps the same phenomenon is also going on with Turbo. I din‘t recall specifics but it was naming inconsistencies with variable names, meaning: same variable name got a typo somewhere, but the typo was close enough - perhaps a space vs. an underscore or something like that. Model could be the same, but maybe some in the infra is different.
- great_psy 2y agoI can’t speak for what OpenAI is doing, but I’ve noticed those types of hallucinations occurring when I quantize a model beyond a certain point. Maybe they are trying to cut down on memory usage ?
- edub 2y agoIs it the same? On the Models page of the API docs it says that GPT-4 is using the June 13th which would be different than the March 23rd.
- labrador 2y agoWaht language? Because I'm guessing they work well for languages with a large amount of training data like Python (in my experience), less well for less used languages like Zig or Clojure (haven't tried them but that's my theory)
- stuckinhell 2y agoI've had way better success with GPT-4o than claude. I wonder why
- netsec_burn 2y agoHave you tried 3 Opus or 3.5 Sonnet? Are you using it for programming, or something else?
- stuckinhell 2y agoeverything really. just opus so far
- simonw 2y agoPersonal prompting style, I imagine,
- Workaccount2 2y agoPeople really, really, underestimate how important prompting is. I would be confident in stating that half the people who complain about a model are actually just suffering from poor prompting.
- SSLy 2y agoare non-snake oil prompting techniques described anywhere?
- simonw 2y agoThose are hard to come by, but the Anthropic prompting documentation is a pretty great source: https://docs.anthropic.com/en/docs/build-with-claude/prompt-engineering/be-clear-and-direct https://docs.anthropic.com/en/docs/build-with-claude/prompt-...
- prng2021 2y agoAnd what makes you so confident that all those people are using different prompt styles when comparing models? You think most people don’t even understand the bare basics of how to compare two products?
- henry_viii 2y ago> Lack of conversation sharing... [there is] no way to even print the conversation because it cuts off on the browser (tested on Firefox). Until they make conversations shareable, in the meantime you can print the whole page in Chrome by: - going to Developer Tools (Ctrl + Shift + I) - opening the Command Palette (Ctrl + Shift + P) - searching for 'screenshot' - selecting Capture full size screenshot
- SubiculumCode 2y agoI do wonder if GPT quality fluctuates seasonally, or with electricity costs, in an engineering effort to balance costs with performance. I agree on all your points, but would like to emphasize that I really do enjoy the voice input voice output thing that chatgpt's app has. Its not how I use it when working, but when commuting, a lot of times, I'll turn on the the chatgpt app and have a conversation with it exploring ideas related to work or side projects. Its better than NPR, and I can't listen to the '3d6 Down the Line' podcast everyday, just once a week. I've been subscribed to PHind, which is a decent service allowing access to their models, chatgpt 4 turbo and o, and claudes. Its been incredibly useful, especially with their search integration. Unfortunately, while chatgpt can be used 500 times a day, Claude is only 10, although I guess it goes into an API like payment mode after that on top of subscription. I sure wish I'd buckle down and calculate my usage to really get an idea of whether subscription is cheaper or more expensive for me compared to API.
- lxgr 2y agoShort of switching between models (which at least OpenAI definitely does for free customers, but I believe they always indicate it), how would that work? Different quantizations?
- SubiculumCode 2y agocaught me speculating. I suppose some mild quanting and/or prompt injection to keep responses smaller unless specifically asked: e.g. use ...
- Powdering7082 2y ago> GPT-5 only began its training several weeks ago Source?
- netsec_burn 2y agohttps://openai.com/index/openai-board-forms-safety-and-security-committee/ https://openai.com/index/openai-board-forms-safety-and-secur... (May 28th) > OpenAI has recently begun training its next frontier model and we anticipate the resulting systems to bring us to the next level of capabilities on our path to AGI.
- cadence- 2y agoBased on other things they said in the last couple of months, it looks like GPT-4.5 is coming this summer, and then GPT-5 in the Fall.
- gagagaga7 2y agoNo doubt openai have been training big models for the last year. If “gpt5” is only just starting it means recent training runs have had disappointing results and have been passed off as “Gpt4o” or whatever. The value of all the AI companies is predicated on high chance of AGI, and gpt5 failing to be revolutionary may pop the whole bubble (+10 trillion of market cap)
- Workaccount2 2y agoSam said on Lex's podcast that people should temper their expectations for GPT-5, not in that it will necessarily suck, but that they want to ramp up ability slowly over time rather than discrete large steps.
- letitgo12345 2y agoSounds like an excuse tbh. Esp when other companies are pushing ahead beyond OAI and open source is close to rivaling them
- mac-attack 2y agoI'm sticking w/ Claude for the foreseeable future as they seem less slimy than OpenAI/Microsoft/Google so far and care about safety. I'm in the same boat waiting for an Android app btw. One other feature that I'm hoping they catch up to others on is a permanent context window so that I can get Claude to stop speaking so formally all the time
- joshstrange 2y agoTo each their own, but I still prefer ChatGPT. The UI for Claude is terrible in my opinion. I had subscriptions for both and I would fire off questions to both of them and see which one I liked more and I consistently liked the ChatGPT ones more. I canceled my subscription last week for Claude. I am super happy that Anthropic continues to push the envelope on this and I hope to re-subscribe to them in the future.
- spidersouris 2y agoIf it's really only the UI that's bothering you, why not use a web UI such as Open WebUI?
- joshstrange 2y agoThe UI wasn’t the only issue, but I will look into that.
- trungdq88 2y agoIf you have an API key, using Opus with a 3rd party UI like typingmind.com solves all of the problems you mentioned (disclaimer: I'm the app developer)
- lannisterstark 2y agoI use LibreChat for this as self hosted UI. Works awesome.
- wonderfuly 2y ago> Lack of conversation sharing You can use my product https://ChatHub.gg https://ChatHub.gg which supports dozens of chatbots including Claude and can share conversations from any of them.
- Alifatisk 2y ago> I had a conversation with Claude where I asked it to reverse engineer some assembly code and it did it perfectly on the first try. I was stunned I share the same experience with you but with Claude 3 Sonnet. I can’t count how many times I’ve shared some code with Claude with barely any hope because other GPTs failed aswell, yet, Claude surprised me and performed the task with success. I’ve actually reached to the point that I expressed my gratitude to Claude because of how well it performs on coding tasks and other tasks in general. I don’t know what Anthropic did, but something did they right. Being able to handle large amounts of tokens, “understand” and perform tasks on it & spit out large amounts of data back with barely any cut-offs (unlike Gemini) has made me feel like Claude is at the moment the best option.