36 ms·
Claude Sonnet 4.5
System card: https://assets.anthropic.com/m/12f214efcc2f457a/original/Claude-Sonnet-4-5-System-Card.pdf https://assets.anthropic.com/m/12f214efcc2f457a/original/Cla...
- jdthedisciple 1y agoWhy the focus on the "alignment"-aspect of safety? Surely there are more pressing issue with LLMs currently...
- smakosh 1y agoAvailable on llmgateway.io already
- vinhnx 1y agoClaude Sonnet 4.5 has landed support in my CLI coding agent VT Code, combining SOTA language model and agentic semantic code understanding github.com/vinhnx/vtcode
- idkmanidk 1y agoPage cannot be found Empty screen mocks my searching Only void responds but: https://imgur.com/a/462T4Fu https://imgur.com/a/462T4Fu
- dbbk 1y agoSo Opus isn't recommended anymore? Bit confusing
- SatvikBeri 1y agoFor now, yeah. Presumably they'll come out with Opus 4.5 soon.
- causal 1y agoDon't think I've ever preferred Opus to Sonnet
- cryptoz 1y agoI've really got to refactor my side project which I tailored to just use OpenAI API calls. I think the Anthropic APIs are a bit different so I just never put in the energy to support the changes. I think I remember reading that there are tools to simpify this kind of work, to support multiple LLM APIs? I'm sure I could do it manually but how do you all support multiple API providers that have some differences in the API design?
- willcodeforfoo 1y agohttps://openrouter.ai/ https://openrouter.ai/?
- pinum 1y agoI use LiteLLM as a proxy.
- dingnuts 1y ago> think I remember reading that there are tools to simpify this kind of work, to support multiple LLM APIs just ask Claude to generate a tool that does this, duh! and tell Claude to make the changes to your side project and then to have sex with your wife too since it's doing all the fun parts
- deleted 1y ago[deleted]
- adidoit 1y agoLiteLLM is your friend.
- adidoit 1y agoor AI SDK
- l1n 1y agohttps://docs.anthropic.com/en/api/openai-sdk https://docs.anthropic.com/en/api/openai-sdk
- 1y ago
- yewenjie 1y agoLooking at the chart here, it seems like Sonnet 4 was already better than GPT-5-codex in the SWE verified benchmark. However, my subjective personal experience was GPT-5-codex was far better at complex problems than Claude Code.
- jjcm 1y agoHow long have you had early access for?
- CuriouslyC 1y agoThe Anthropic models have been vibe-coding tuned. They're beasts at simple python/ts programs, but they definitely fall apart with scientific/difficult code and large codebases. I don't expect that to change with the new Sonnet.
- patates 1y agoIn my experience Gemini 2.5 Pro is the star when it comes to complex codebases. Give it a single xml from repomix and make sure to use the one at the aistudio.
- CuriouslyC 1y agoYup. In fact every deep research tool on the market is just a wrapper for gemini, their "secret sauce" is just how they partition/pack the codebase to feed it into gemini.
- Workaccount2 1y agoIts mostly because it is so damn good with long contexts. It can stay on the ball even at 150k whereas other models really wilt around 50-75k.
- garciasn 1y agoIn my experience, G2.5P can handle so much more context and giving an awesome execution plan that is implemented by CC so much better than anything G2.5P will come up with. So; I give G2.5P the relevant code and data underneath and ask it to develop an execution plan and then I feed that result to CC to do the actual code writing. This has been outstanding for what I have been developing AI assisted as of late.
- chipgap98 1y agoInteresting that this is better than Opus 4.1. I want to see how this holds up under real world use, but if that's the case its very impressive. I wonder how long it will be before we get Opus 4.5
- FergusArgyll 1y agoIIRC sonnet 3.5 (and definitely 3.5-new aka 3.6) was better than opus 3. There's still a lot of low hanging fruit apparently
- kixiQu 1y agoLots of feature dev here – anyone have color on the behavior of the model yet? Mouthfeel, as it were.
- meetpateltech 1y agoSeeing the progress of the Claude models is really cool! Charting Claude's progress with Sonnet 4.5: https://youtu.be/cu1iRoc1wBo https://youtu.be/cu1iRoc1wBo
- clueless 1y agowould love to see the prompt they used and the final code of the Claude.ai clone it generated
- mohsen1 1y agoPrice is playing a big role in my AI usage for coding. I am using Grok Code Fast as it's super cheap. Next to it GPT-5 Codex. If you are paying for model use out of pocket Claude prices are super expensive. With better tooling setup those less smart (and often faster) models can give you better results. I am going to give this another shot but it will cost me $50 just to try it on a real project :(
- muttantt 1y agohow are you using grok code fast? what tooling/cli/etc?
- deleted 1y ago[deleted]
- rafaquintanilha 1y agoIt’s currently free in OpenRouter.
- esafak 1y agoThrough Opencode.
- xwowsersx 1y agoSame
- hu3 1y agofree in GitHub copilot atm
- _joel 1y agoI'm paying $90(?) a month for the Max and it holds up for about an hour or so of in depth coding before it kicks in the 5-hour window lockout (so effectively about 4 hours of time when I can't run it). Kinda frustrating, even with efficient prompt and context length conservation techniques. I'm going to test this new sonnet 4.5, now but it'll probably be just as quick to gobble my credits.
- greenfish6 1y agoAs the rate of model improvement appears to slow, the first reactions seem to be getting worse and worse, as it takes more time to assess the model's quality and understand the nuances & subtler improvements
- alach11 1y agoI'm really interested in the progress on computer use. These are the benchmarks to watch if you want to forecast economic disruption, IMO. Mastery of computer use takes us out of the paradigm of task-specific integrations with AI to a more generic interface that's way more scalable.
- mrshu 1y agoWhat are some standard benchmarks you look at in this space?
- sipjca 1y agoMaybe this is true? But it's not clear to me this methodology will ever be quite as good as native tool calling. Or maybe I don't know the benchmark well enough, I just assume it's vision based Perhaps Tesla FSD is a similar example where in practice self driving with vision should be possible (humans), but is fundamentally harder and more error prone than having better data. It seems to me very error prone and expensive in tokens to use computer screens as a fundamental unit. But at the same rate, I'm sure there are many tasks which could be automated as well, so shrug
- simianwords 1y agoLooks like RPA vs API debate all over again
- cantor_S_drug 1y agoDo you think a Genie like model specifically trained on data consisting of interacting with application interfaces would be good on computer use tasks?
- _joel 1y ago`claude model claude-sonnet-4-5-20250929` for cli users
- mohsen1 1y agoThat's a pretty pelican on a bicycle! https://jsbin.com/hiruvubona/edit?html,output https://jsbin.com/hiruvubona/edit?html,output https://claude.ai/share/618abbbf-6a41-45c0-bdc0-28794baa1b6c https://claude.ai/share/618abbbf-6a41-45c0-bdc0-28794baa1b6c
- greenfish6 1y agopelican on a bicycle benchmark probably getting saturated... especially as it's become a popular way to demonstrate model ability quickly
- AlecSchueler 1y agoBut where is the training set of good pelicans on bikes coming from? You think they have people jigging them up internally?
- eli 1y agoAssuming they updated the crawled training data, just having a bunch of examples of specifically pelicans on bicycles from other models is likely to make a difference.
- AlecSchueler 1y agoBut then how does the quality increase? Normally we hear that when models are trained on the output of other models the style becomes very muted and various other issues start to appear. But this probably the best pelicans on a bicycle I've ever seen, by quite some margin.
- Kuinox 1y agoJust compare it with a human on a bicycle, you would see that LLMs are weirdly good at drawing pelicans in SVG but not humans.
- deleted 1y ago[deleted]
- atemerev 1y agoAh, the company where the models are unusable even with Pro subscription (start to hit the limit after 20 minutes of talking), and free models are not usable at all (currently can't even send a single message to Sonnet 4.5)...
- usr19021ag 1y agoTheir benchmark chart doesn't match what's published on https://www.swebench.com/ https://www.swebench.com/. I understand that they may have not published the results for sonnet 4.5 yet, but I would expect the other models to match...
- zurfer 1y agoSame price and a 4.5 bp jump from 72.7 to 77.2 SWEBench Pretty solid progress for roughly 4 months.
- zurfer 1y agoAlso getting a perfect score on AIME (math) is pretty cool. Tongue in cheek: if we progress linearly from here software engineering as defined by SWE bench is solved in 23 months.
- wohoef 1y agoJust a few months ago people were still talking about exponential progress. The fact that we’re already going for just linear progress is not a good sign
- falcor84 1y agoLinear growth on a 0-100 benchmark is quite likely an exponential increase in capability.
- usaar333 1y agoExcept it is sublinear. Sonnet 4 was 10.2% above sonnet 3.7 after 3 months.
- GoatInGrey 1y agoWe should all know that in the software world, the last 10% requires 90% of the effort!
- baq 1y agoSublinear as demonstrated on a sigmoid scale is quite fast enough for me thank you.
- falcor84 1y agoThis got me thinking - is there any reasonable metric we could use to measure the intellectual capabilities of the most capable species on Earth that had evolved at each point in time? I wonder what kind of growth function we'd see. Silly idea - is there an inter-species game that we could use in order to measure ELO?
- schmorptron 1y agoOh wow, a lot of focus on code from the big labs recently. In hindsight it makes sense that the domain the people building it know best is the one getting the most attention, and it's also the one the models have seen the most undeniable usefulness in so far. Though personally, the unpredictability of the future where all of this goes is a bit unsettling at the same time...
- modeless 1y agoOpenAI and Anthropic are both trying to automate their own AI research, which requires coding.
- martinald 1y agoThing is though if you are good at code it solves many other adjacent tasks for LLMs, like formatting docs for output, presentations, spreadsheet analysis, data crawling etc.
- doctoboggan 1y agoAlong with developers wanting to build tools for developers like you said, I think code is a particularly good use case for LLMs (large language models), since the output product is a language.
- fragmede 1y agoIt's because the output is testable. If the model outputs a legal opinion or medical advice, a human needs to be looped in to verify that the advice is not batshit insane. Meanwhile, if the output is code, it can be run through a compiler and (unit) tests run to verify that the generated code is cromulent without a human being in the loop for 100% of it, which means the supercomputer can just go off and do it a thing with less supervision.
- neuronexmachina 1y agoI think coding is also the area where companies are most likely to buy large team licenses.
- 1y ago
- fibers 1y agoThis looks exciting. I hope they add this to Windsurf soon.
- pzo 1y agoit looks like its already there
- simianwords 1y agoIt’s stupid… like just have a registry of models and let people automatically use them. It’s silly to wait for manual whitelisting each time for every app
- ReverseCold 1y agoIt was there a few (<5? I think?) minutes after the Anthropic post went out. If you look at Windsurf's web traffic it looks like they did a thing (model is an int) to make it so the IDE doesn't need to update to get new models.
- fibers 1y agoI agree, I use Windsurf for personal projects and I think the pricing model is a bit better than what a professional dev would be using on cursor or something like that.
- cube2222 1y agoSo… seems like we’re back to Sonnet being better than Opus? At least based on their benchmarks. Curious to see that in practice, but great if true!
- catigula 1y agoI happened to be in the middle of a task in a production codebase that the various models struggled on so I can give a quick vibe benchmark: opus 4.1: made weird choices, eventually got to a meh solution i just rolled back. codex: took a disgusting amount of time but the result was vastly superior to opus. night and day superiority. output was still not what i wanted. sonnet 4.5: not clearly better than opus. categorically worse decision-making than codex. very fast. Codex was night and day the best. Codex scares me, Claude feels like a useful tool.
- poisonborz 1y agoThese reviews are pretty useless to other developers. Models perform vastly differently with each language, task type, framework.
- MichealCodes 1y agoI really hope benchmarking improves soon to monitor the model in the weeks following the announcement. It really seems like these companies introduce a new "buffed" model and then slowly nerf the intelligence through optimizations. If we saw task performance week 1 vs week 8 on benchmarks, this would at least give us more insight into the loop here. In an environment lacking true progress a company could surely "show" it with this strategy.
- SubiculumCode 1y agoI do wonder about this. I just don't know if it real or in our heads
- beefnugs 1y agoCapitalism is pure scam now on every level: they did this with nvme drives in the last couple years. Sending out perfect hardware to reviewers then rug pulling trash to ship to the world
- commakozzi 1y agoIt does feel like it has to be real. I've noticed it since chatGPT with GPT-3.5, once it hit big news publicly and demands were made on "censoring" its output to limit biases, etc. (not inherently a problem to do this with LLMs as a society, but it does affect the output for obvious reasons). Whatever workflow OpenAI and others have applied, seems to be post-release somehow? i'm ignorant and just speculating, but literally every model release i've noticed it. Starts strong, ends up feeling less capable days, weeks, months after. I'm sure some of it could be in the parallelization of processing that has to occur to service the large amount of requests. and more and more traffic are spreading it thin?
- MichealCodes 1y ago> I'm sure some of it could be in the parallelization of processing that has to occur to service the large amount of requests. and more and more traffic are spreading it thin? Even if this is the case, benchmarks should be done at scale too if the models suffer from symptoms of scale. Otherwise the benchmarks are just a lie unless you have access to an unconstrained version of the model.
- scosman 1y agoInteresting quirk on first use: "`temperature` and `top_p` cannot both be specified for this model. Please use only one."
- zora_goron 1y agoWhy might this be, does anyone know?
- epolanski 1y agoThis isn't new to other models, and it doesn't make much sense to specify both.
- seaal 1y agoThey really had to release an updated model, I can only imagine how many people cancelled their plans and switched over to Codex over the past month. I'm glad they at least gave me the full $100 refund.
- GenerWork 1y agoI'm one of them, but I'm just a product designer who likes to jump between various AI tools to get experience with them. Once my month with OpenAI is up, I may jump back to CC as I liked some of the non-coding features more, specifically plan mode.
- epolanski 1y agoGoing from pro to Max was a giant let down. Then they even started sending me marketing emails which was the straw that broke the camel's back, I use to cancel subscriptions of companies spamming my email.
- user1999919 1y agoits time to start benchmarking benchmarks. im pretty sure they are bmw levels doping the game here
- user1999919 1y ago*vw (volkswagen)
- trevin 1y agoI’m always fascinated by the fine-tuning of LLM personalities. Might we finally get less of the reflexive “You’re absolutely right” with this one? Maybe we’re entering the Emo Claude era. Per the system card: In 250k real conversations, Claude Sonnet 4.5 expressed happiness about half as often as Claude 4, though distress remained steady.
- fnordsensei 1y agoI personally enjoy the “You’re absolutely right!” exclamation. It signals alignment with my feedback in a consistent manner.
- transcriptase 1y agoYou’re overlooking the fact that it still says that when you are, in reality, absolutely wrong.
- podgietaru 1y agoAnd that it often spits out the exact same wrong answer in response.
- fnordsensei 1y agoThat’s not the purpose of it, as I understand it; it’s a token phrase generated to cajole it down a particular path.[1] An alignment mechanism. The complement appears to be, “actually, that’s not right.”, a correction mechanism. 1: https://news.ycombinator.com/item?id=45137802 https://news.ycombinator.com/item?id=45137802
- baobabKoodaa 1y agoHmmh. I believe your explanation, but I don't think that's the full story. It's also a sycophancy mechanism to maximize engagement from real users and reward hack AI labelers.
- 1y ago
- rudedogg 1y agoI just ran this through a simple change I’ve asked Sonnet 4 and Opus 4.1, and it fails too. It’s a simple substitution request where I provide a Lint error that suggests the correct change. All the models fail. I could ask someone with no development experience to do this change and they could. I worry everyone is chasing benchmarks to the detriment of general performance. Or the next token weight for the incorrect change outweigh my simple but precise instructions. Either way it’s no good Edit: With a followup “please do what I asked” sort of prompt it came through, while Opus just loops. So theres that at least
- darksaints 1y ago> I worry everyone is chasing benchmarks to the detriment of general performance. I've been worried about this for a while. I feel like Claude in particular took a step back in my own subjective performance evaluation in the switch from 3.7 to 4, while the benchmark scores leaped substantially. To be fair, benchmarking has always been the most difficult problem to solve in this space, so it's not surprising that benchmark development isn't exactly keeping pace with all of the modeling/training development happening.
- GoatInGrey 1y agoNot that it was better at programming, but I really miss Sonnet 3.5 for educational discussions. I've sometimes considered that what I actually miss was the improvement 3.5 delivered over other models at that time. Though since my system message for Sonnet since 3.7 has been primarily instructing it to behave like a human and have a personality, I really think we lost something.
- walthamstow 1y agoI still use 3.5 today in Cursor. It's still the best model they've produced for my workflow. It's twice as fast as 4 and doesn't vomit pointless comments all over my code.
- MichealCodes 1y agoMore like churning benchmarks... Release new model at max power, get all the benchmark glory, silently reduce model capability in the following weeks, repeat by releasing newer, smarter model.
- sberens 1y agoIs "parallel test time compute" available in claude code or the api? Or is it something they built internally for benchmark scores?
- ancorevard 1y agoCan't use Anthropic models in Cursor. Completely cost prohibitive compared to gpt-5 and grok models. Why is this? Does Anthropic have just higher infrastructure costs compared to OpenAI/xAI?
- doctoboggan 1y agoPossibly, or they are pricing for sustainability and OpenAI/xAI are just burning through VC money.
- acchow 1y agoThe Anthropic models are also better at coding. Why wouldn’t they price it higher?
- dbbk 1y agoIt's meant to be used with the Max subscription
- wohoef 1y agoAnd Sonnet is again better than Opus. I’d love to see simultaneous release dates for Sonnet and Opus one day. Just so that Opus is always better than Sonnet
- cloverich 1y agoPlease y'all, when you list supportive or critical complaints based on your actual work, include some specifics of the task and prompt. Like actual prompt, actual bugs, actual feature, etc. I've had great success with both ChatGPT and Claude for years, am around 3x sustained output increase in my professional work, and kicking off and finishing new side projects / features that I used to simply not ever finish. BUT there's some tasks I run into where it's god awful. Because I have enough good experience, I know how to work around, when to give up, when to move on, etc. I am still surprised at things it cannot do, for example Claude code could not seem to stitch together three screens in an iOS app using the latest SwiftUI (I am not an iOS dev). IMHO for people using it off and on or sparingly, it's going to seem either incredible or worthless depending on your project and prompt. Share details, it's so helpful for meaningful conversation!
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- emil-lp 1y agoHow do you measure 3x sustained output increase? Is it number of lines? Tickets closed? PRs opened or merged? Number of happy customers?
- inopinatus 1y agoIt is undoubtedly 3x as many bugs.
- _alternator_ 1y agoThis would be a win. Professionals make about 1 bug for every 100 loc. If you get 3x the code with 3x the bugs, this is the definition of scaling yourself.
- senordevnyc 1y agoOh good, a new discussion point that we haven't heard 1000x on here. Have you heard of that study that shows AI actually makes developers less productive, but they think it makes them more productive?? EDIT: sorry all, I was being sarcastic in the above, which isn't ideal. Just annoyed because that "study" was catnip to people who already hated AI, and they (over-) cite it constantly as "evidence" supporting their preexisting bias against AI.
- marginalia_nu 1y agoIs there some accessible explainer for what these numbers that keep going up actually mean? What happens at 100% accuracy or win rate?
- lukev 1y agoIt means that the benchmark isn't useful anymore and we need to build a harder one. edit: as far as what the numbers mean, they are arbitrary. They are only useful insofar as you can run two models (or two versions of the same model) on the same benchmark, and compare the numbers. But on an absolute scale the numbers don't mean anything.
- typpilol 1y agoI thought the percentage was how many problems it successfully solved
- baq 1y agoTechnically correct, but not helpful nor actionable.
- marginalia_nu 1y agoIt was actually very helpful as it answered my question about what the benchmark numbers are. It wasn't a request for advice, but I'm merely looking to understand the article, which doesn't really elaborate on what they are presenting; either assuming an audience that is very familiar with these benchmarks prior, or so dazzled by number going up they forget to ask what number is.
- asadm 1y agothen we need new bench.
- unshavedyak 1y agoInteresting, in the new 2.0.0 claude code they got rid of the "Plan with Opus then switch to Sonnet" feature. I hope they're correct in Sonnet being good enough to Plan too, because i quite preferred Opus planning. It wasn't necessarily "better", just more predictable in my experience. Also as a Max $200 user, feels weird to be paying for an Opus tailored sub when now the standard Max $100 would be preferred since they claim Sonnet is better than Opus. Hope they have Opus 4.5 coming out soon or next month i'm downgrading.
- Implicated 1y agoI'm also a max user and I just _leave_ it on Opus 4.1 - I've never hit a rate limit.
- danielbln 1y agoI'm on the 25x MAX plan and if I go full hog on multiple projects I might see the yellow "Approaching Opus limits" message in Claude Code, but I have yet to have it lock me down, I usually slip right into the next 5h block and the message vanishes.
- stavros 1y agoSame, it very quickly says "approaching rate limits", and then just keeps going forever.
- asar 1y agoIn the same boat and ready to downgrade. But this must be on their radar, or they were/are losing money with opus...
- vb-8448 1y agoclaims against gpt-5 are huge! I used to use cc, but I switched to codex (and it was much better) ... no I guess I have to switch batch to CC, at least to test it
- bradley13 1y agoI need to try Claude - haven't gotten to it. I use AI for different things, though, including proofreading posts on political topics. I have run into situations where ChatGPT just freezes and refuses. Example: discussing the recent rape case involving a 12-year-old in Austria. I assume its guardrails detect "sex + kid" and give a hard "no" regardless of the actual context or content. That is unacceptable. That's like your word processor refusing to let you write about sensitive topics. It's a tool, it doesn't get to make that choice.
- Implicated 1y agoI'd imagine that the proportion of "legit" conversations around these topics and those that they're intending to not allow is large enough that it doesn't make sense for them to even entertain the idea of supporting those conversations. As a rather hilarious and really annoying related issue - I have a real use where the application I'm working on is partially monitoring/analyzing the bloodlines of some rather specific/ancient mammals used in competition and... well.. it doesn't like terms like "breeders" and "breeding"
- user34283 1y agoThis is the result of Anthropic and others focusing on imaginary threats about things the model cannot realistically do - such as engineer bio weapons. To guard against the imaginary threats, they compromise real use cases.
- jjordan 1y agoThis is why eventually, the AI with the fewest guardrails will win. Grok is currently the most unguarded of the frontier models, but it could still use some work on unbiased responses.
- beefnugs 1y agoStill has to be a local model too. Arbitrary government censorship on top of arbitrary corporate censorship is a hell no for me forever into the future
- catigula 1y agoI'm still absolutely right constantly, I'm a genius. I also make various excellent points.
- hu3 1y agoI wonder if/when this will be available to GitHub Copilot in VSCode.
- Osyris 1y agoWonder no more: https://github.blog/changelog/2025-09-29-anthropic-claude-sonnet-4-5-is-in-public-preview-for-github-copilot/ https://github.blog/changelog/2025-09-29-anthropic-claude-so...
- aliljet 1y agoThese benchmarks in real world work remain remarkably weak. If you're using this for day-to-day work, the eval that really matters is how the model handles a ten step action. Context and focus are absolutely king in real world work. To be fair, Sonnet has tended to be very good at that... I wonder if the 1m token context length is coming for this ride too?
- data-ottawa 1y agoAnecdotally this new Sonnet model is massively falling apart on my tool call based workflows. I’m having to handhold it through analysis tasks. At one point it wrote a python script that took my files it needed to investigate and iterated through them and ran `print(f”{i}. {file}”)` then printed “Ready to investigate files…” And that’s all the script did. I have no idea what’s going on with those benchmarks if this is real world use.
- edude03 1y agoAh, I figured something was up - I had sonnet 4 selected but it changed to "Legacy Model" while I was using the app.
- peterdstallion 1y agoI am a paying subscriber to Gemini, Claude and OpenAI. I don't know if it's me, but over the last few weeks I've got to the conclusion ChatGPT is very strongly leading the race. Every answer it gives me is better - it's more concise and more informative. I look forward to testing this further, but out of the few runs I just did after reading about this - it isn't looking much better
- yepyip 1y agoWhat about Grok, are they catching up?
- jjordan 1y agoGrok has been free for over a month now and for me it has certainly proven itself competent at most tasks that you would otherwise have to pay for with Claude, ChatGPT, etc.
- ethmarks 1y agoI've only tried Grok Code Fast 1, so I can't speak for any of the other models. In my experience, Grok is very fast and very cheap, but only moderately intelligent. It isn't stupid, but it rarely does anything that impresses me. The reason it's a useful model is that it is very, very fast (~90 tokens per second) and is very competitively priced.
- conception 1y agoYou should try cerebras with qwen. 2000 tokens/sec. It’s like chatting with the future usually- just an instant response.
- porphyra 1y agoGrok 4 is extremely capable, but for everyday chatting, Grok kinda sucks since it keeps repeating what you told it, and saying the current timestamp for some reason. ChatGPT is much better with its post training and prompt I feel like.
- 1y ago
- Bjorkbat 1y ago> Practically speaking, we’ve observed it maintaining focus for more than 30 hours on complex, multi-step tasks. Really curious about this since people keep bringing it up on Twitter. They mention it pretty much off-handedly in their press release and doesn't show up at all in their system card. It's only through an article on The Verge that we get more context. Apparently they told it to build a Slack clone and left it unattended for 30 hours, and it built a Slack clone using 11,000 lines of code (https://www.theverge.com/ai-artificial-intelligence/787524/anthropic-releases-claude-sonnet-4-5-in-latest-bid-for-ai-agents-and-coding-supremacy https://www.theverge.com/ai-artificial-intelligence/787524/a...) I have very low expectations around what would happen if you took an LLM and let it run unattended for 30 hours on a task, so I have a lot of questions as to the quality of the output
- sigmoid10 1y agoThis is obviously much more than just taking an LLM an letting it run for 30 hours. You have to build a whole environment together with external tool integration and context management and then tune the prompts and perhaps even set up a multi-agent system. I believe that if someone puts a ton of work into this you can have an LLM run for that long and still produce sellable outputs, but let's not pretend like this is something that average devs can do by buying some API tokens and kicking off a frontier model.
- Philpax 1y agoWell, yes, that's Claude Code. And OpenAI Codex. And Google Gemini CLI. Your average dev can just use those.
- ewoodrich 1y agoBut then that goes back to the original question, considering my own experiences observing the amount of damage CC or Codex can do in a working code base with a couple tiny initial mistakes or confusion about intent while being left unattended for ten minutes, let alone 30 hours....
- 1y ago
- asdev 1y agohow do claude/openai get around rate limiting/captcha with their computer use functionality?
- chrisford 1y agoThe vision model has consistently been degraded since 3.5, specifically around OCR, so I hope it has improved with Claude Sonnet 4.5!
- nickphx 1y agoIt will be great when the VC cash runs out, the screws tighten, and finally an end to the incessant misleading marketing claims.
- simonw 1y agoI had access to a preview over the weekend, I published some notes here: https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ https://simonwillison.net/2025/Sep/29/claude-sonnet-4-5/ It's very good - I think probably a tiny bit better than GPT-5-Codex, based on vibes more than a comprehensive comparison (there are plenty of benchmarks out there that attempt to be more methodical than vibes). It particularly shines when you try it on https://claude.ai/ https://claude.ai/ using its brand new Python/Node.js code interpreter mode. Try this prompt and see what happens: Checkout https://github.com/simonw/llm and run the tests with pip install -e '.[test]' pytest I then had it iterate on a pretty complex database refactoring task, described in my post.
- kurtis_reed 1y agoWhy did you have access to a preview?
- simonw 1y agoI get access to previews from OpenAI, Anthropic and Gemini pretty often. They're usually accompanied by an NDA and an embargo date - in this case the embargo was 10am Pacific this morning. I won't accept preview access if it comes with any conditions at all about what I can say about the model once the embargo has lifted.
- poopiokaka 1y ago[flagged]
- 0x696C6961 1y agoWhy do you even care?
- dboreham 1y agoNDAs often prohibit publishing...the NDA.
- jonathanstrange 1y agoI would like to see completely independent test results of these companies' products. I'm skeptical because every AI company claims their new product is the best.
- AtNightWeCode 1y agoSonnet is just so expensive comparing to other competitors. Have they fixed this?
- ripped_britches 1y agoPricing is the same as sonnet 4
- coconut08 1y agowhich was expensive compared to it's competitors
- AtNightWeCode 1y agoexactly, cost per token is higher but it also uses tokens like a chipmunk on steroids
- pembrook 1y agoIf they stopped the automatic "You're absolutely right!" responses after the model fails to fix something 20 times in a row, then that alone will be worth the upgrade. Me: "You just burned my house down" Claude: "You're absolutely right! I burned your house down, I need to revert the previous change and..." Me: "Now you rebuilt my house with a toilet in the living room" Claude: "You're absolutely right! I put a toilet in your living room..." Etc.
- croemer 1y agoIt's not yet on LMarena: https://lmarena.ai/leaderboard/text https://lmarena.ai/leaderboard/text
- nickstinemates 1y agoI gave it a quick spin with System Initiative[1]. The combination solved a 503 error in our infrastructure in 15 minutes that took over 2 hours to debug manually. It's pretty good! I wrote about a few other use cases on my blog[2] 1: https://systeminit.com https://systeminit.com 2: https://keeb.dev/2025/09/29/claude-sonnet-4.5-system-initiative/ https://keeb.dev/2025/09/29/claude-sonnet-4.5-system-initiat...
- iagooar 1y agoAnecdotal evidence. I have a fairly large web application with ~200k LoC. Gave the same prompt to Sonnet 4.5 (Claude Code) and GPT-5-Codex (Codex CLI). "implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s delay" Sonnet 4.5 went really fast at ~3min. But what it built was broken and superficial. The code did not even manage to reuse already existing auth and started re-building auth server-side instead of looking how other API endpoints do it. Even re-prompting and telling it how it went wrong did not help much. No tests were written (despite the project rules requiring it). GPT-5-Codex needed MUCH longer ~20min. Changes made were much more profound, but it implemented proper error handling, lots of edge cases and wrote tests without me prompting it to do so (project rules already require it). API calls ran smoothly. The entire feature worked perfectly. My conclusion is clear: GPT-5-Codex is the clear winner, not even close. I will take the 20mins every single time, knowing the work that has been done feels like work done by a senior dev. The 3mins surprised me a lot and I was hoping to see great results in such a short period of time. But of course, a quick & dirty, buggy implementation with no tests is not what I wanted.
- teekert 1y ago[flagged]
- iagooar 1y agoI even added a disclaimer "anecdotal evidence". Believe me, I am not the biggest fan of Sam. I just happen to like the best tools available, have used most of the large models and always choose the one that works best - for me.
- Implicated 1y agoI'm not trying to be offensive here, feel the need to indicate that. But that prompt leads me to believe that you're going to get rather 'random' results due to leaving SO much room for interpretation. Also, in my experience, punctuation is important - particularly for pacing and grouping of logical 'parts' of a task and your prompt reads like a run on sentence. Making a lot of assumptions here - but I bet if I were in your shoes and looking to write a prompt to start a task of a similar type that my prompt would have been 5 to 20x the length of yours (depending on complexity and importance) with far more detail, including overlapping of descriptions of various tasks (ie; potentially describing the same thing more than once in different ways in context/relation to other things to establish relation/hierarchy). I'm glad you got what you needed - but these types of prompts and approaches are why I believe so many people think these models aren't useful. You get out of them what you put into them. If you give them structured and well written requirements as well as a codebase that utilizes patterns you're going to get back something relative to that. No different than a developer - if you gave a junior coder, or some team of developers the following as a feature requirement: `implement a fuzzy search for conversations and reports either when selecting "Go to Conversation" or "Go to Report" and typing the title or when the user types in the title in the main input field, and none of the standard elements match, a search starts with a 2s delay` then you can't really be mad when you don't get back exactly what you wanted. edit: To put it another way - spend a few more minutes on the initial task/prompt/description of your needs and you're likely to get back more of what you're expecting.
- andrewstuart 1y agoStill waiting to be able to upload zip files to Claude, which Gemini and ChatGPT have had for ages. ChatGPT even does zip file downloads, packaging up all your files.
- rishabhaiover 1y agohn displays a religious hatred towards ai progress
- epolanski 1y agoMost people here use these models as you can see from the comments. But we can also see that we're one of the few sane skeptical places in a world that is making the most diverse claims about AI.
- rishabhaiover 1y agoFair.
- deleted 1y ago[deleted]
- dr_dshiv 1y agoAnyone try the Imagine with Claude yet? How does it work?
- lexarflash8g 1y agoJust tested this on a rather simple issue. Basically it falls into rabbits holes just like the other models and tries to brute force fixes through overengineering through trial and error. It also says "your job should now pass" maybe after 10 prompts of roughly doing the same thing stuck in a thought loop. A GH actions pipeline was failing due to a CI job not having any source code files -- error was "No build system detected". Using Cursor agent with Sonnet 4.5, it would try to put dummy .JSON files and set parameters in the workflow YAML file to false, and even set parameters that don't exist. Simple solution was to just override the logic in the step to "Hello world" to get the job to pass. I don't understand why the models are so bad with simple thinking outside the box solutions? Its like a 170 iq savant who can't even ride public transporation.
- mirsadm 1y agoThey're very good at things have been done a million times before. I use both Claude and Gemini and they are pretty terrible at writing any kind of Vulkan shader but really good for spitting out web pages and small bits of code here and there. For me that's enough to make them useful.
- baq 1y ago> why the models are so bad with simple thinking outside the box solutions There is no outside the box in latent space. You want something a plain LLM can’t do by design - but it isn’t out of question that it can step outside of its universe by random chance during the inference process and thanks to in-context learning.
- i-chuks 1y agoAI companies really need to consider regional pricing. Huuuuge barrier!
- Jcampuzano2 1y agoThe price of training and running the models doesn't really change much no matter which region you're hosting/making requests from. Regional pricing unfortunately doesn't really make much sense for them unless they're willing to take even larger losses, even if it is a barrier to lower income countries/regions.
- pants2 1y agoUnfortunately also disappointed with it in Cursor vs GPT-5-Codex. I asked it to add a test for a specific edge case, it hallucinated some parameters and didn't use existing harnesses. GPT-5-Codex with the same prompt got everything right.
- baobabKoodaa 1y agoHere's an anecdata. I have a real-world use case financial dataset where I have created benchmarks. Sonnet 4.5 provides no measurable improvement on these benchmarks over Sonnet 4. This is a bit surprising to me, especially when considering that the benchmark results published by Anthropic indicate that Sonnet 4.5 should be better than Sonnet 4 specifically on financial data analysis.
- siva7 1y agoDoes 4.5 still answer everything with "You're absolutely right!" or is it now able to communicate like a real programmer?
- simonw 1y agoIt still says "Perfect!" about its own work far too often.
- kenjackson 1y agoIn fairness that sounds like me when I code. It's either "Perfect!" or "Genius!". Or conversely "I'm a complete idiot!"
- onraglanroad 1y agoFor me, all three tend to follow in rapid succession.
- neutronicus 1y agoI'm more of a "Kneel before Zod" kind of guy Wonder if I could Claude to do that
- anshumankmr 1y agohttps://www.youtube.com/watch?v=fXW02XmBGQw https://www.youtube.com/watch?v=fXW02XmBGQw You have to be lying if you haven't felt about your own work like that guy from the Bond movie.
- inopinatus 1y agoI won’t be satisfied until I get a Linus Torvalds mode. “Your idea is shit because you are so fucking stupid” “Please stop talking, it hurts my GPUs thinking down to your level” “I may seem evil but at least I’m not incompetent”
- atonse 1y agoWhy is this getting downvoted? It was hilarious! I actually added a fun thing to my user-wide CLAUDE.md, basically saying that it should come up with a funny insult every time I come up with an idea that wasn't technically sound (I got the prompt from someone else). It seems to be disobeying me, because I refuse to believe that I don't have bad ideas. Or some other prompt is overriding it.
- Aflynn50 1y agoWhen I see how much the latest models are capable of it makes me feel depressed. As well as potentially ruining my career in the next few years, its turning all the minutiae and specifics of writing clean code, that I've worked hard to learn over the past years, into irrelivent details. All the specifics I thought were so important are just implementation details of the prompt. Maybe I've got a fairly backwards view of it, but I don't like the feeling that all that time and learning has gone to waste, and that my skillset of automating things is becoming itself more and more automated.
- elAhmo 1y agoDon't be so grim! This will just give you access to not worry about writing clean code as much as you did in the past - you can focus on other parts of the development lifecycle. The skill of writing good quality code is still going to be beneficial, maybe less emphasized on writing side, but critical of shipping good code, even when someone (something) else wrote it.
- FridgeSeal 1y ago“Do t worry about the fit and finish in your craftsmanship anymore, just bolt everything together and move on to other woodworking” Is how that argument comes across.
- jaggederest 1y agoAnd contrariwise, the argument against tools like these sounds like: "I never use power tools or CNC, I only use hand tools. Even if they would save me an incredible amount of time and let me work on other things, I prefer to do it the slow and painstaking way, even if the results are ultimately almost identical." Sure, you can absolutely true up stock using a jointer plane, but using a power jointer and planer will take about 1/10th of the time and you can always go back with a smoothing plane to get that mirror finish if you don't like the machine finish. Likewise, if your standards are high and your output indistinguishable, but the AI does most of the heavy lifting for the rough draft pass, where's the harm? I don't understand everyone who says "the AI only makes slop" - if you're responsible for your commits and you do a good job, it's indistinguishable.
- iFire 1y agoIs it 15x cheaper like Grok?
- typpilol 1y agoI heard 5x cheaper then opus
- Attummm 1y agoAnthropic really nailed this release. There had been a trend where each new model released from OpenAI, Anthropic, etc. felt like a letdown or worse a downgrade. But the release of 4.5 break that trend, And is a pleasant surprise on day one. Well done! :)
- rtp4me 1y agoJust updated to Sonnet 4.5 and Claude Code 2.0 this afternoon. I worked on a quick project (creating PXE bootable files) using the updates and have to say, this new version seems much faster and more accurate than before. I did not go round-and-round trying to get good output and Claude did not go down rabbit holes like before. So far, so good.
- mchusma 1y agoFor me, Opus 4.1 was so much better than Sonnet 4.0 that I used it exclusively in Claude Code and cancelled Cursor. I'm a bit skeptical that Sonnet 4.5 will be in practice better, but will test with it and see! Hopefully we get Opus 4.5 soon.
- throwaway638637 1y agoIsn't Opus much slower than Sonnet? I haven't been using Opus for that reason
- ojosilva 1y agoTo @simonw and all the coding agent and LLM benchmarkers out there: please, always publish the elapsed time for the task to complete successfully! I know this was just a "it works straight in claude.ai" post, but still, nowhere in the transcript there's a timestamp of any kind. Durations seem to be COMPLETELY missing from the LLM coding leaderboards everywhere [1] [2] [3] There's a huge difference in time-to-completion from model to model, platform to platform, and if, like me, you are into trial-and-error, rebooting the session over and over to get the prompt right or "one-shot", it's important how reasoning efforts, provider's tokens/s, coding agent tooling efficiency, costs and overall model intelligence play together to get the task done. Same thing applies to the coding agent, when applicable. Grok Code Fast and Cerebras Code (qwen) are 2 examples of how models can be very competitive without being the top-notch intelligence. Running inference at 10x speed really allows for a leaner experience in AI-assisted coding and more task completion per day than a sluggish, but more correct AI. Darn, I feel like a corporate butt-head right now. 1. https://www.swebench.com/ https://www.swebench.com/ 2. https://www.tbench.ai/leaderboard https://www.tbench.ai/leaderboard 3. https://gosuevals.com/agents.html https://gosuevals.com/agents.html
- fabmilo 1y agoYeah I totally agree, we need time to completion of each step and the number of steps, sizes of prompts, number of tools, ... and better visualization of each run and break down based on the difficulty of the task
- simonw 1y agoThat's a good call, I'll try to remember that for next time.
- Imustaskforhelp 1y agoI just wanted to say that I really liked your this comment which just showed professionalism and just learning from your mistakes/improving yourself. I definitely consider you to be an AI influencer, especially in hackernews communities and so I wanted to say that I see influencers who will double down,triple down on things when in reality, people just wanted to help them in the first place. I just wanted to say thanks with all of this in mind, also that your generate me a pelican riding a bicycle has been a fun ride and is always going to be interesting, so thanks for that as well. I just wanted to share my gratitude with ya.
- AbuAssar 1y agoI used to treat writing code as a form of art, with attention to details and best practices, and using design patterns whenever possible. but it seems this will come to an end eventually as these agents become more stronger and capable each day, and will be better and faster than human coders.
- nvarsj 1y agoYup, we're headed to the robot assembly line, with a few experts making sure it all works correctly. Craftsmen will remain, but it will be niche (and probably not pay anything unless you are a true master).
- labrador 1y agoI'm sympathetic, but it occcured to me that ccording to my amatuer studies, Germany lost WW2 in part because it had a craftsman mentality to manufacture war machines and ended up with a bazillion different part requirments and a shortage of skilled craftsmen, while America used Henry Ford's assembly line process to stamp out hundreds of thousands of identical machines sharing the same parts. Now we are at the assembly line stage of software production with AI. Us craftsmen will have to find other ways to enjoy our crafts.
- mirsadm 1y agoThis has been said about the release of every model in the last couple of years. Personally I can't even tell if this one is better than 3.5. I actually found that one more useful.
- devinprater 1y agoI hope that one day Anthropic work on making Claude more accessible to screen reader users. ChatGPT is currently the only AI that I know of that, when it's thinking, sends that status to the screen reader, and then sends the response to the screen reader to be spoken as well, like any other good chat app does.
- ionwake 1y agoDo we have a pelican for it yet ?
- 0xbadcafebee 1y agoClaude doesn't know how to calculate realistic minimum voltages for solar arrays w/MPPT chargers. ChatGPT does. Prompt: "Can I use two strings of four Phono Solar PS440M8GFH solar panels with a EG4 12kPV Hybrid Inverter? I want to make sure that there will not be an issue any time of year. New York upstate." Claude 4.5: Returns within a few seconds. Does not find the PV panel specs, so it asks me if I want it to search for them. I say yes. Then it finally comes up with: "YES, your configuration is SAFE [...] MPPT range check: Your operating voltage of 131.16V fits comfortably in the 120-500V MPPT operating range". ChatGPT 5: Returns after 78 seconds. Says: "Hot-weather Vmpp check: Vmpp_string @ STC = 4 × 32.79 = 131 V (inside 120–500 V). Using the panel’s NOCT point (31.17 V each), a typical summer operating point is ~125 V — still OK. But at very hot cell temps (≈70 °C is possible), Vmpp can drop roughly ~13% from STC → ~114 V, which is below the EG4’s 120 V MPPT lower limit. That can cause the tracker to fall out of its optimal range and reduce harvest during peak heat." ChatGPT used deeper thinking to determine that the lowest possible voltage in the heat would be below the MPPT's minimum operating voltage. It doesn't indicate that in reality it might not charge at all at that point... but it does point out the risk, whereas Claude says everything is fine. I need about 5 back-and-forths with Claude to get it to finally realize its mistake.
- ranguna 1y agoThis HN post is about claude 4.5 and you come here speaking about how "claude" does not give you satisfactory answer when, most likely, you didn't even try claude 4.5 in the first place. Claude 4.5 after a few web searches and running a couple python scripts for analysis: Yes, your configuration should work! Based on my analysis, two strings of four Phono Solar PS440M8GFH panels will be compatible with the EG4 12kPV inverter for upstate New York conditions. Key Findings: Voltage Safety: Cold weather maximum (-25°C/-13°F): 182V - well below the 600V limit (only 30% of maximum) Standard operating voltage: 128V - comfortably within the 120-500V MPPT range Hot weather minimum (40°C/104°F panel temp): 121V - just above the 120V MPPT minimum Current: Operating current: ~13.8A per string - well within the 25A MPPT limit (55% of capacity) Total System: 8 panels × 440W = 3,520W (3.5kW) - well below the 12kW inverter rating Important Considerations: Hot weather margin is tight: At extreme hot temperatures, the voltage drops to about 121V, which is only 1V above the MPPT minimum. This means: The system will work, but efficiency might be slightly reduced on the hottest days The MPPT controller should still track power effectively More robust alternative: If you want more safety margin, consider 5 panels per string instead: Cold: 228V (still safe) Hot: 151V (much better margin above 120V minimum) Total: 10 panels = 4.4kW Wire each string to a separate MPPT on the EG4 12kPV (it has 2 MPPTs), which is perfect for your 2-string configuration. Bottom Line: Your planned configuration of 2 strings × 4 panels will work year-round in upstate New York without safety issues. The system is conservatively sized and should perform well!
- tresil 1y agoI'll add another really positive review here. Sonnet 4.0 had been really struggling to implement an otel monitoring solution using grafana's lgtm stack. Sonnet 4.0 had 4 or 5 different attempts - some of them longer than 10 min - troubleshooting why metrics were supposedly being emitted from the api, but not showing up in Prometheus. Sonnet 4.5 correctly diagnosed and fixed the real issue within about 5 min. Not sure if that's the model being smarter, but I definitely saw the agent using some new approaches and seemingly managing it's context better.
- deleted 1y ago[deleted]
- mccoyb 1y agoCongratulations: it's faster, but worse, with a larger context window.
- j45 1y agoA question I have for anyone is -- has Claude Max returned to or repaired the response quality and service issues between the usage limits and performance of the model for coding and non-coding tasks? Anecdata is welcome as it seems like it's the only thing available sometimes.
- manofmanysmiles 1y agoI haven't shouted into the void for a while. Today is as good a day as any other to do so. I feel extremely disempowered that these coding sessions are effectively black box, and non-reproducible. It feels like I am coding with nothing but hopes and dreams, and the connection between my will and the patterns of energy is so tenuous I almost don't feel like touching a computer again. A lack of determinism comes from many places, but primarily: 1) The models change 2) The models are not deterministic 3) The history of tool use and chat input is not availabler as a first class artifact for use. I would love to see a tool that logs the full history of all agents that sculpt a codebase, including the inputs to tools, tool versions and any other sources of enetropy. Logging the seed into the RNGs that trigger LLM output would be the final piece that would give me confidence to consider using these tools seriously. I write this now after what I am calling "AI disillusionment", a feel where I feel so disconnected from my codebase I'd rather just delete it than continue. Having a set of breadcrumbs would give me at least a modicum of confidence that the work was reproducible and no the product of some modern ghost, completely detached from my will. Of course this would require actually owning the full LLM.
- johnfn 1y agoIf you care about this so much why don't you use one of the open source OpenAI models? They're pretty good and give you the guarantees you want.
- int_19h 1y agoNone of the open weight models are really as good as SOTA stuff, whatever their evals says. Depending on the task at hand this might not actually manifest if the task is simple enough, but once you hit the threshold it's really obvious.
- alex77456 1y agoI share the sentiment. I would add that people I would like to see use LLMs for coding (and other technical purposes) tend to be jaded like you, and people I personally wouldn't want to see use LLMs for that, tend to be pretty enthusiastic
- 1y ago
- lihaciudanieljr 1y ago[dead]
- niyazpk 1y agoDoes anyone know whatever happened to the Haiku family of models? They've not been updated since 3.5! Did Anthropic give up on them?
- Galaco 1y agoIf you pause your subscription, Claude.ai breaks. I paused my subscription, and my account immediately transitioned to free. It has removed my invoice history, and attempts to upgrade again fail with an internal error. Their chatbot is telling me to navigate to UI elements that don't exist, and free users do not have the option of human support. So I'm stuck; my sub is paused, and I cannot either cancel, or unpause and cannot speak to a human to solve this because the pause process took away all possibility of human interaction. This is the future we live in.
- labrador 1y agoDid you use Google Play to pause subscription? Because Claude Pro says there is no pause subscription except on Google Play and then goes on to explain your problem if that's the case.
- Galaco 1y agoThanks for the info, but I did pause this, not through Google Play, it was via the UI. I received an automated email from them that my subscription had been paused as I expected and will resume in 1 month unless I cancel (I can’t cancel because the cancel UI doesn’t exist in whatever status my account is somehow in). It’s funny that Claude Pro says this isn’t a feature, because their chatbot gave me instructions on how to unpause via the UI (although said UI does not exist) so the bot seems to know it’s a feature.
- deleted 1y ago[deleted]
- zulban 1y ago> This is the future we live in. It's just a bug. Chill. Wait a business day and try again. You write as if you've never experienced a bug before.
- simondotau 1y agoIf you’re being sarcastic, you might want to edit your post to make that clearer.
- jatins 1y agoI tested this on some day to day pattern matching kind of tasks and it didn't do well. Still the same over eagerness to make wild code changes instead of "reasoning" about the error
- hsn915 1y agoIt is time to acknowledge that AI coding does not actually work. ok, you think it's a promising field and you want to explore it, fine. Go for it. Just stop pretending that what these models are currently doing is good enough to replace programmers. I use LLMs a lot, even for explaining documentation. I used to use them for writing _some_ code, but I have never ever gotten a code sample over 10 lines that was not in need of heavy modifications to make it work correctly. Some people are pretending to write hundreds of lines of code with LLMs, even entire applications. All I have to say is "lol".
- gabriel-uribe 1y agoVery interesting observation. I haven’t written a function by hand in 18 months.
- sneilan1 1y agoSame. I haven't written any code by hand in some time. Oh well. I guess I'm just doing it wrong.
- BoorishBears 1y agoHave you built anything public that folks can try out? Not doubting but it helps to contextualize things
- sneilan1 1y agohttps://app.grantpuma.com/ https://app.grantpuma.com/ It's a startup for finding grants. We have california state, federal, non-profit and california city/county grants. My landing page absolutely sucks but if you sign up / upload some papers or make some search cards you'll like the experience. I'm very excited to try out the new Qwen XL that came out recently for visual design. I could really use some better communication to users of the capabilities of the platform.
- 1y ago
- system2 1y agoI didn't try the checkpoints, I use local git + /resume from a chat that I pick closer to the git version I restore if Claude screws up. Will this checkpoint help with chat memory and disregard the latest chat's info? I use WSL under Windows, VSCode with the WSL plugin, and Claude-Code installed on Ubuntu 24. It is generally solid and has no issue with this setup.
- cmrdporcupine 1y agoSo far I'm liking that it seems to follow my CLAUDE.md instructions better, doing more frequent checkins with me to ask me to review what it's done, etc, and taking my advice more. What I'm not liking is it seems even... lazier... than previously. By which I mean the classic "This is getting complicated so..." (followed by cop-out, dropping the original task and motivation). There's also a bug where compaction becomes impossible. ("conversation too long" and its advice on how to fix doesn't work)
- bdangubic 1y ago> There's also a bug where compaction becomes impossible. ("conversation too long" and its advice on how to fix doesn't work) I have seen this issue with every model so far
- ChaoPrayaWave 1y agoWhat impressed me most about Claude Sonnet 4.5 is that its output structure is more stable than many other models and less prone to crashes. I ran some real world scripts from my own projects, and it exhibited fewer hallucinations than GPT-4 and performed more faithfully on code interpretation tasks. However, it can be a bit slow to warm up, and sometimes I needed more prompts in the first few rounds.
- mattlangston 1y agoIt does well with screenshot-calculus for me. For example, I pasted a screenshot of the Layer Norm equation into Claude Code 2 and asked: "Differentiate y(x) w.r.t x, gamma and beta." It not only produced the correct result, but it understood the context - I didn't tell it the context was layer norm, back-propagation and matrices. This release is a step function for my use cases. My screenshot came from here: https://docs.pytorch.org/docs/stable/generated/torch.nn.LayerNorm.html https://docs.pytorch.org/docs/stable/generated/torch.nn.Laye...
- StarterPro 1y agoOnce the bottom falls out of ai, will programming be seen as a marketable skill again?
- hu3 1y agoI think so. Because systems require a lot of knowledge to create and maintain without breaking. How many years till AI can be trusted to deploy changes to production without supervision? Maybe never.
- techpression 1y agoIt took me one question to have it spit out a completely dreamt up codebase, complete with emojis, promises of solutions and fixing all my problems, and of course nothing of it worked. It was a very simple question about something very well documented (Oban timeouts). I doubt LLM benchmarks more and more, what are they even testing?
- nakamoto_damacy 1y ago> what are they even testing? How well the LLM does on the benchmarks. Obviously. :P
- techpression 1y agoIs there some kind of conversion ratio to actual value? ;)
- ileonichwiesz 1y agoSure there is. It’s called “higher numbers = more investor money”. Any improvement in actual utility is purely coincidental.
- doix 1y ago> It was a very simple question about something very well documented (Oban timeouts). It's some 3rd party thing for Elixir, a niche within a niche. I wouldn't expect an LLM to do well there. > I doubt LLM benchmarks more and more, what are they even testing? Probably testing by asking it to solve a problem with python or (java|type)script. Perhaps not even specifying a language and watching it generate a generic React application.
- techpression 1y agoSomething that is well documented should still perform well, there’s few places to go wrong, compared with something like React where the training data seems to be a cesspool of the worst code imaginable, at least that’s my experience using it for React.
- deviation 1y agoInteresting. In a thought process while editing a PDF, Claude disclosed the folder hierarchy for it's "skills". I didn't know this was available to us: > Reading the PDF skill documentation to create the resume PDF > Here are the files and directories up to 2 levels deep in /mnt/skills/public/pdf, excluding hidden items and node_modules:
- vbtechguy 1y agoClaude Sonnet 4.5 definitely the best model I've tried to date - my evaluation rankings against 23 AI models at https://github.com/centminmod/claude-sonnet-4.5-evaluation https://github.com/centminmod/claude-sonnet-4.5-evaluation :)
- drbojingle 1y agoImo we're going to start needing more examples of where the successor is better than what came before, and not just benchmarks.
- oscord 1y agoSonnet 4 had turned to shit recently (about 2.5 months according to my observations). It hallucinated on 3 questions in a row while looking at a simple bash script. Was enough for me to cancel. Claude biz is killing Claude dev. It was good while they were not so stingy on GPU.
- miletus 1y agowe at agentsea.com have been playing with it for a while. here's what we think about it: - still sucks at generating pretty ui - great for creative writing and long-form planning - it’s really fast but not smarter than gpt-5 - pairs well with external tools/agents for research and automation - comes with a 1m token context window, so you can feed it monstrous codebases or giant docs - still hallucinates or stumbles on complex requests
- jdlyga 1y agoCompared to Claude Sonnet 4, anecdotal evidence. But I'm noticing very little difference.
- cwoolfe 1y agoI've been really impressed with how good Cursor is at coding. I threw it a standard backend api endpoint and database task yesterday and it generated 4 hours of code in 2 minutes. It was set to Auto which I think uses some Claude model.
- mutant 1y agoDidn't they promise a 1m token input ? I don't see that here.
- MarcelOlsz 1y agoTerrible. It can't even do basic scaffolding which is all it was good for, now it can't even do that. You can wrangle it with taskmaster or bmadcode or whatever but at that point I'd rather just write it myself. Writing English to build things is goofy. Unsubscribed.
- n8m8 1y agoSo far the only thing I’ve noticed is that it made me confirm that it should do a 10 minute task “manually” because it “would take 2 or 3 hours” It was a context merging task for my unorganized collection of agents… it sort of made sense, but was the exact reason I was asking it to do it… like you’re the bot, lol
- n8m8 1y agoExcited to try Claude agents sdk though
- virtualritz 1y agoSo I was using Opus exclusively (Max plan) to write Rust since June. CC switched to Sonnet 4.5 by default yesterday, I'm just very unimpressed. It seems like a considerable regression. Probably this is related to me using it to write Rust and not Python or JS/TS? Example: I asked it to refactor a for loop to functional code with rayon, compiler barfs about mutation (it used par_iter()). It rolls back the code to what it was before. Then this happens: Me: There is par_iter_mut(). Sonnet: Ah yes, but we can't use par_iter_mut() on self.vertices because we're calling self.set_vertex_position() which needs mutable access to the entire mesh (not just the vertices map). However, self is &mut. Wtf? This would have never happened with Opus, i.e. Opus would have used par_iter_mut() to start with (or applied the right fix w/o hand-holding after the compile failed with par_iter()). I had just a bunch of those wtfs since yesterday from more or less zero before. I.e. it doesn't feel like coincidence.
- Kim_Bruning 1y agoI think it's to do with system prompting too, but Sonnet 4.5 actually pushes back at times. And it tries to keep me on topic. It's refreshing!
- ath3nd 1y ago[dead]