5 ms·
It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in
by dpe82 8mo ago
It's wild that Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks. It will be interesting to see if that's the case in real, practical, everyday use. The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s.
- iLoveOncall 8mo agoGiven that users prefered it to Sonnet 4.5 "only" in 70% of the cases (according to their blog post) makes me highly doubt that this is representative of real-life usage. Benchmarks are just completely meaningless.
- jwolfe 8mo agoFor cases where 4.5 already met the bar, I would expect 50% preference each way. This makes it kind of hard to make any sense of that number, without a bunch more details.
- gnatolf 8mo agoGood point. So much functionality gets commoditized, we have to move goalposts more or less constantly.
- dpe82 8mo agosimonw hasn't shown up yet, so here's my "Generate an SVG of a pelican riding a bicycle" https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5cb2d5b8f732 https://claude.ai/public/artifacts/67c13d9a-3d63-4598-88d0-5...
- coffeebeqn 8mo agoWe finally have AI safety solved! Look at that helmet
- 1f60c 8mo ago"Look ma, no wings!" :D
- AstroBen 8mo agoif they want to prove the model's performance the bike clearly needs aero bars
- thinkling 8mo agoFor comparisonI think the current leader in pelican drawing is Gemini 3 Deep Think: https://bsky.app/profile/simonwillison.net/post/3meolxx5s7227 https://bsky.app/profile/simonwillison.net/post/3meolxx5s722...
- konart 8mo agoMy take (also Gemini 3 Deep Think): https://gemini.google.com/share/12e672dd39b7 https://gemini.google.com/share/12e672dd39b7 Somehow it's much better now.
- jazzyjackson 8mo agoI’m not familiar with Gemini, isn’t this just a diffusion model output? The Pelican test is for the llm to produce SVG markup.
- konart 8mo agoYeah, I was so amazed by the result I didn't even realize Gemini used Nano Banana while producing the result.
- kingbob000 8mo agoIs that actually better? That pelican has arms sprouting out of its wings
- badc0ffee 8mo agoThe point of the penny-farthing is that you drive the front wheel directly with the pedals, but this seems to have the pedals in a spot where they would drive a chain, although there is no chain?
- dyauspitr 8mo agoCan’t beat Gemini’s which was basically perfect.
- estomagordo 8mo agoWhy is it wild that a LLM is as capable as a previously released LLM?
- simianwords 8mo agoIt means price has decreased by 3 times in a few months.
- Retr0id 8mo agoBecause Opus 4.5 inference is/was more expensive.
- crummy 8mo agoOpus is supposed to be the expensive-but-quality one, while Sonnet is the cheaper one. So if you don't want to pay the significant premium for Opus, it seems like you can just wait a few weeks till Sonnet catches up
- ceroxylon 8mo agoStrangely enough, my first test with Sonnet 4.6 via the API for a relatively simple request was more expensive ($0.11) than my average request to Opus 4.6 (~$0.07), because it used way more tokens than what I would consider necessary for the prompt.
- svachalek 8mo agoThis is an interesting trend with recent models. The smarter ones get away with a lot less thinking tokens, partially to fully negating the speed/price advantage of the smaller models.
- smartbit 8mo agoJust like humans :-) Eg a smart person will automate a task instead of executing the task repeatedly.
- 8mo ago
- simlevesque 8mo agoThe system card even says that Sonnet 4.6 is better than Opus 4.6 in some cases: Office tasks and financial analysis.
- justinhj 8mo agoWe see the same with Google's Flash models. It's easier to make a small capable model when you have a large model to start from.
- karmasimida 8mo agoFlash models are nowhere near Pro models in daily use. Much higher hallucinations, and easy to get into a death sprawl of failed tool uses and never come out You should always take those claim that smaller models are as capable as larger models with a grain of salt.
- justinhj 8mo agoFlash model n is generally a slightly better Pro model (n-1), in other words you get to use the previously premium model as a cheaper/faster version. That has value.
- karmasimida 8mo agoThey do have value, because they are much much cheaper. But no, 3.0 flash is not as good as 2.5 pro, I use both of them extensively, especially in translation. 3.0 flash will confidently mistranslate some certain things, while 2.5 pro will not.
- justinhj 8mo agoTotally fair. Translation is one of those specific domains where model size correlates directly with quality, and no amount of architectural efficiency can fully replace parameter count.
- madihaa 8mo agoThe most exciting part isn't necessarily the ceiling raising though that's happening, but the floor rising while costs plummet. Getting Opus-level reasoning at Sonnet prices/latency is what actually unlocks agentic workflows. We are effectively getting the same intelligence unit for half the compute every 6-9 months.
- mooreds 8mo ago> We are effectively getting the same intelligence unit for half the compute every 6-9 months. Something something ... Altman's law? Amodei's law? Needs a name.
- merlindru 8mo agoHow about More's law - because we keep getting "more" compute at a lower cost?
- turnsout 8mo agoThis is what excited me about Sonnet 4.6. I've been running Opus 4.6, and switched over to Sonnet 4.6 today to see if I could notice a difference. So far, I can't detect much if any difference, but it doesn't hit my usage quota as hard.
- nimonian 8mo agoMoore's law lives on!
- scottmf 8mo ago2024: Intelligence too cheap to meter 2026: Everyone is spending $500/month on LLM subscriptions
- qingcharles 8mo agoMy Dad used to make the same joke in the 1980s about how they'd told him in the 1950s that nuclear power would be "too cheap to meter" which I assume is probably where the trope originated.
- amelius 8mo ago> The speed at which this stuff is improving is really remarkable; it feels like the breakneck pace of compute performance improvements of the 1990s. Yeah, but RAM prices are also back to 1990s levels.
- mrcwinn 8mo agoRelief for you is available: https://computeradsfromthepast.substack.com/p/connectix-ram-doubler https://computeradsfromthepast.substack.com/p/connectix-ram-...
- deleted 8mo ago[deleted]
- isoprophlex 8mo agoYou wouldn't download a RAM
- MarsIronPI 8mo agohttps://downloadmoreram.com https://downloadmoreram.com Yes I would.
- Rapzid 8mo agoWe don't rent RAMs!
- mikkupikku 8mo agoI knew I've been keeping all my old ram sticks for a reason!
- ge96 8mo agoI sent Opus a photo of NYC at night satellite view and it was describing "blue skies and cliffs/shore line"... mistral did it better, specific use case but yeah. OpenAI was just like "you can't submit a photo by URL". Was going to try Gemini but kept bringing up vertexai. This is with Langchain
- danielbln 8mo agoI just sent Opus a NYC night satellite view and it described it just as expected. Seems like you have a tooling problem, not a model problem.
- ge96 8mo agoWould be curious your setup this was mine. satellite_imagery_analysis_agent = create_agent( model="claude-opus-4-6", system_prompt="your task is to analyze satellite images" ) response = satellite_imagery_analysis_agent.invoke({ "messages": [ { "role": "user", "content": "What do you see in this satellite image? https://images.unsplash.com/photo-1446776899648-aa78eefe8ed0?q=80&w=1744&auto=format&fit=crop&ixlib=rb-4.1.0&ixid=M3wxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8fA%3D%3D https://images.unsplash.com/photo-1446776899648-aa78eefe8ed0..." } ] }) With this output: # Satellite Image Analysis I can see this image shows an *aerial/satellite view of a coastline*. Here are the key features I can identify: ## Geographic Features - *Ocean/Sea*: A large body of deep blue water dominates a significant portion of the image - *Coastline*: A clearly defined boundary between land and water with what appears to be a rugged or natural shoreline - *Beach/Shore*: Light-colored sandy or rocky coastal areas visible along the water's edge ## Terrain - *Varied topography*: The land area shows a mix of greens and browns, suggesting: - Vegetated areas (green patches) - Arid or bare terrain (brown/tan areas) - *Possible cliffs or elevated terrain* along portions of the coast ## Atmospheric Conditions - *Cloud cover*: There appear to be some clouds or haze in parts of the image - Generally clear conditions allowing good visibility of surface features ## Notable Observations - The color contrast between the *turquoise/shallow nearshore waters* and the *deeper blue offshore waters* suggests varying ocean depths (bathymetry) - The coastline geometry suggests this could be a *peninsula, island, or prominent headland* - The landscape appears relatively *semi-arid* based on the vegetation patterns --- Note: Without precise geolocation metadata, I'm providing a general analysis based on visible features. The image appears to capture a scenic coastal region, possibly in a Mediterranean, subtropical, or tropical climate zone. Would you like me to focus on any specific aspect of this image?
- satvikpendem 8mo ago> Sonnet 4.6 is roughly as capable as Opus 4.5 - at least according to Anthropic's benchmarks Yeah it's really not. Sonnet still struggles while Opus, even 4.5 succeeds (and some examples show Opus 4.6 is actually even worse than 4.5, all while being more expensive and taking longer to finish).