6 ms·
[SWE-bench co-author here] It seems like they run this test on a subset of 50 tasks, and that they only run the test once per day. So a lot of the movement in a
by ofirpress 8mo ago
[SWE-bench co-author here]
It seems like they run this test on a subset of 50 tasks, and that they only run the test once per day. So a lot of the movement in accuracy could be attributed to that.
I would run on 300 tasks and I'd run the test suite 5 or 10 times per day and average that score. Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded.
- dana321 8mo ago"Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded" Aha, so the models do degrade under load.
- mohsen1 8mo agoHope you don't mind the unrelated question: How do you pay for those SWE-bench runs? I am trying to run a benchmark but it is too expensive to run enough runs to get a fair comparison. https://mafia-arena.com https://mafia-arena.com
- ofirpress 8mo agoBenchmarks can get costly to run- you can reach out to frontier model creators to try and get them to give you free credits, but usually they'll only agree to that once your benchmark is pretty popular.
- mohsen1 8mo agoyes I reached out to them but as you say it's a chicken-and-egg problem. Thanks!
- Dolores12 8mo agoso basically they know requests using your API key should be treated with care?
- Deklomalo 8mo ago[dead]
- swyx 8mo agothey could but you can also have some trust in anthropic to have some integrity there, these are earnest people. "trust but verify" ofc . https://latent.space/p/artificialanalysis https://latent.space/p/artificialanalysis do api keys but also mystery shopper checks
- mrandish 8mo ago> these are earnest people. I agree. I'll also add that when my startup got acquired into a very large, well-known valley giant with a sterling rep for integrity and I ended up as a senior executive - over time I got a first-hand education on the myriad ways genuinely well-intentioned people can still end up being the responsible party(s) presiding over a system doing net-wrong things. All with no individual ever meaning to or even consciously knowing. It's hard to explain and I probably wouldn't have believed myself before I saw and experienced it. Standing against an overwhelming organizational tide is stressful and never leads to popularity or promotion. I think I probably managed to move on before directly compromising myself but preventing that required constant vigilance and led to some inter-personal and 'official' friction. And, frankly, I'm not really sure. It's entirely possible I bear direct moral responsibility for a few things I believe no good person would do as an exec in a good company. That's the key take-away which took me a while to process and internalize. In a genuinely good organization with genuinely good people, it's not "good people get pressured by constraints and tempted by extreme incentives, then eventually slip". I still talk with friends who are senior execs there and sometimes they want to talk about whether something is net good or bad. I kind of dread the conversation going there because it's inevitably incredibly complex and confusing. Philosopher's trolley car ethics puzzles pale next to these multi-layered, messy conundrums. But who else are they going to vent to who might understand? To be clear, I still believe that company and its leadership to be one of the most moral, ethical and well-intentioned in the valley. I was fortunate to experience the best case scenario. Bottom line: if you believe earnest, good people being in charge is a reliable defense against the organization doing systemically net-wrong things - you don't comprehend the totality of the threat environment. And that's okay. Honestly, you're lucky. Because the reality is infinitely more ambiguously amoral than white hats vs black hats - at the end of the day the best the 'very good people' can manage is some shade of middle gray. The saddest part is that good people still care, so they want to check the shade of their hat but no one can see if it's light enough to at least tell yourself "I did good today."
- epolanski 8mo agoThe last thing a proper benchmark should do is reveal it's own API key.
- sejje 8mo agoThat's a good thought I hadn't had, actually.
- plagiarist 8mo agoIMO it should need a third party running the LLM anyway. Otherwise the evaluated company could notice they're receiving the same requests daily and discover benchmarking that way.
- jabedude 8mo agoBut that's removing a component that's critical for the test. We as users/benchmark consumers care that the service as provided by Anthropic/OpenAI/Google is consistent over time given the same model/prompt/context
- plagiarist 8mo agoMight as well have the free tokens, then, especially if it is an open benchmark they are already aware of. If they want to game it they cannot be stopped from doing so when it's on their infra.
- mrandish 8mo agoWith the insane valuations and actual revenue at stake, benchmarkers should assume they're assessing in an adversarial environment. Whether from intentional gaming, training to the test, or simply from prioritizing things likely to make results look better, targeting benchmarks will almost certainly happen. We already know large graphics card manufacturers tuned their drivers to recognize specific gaming benchmarks. Then when that was busted, they implemented detecting benchmarking-like behavior. And the money at stake in consumer gaming was comparatively tiny compared to current AI valuations. The cat-and-mouse cycle of measure vs counter-measure won't stop and should be a standard part of developing and administering benchmark services. Beyond hardening against adversarial gaming, benchmarkers bear a longer term burden too. Per Goodhart's Law, it's inevitable good benchmarks will become targets. The challenge is the industry will increasingly target performing well on leading benchmarks, both because it drives revenue but also because it's far clearer than trying to glean from imprecise surveys and fuzzy metrics what helps average users most. To the extent benchmarks become a proxy for reality, they'll bear the burden of continuously re-calibrating their workloads to accurately reflect reality as user's needs evolve.
- cedws 8mo agoAgreed, this benchmark would be much more useful ran multiple times a day. That could reveal degredation in line with load patterns.
- bredren 8mo agoFor CC, I suspect it also need to be testing and labeling separate runs against subscription, public API and Bedrock-served models? It’s a terrific idea to provide this. ~Isitdownorisitjustme for LLMs would be the parakeet in the coalmine that could at least inform the multitude of discussion threads about suspected dips in performance (beyond HN). What we could also use is similar stuff for Codex, and eventually Gemini. Really, the providers themselves should be running these tests and publishing the data. The availability status information is no longer sufficient to gauge the service delivery because it is by nature non-deterministic.
- swyx 8mo agoi recall another project here on HN maybe 4-6 months ago that would run tests 4x a day or something. not sure how to find them again
- Davidzheng 8mo agobut degradation from servers being overloaded would be the type of degradation this SHOULD measure no? Unless it's only intended for measuring their quietly distilling models (which they claim not to do? idk for certain)
- cmrdporcupine 8mo agoI've personally witnessed large variability in behaviour even within a given session -- which makes sense as there's nothing stopping Anthropic from shuttling your context/session around load balanced through many different servers, some of which might be quantized heavily to manage load and others not at all. I don't know if they do this or not, but the nature of the API is such you could absolutely load balance this way. The context sent at each point is not I believe "sticky" to any server. TLDR you could get a "stupid" response and then a "smart" response within a single session because of heterogeneous quantization / model behaviour in the cluster.
- epolanski 8mo agoI've defended opus in the last weeks but the degradation is tangible. It feels like it degraded by a generation tbh.
- cmrdporcupine 8mo agoit's just extremely variable
- megabless123 8mo agonoob question: why would increased demand result in decreased intelligence?
- vidarh 8mo agoIt would happen if they quietly decide to serve up more aggressively distilled / quantised / smaller models when under load.
- epolanski 8mo agoStilll relevant over time.
- seunosewa 8mo agoThe degradation may be more significant within the day than at the same time every day.
- GoatInGrey 8mo agoSure, but it's still useful insight to see how it performs over time. Of course, cynically, Anthropic could game the benchmark by routing this benchmark's specific prompts to an unadulterated instance of the model.
- chrisjj 8mo ago> Lots of variance in the score can come from random stuff like even Anthropic's servers being overloaded. Are you suggesting result accuracy varies with server load?
- rootnod3 8mo agoSorry what? "You can't measure my Cloud Service's performance correctly if my servers are overloaded"? "Oh, you just measured me at bad times each day. On only 50 different queries." So, what does that mean? I have to pick specific times during the day for Claude to code better? Does Claude Code have office hours basically?
- copilot_king 8mo ago[flagged]
- rootnod3 8mo agoVerily, my vichyssoise of verbiage veers most verbose, so let me run that thing out of tokens fast.
- johnsmith1840 8mo agoThis has been happening for years. Tgere's a great paper from microsoft on Deepspeed AI inference. Basically the paper showed methods for how to handle heavy traffic load by changing model requirements or routing to different ones. This was awhile ago and I'm sure it's massively more advanced now. Also why some of AI's best work for me is early morning and weekends! So yes, the best time to code with modern LLM stacks is when nobody else is. It's also possibly why we go through phases of "they neutered the model" some time after a new release.
- swyx 8mo agochill out, ofir does not work for anthropic. he's just saying there's inherent variability in LLMs and you need to at least 30x the samples that OP is doing in order to make any form of statistically significant conclusions.
- kuboble 8mo agoI wonder if my great experience with claude are partly due to the fact that my working hours don't overlap with the US west coast
- bhk 8mo agoAccording to Anthropic: "We never reduce model quality due to demand, time of day, or server load." https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues https://www.anthropic.com/engineering/a-postmortem-of-three-...
- embedding-shape 8mo agoThey've had issues before with things like "TPU top-k error - Claude sometimes dropped the best next token" (https://www.anthropic.com/engineering/a-postmortem-of-three-recent-issues https://www.anthropic.com/engineering/a-postmortem-of-three-...) so what's going on might not be intentional even.
- deleted 8mo ago[deleted]
- mgraczyk 8mo agoThat issue did not have any time of day dependence
- nikcub 8mo ago> I would run on 300 tasks and I'd run the test suite 5 or 10 times per day and average that score. assume this is because of model costs. anthropic could either throw some credits their way (would be worthwhile to dispel the 80 reddit posts a day about degrading models and quantization) or OP could throw up a donation / tip link
- phist_mcgee 8mo agoThen you'd get people claiming that the benchmarks were 'paid for' by anthropic
- nikcub 8mo agoone thing you learn from being on the internet is that you're never going to satisfy everybody
- simsla 8mo agoProbably, but with a small sample size like that, they should probably be taking the uncertainty into account, because I wouldn't be surprised if a lot of this variation falls within expected noise. E.g. some binomial interval proportions (aka confidence intervals).
- sjtgraham 8mo agoWhy should users care about Anthropic's servers being overloaded?