4 ms·
Bad benchmarks and evals: Senior SWE-Bench, napkin math, and winter tires
- alexpran 24d ago[flagged]
- stephantul 21d agoThe (relevant) segue into tires elevates this post so much. I don’t understand why, but it does
- vintagedave 21d agoMaybe that it's an example of the same pattern in a real-world, non-digital area? It makes it feel more grounded.
- daveguy 21d agoIt's something an LLM would never be able to come up with because it requires a depth of experience that LLMs just don't have. It's authentically human.
- IsTom 21d agoI'm a little bit confused about these tire claims as > down to 0C / 32 F (he didn't test colder conditions) I get that people live in different places, but that's a huge caveat. How's that winter if you're not below 0°C? That sounds like "winter tires are worse than all-seasons tires in winter if you exclude winter".
- pm215 21d agoI agree that testing colder conditions would be useful, but I guess you get into problems with it being ice and snow that you're testing, not the "summer tyres get too hard in the cold" hypothesis. It's winter if it's noticeably cooler than summer. For instance London's climate has a January average minimum temperature of just under 3C, which is very clearly winter compared to the July minimum of 14C. (figures from Met office site on their Heathrow station.) While the temperature does drop below 0C sometimes, it is not consistently below zero. In southern England pretty much nobody changes tyres for winter -- you just use the same set all year. Optimising for "2C in the wet" seems about right...
- IsTom 21d ago> but I guess you get into problems with it being ice and snow that you're testing If it's not actively snowing for a few days roads get clear by thawing during day because of salt/cars being hot/sun shining. The "summer tyres get too hard in the cold" idea could very well come from places where the typical winter is at least < -5°C. It's not uncommon in central/eastern/northern europe for temperatures to be < -10°C for extended periods of time.
- rjsw 21d agoThe crossover point where winter tyres work better is < +7°C, it doesn't need to be below freezing.
- pm215 21d agoThe linked article claims "in dry conditions, summer tires have the best grip down to 0C", though, which doesn't seem to match where you suggest the crossover point is.
- attila-lendvai 21d agowhich is my experience, too. winter tires, especially the snow versions, are straight out a source of danger here in central europe. i just put some snow chains in for the winter surptises, and ride with my summer tires because its grip is clearly better on dry and wet bithumen, even in near zero temps, which is most of the season.
- sgerenser 21d agoDid you read the article? I thought his whole point is that little piece of folk wisdom didn't turn out to be true in actual testing.
- TacticalCoder 21d ago> How's that winter if you're not below 0°C? Seasons are called summer, autumn, winter and spring and each last three months and that bears no relation to whether or not there's snow or sub-zero temperatures. In the country I live in atm the rules regarding winter tires are not the absolute dumbest but they're still very dumb: you need to have either winter tires or all-seasons tires "if the conditions are winter'y". Which means, basically, both sub-zero AND either wet or icy. Sub-zero and all sunny means winter tires aren't mandatory. The reason it's still dumb it's that that correspond to, at most, 10 days per year. And this forces a lot of people to have worse performing tires during much more than 10 days. Which is probably the cause for a lot of accidents (e.g. people on days where it's + 3 C would be safer with summer tires, that do perform way better than "I've got winter tires because tomorrow at 7am it may or may not be -1 C and it may or may not be raining"). Not that's of course dumbtardation but there's worse: there are countries where from that month to that month of winter, no matter the temperature, you must have winter (or all-seasons) tires. And at times you'll have an entire winter without freezing temperatures. So politicians who voted these laws are basically creating more accidents due to cars having inferior tires (the tires lobby does love it though). It's sad but it's how it is.
- kuerbel 21d agoA lot of drivers don't want it to be true but the quality of the tires absolutely matters. E.g. the usual ADAC test winner are actually good in dry conditions, but they are also expensive. The worst are just bad, but cheap. You get what you pay for. The best winter tires outclass bad summer tires, even up to lower middle end. The worst tires however are all-seasons.
- 27183 21d agoLately I've just been leaving my snow tires on year round (Bridgestone Blizzaks). It gets cold here, between November and April the number of days with daytime high above freezing is small. I used to run all seasons in the summer and snows in the winter but they were lasting too long that way. It's better to wear them out within ~5yr.
- threetonesun 21d agoI've done that with "mild" Winter tires, really the thing full Summer tires excell at is removing water, which requires a center tread that's useless in snow. So if you don't get a lot of rain and it doesn't get too hot (some Winter compounds will get mushy and wear very quickly) they're fine. But really I've just described a Winter rated All-Season at this point, and you should buy those.
- 27183 21d agoAll season tires are useless on snow and ice unless they're brand new. They're OK the first winter, after that it's bad news. On heavier vehicles I run BFGoodrich All Terrain T/A tires year round. They have good siped treads which grip on ice. For the car, studless Blizzaks have been holding up well year round. No abnormal wear so far, and they do fine in the wet, dry, heat, etc. I probably only drive 1-2 days/yr over 90°F though, in a hotter climate they might not be so great. [edit] obviously this is a tradeoff--I'm trading slightly reduced hot/dry/wet performance for massively increased winter performance. The reduction in summer performance is small enough to not be noticeable, whereas the increase in winter performance is large. On all season tires I would have to chain up a dozen or so times per year, often just to move the car like 3 car lengths out of a parking spot. I've only ever had to put the chains on once with snow tires on the car, and that was bashing through 6" of unplowed crusty icy stuff up a steep driveway.
- dgacmu 21d agoI dunno about this - I leave my crossclimate 2 aw's on year round and they're pretty fantastic in snow. We had a pretty good winter this year (44" from dec-feb) and they just worked. I didn't go meandering around any mountains on them, mind you, but they're so far ahead of most typical all-season tires it's almost not fair to compare. For -most- lazy people who don't want to swap tires or rims, they seem a better option than leaving true winter tires on year round. Obviously, there are exceptions depending on where you live.
- mplanchard 21d agoYes, this stuck out to me, too. I use winter tires in the winter specifically for the snow and consistently below-freezing temps. Are there places people bother with winter tires that aren’t actually cold in the winter?
- menaerus 21d agoIf warranted by the law then yes, you have to abide to it.
- mplanchard 21d agoAre there any places that require winter tires by law that don’t have cold winters? I don’t know of anywhere in the US that even requires winter tires, although I think some states require you to have winter tires OR chains.
- menaerus 21d agoSometimes the winters may be cold, some other times they may be exceptionally warm, and sometimes they're a mix of both. The law remains the same regardless what the winter was or is like. Countries that have exceptionally warm winters do not really have a winter conditions so I'd guess their law wouldn't have a requirement for winter tires or chains
- mplanchard 21d agoYes of course, anyone who lives somewhere with winter knows that some winters are colder than others. In the late fall when it is time to swap the tires, you don’t know in advance whether the winter will be harsh or mild, but people put snow tires on anyway in places where harsh winters are common. It would only make sense to require them every winter if the large majority of winters were quite cold. As such, nowhere in the US requires winter tires, but someone else noted that Quebec does. This makes sense, since even a mild Quebec winter will spend most of the time getting below freezing. I would be very surprised if there were a place that required winter tires by law where the winter does not always get and stay at or below freezing for significant periods of time almost every winter.
- quietraster 21d ago[dead]
- jdw64 21d agoI feel like AI development is hitting a wall now. Evaluations are driven by benchmarks, but I have no idea who is even evaluating the validity of these benchmarks. The idea that continuously training and scaling models will automatically yield broad capabilities across multiple domains hardly seems true anymore. Benchmarks like Senior SWE-Bench apply discontinuous and arbitrary thresholds to evaluate results. If the generated code is semantically identical to the reference solution but even slightly exceeds a length cutoff, it fails? That seems like a genuinely flawed criterion, and the LLM-based graders themselves feel highly unstable. Honestly, my impression is that LLM advancement is now completely dictated by benchmarks. But looking at recent trends, token prices are skyrocketing while coding capabilities haven't shown any massive improvements beyond a certain model generation. In fact, comparing GPT-6 Astra to 5.6 SOL, SOL often writes better code. Considering all this, while training across diverse domains can pack a model with various pieces of knowledge, ultimately, it feels impossible for this approach to do something like derive the theory of relativity from medieval knowledge.
- lh712 21d agoNitpicking: It is not possible to derive the theory of relativity from medieval knowledge. At least not in any way resembling how the development actually happened, that is, heavily influenced by observations (or rather an iteration of speculation and observation). [Unless one counts both Galilean relativity and some parts of the theory of electromagnetism as being contained in medieval knowledge.] Perhaps it would be possible if the classically observed reality turned out to be possible (consistent under some reasonable conditions) only as a consequences of some deeper, sufficiently determined theory; but that would go far beyond our current level of knowledge, at least as far as I can say.
- wongarsu 21d agoThe earliest you could reasonably derive special relativity is probably 1881 with the first version of the Michelson–Morley experiment, which showed that speed of light is the same no matter how you move through space. Which, on the scale of discoveries, is a pretty short timeline to the 1905 publication of special relativity. Which doesn't really stop you from training an LLM on 1904 knowledge and having it derive special relativity. LLMs trained with knowledge cutoffs that far back is a fun exercise, and one I've also dabbled in in the past. But you don't have anywhere near enough training data to reach SotA levels of intelligence, even if you had the money for those training runs. And lobotomizing all modern information out of LLMs doesn't seem viable either. So I don't see how you could ever turn that into a viable benchmark
- dom96 21d agoThis is great. I've been building my own model benchmark lately and it has indeed been so easy to mess up the scoring. It's simply much harder to come up with an algorithm that combines all your individual scores into something that isn't broken in some special circumstances. That's why I think many just start capping the results.
- cb321 21d agoSomething Dan does not observe in his article (perhaps Jamie does elsewhere? edit: or even Dan elsewhere) is that the same problem which makes the memory latency benchmark unrealistic (or at least misleading) often impacts hash table lookup benchmarks as mentioned at https://github.com/c-blake/bu/blob/main/doc/memlat.md https://github.com/c-blake/bu/blob/main/doc/memlat.md and probably many other benchmarks. Essentially, CPU work prediction/speculative execution has become so good that much care is often required to measure latency rather than reciprocal throughput. This all started in the 1990s (or probably earlier with Cray), but I guess there's been an ongoing educational failure/oversimplification tendency. Of course, throughput at one "level of work" is often latency at a higher level (like command-to-command execution, for example). So, "what matters" might be throughput or "latency". It all depends. :-) People often fiddle with such semantics to market methods, products, ideas, ... and marketing is often at cross purposes with understanding.
- jbellis 21d agoWhile I'm generally sympathetic to the idea that the public coding benchmarks are inadequate, to the point that I've written my own in the past and will likely do so again, the complaint here is that "these tasks don't match what I do in my day job" which is ~always going to be the case. The hope with benchmarks is that you can capture properties that generalize, from examining performance against small set of tasks, and I do think that this is at least directionally true for well-designed evals. (There's https://withspecific.com/benchmarks/real-swe https://withspecific.com/benchmarks/real-swe but since the tasks are private we still don't really know what they're representative of.)
- Zigurd 21d agoI don't mean to make you write a dissertation but to say that AI benchmarks can "capture properties that generalize from examining performance against a small set of tasks" is a bald assertion. It's a hypothesis without a theory behind it. I can profile some code and then tell you where the slow parts are. Unless there's some explanatory power to an AI benchmark, it's a bit like benchmarking pillows visually.
- menaerus 21d agoI started using gemini with caution given the "much worse" benchmarking points it has gotten and still does but in practice there's very little evidence I found in comparison to claude models. It performs really well on non trivial tasks.
- ennepoai 21d ago[flagged]
- yg2k 19d ago[flagged]