4 ms·
Browser Agent Benchmark: Comparing LLM models for web automation
- pixel_popping 8mo agoIt's lacking the best model (Opus 4.5) on the benchmark tho.
- djohnston 8mo agoYeah but then their own product might not score the highest.
- pixel_popping 8mo agoExactly why I'm pointing it out, which feels a bit corrupt, but understandable.
- djohnston 8mo agotbh i was a bit cranky yesterday - even if they are #2 on a legit benchmark that would be impressive
- wiradikusuma 8mo agoSince we're in this topic, can anyone suggest good AI-based tool for exploratory (fuzzy?) web testing?
- MagMueller 8mo ago[dead]