6 ms·
> Extraordinary claims require extraordinary evidence so see below for the receipts. Yes, that’s the kind of attitude I want to see in these model releases
by pennomi 24d ago
> Extraordinary claims require extraordinary evidence so see below for the receipts.
Yes, that’s the kind of attitude I want to see in these model releases
- ramon156 24d agoBut the evidence is not there...
- pennomi 24d agoIndeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.
- simianwords 24d agoThey gesture at not using benchmarks for some reason...
- meric_ 24d agohttps://typesafe.ai/blog/antibenchmaxxing https://typesafe.ai/blog/antibenchmaxxing But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do