4 ms·
I need someone to run actual benchmarks between the two.
by jadbox 1mo ago
I need someone to run actual benchmarks between the two.
- swatcoder 1mo agoBenchmarks are the BMI of model evaluation. They may have utility in trying to look at the whole landscape of models, but are very misleading when it comes to making 1:1 comparisons or in developing confidence at to how a given model will deliver on your workflow.
- dofm 1mo ago> Benchmarks are the BMI of model evaluation. That is such an elegant way to put it.
- NitpickLawyer 1mo agoOnly relevant benchmarks are those you make yourself, targeted specifically for your workflows. Anything else is just number go up on a pretty graph, and every model out there is probably benchmaxxed to hell on the public ones anyway. Keep yours private.
- gertlabs 1mo agoThese models have gotten a fair amount of attention -- we're hoping it's enough to get them added to some reliable inference providers and OpenRouter, at which point we'll run them on our full benchmark suite.
- RomulusHill 26d agoHi Gertlabs, my inference company has actually started offering Ornith1.5 today for 9B and 35B A3B. If you are still interested, send me a message on X and I'll help you get started! https://x.com/romulushill https://x.com/romulushill https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-9b/ https://scalattice.com/models/ornith-1.5-35b-a3b/ https://scalattice.com/models/ornith-1.5-35b-a3b/ https://scalattice.com/developers/ https://scalattice.com/developers/