4 ms·
* Release new model that scores an arbitrary 100 on a benchmark * Get everyone to talk about you as the first model to ever score 100 on the 100benchmark. * T
by well_ackshually 12d ago
* Release new model that scores an arbitrary 100 on a benchmark
* Get everyone to talk about you as the first model to ever score 100 on the 100benchmark.
* Tune it down over time so that you end up only scoring 75 on the benchmark and people get used to it, gaslight them into thinking it never changed or that it's just a harness problem, they can't run the old version locally anyways to verify. This also cuts your costs in half. Your gross margin on API calls goes from 70% to 150%.
* Release new model that scores 120 on the benchmark and advertise it as 50% better than the current model, while it's only in practice a minor increment. Everyone praises it as the second coming of Jesus Christ.
* Get everyone to talk about you as the first model to ever score 120 on the 100benchmark.
Bis repetitae.
- scrollop 12d agoWhy can't the models be benchmarked again after a few weeks/months to confirm this (likely true) theory? I imagine some people have their own personal in depth benchmarks they could do this for.
- Aurornis 12d ago> * Tune it down over time so that you end up only scoring 75 on the benchmark Where? I see so many accusations of this happening and it's so easy to check, but nobody ever proves it.