3 ms·
I’m struggling to build my own evaluation bench for local models against my own (scientific coding) use cases, and realising it’s quite hard. Good coding has ma
by djc404 26d ago
I’m struggling to build my own evaluation bench for local models against my own (scientific coding) use cases, and realising it’s quite hard. Good coding has many dimensions and it varies depending on the need. Nothing seems to collapse cleanly to a few numbers.
Anyone played with this and have examples I can steal from?