3 ms·
Yep, exactly! My approach is similar - closed source benchmarks with prompts and tests from real LLM-driven products (mostly around boring business automation
by abdullin 2y ago
Yep, exactly!
My approach is similar - closed source benchmarks with prompts and tests from real LLM-driven products (mostly around boring business automation and enterprise workflows).
Although it would be neat to upgrade the setup to work on the synthetic data. This will at least make benchmarks shareable publicly (not just the results)