3 ms·
just ran the LLM to SQL benchmark over opus-4.1 and it didn't top previous version :thinking: => https://llm-benchmark.tinybird.live/ https://llm-benchmark.tiny
by alrocar 1y ago
just ran the LLM to SQL benchmark over opus-4.1 and it didn't top previous version :thinking: => https://llm-benchmark.tinybird.live/ https://llm-benchmark.tinybird.live/
- epolanski 1y agoHow does running it multiple times performs? LLMs are non-deterministic, I think benchmarks should be more about averages of N runs, rather than single shot experiments.