4 ms·
I think this says more about the benchmark than the capabilities of the model. If it were the case that 90th percentile performance on the bar exam mean that a
by Imnimo 2y ago
I think this says more about the benchmark than the capabilities of the model. If it were the case that 90th percentile performance on the bar exam mean that a model was a 90th percentile lawyer, and we've had these models for 15 months (in fact longer), where are all the LLM lawyers? The lesson here is that a test designed for humans may not be equally representative of capabilities when given to an LLM.
- pas 2y agonah, it's probably an okay-ish benchmark for humans, but it's just that, a filter to weed out those who can't learn hundreds of pages of legal trivia. the model is great at this, because the training set is full of this stuff. the questions are static and simple. (the actual text of the questions change, of course, but the format and the answers are from a fixed set. and the LLM doesn't need to generate text basically, just one token. A B C or D ... that said I'm curious how it's administered to the LLMs and how much that influences their performance.)