4 ms·
> 15 months ago, general-purpose LLMs that have not been specifically trained on legal reasoning could score better than 90% of humans on the multistate bar exa
by elicksaur 2y ago
> 15 months ago, general-purpose LLMs that have not been specifically trained on legal reasoning could score better than 90% of humans on the multistate bar exam
The claims of similar performance on coding problems were shown to be due to contamination of the training data on the tested problems. It did abysmal on problems made public after the model training cutoff.
I don’t think anyone has tested contamination for the MBE claims, but I would lean toward assuming the same issue exists for that assessment until proven otherwise.
- coredog64 2y agoYou can see similar (poor) results if you give it a slightly tweaked puzzle that it’s seen before. The most recent example on Twitter was the farmer with the animals and the boat. If there are no constraints, it will give you an answer based on the tricks for the original unless you harass it.