3 ms·Measuring What Matters: Construct Validity in Large Language Model Benchmarks1 points by Cynddl 11mo ago