4 ms·
AA-Omniscience: Evaluating Cross-Domain Knowledge Reliability in Language Models
- karty1 11mo agoI really like this. One of the most realistic evals I've seen, finally quantifying the good vibes many of us feel from the Claude models. Also lmao at CapGPT(-oss).