4 ms·
Sure, it did good in frontiermath. That's not what this thread is about. Your comment isn't relevant at all
by MattDaEskimo 2y ago
Sure, it did good in frontiermath. That's not what this thread is about.
Your comment isn't relevant at all
- whimsicalism 2y agothis thread is about math LLM capability, it’s a bit ridiculous to say that mentioning frontiermath is off topic but that’s just me
- MattDaEskimo 2y agoJust because you can generalize the topic doesn't mean you can ignore the specific conversation and choose your hill to argue. Additionally, the conversation of this topic is about the model's ability to generalize and it's potential overfitting, which is arguably more important than parroting mathematics.
- whimsicalism 2y agoperformance on a held-out set (like frontiermath) compared to putnam (which is not held out) is obviously relevant to a model's potential overfitting. i'm not going to keep replying, others can judge whether they think what i'm saying is "relevant at all."
- MattDaEskimo 2y agoAgain, you set your own goal posts and failed to add any insights. The topic here isn't "o-series sucks", it's addressing a found concern.