4 ms·
Our internal blinded human evals for summarization/creative work have always preferred Claude 3.0 Opus by a huge margin, so we've been using it for months - GPT
by icelancer 2y ago
Our internal blinded human evals for summarization/creative work have always preferred Claude 3.0 Opus by a huge margin, so we've been using it for months - GPT-4o didn't unseat it either.
GPT-4o IMO was better for coding (still using GPT-4 original w/ Cursor, but long-form stuff GPT-4o seemed better) but with this new launch, will definitely have to retest.
Pretty big news.