3 ms·
grok is 17%? And that's the lowest, most models are like 80%+? While hallucination is probably closer to 100% depending on the question. This benchmark makes n
by dubcanada 5mo ago
grok is 17%? And that's the lowest, most models are like 80%+?
While hallucination is probably closer to 100% depending on the question. This benchmark makes no sense.
- elAhmo 5mo agoNo one serious uses grok.
- ajdegol 5mo ago@grok is this true?
- NamlchakKhandro 5mo agono
- for_i_in_range 5mo agoThis comment deserves more love
- RALaBarge 5mo agoYMMV but Grok 4.1 Fast can usually find via static analysis a few things that other models dont seem to catch with the same prompt
- d0gsg0w00f 5mo agoWhy not? Honest question.
- phillipcarter 5mo agoBecause the Grok models offer nothing different in serious contexts from the other leading models which don't come with a heaping pile of bad baggage.
- Jensson 5mo ago> While hallucination is probably closer to 100% depending on the question. But the benchmark didn't ask those questions, and it seems grok is very well at saying it doesn't know the answer otherwise.
- MagicMoonlight 5mo agoIt makes sense. Grok is taught to answer the question, regardless of how explicit or extreme it is. These other models are taught to suppress any wrongthink. That's going to make it hard to answer things correctly. If you've been told to answer something incorrectly because it's wrong, then you'll have to make up an answer.