3 ms·
Good one; reading this briefly took me back to the days of “Needle in a Haystack” being super challenging for LLMs. Maybe there needs to be a benchmark of “Rule
by mehmetoguzderin 24d ago
Good one; reading this briefly took me back to the days of “Needle in a Haystack” being super challenging for LLMs. Maybe there needs to be a benchmark of “Rule in a Haystack” (similar to information one, testing not only independent rules also the ones that need hops) to clarify model performance regarding this. Thank you for the resource.
- oofbaroomf 23d agoI believe you're looking for IFBench (https://arxiv.org/abs/2507.02833 https://arxiv.org/abs/2507.02833, https://artificialanalysis.ai/evaluations/ifbench https://artificialanalysis.ai/evaluations/ifbench, https://github.com/allenai/IFBench https://github.com/allenai/IFBench)