3 ms·
This paper wasn't for SKILLS per se, but for LLMs, and suggests that LLMs can't follow too many constraints at one time very well https://arxiv.org/abs/2608.124
by MrCoffee7 13d ago
This paper wasn't for SKILLS per se, but for LLMs, and suggests that LLMs can't follow too many constraints at one time very well https://arxiv.org/abs/2608.12426 https://arxiv.org/abs/2608.12426 . Other papers suggest that LLMs show more information sparsity in the results with negative constraints.
- mehmetoguzderin 13d agoGood one; reading this briefly took me back to the days of “Needle in a Haystack” being super challenging for LLMs. Maybe there needs to be a benchmark of “Rule in a Haystack” (similar to information one, testing not only independent rules also the ones that need hops) to clarify model performance regarding this. Thank you for the resource.
- oofbaroomf 11d agoI believe you're looking for IFBench (https://arxiv.org/abs/2507.02833 https://arxiv.org/abs/2507.02833, https://artificialanalysis.ai/evaluations/ifbench https://artificialanalysis.ai/evaluations/ifbench, https://github.com/allenai/IFBench https://github.com/allenai/IFBench)