3 ms·
If you want to disabuse yourself of your notions annd intuitions on how LLMs work, run safety models and tests. On a hate speech policy test for a lightweight
by intended 12d ago
If you want to disabuse yourself of your notions annd intuitions on how LLMs work, run safety models and tests.
On a hate speech policy test for a lightweight LLM, the presence or absence of the last full stop on the last sentence would cause the model to flip its decisions.
- latentsea 12d agoI blame the $slur
- user43928 12d agoI'm not sure how crappy small models behaving unreliably is relevant here, when a large SOTA model does presumably not produce the same issue.
- intended 12d agoIts relevant because those same issues occur with large models. Also, the "crappy small model", was a model trained for safety tasks, and outperformed the frontier lab safety models.
- user43928 12d agoAnd you tested this, that the presence of the last full stop flips the outcome with a large SOTA model? Small models are notoriously unreliable and prone to hallucination in my experience, so that would not surprise me to be an issue there.