4 ms·
It's absolutely not reliable, and we have opened a few issues for that. For example: our instructions (which are read by the model and classifier) include "do
by whstl 1mo ago
It's absolutely not reliable, and we have opened a few issues for that.
For example: our instructions (which are read by the model and classifier) include "do not use sed/python/perl/etc, always use the edit tool for editing", and this only gets followed for a few messages. We have introduced scripts to block those ourselves, since the classifier doesn't care.
Because of those problems, my team is currently testing OpenAI after about a year of Anthropic.