3 ms·
That's the result of guard rails on the hosted service. They run checks on the query before it even hits the LLM as well as ongoing checks at the LLM generates
by RevEng 2y ago
That's the result of guard rails on the hosted service. They run checks on the query before it even hits the LLM as well as ongoing checks at the LLM generates output. If at any moment it detects something in its rules, it immediately stops generation and inserts a canned response. A model alone won't do this.
- aussieguy1234 2y agoFor these tests, I self hosted the 14b version of R1 and ran it on my gaming gpu with ollama.