3 ms·
Stick those same classifiers (that you admit dont seem to work) on the open models, and done.
by bcjdjsndon 3mo ago
Stick those same classifiers (that you admit dont seem to work) on the open models, and done.
- eddyg 3mo agoClassifiers are policy enforced by the process serving the model. Input classifiers get applied before it reaches the model so somebody hacking an open-weight model would skip this. Streaming classifiers get polled during decoding; hackers delete this check in the sampling loop. But both are always applied in closed weight models. Set Llama Guard to 1.0 and nothing is ever unsafe.
- bcjdjsndon 3mo agoMore than one way to do guard rails, slopboy
- eddyg 3mo agoAll of which are easily bypassed in open weight models; see my previous comment about K2.5.
- bcjdjsndon 3mo agoUsing this as your argument is dangerous. You're saying AI is inherently dangerous and an emergency stop button is actually what you need to be safe from the dangers... Forget the fact your open weighted model could be ran from a central location