5 ms·
Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think. https://github.com/BlueFalconHD/apple_generative_model_safety_decrypted
by binarymax 1y ago
Wow, this is pretty silly. If things are like this at Apple I’m not sure what to think.
https://github.com/BlueFalconHD/apple_generative_model_safety_decrypted/blob/main/decrypted_overrides/com.apple.gm.safety_deny.input.mail_reply.long_form_basic.generic/AssetData/metadata.json https://github.com/BlueFalconHD/apple_generative_model_safet...
EDIT: just to be clear, things like this are easily bypassed. “Boris Johnson”=>”B0ris Johnson” will skip right over the regex and will be recognized just fine by an LLM.
- deepdarkforest 1y agoIt's not silly. I would bet 99% of the users don't care that much to do that. A hardcoded regex like this is a good first layer/filter, and very efficient
- BlueFalconHD 1y agoYep. These filters are applied first before the safety model (still figuring out the architecture, I am pretty confident it is an LLM combined with some text classification) runs.
- brookst 1y agoAll commercial LLM products I’m aware of use dedicated safety classifiers and then alter the prompt to the LLM if a classifier is tripped.
- latency-guy2 1y agoThe safety filter appears on both ends (or multi-ended depending on the complexity of your application), input and output. I can tell you from using Microsoft's products that safety filters appears in a bunch of places. M365 for example, your prompts are never totally your prompts, every single one gets rewritten. It's detailed here: https://learn.microsoft.com/en-us/copilot/microsoft-365/microsoft-365-copilot-architecture https://learn.microsoft.com/en-us/copilot/microsoft-365/micr... There's a more illuminating image of the Copilot architecture here: https://i.imgur.com/2vQYGoK.png https://i.imgur.com/2vQYGoK.png which I was able to find from https://labs.zenity.io/p/inside-microsoft-365-copilot-technical-breakdown https://labs.zenity.io/p/inside-microsoft-365-copilot-techni... The above appears to be scrubbed, but it used to be available from the learn page months ago. Your messages get additional context data from Microsoft's Graph, which powers the enterprise version of M365 Copilot. There's significant benefits to this, and downsides. And considering the way Microsoft wants to control things, you will get an overindex toward things that happen inside of your organization than what will happen in the near real-time web.
- twoodfin 1y agoEfficient at what?
- miohtama 1y agoSounds like UK politics is taboo?
- immibis 1y agoAll politics is taboo, except the sort that helps Apple get richer. (Or any other company, in that company's "safety" filters)
- tpmoney 1y agoI doubt the purpose here is so much to prevent someone from intentionally side stepping the block. It's more likely here to avoid the sort of headlines you would expect to see if someone was suggested "I wish ${politician} would die" as a response to an email mentioning that politician. In general you should view these sorts of broad word filters as looking to short circuit the "think of the children" reactions to Tiny Tim's phone suggesting not that God should "bless us, every one", but that God should "kill us, every one". A dumb filter like this is more than enough for that sort of thing.
- XorNot 1y agoIt would also substantially disrupt the generation process: a model which sees B0ris and not Boris is going to struggle to actually associate that input to the politician since it won't be well represented in the training set (and on the output side the same: if it does make the association, a reasoning model for example would include the proper name in the output first at which point the supervisor process can reject it).
- quonn 1y agoI don‘t think so. My impression with LLMs is that they correct typos well. I would imagine this happens in early layers without much impact on the remaining computation.
- lupire 1y ago"Draw a picture of a gorgon with the face of the 2024 Prime Minister of UK."
- chgs 1y agoThere were two.
- binarymax 1y agoNo it doesn't disrupt. This is a well known capability of LLMs. Most models don't even point out a mistake they just carry on. https://chatgpt.com/share/686b1092-4974-8010-9c33-86036c88e789 https://chatgpt.com/share/686b1092-4974-8010-9c33-86036c88e7...
- bigyabai 1y ago> If things are like this at Apple I’m not sure what to think. I don't know what you expected? This is the SOTA solution, and Apple is barely in the AI race as-is. It makes more sense for them to copy what works than to bet the farm on a courageous feature nobody likes.
- stefan_ 1y agoWhy are these things always so deeply unserious? Is there no one working on "safety in AI" (oxymoron in itself of course) that has a meaningful understanding of what they are actually working with and an ability beyond an interns weekend project? Reminds me of the cybersecurity field that got the 1% of people able to turn a double free into code execution while 99% peddle checklists, "signature scanning" and deal in CVE numbers. Meanwhile their software devs are making GenerativeExperiencesSafetyInferenceProviders so it must be dire over there, too.
- Aeolun 1y agoThe LLM will. But the image generation model that is trained on a bunch of pre-specified tags will almost immediately spit out unrecognizable results.
- Lockal 1y agoWhat prevents Apple from applying a quick anti-typo LLM which restores B0ris, unalive, fixs tpyos, and replaces "slumbering steed" with a "sleeping horse", not just for censorship, but also to improve generation results?
- the_mar 1y agowhy do you think this doesn't already exist?