4 ms·
Random thought: why isn't LLM-based spam filtering ubiquitous yet? Some of the obvious spam that lands into my inbox would be caught easily even by the tiniest,
by fph 1y ago
Random thought: why isn't LLM-based spam filtering ubiquitous yet? Some of the obvious spam that lands into my inbox would be caught easily even by the tiniest, cheapest models.
- dpcx 1y agoSeems like an llm keeping up with the amount of email that people receive would be cost prohibitive, either in dollars or cpu time.
- torton 1y agoA simple, open-source spam filtering approach gets rid of 99.99% of spam. My total filtered email volume in a day is in the single to low double digits in the personal account and double digits at work. This is very much in range for LLM filtering of mail that passes the mechanical spam filter.
- robterrell 1y agoI have an appscript that uses Gemini for this. Works great. Usage is in in the free tier. I even had Gemini write the appscript.
- mthoms 1y agoPlease share :-)
- bityard 1y agoI've heard multiple opinions that traditional spam filtering techniques work well enough. And as someone who has implemented many of them for work and personal email systems, that is true only if you consider low-effort spam that comes from dodgy networks and domains, doesn't get past a graylist, and fails simple content analysis. High-effort spam is quickly becoming the norm and comes from spammers who take the time and expense to warm up their domain, correctly configure their DNS records, use TLS certificates and ports, are careful not to send too many messages to one provider at a time, address and format their emails to make them look like newsletters, and so on. Traditional spam mitigation system are not very effective against these techniques because it was always believed that spammers could never justify the expense of high-effort spam. That may have been true at one point, but it was silly to think it would always be true. An LLM that can read an email and quickly perform sentiment/classification analysis on the content itself is the only thing that I can see which will fill that gap. (Aside from just choosing to live with it, of course.) I'm just getting into this myself but my understanding is that very small models can do this surprisingly well. As to why it's not ubiquitous yet, I suspect it's only a matter of time. My Outlook (for work) has been automatically summarizing all of my emails for months now (something I can't opt out of, by the way). Hooking that up to an action that sends spam messages into the Junk Email folder is just a few lines of code on Microsoft's part.