8 ms·
The funny rules of SpamAssassin in 2023
- lloydatkinson 3y agoSome of those highlighted rules, such as using CC or having the string “can help” being used to decide if something is spam or not is so absurd I’ll make sure to never use SpamAssasin.
- jeffbee 3y agoIt is not and has never been a good classifier. If open AI fans want to contribute something of value to society, they would train a spam classifier on a large, manually-labeled corpus of mail, where the features include envelope data. That would get open source maybe 10% of the way to Gmail quality, or 100x better than SA.
- CamperBob2 3y agoDoes SpamBayes still work?
- dugite-code 3y agoThe Python project that's clent side? Best of my knowledge it hasn't been updated in years, spamassasin's in built Bayesian classifier works just fine.
- jepler 3y agoI still used spambayes up until I basically abandoned my self-hosted e-mail setup (in favor of an @gmail). It occurred to me recently that LLM-style tokenization + bayesian classification would be a sweet upgrade for spambayes, which always struggled with ad-hoc tokenization rules. (I don't think of it as "client side"; it was integrated with my system via procmail on what I'd call the "server side". You could use it in other ways, including as an Outlook plugin, way back in the day. Or it could connect to an IMAP mailbox and filter messages it found, etc. Really versatile tool for its time)
- acidburnNSA 3y agoFWIW I've been using SpamAssassin for over a decade personally (partly to avoid Google dependence), and it's been pretty darn good once I ran the Bayesian learning thing a few times many years ago. I get like 3-5 spams per week in my inbox. Do others really consider SA that bad?
- brokensegue 3y agoOn Gmail I get maybe 1 spam a month max in my inbox (and it blocks many per day)
- acidburnNSA 3y agoThat's pretty good. I have no doubt that Gmail spam protection is better than my self-hosted SA protection. For me, independence from a somewhat suspicious large company for something as important as e-mail is worth it.
- dugite-code 3y agoI have a gmail honeypot where I fetchmail junk email straight to my junk folder and have a scheduled sa-learn cronjob. Ever since I started this I essentially stopped getting junk email in my selfhosted inbox. I also have dovecot set to learn Ham every time I file an email from the inbox to a folder for good measure.
- kjs3 3y agoSo...statistically insignificant difference from SA for most mail users.
- brokensegue 3y ago3-5 spams per week vs. 1 per month is a big difference
- kjs3 3y ago
- jerf 3y agoWe'll get there eventually, but it will be a bit. Spam classification at scale is already a compute-bound, or at least compute-starved, operation. Spam classification systems already do what they can to avoid so much as invoking a virus scanner if they can avoid it, because at scale it's so expensive. LLM-based spam classification is another order of magnitude more expensive and would require hardware that current spam systems do not have. But that's a problem that will resolve itself over time, in a variety of ways. And the spam systems can play the same tricks with only invoking it on a fraction of emails too, of course. It's just at current expense levels, that would be a very small fraction indeed. I'd hazard that trying to use modern AI on spam classification at scale could easily consume 10x-100x of all current AI hardware and still make less of a dent than you'd hope.
- jeffbee 3y agoIt doesn't need to be computationally costly because, as you seem to imply, there are tiers of cost tradeoffs. You can invoke a very cheap classifier at SMTP time, that is biased to have few false positives, that will temporarily reject all that which is highly likely to be spam. You can do this without even glancing at the body. Of course, having signals about peer reputation is the strong suit of Gmail or Microsoft, and the distributed, open community would need to solve the problem of promptly updating and distributing such reputation signals. And by "promptly" I mean within seconds of the leading edge of an attack. Then there are increasing tiers of cost that you would only run after it becomes likely that the message is acceptable. As you say, you would only run an antivirus on a message on the verge of delivery, because decoding the attachment and running the AV (in an expensive sandbox) is so costly.
- dugite-code 3y agoReally any of the open source language models might work well enough for the job. If you could manage to get a classifier that runs with tensorflow to take advantage of a coral tpu it would certainly be a major step up with managable performance.
- thesuitonym 3y agoI hope against hope that AI spam detection never becomes a thing. At least with today's methods, I can tell a person why their message was marked as spam. If AI detection becomes the norm, all I can do is shrug and say, "Sorry, it's the algorithm."
- potatoman22 3y agoHow do you think SpamAssassin/gmail/outlook created their spam rules?
- jeffbee 3y agoGmail has used machine learning to classify spam since its creation. https://workspace.google.com/blog/identity-and-security/an-overview-of-gmails-spam-filters https://workspace.google.com/blog/identity-and-security/an-o...
- Mailtemi 3y agoGmail is successful because it naturally is the biggest honeypot. Most antispam API filters are like accumulators. When a trend is detected, the rest are protected. But overall, it's about scale.
- layer8 3y agoSpamAssassin has a Bayes filter that you can train with ham and spam. This has basically been a thing since forever.
- andrewf 3y ago"A plan for Spam" (2002) - https://paulgraham.com/spam.html https://paulgraham.com/spam.html
- El_RIDO 3y agoCorrect, Spamassassin will not classify emails be default. But it does include a Bayesian classifier and tooling that can be used to train it with curated ham and spam emails. It does require extra steps to set this up and feed it selected maildirs (for example, exclude inbox, spam and trash for training ham). https://spamassassin.apache.org/full/3.0.x/dist/doc/sa-learn.html#description https://spamassassin.apache.org/full/3.0.x/dist/doc/sa-learn... https://cwiki.apache.org/confluence/display/spamassassin/BayesInSpamAssassin https://cwiki.apache.org/confluence/display/spamassassin/Bay...
- creeble 3y agoI find it far more likely that AI will be used (indeed, is used already) to generate spam, rather than filter it.
- throw0101d 3y ago> being used to decide if something is spam or not Each rule has a score associated with it. By default a message needs to reach 5.0 to be marked as "spam": * https://spamassassin.apache.org/full/3.0.x/dist/doc/Mail_SpamAssassin_Conf.html#item_nn https://spamassassin.apache.org/full/3.0.x/dist/doc/Mail_Spa... The threshold is configurable. An header is added post-processing, e.g.: X-Spam-Status: Yes, score=21.6 required=4.0 […] * https://cwiki.apache.org/confluence/display/SPAMASSASSIN/X+Spam+Status https://cwiki.apache.org/confluence/display/SPAMASSASSIN/X+S... One can then choose what do to with this information (via procmail or Sieve). There is another header as well: > X-Spam-Level: This displays your spam level with asterisks, with one asterisk displayed per point, rounded down. For example, if your overall SpamAssassin score is 4.3, it will display ****. If you score less than 1, for example, 0.5, it will display nothing. * https://www.mailercheck.com/articles/spamassassin-score https://www.mailercheck.com/articles/spamassassin-score
- londons_explore 3y agoHaving the rules public seems to take away most of the benefits... Any smart spammer will just tweak his spam to not hit these rules... And if he hasn't, it's because the vast majority of people don't use SpamAssassin
- Hnrobert42 3y ago>smart spammer I am sure there are plenty of smart spammers, but it also seems like a lot of spam comes from folks using scripts and email lists they use without fully understanding. It appears SpamAssassin would help with those operations.
- bombcar 3y agoI'm starting to think the smart spammers are the ones selling worthless spam tools to the dumb spammers, because so much spam is entirely unactionable.
- kjs3 3y ago[dead]
- brightball 3y agoPart of the smart spammer approach is to condition people to what spam looks like, so you're more likely to let through the ones they really care about.
- adrienjarthon 3y agoHi, Adrien here author of this article (and of updown.io). That is true and I actually hesited to write the article for this reason, because it could make the spammer life easier. But after seing some of the legacy and nonsense in here I though it's still worth it so people at least understand what they are using.
- bell-cot 3y ago> Any smart spammer will just... Spam is all about high-volume/~no-cost delivery of crap. Time spent tweaking the spam - to evade $Defense_1, $Defense_2, etc. - is added cost. Especially if $Defense_n is only used by a few of the prospective victims (folks too savvy or paranoid to be suckered do not count), then tweaking to get around $Defense_n is a losing strategy for the spammer.
- srmarm 3y agoI've been using SpamAssassin for at least 15 years and it's sadly gotten less useful as the spam arms race has moved on. We regularly see people on here post about deliverability issues with Gmail/Outlook but the truth is that sender reputation is by far the biggest indicator of whether a message will be spam - these type of rules are just counting deckchairs on the titanic in comparison. And this plays into the strengths of the big mail networks in detection. It's a bonus to them that every time they block a smaller host there is a good chance that sender will consider a move to office365 or Google Workspace for their mail. As an aside, not sure if OP is related to them but updown.io is a nice service and I appreciate the simple PAYG pricing! For what it's worth their mails seem to get through successfully to me too. Also for those facing mail delivery issues (or just practicing good email hygiene) - I recommend www.mail-tester.com - they give you an email address to send a mail to and carry out a heap of tests - including checking against SpamAssassin + blacklists, SPF/DNS/etc testing.
- andrewfromx 3y agothere needs to be like a mozillia vs chrome thing here no? What's the best try so far for something like letsencrypt or mozilla foundation for not owned by big tech email so "will consider a move to office365 or Google Workspace for their mail" the sender has this other awesome option?
- jeffbee 3y agoIf you wanted to operate a haven for independent email hosting, where you want to assure deliverability in the face of Gmail's sender reputation system, you would need to classify your outbound traffic, and have a death penalty for spammers. If you tolerate any activity that peers classify as spam, that would tank your reputation.
- BSDobelix 3y agoI like rspamd much more (performance and redis) than SpamAssassin, and as you mentioned: -https://www.mail-tester.com https://www.mail-tester.com -https://www.learndmarc.com https://www.learndmarc.com -https://mecsa.jrc.ec.europa.eu/en/ https://mecsa.jrc.ec.europa.eu/en/ Are exellent tool's to check your "deliverability".
- matthews2 3y agoPutting your outbound emails through SpamAssassin as part of a regression test sounds like a really good idea - would have never thought of doing that myself!
- uean 3y agoI love the analysis. But I hate that the 'fixed' email ends up being wordier for no reason at all. Brevity has value. Having to bloat content (an email to get past anti-spam; a cooking blog to rank better within Google SEO; ...) brings back memories of high-school english papers, or the modern equivalent ChatGPT.
- adrienjarthon 3y ago100% agree, I also hate that I had to do this.
- layer8 3y agoCouldn’t you add some “hidden” text instead, e.g. white on white or display:none?
- adrienjarthon 3y agoI could but I don't want to, it's even more of a dark pattern and looks way too "spammish" IMO. I don't want my users to find this in their email and think that I'm trying to trick their system. Also I wouldn't be suprised if some antispam tries to detect this as a spam criteria.
- chrismorgan 3y agoAnother piece of feedback: the link doesn’t look like a link any more. It wasn’t great before, but the verbiage made it adequately clear. But now it’s terrible, because the wording doesn’t suggest an action, and it doesn’t look like a link or a button. You should either restore its underline and lean into “link”, or give a background colour or (generally better) gradient and lean into “button”. But when it’s just a border, it doesn’t look like a button, especially when there’s a tick after it. And change the wording again.
- adrienjarthon 3y agoThanks for this feedback, I actually changed this because some of my clients complained of the opposite, that the link was a bit too "dim" and didn't look like the the obvious Call To Action in the email. But it's all very debatable I agree and I may change this again in the future.
- jwr 3y agoI've been using SpamAssassin since, well, forever, in internet terms. My recent facepalm moment was when I noticed that E-mails from the Playdate developer forum (Playdate is a really cool tiny gaming console) land in my spam folder, because anything in the .date domain (and the forum uses play.date as the domain) is assumed to be "dating spam".
- layer8 3y agoGiven the double meaning of “play date”, it’s not surprising that it would cause a higher score, even if it used a different TLD.
- loloquwowndueo 3y agoI have deny listed most of those new funny tld as it’s indeed a good indication of spam. Here the face palm should be playdate’s because they have realized their domain looks like a spam domain.
- JohnFen 3y agoYes, I do the same. It's very useful to me because there is no non-generic TLD that I would be getting legitimate email from, but it may not work well for people who do want to get emails from such TLDs.
- linsomniac 3y agoIt's been a very long time since I ran a mail server, but for a decade or more I pumped all our outgoing mail through Hashcash because it gave a good boost to the Spam Assassin score. We'd crank it through the largest one, and it would add ~60sec to the mail delivery, unless we had a bunch of outgoing mail, but it was worth it I felt.
- creeble 3y agoI tried using SpamAssassin (via Proxmox Mail Gateway, which makes it much easier to set up) to replace a Barracuda email appliance (it was destined to get a *6x* service price increase in 2024!), and after several months of trying to get the number of FPs down, I gave up. The problem wasn't just the number of FPs (which were much higher than the 'Cuda) -- it was that they came from real people, who were often common senders. This is not corporate email, or anything that was even remotely spam (except as SA's crazy ruleset determined). These all required whitelisting, and it became a real chore for all my users to keep up with all the whitelisting. So back to the Barracuda for another year. It lets a little more spam through, but virtually no FPs. I just couldn't make SA get the same performance, even with many tweaks to the weights and rulesets.
- bschne 3y agoI couldn't help but think about mechanistic interpretability research on large neural models reading this — I guess this is what happens when humans do something similar, adding and removing tweaks here and there to better fit this or that case, over a long period of time.
- jaimex2 3y agoI won the war a while ago now. I basically trash all emails not in my contact lists. Easy.
- TwoNineFive 3y agoSpamassassin is doing it's job here, and doing a good job! Most spammers and marketing/sales sleezoids never think they are doing anything wrong. They are totally empathy incapable. Or they know they are scum and don't care. Either way. OP talks about adding "invisible text" and other such common spammer tactics to get around some of the rules. Zero self-awareness. At no point did this person ever think "did I do something wrong?". No, it's that shitty Spamassassin!