3 ms·
It's using base64 encoded strings to deliver the initial stage. Can this be avoided/flagged more easily if by adding a scan of statements featuring base64 or im
by sigg3 4y ago
It's using base64 encoded strings to deliver the initial stage. Can this be avoided/flagged more easily if by adding a scan of statements featuring base64 or import?
- banana_giraffe 4y agoThat'd catch a ton of valid packages. Right now on my random collection of packages in site-packages I have ~60 packages that have 'import base64' in them.
- sigg3 4y agoYeah, it would probably create more manual work if you have too many false positives. I have maybe six base64 strings in the code I'm working on, so it might be worthwhile looking into provided my legitimate imports don't have any.
- louislang 4y agoYes, this works really well. But as soon as you deploy it, the actors change tactics. We've had to build a defense in depth approach to discovering malicious packages as they are introduced into the system.
- csunbird 4y agoit is probably very easy to bypass, by creating a sub package that can do the decoding via proxy functions, which is not evil at all, and depending on that package on the evil one. It won’t trigger the alarm, as it is indirectly depending on the base64 :)
- louislang 4y agoExactly, I don't think you can rely on a naïve import to determine maliciousness. We've basically had to build out heuristics that are capable of walking function calls for this exact reason. Otherwise things are just too noisy.
- woodruffw 4y agoWe tried doing this on PyPI a couple of years ago, and it produced a large number of false positives (too many to manually review). You can see the rules we tried here[1]. [1]: https://github.com/pypi/warehouse/blob/main/warehouse/malware/checks/setup_patterns/setup_py_rules.yara https://github.com/pypi/warehouse/blob/main/warehouse/malwar...
- deleted 4y ago[deleted]
- fabioz 4y agoThe way I'd go about this is probably starting a VM, installing the package and seeing what in the filesystem is affected by it rather than trying to do static analysis (which becomes a cat and mouse game as detection heuristics improve so do the stealth heuristics). The attack surface area is too big when random python code is executed, which is the case for `setup.py`, but even if there wasn't code executed there, as soon as you import the package and use it, you'd have the same issue.
- bigDinosaur 4y agoUnless you can hide the fact that it's running in a VM, I don't see why code couldn't act normally if it thought it was being analysed like this. Or what about some kind of payload that executes after a long delay, and would become visible for long running programs but not short tests? and so on.
- ashishbijlani 4y agoThis is exactly what Packj [1] scans packages for (30+ such risky attributes). Many packages will use base64 for benign reasons, this is why no fully-automated tool could be 100% accurate. Manual auditing is impractical, but Packj can quickly point out if a package accesses sensitive files (e.g., SSH keys), spawns shell, exfiltrates data, is abandoned, lacks 2FA, etc. Alerts could be commented out if don't apply. 1. https://github.com/ossillate-inc/packj https://github.com/ossillate-inc/packj Disclaimer: I developed this.