19 ms·
Curl-Impersonate
- londons_explore 2y ago> The resulting curl looks, from a network perspective, identical to a real browser. How close is it? If I ran wireshark, would the bytes be exactly the same in the exact same packets?
- dchest 2y agoWhat else could "identical" mean?
- londons_explore 2y agoIt could be that the TCP streams are the same, but packetiation is different. It could mean that the packets are the same, but timing is off by a few milliseconds. It could mean a single HTTP request exactly matches, but when doing two requests the real browser uses a connection pool but curl doesn't. Or uses HTTP/3's fast-open abilities, etc. etc.
- zlagen 2y agoIt replicates the browser at the HTTP/SSL level, not TCP. From what I know this is good enough to bypass cloudflare's bot detection.
- Retr0id 2y agoTwo TLS streams are never byte-identical, due to randomness inherent to the protocol. Identical here means having the same fingerprint - i.e. you could not write a function to reliably distinguish traffic from one or the other implementation (and if you can then that's a bug).
- jsnell 2y agoThe packets from Chrome wouldn't be exactly the same as packets sent by Chrome at a different time either. "The exact same packets" is not a viable benchmark, since both the client and the server randomize the payloads in various ways. (E.g. key exchange, GREASE).
- peetistaken 2y agoYou can check your fingerprint on https://tls.peet.ws https://tls.peet.ws
- peetistaken 2y agohttps://github.com/bogdanfinn/tls-client https://github.com/bogdanfinn/tls-client is the go-to package for the go world, it does the same thing
- zlagen 2y agoIn case anyone is interested, I created something similar but for python(using chromium's network stack) https://github.com/lagenar/python-cronet https://github.com/lagenar/python-cronet I'm looking for help to create the build for windows.
- hk__2 2y agoAny reason you didn’t use https://github.com/lexiforest/curl_cffi https://github.com/lexiforest/curl_cffi?
- zlagen 2y agoI wanted to try a diffent approach which is to use chromium's network stack directly instead of patching curl to impersonate it. In this case you're using the real thing so it's a bit easier to maintain when there are changes in the fingerprint.
- Klonoar 2y agoSimilar projects exist for C# (https://github.com/sleeyax/CronetSharp https://github.com/sleeyax/CronetSharp), Go (https://github.com/sleeyax/cronet-go https://github.com/sleeyax/cronet-go) and Rust (https://github.com/sleeyax/cronet-rs https://github.com/sleeyax/cronet-rs). These can work well in some cases but it's always a tradeoff.
- thrdbndndn 2y agoAny plan to offer a sync API?
- Retr0id 2y agoI recently used ja3proxy, which uses utls for the impersonation. It exposes an HTTP proxy that you can use with any regular HTTP client (unmodified curl, python, etc.) and wraps it in a TLS client fingerprint of your choice. Although I don't think it does anything special for http/2, which curl-impersonate does advertise support for. https://github.com/LyleMi/ja3proxy https://github.com/LyleMi/ja3proxy https://github.com/refraction-networking/utls https://github.com/refraction-networking/utls
- TekMol 2y agoWhat is the use case? If you have to read data from one specific website which uses handshake info to avoid being read by software? When I have to do HTTP requests these days, I default to a headless browser right away, because that seems to be the best bet. Even then, some website are not readable because they use captchas and whatnot.
- mschuster91 2y ago> What is the use case? If you have to read data from one specific website which uses handshake info to avoid being read by software? Evade captchas. curl user agent / heuristics are blocked by many sites these days - I'd guess many popular CDNs have pre-defined "block bots" stuff that blocks everything automated that is not a well-known search engine indexer.
- adastral 2y ago> I default to a headless browser Headless browsers consume orders of magnitude more resources, and execute far more requests (e.g. fetching images) than a common webscraping job would require. Having run webscraping at scale myself, the cost of operating headless browsers made us only use them as a last resort.
- TekMol 2y agoSo you maintain a table of domains and how to access them? How do you build that table and keep it up to date? Manually?
- at0mic22 2y agoBlocking all image/video/CSS requests is the rule of thumb when working with headless browsers via CDP
- sangnoir 2y agoSpeaking as a person who has played on both offense and defense: this is a heuristic that's not used frequently enough by defenders. Clients that load a single HTML/JSON endpoint without loading css or image resources associated with the endpoints are likely bots (or user agents with a fully loaded cache, but defenders control what gets cached by legit clients and how). Bot data thriftiness is a huge signal.
- jollyllama 2y ago>The Client Hello message that most HTTP clients and libraries produce differs drastically from that of a real browser. Why is this?
- throwaway99210 2y agoBased on what I've seen, most command-line clients and basic HTTP libraries typically ship with leaner, more static configurations (e.g., no GREASE extensions in the Client Hello, limited protocols in the ALPN extension header, smaller number of Signature Algorithms). Mirroring real browser TLS fingerprints is also more difficult due to the randomization of the Client Hello parameters (e.g., current versions of Chrome)
- Retr0id 2y agoThe protocols are flexible and most browsers bring their own HTTP+TLS clients
- zlagen 2y agoThey use different SSL libraries/configuration. Chrome uses BoringSSL and other libraries may use OpenSSL or some other library. Besides that the SSL library may be configured with different cipher suites and extensions. The solution these impersonators provide is to use the same SSL library and configuration as a real browser.
- cle 2y agoThe same author also makes a Python binding of this which exposes a requests-like API in Python, very helpful for making HTTP reqs without the overhead of running an entire browser stack: https://github.com/lexiforest/curl_cffi https://github.com/lexiforest/curl_cffi I can't help but feel like these are the dying breaths of the open Internet though. All the megacorps (Google, Microsoft, Apple, CloudFlare, et al) are doing their damndest to make sure everyone is only using software approved by them, and to ensure that they can identify you. From multiple angles too (security, bots, DDoS, etc.), and it's not just limited to browsers either. End goal seems to be: prove your identity to the megacorps so they can track everything you do and also ensure you are only doing things they approve of. I think the security arguments are just convenient rationalizations in service of this goal.
- throwaway99210 2y ago> I can't help but feel like these are the dying breaths of the open Internet though I agree with the over zealous tracking by the megacorps but this is also due to bad actors, I work for a financial company and the amount of API abuse, ATO, DDoS, nefarious bot traffic, etc. we see on a daily basis is absolutely insane
- berkes 2y agoBut how much of this "bad actor" interaction is countered with tracking? And how many of these attempts are even close to successfull with even the simplest out of the box security practices set up? And when it does get more dangerous, is over zealous tracking the best counter for this? I've dealt with a lot of these threats as well, and a lot are countered with rather common tools, from simple fail2ban rules to application firewalls and private subnets and whatnot. E.g. a large fai2ban rule to just ban anything that attempts to HTTP GET /admin.php or /phpmyadmin etc, even just once, gets rid of almost all nefarious bot traffic. So, I think the amount of attacks indeed can be insane. But the amount that need over zealous tracking is to be countered, is, AFAICS, rather small.
- throwaway99210 2y ago
- oefrha 2y agoWhat are some example sites where this is both necessary and sufficient? In my experience sites with serious anti-bot protection basically always have JavaScript-based browser detection, and some are capable of defeating puppeteer-extra-plugin-stealth even in headful mode. I doubt sites without serious anti-bot detection will do TLS fingerprinting. I guess it is useful for the narrower use case of getting a short-lived token/cookie with a headless browser on a heavily defended site, then performing requests using said tokens with this lightweight client for a while?
- jonatron 2y agoThere are sites that will block curl and python-requests completely, but will allow curl-impersonate. IIRC, Amazon is an example that has some bot protection but it isn't "serious".
- ekimekim 2y agoIn most cases this is just based on user agent. It's widespread enough that I just habitually tell requests not to set a User Agent at all (these aren't blocked, but if the UA contains "python" it is).
- Retr0id 2y agoA lot of WAFs make it a simple thing to set up. Since it doesn't require any application-level changes, it's an easy "first move" in the anti-bot arms race. At the time I wrote this up, r1-api.rabbit.tech required TLS client fingerprints to match an expected value, and not much else: https://gist.github.com/DavidBuchanan314/aafce6ba7fc49b19206bd2ad357e47fa https://gist.github.com/DavidBuchanan314/aafce6ba7fc49b19206... (I haven't paid attention to what they've done since so it might no longer be the case)
- oefrha 2y agoMakes sense, thanks.
- Avamander 2y ago
- ape4 2y agoI like this project! Is there a way to request impersonization of the current version of Chrome (or whatever)?
- jakeogh 2y agoThe latest version is a moving target, currently you get the following chrome versions: $ curl_chrome <TAB><TAB> curl_chrome100 curl_chrome101 curl_chrome104 curl_chrome107 curl_chrome110 curl_chrome116 curl_chrome119 curl_chrome120 curl_chrome123 curl_chrome124 curl_chrome131 curl_chrome131_android curl_chrome99 curl_chrome99_android
- ape4 2y agoPerhaps plain `curl_chrome` could use the latest available `curl_chromeNNN`
- aninteger 2y agoI think we should list the sites where this fingerprinting is done. I have a suspicion that Microsoft does it for conditional access policies but I am not sure of other services.
- Galanwe 2y agoWe cannot really list them, as 90% of the time, it's not the websites themselves, it's their WAF. And there is a trend toward most company websites to be behind a WAF nowadays to avoid 1) annoying regulations (US companies putting geoloc on their websites to avoid EU cookie regulations) and 2) DDoS. It's now pretty common to have cloudflare, AWS, etc WAFs as main endpoints, and these do anti bots (TLS fingerprinting, header fingerprinting, Javascript checks, capt has, etc).
- pixelesque 2y agoCloudflare (which seems to be fronting half the web these days based off the number of cf-ray cookies that I see being sent back) does this with bot protection on, and Akamai has something similar I think.
- Sytten 2y agoThankfully only a small fraction of website does JA3/JA4 fingerprinting. Some do more advanced stuff like correlating headers to the fingerprint. We have been able to get away without doing much in Caido for a long time but I am working on an OSS rust based equivalent. Neat trick, you can use the fingerprint of our competitor (Burp Suite) since it is whitelisted for the security folks to do their job. Only time you will not hear me complain about checkbox security.
- jandrese 2y agoThe build scripts on this repo seem a bit cursed. It uses autotools but has you build them in a subdirectory. The default built target is a help text instead of just building the project. When you do use the listed build target it doesn't have the dependencies set up correctly so you have to run it like 6 times to get to the point where it is building the application. Ultimately I was not able to get it to build because the BoringSSL disto it downloaded failed to build even though I made sure all of the dependencies the INSTALL.md listed are installed. This might be because the machine I was trying to build it on is an older Ubuntu 20 release. Edit: Tried it on Ubuntu 22, but BoringSSL again failed to build. The make script did work better however, only requiring a single invocation of make chrome-build before blowing up. Looks like a classic case of "don't ship -Werror because compiler warnings are unpredictable". Died on: /extensions.cc:3416:16: error: ‘ext_index’ may be used uninitialized in this function [-Werror=maybe-uninitialized] The good news is that removing -Werror from the CMakeLists.txt in BoringSSL got around that issue. Bad news is that the dependency list is incomplete. You will also need libc++-XX-dev and libc++abi-XX-dev where the XX is the major version number of GCC on your machine. Once you fix that it will successfully build, but the install process is slightly incomplete. It doesn't run ldconfig for you, you have to do it yourself. On a final note, despite the name BoringSSL is huge library that takes a surprisingly long time to build. I thought it would be like LibreSSL where they trim it down to the core to keep the attack surface samll, but apparently Google went in the opposite direction.
- at0mic22 2y agoPlayed this game and switched to prebuilt libraries. Think builder docker images have also been broken for a while.
- 38 2y agothat's exactly why I stopped using C/C++. building is many times a nightmare, and the language teams seems to have no interest in improving the situation
- userbinator 2y agoLook on the bright side: the harder it is to build and use correctly, the harder it is for the enemy to analyse and react.
- ElizabethVio52 2y ago[dead]
- kerblang 2y agoInteresting in light of another much-discussed story about AI scraper farms swamping/DDOSing sites https://news.ycombinator.com/item?id=42549624 https://news.ycombinator.com/item?id=42549624
- userbinator 2y agoI can't help but think that projects like these shouldn't be posted here, since the enemy is among us. Prodding the bear even more might lead to an acceleration towards the dystopia that others here have already prophesised. The following browsers can be impersonated. ...unfortunately no Firefox to be seen. I've had to fight this too, since I use a filtering proxy. User-agent discrimination should be illegal. One may think the EU could have some power to change things, but then again, they're also hugely into the whole "digital identity" thing.
- crtasm 2y agoIt says "Firefox(In progress)", and the original project this was forked from has it: https://github.com/lwthiker/curl-impersonate https://github.com/lwthiker/curl-impersonate
- ospider 2y agoMaintainer here. Curl drops NSS support since like a year ago, which is the SSL engine firefox uses. Without NSS, two special extensions can not be added. And that's why only webkit-based browsers are left. You can find support for old firefox versions in the original repo.
- jakeogh 2y ago(very rough) ebuild: https://github.com/jakeogh/jakeogh/blob/master/net-misc/curl-impersonate/curl-impersonate-9999.ebuild https://github.com/jakeogh/jakeogh/blob/master/net-misc/curl...
- 0x676e67 2y agoI think someone should need this. It is based on boring tls and makes some fake extensions similar to utls to support Firefox TLS fingerprint imitation repo: https://github.com/penumbra-x/rquest https://github.com/penumbra-x/rquest