6 ms·
Show HN: Use cookies from Chrome (CDP) in cURL without copy pasting
- juujian 4y agoI have done a lot of scraping in the past. Cookies are a pain, this is a really elegant solution. Of course the biggest problem is that everything interesting is hidden away behind JavaScript these days and then you have to resort to Selenium and the whole thing just spirals out of control. But I'm looking forward to giving this a shot for non-JavaScript content in the future. edit: JavaScript not Java
- mkl 4y agoDo you mean JavaScript? I have never run into content hidden by Java, but many pages load content dynamically using JavaScript. I have found it's quite easy to snoop on those JavaScript API requests using the Network tab of Chrome Devtools, then copy the network request as a curl command for bash scripts or as JavaScript for browser extensions.
- ghqst 4y agoYeah, you can sometimes find the API or find data sent in JavaScript but not in prerendered HTML, which can save you the pain of headless scraping.
- juujian 4y agoI do mean JavaScript. Not sure how many times I have made that mistake... And great advice, that sounds like a neat approach.
- tomashubelbauer 4y ago> I have never run into content hidden by Java Tongue in cheek: You'd never know - servers running Java code generating HTML pages have probably conditionally not-rendered many pieces of HTML that you've never come across in your browsing :)
- totetsu 4y agoThere are python libraries you can use that import cookies directly from wherever your browsers stores them to use in selenium projects.
- berkle4455 4y agoJavascript is delivered as text and sends text-based HTTP calls to the server to fetch more data. Why do you need selenium?
- rhd 4y agoI've once used Selenium to run javascript in the webpage to steal a few dynamic tokens required by the sites API to reuse in my more well-trodden python-requests workflow.
- LelouBil 4y agoif you don't want to reverse engineer the javascript
- KomoD 4y agoMost of the time you don't need to, just open up devtools, look at the network tab, locate the right request(s).
- deleted 4y ago[deleted]
- bdcravens 4y agoIf you'd be standing up CDP to grab the cookies, you'd probably use Puppeteer or Playwright instead of Selenium.
- juujian 4y agoAppreciate the recommendation, I just used whatever python had to offer, Puppeteer looks promising though!
- bdcravens 4y agoUsing the tools at hand is often the best approach. That said, I've spent most of the last 13 years of my career automating browsers. For years, I used Selenium with a variety of libraries. After switching to Puppeteer/Playwright, I have zero interest in going back lol. Playwright actually has first party Python support. (Puppeteer has a port called Pyppeteer, but it's no longer maintained and the author recommends using Playwright) https://playwright.dev/python/ https://playwright.dev/python/
- rgrieselhuber 4y agoI second Playwright, it's amazing.
- robertlagrant 4y agoThird.
- 1vuio0pswjnm7 4y agoThe term "everything interesting" is of course subjective. What is interesting to person A might not be interesting to person B. I never use Selenium and I generally have no problem acessing "everything interesting". The simplest example is reading and submitting HN comments. Presumably we all find this interesting enough. Javascript is neither required to read, vote nor submit to HN. What if the phrase "everything interesting" was replaced with specific examples and questions. Something like, "I cannot access X without Javascript. How do I access X without using Javascript."
- RockRobotRock 4y agoHN is in the minority of websites in that it works completely without JS. Surely you're aware of this, right?
- 1vuio0pswjnm7 4y ago1. Define "works". 2. Provide examples of sites that do not "work". It's possible that people might disagree on the definition of "works". For example, perhaps web developers might be biased toward a definition that puts them in control instead of the user. If I can retrieve information from a server with HTTP requests then the website "works" for me. As a user, I certainly do not need to use Javascript to make HTTP requests. Nor do I need to use a particular client. One could argue that even HN does not "work" completely without Javascript. For example, the script at https://news.ycombinator.com/hn.js https://news.ycombinator.com/hn.js will not run.
- thrdbndndn 4y agoFeel like you can just read Chrome's cookie from the file (and filter out the ones you need by site, of course) so you don't need to bother run chrome in debugging mode? Like https://github.com/borisbabic/browser_cookie3 https://github.com/borisbabic/browser_cookie3
- toomuchtodo 4y agoyt-dlp does this also. https://news.ycombinator.com/item?id=28320666 https://news.ycombinator.com/item?id=28320666
- thrdbndndn 4y agoThanks for the link. I know yt-dlp does, but from your link I found another library (https://github.com/n8henrie/pycookiecheat https://github.com/n8henrie/pycookiecheat) that can do that and it seems more popular than browser_cookie3. (browser_cookie3 works totally fine last time I tried).
- fipso 4y agoThis is awesome. I did not know decrypting chrome's password db is still that easy.
- paulirish 4y agoCookies != Passwords.. But anyway... You know this is also easily accessible within DevTools, yah? https://umaar.com/dev-tips/3-copy-as-curl/ https://umaar.com/dev-tips/3-copy-as-curl/
- eurasiantiger 4y agoOne could argue that cookies need to be more securely stored than passwords, because they can allow an attacker to bypass passwords and all other authentication factors.
- djbusby 4y ago
- 2h 4y ago> Tired of copy pasting cURL commands from chrome to your terminal ? FYI for anyone that does this, MITM Proxy is usually a better option for this type of stuff. Not sure about Chrome, but especially with Firefox, you have no way of getting the full raw request on anything with a request body like POST. You have to Copy Request Headers, then Copy POST Data. With MITM Proxy or similar you can just get the full request at once. Also you can inject headers like X-Forwarded-For into all or specific requests.
- folmar 4y ago> you have no way of getting the full raw request on anything with a request body like POST On FF right click on Request -> Copy Value -> As cURL This gives everything and works with POST since a few years at least.
- cookiengineer 4y agoI've had kind of the same problem in the past. For me I built a cookiejar textfile generating chrome extension, because it turns out most relevant tracking or session cookies are on external domains or oauth provider domains. [1] You just need to copy/paste the generated text content to a cookies.txt and you're set, so it worked for my workflow in the terminal. [1] https://github.com/cookiengineer/me-want-cookies https://github.com/cookiengineer/me-want-cookies
- natpalmer1776 4y agoRepo has a single open pull request where someone (by their own admission) just fed the code into ChatGPT then pasted the results into a pull request without actually testing it. Has anyone seen this happening anywhere else?
- fipso 4y agoHonestly thought it's troll cuz April fools or something, but the guy actually has like 10% usefull suggestions in the PR. Will probably close that one tho.
- crimsonraven 4y agoWoah! I didn't realise it was 1st April when I raised that PR. Sorry if it isn't of much use (as can be expected of straight GPT responses). I didn't mean to come out as a troll. I agree; should've at least tested the change before raising a PR, but I would've just dropped the idea, so felt it best to share my findings at least.
- fipso 4y agoAll good there were some good points in the PR which I will likely adapt into the script. But when you wrote that you ported my 3 line bash script to rust I thought you were kidding.
- malborodog 4y agoHow could this underlying idea be implemented in a python script from a running selenium driver?
- 1vuio0pswjnm7 4y agoNB. Cannot enable remote debugging if using Chrome in Guest mode on ChromeOS. Why is left as question for the reader. A more universal solution, one that does not require enabling websockets, is a localhost-bound forward proxy where HTTP traffic, including cookies, is saved in log files. No need to copy/paste with a GUI or mouse. Can use standard UNIX utilities to work with log files.
- GauntletWizard 4y agoI used to use an extension called "Get Cookies.txt" to do this. Then several more by similar names with the same codebase. I looked at the code myself and didn't see anything up, but they each got removed from the chrome store for security violations.