7 ms·
I'm curious - What is the utility of headless browsers? Are there people who earn money by getting it to automatically fill out forms, enter competitions etc?
by redsummer 9y ago
I'm curious - What is the utility of headless browsers?
Are there people who earn money by getting it to automatically fill out forms, enter competitions etc?
- kasparsklavins 9y agoAutomated tests.
- askmike 9y agoAs well as scraping more JS heavy stuff.
- kowdermeister 9y agoServer side rendering, testing, automation just to name a few.
- praptak 9y agoWeb scraping for sites that discourage simple robots by checking for JS or serving content via JS. Also, the darker stuff like click fraud and all the other kinds of fraud where you pretend there are humans doing something when in fact there's just a bot.
- ajholyoake 9y agoI have used them in the past to convert graphs and reports to JPGs and PDFs so that I can automatically email them to people in the company who are unable (or unwilling) to use a web page.
- d--b 9y agoI use it to automate turning webpages in pdf.
- avaer 9y agoTesting, rendering, scraping, streaming, botting, and many more. Think "what could I accomplish with a browser, with the slow human replaced by a fast program" and let your imagination run wild. The space is interesting enough that people have jumped through a lot of hoops to make it work in the past; this makes it one less hoop. Oh, and if you ever wondered why web captchas are a thing, one of the reasons is headless browsers.
- mvitorino 9y agoServer side automation of multiple kinds (not just tests): screenshots, advanced crawlers, etc.
- demonshalo 9y agoA great example would be PDF generation for things like invoices. Rather than generating a pdf with something like PHP or Java. Render a regular html page with all the css you want (super easy compared to drawing a PDF using PHP) and then proceed to use a pdf printer on that page. You could run such a thing as a microservice using a headless browser or PhantomJS. There are probably better ways to do this but that's one of the first things that popped into my head!
- tjelen 9y agoThe Webkit/PhantomJS PDF export actually supports SVG embedding (as real vectors), Webfonts and many other things. It's possible to create pretty advanced layouts with maps, graphs using that. Even embedding IFrames with e.g. Google Maps works.
- asdfgadsfgasfdg 9y agoThis sounds horrible.. In general all these html->pdf ways of generating pdfs sounds horrible -- why don't people use latex for this?
- demonshalo 9y agoIt is NOT horrible. However, Latex is another way of doing things. Perhaps your requirements should dictate which of the two you should go for!
- SigmundA 9y agoBecause you probably already did the layout work in HTML to display on screen to the user and now just want a PDF version of it. Or you can redo the layout in latex and maintain two layouts. The full print css is actually pretty complete, problem is the only browser that fully supports it is PrinceXML. None of the major browsers seem to care much about print layout.
- asdfgadsfgasfdg 9y agoBut HTML layout is very different to page based layout. HTML is responsive and has no concept of pagination. PDF is paginated and has no concept of responsiveness.
- etatoby 9y agoDownloading any kind of web page where a simple wget or curl turns out empty, for instance anything made with React or other advanced JS frameworks.
- hakanensari 9y agoWe automate buying things on sites like Amazon.
- redsummer 9y agoCan you give me more details? Is it like some automated stock control system which orders from Amazon when stock is low?
- wfunction 9y agoWhile we're on the topic, does anyone know where one might find scripts to scrape bank statements so you don't have to download them manually every month? (This is one thing I would find headless browsers useful for...)
- pricechild 9y agoEvery UK bank I've used allows some format of csv/qif/ofx export. Are you in the US?
- wfunction 9y agoThe format isn't the issue. The problem is I want it to be API-friendly so I don't even have to think about it; my system should download it automatically. But yes, I'm talking about the US.
- vidarh 9y agoThey generally do. But many UK banks make it hard to automate things. Either insisting on 2FA in all cases, or having a secondary login without 2FA that only gives very limited access. Some way of authorising API access to read-only access to things like statements would be fantastic, to the extent that I'd consider changing banks over it, if you know of any UK banks that offer it.
- thibaut_barrere 9y agoAs well, most bank OFX/CSV exports I've dealt with are truncated in some ways (e.g. truncated labels), which make it harder to really leverage sometimes.
- kingosticks 9y agoMonzo will be offering current accounts soon. https://monzo.com/blog/2017/04/05/banking-licence/ https://monzo.com/blog/2017/04/05/banking-licence/
- pricechild 9y agoAhh of course. Apologies for missing that.
- avip 9y ago>Are there people who earn money by getting it to automatically fill out forms, enter competitions etc? Yes. Tiny ex. https://news.ycombinator.com/item?id=1165680 https://news.ycombinator.com/item?id=1165680
- LoSboccacc 9y agoWe have a complex matrix of layouts and styles an user can chose from and we need to test them across all browser to make sure any improvement doesn't mess with others. It's way cheaper to launch a browser headless at all the resolution we need and grab screenshot to visually compare them at glance, instead of goin one by one at hand
- vishbar 9y agoWe used it a lot for full automation tests for the UI. It's nice being able to interface with a full-featured browser that can run javascript, etc. And take screenshots when things go wrong.
- tomp 9y agoI wrote a PhantomJS script to download data from my bank accounts. They offer no API and disfunctional text-based exports, their websites are ridden with "good" (=terrible) web & security practices, like 3-characters-of-the-whole-password authentication, single-tab sessions, frames, etc. that makes it pretty much impossible to scrape with Python but relatively easy with a fully-fledged browser (although that still requires a lot of bank-specific boilerplate code).
- slig 9y agoIf your bank has a mobile app, it might be easier to MITM and figure out their API and use it directly.
- cr0sh 9y agoThat actually sounds like it could run you into legal issues (or worse), depending on your location (ie - access to a computer system without permission; they give you permission to use the app on the phone, but maybe not to use the API directly). YMMV.
- sjtgraham 9y agoWho do you bank with?
- tomp 9y agoNatWest, Halifax and AMEX (not a bank, but I want my account data from there as well).
- sjtgraham 9y agoI can't help you with Halifax or AMEX (yet), but my company (https://teller.io/ https://teller.io/) has a Natwest API in production (private beta). If you would like access, please ping me. sg -at- teller.io
- camus2 9y agoautomated testing and web-scraping mostly.
- thibaut_barrere 9y agoMy favorite use is automated browser testing. Examples: - in Ruby, poltergeist (https://github.com/teampoltergeist/poltergeist https://github.com/teampoltergeist/poltergeist) - in Elixir, hound (https://github.com/HashNuke/hound https://github.com/HashNuke/hound)
- acdha 9y agoI use https://github.com/bfirsh/needle/blob/master/README.md https://github.com/bfirsh/needle/blob/master/README.md for automated UI regression testing. Using a headless browser means your test suite can run faster and with fewer dependencies.
- tzs 9y agoA few little personal things I've done with PhantomJS: • A script that would go to Comcast's TV schedule for my area and make a list of all movies upcoming in the next two weeks on all channels that are included in my subscription. I could then grep that for a list of movies I've been looking for. I couldn't just grab the page with curl and parse it, because JavaScript does most of the work. JavaScript fetches the listings, and when you advance the listings it it fetches the new listings and replaces the old ones on the page. • A script that goes to the FCCs license information site and gets a list of all ham radio callsigns issued recently [1]. • A script that given a URL to a tactics problem on lichess gets the FEN for the position. I'd use this if I was doing tactics training there on my iPad and did not understand why my answer was wrong or why their answer was right. I'd mail myself a link to the problem, and then later on my desktop I'd give that URL to this script and it would go to lichess to the problem page, and then from there to the board editor page for that position, and grab and give me the FEN, which I could then use to set up the position in Stockfish to analyze. (This is no longer useful. They have made some changes at lichess and now they have a browser-based version of Stockfish on the problem pages, so I can answer my questions right there). • A script that goes to everquest.com and gets the server population levels from the server population display on that page. I don't think that there was anything in this one that actually needed a headless browser. As far as I recall it could have all been done with getting the page with curl and parsing it. It was just easier to do it in JavaScript using the DOM. (The lichess one may also have been that way). [1] https://github.com/tzs/todays_hams https://github.com/tzs/todays_hams
- dec0dedab0de 9y agoAt work I was given permission by a vendor to screen scrape their site while they worked on building a real API. This site was extremely dependent on javascript. Including doing some really complex token passing between multiple domains that the company owned. Not to mention all of their js was minified and uglified so I had a very hard time understanding what it was doing. It was the first time I wasn't able to successfully reverse engineer a site enough to scrape what I needed with just requests/beautifulsoup. I was however able to get it working just fine using phantomjs via selenium via splinter. It was a fun exercise, but part of me still feels like it was cheating.
- holtalanm 9y agoAutomated browser testing on a CI server that doesn't have a gui.
- ble 9y agoI used PhantomJS as part of a report generation pipeline which served no HTTP requests and contacted no outside servers. We made PDFs and ready-to-email, single-file HTML reports with some minimal interactive features. (Ready-to-email, single file == all images turned into data URIs, styles inlined, for HTML files sent as attachments) PhantomJS loaded up an HTML file written earlier in the pipeline. The HTML consisted of a big slug of JSON containing all the relevant data (which would vary from one run to another) and a bunch of scripts and templates (which were fixed for any given report type). The scripts built into the HTML file would chew up the JSON slug and build up the DOM required for the report. Then the PhantomJS script would identify all the images in the DOM and replace all of them with data URIs, strip out the JSON slug to prevent giving away more data than contained in the DOM, and strip out all of the templating JavaScript, leaving behind only the JavaScript needed for the interactive features, which was inlined. We went with PrinceXML for PDF generation. I was briefly nervous because I saw people praising PhantomJS' pdf generation capabilities... but then I saw the people saying, "we used PhantomJS for pdf generation, we used wkhtmltopdf, then we just paid some money to get something that wouldn't produce weird output some of the time." CSS Paged Media Module FTW, y'all.