20 ms·
Show HN: Flyscrape – A standalone and scriptable web scraper in Go
- unixhero 3y agoI will test this, great stuff
- irishgeoff22 3y ago[dead]
- lucgagan 3y agoThis looks great. I wish I had this a few months ago! Giving it a try.
- philippta 3y agoGlad to hear! You’re welcome to leave any feedback on Github (as an Issue) or right in here.
- bryanrasmussen 3y agoLooks like it doesn't have the possibility of running it as a particular browser etc. Which I guess makes it fine for a lot of pages, but also a lot of scraping tasks would be affected. Am I right or did I miss something?
- philippta 3y agoYes, this is correct. As of right now there is no built-in support for running as a browser. What is possible though, is to use a service like ScrapingBee (not affiliated) and set it as the proxy. This would render the page on their end, in a browser.
- _lvbh 3y agoTry tls-client. It gets around TLS fingerprinting by Cloudflare
- snake117 3y agoLooks interesting, and thank you for sharing this! One common issue with scraping web pages is dealing with data that is dynamically loaded. Is there a solution for this? For example, when using Scrapy, you can have Splash running in Docker via scrapy-splash (https://github.com/scrapy-plugins/scrapy-splash https://github.com/scrapy-plugins/scrapy-splash).
- figmert 3y agoCan't you load the URL that is being dynamically loaded directly within your scraper?
- mdaniel 3y agoNot only can you, in my experience it is substantially less drama and arguably less load on the target system since the full page may make many many other requests that a presentation layer would care about that I don't The trade-offs usually fall into: - authing to the endpoint can sometimes be weird - it for sure makes the traffic stand out since it isn't otherwise surrounded by those extraneous requests - it, as with all good things scraping, carries its own maintenance and monitoring burden However, similar to those tradeoffs, it's also been my experience that a full page load offers a ton more tracking opportunities that are not present in a direct endpoint fetch. I mean, look how many "stealth" plugins out there designed to mask the fact that a headless browser is headless But, having said all of that: without question the biggest risk to modern day scraping is Cloudflare and Akamai gatekeeping. I do appreciate the arguments of "but ddos!11" and yet I would rather only actors that are actually exhibiting bad behavior[1] be blocked instead of everyone trying with a copy of python who have set reasonable rate limits 1 = this setting aside that "bad behavior" can be defined as "downloading data that the site makes freely available to Chrome but not freely available to python"
- philippta 3y agoThanks! As mentioned in another comment, currently there is no build in support for this yet. As a workaround one could use a service like ScrapingBee (not affiliated) as a proxy, that renders the page in a browser for you. Surely, relying on a service for this is not always ideal. I am also working on a small wrapper that turns Chrome into an HTTPS proxy, which you could plug right into flyscrape. Unfortunately it is very experimental still and not public yet. I have not yet decided if I release it as part of flyscrape or as a separate project.
- xyzzy_plugh 3y agoI like web scraping in Go. The support for parsing HTML in x/text/html is pretty good, and libraries like github.com/PuerkitoBio/goquery go a long way to matching ergonomics in other tools. This project uses both, but then also goes on to use github.com/dop251/goja, which is a JavaScript VM and it's accompanying nodejs compatability layer and even esbuild, in order to interpret scraping instruction scripts. I mean, at this point I am not sure Go is the right tool for the job (I am actually pretty confident that it is not). A pretty neat stack of engineering, sure! This is cool, niely done. But I can't help but feel disturbed.
- deleted 3y ago[deleted]
- cxr 3y agoYour comment was posted 4 minutes ago. That means you still have enough time to edit your comment to change it so it contains real URLs that link to the project repos for the packages mentioned: <https://github.com/PuerkitoBio/goquery https://github.com/PuerkitoBio/goquery> <https://github.com/dop251/goja https://github.com/dop251/goja> (Please do not reply to this comment of mine—if you do, I won't be able to delete it once the previous post is fixed, because the existence of the replies will prevent that.)
- cheapgeek 3y ago[flagged]
- xyzzy_plugh 3y agoEven if I saw this post in time, I wouldn't have edited it. They are all proper Go package names.
- deleted 3y ago[deleted]
- deleted 3y ago[deleted]
- sunshadow 3y agoThese days, I'm not even using Go for scraping that much, as the webpage changes makes me crazy and JS code evaluation is a lifesaver, so I moved to Typescript+Playwright. (Crawlee framework is cool, while not strictly necessary). Its been 8+ years since i started scraping. I even wrote a popular Go web scraping framework previously: (https://github.com/geziyor/geziyor https://github.com/geziyor/geziyor). My favorite stack as of 2023: TypeScript+Playwright+Crawlee(Optional) If you're serious in scraping, you should learn javascript, thus, playwright should be good. Note: There are niche cases where lower-level language would be required (C++, Go etc), but probably only <%5
- mikercampbell 3y agoHave you seen Crul?? I love the JS flow, but I thought crul was an interesting newer tool!! But I agree, you gotta get in there and it’s easier with JS
- reyostallenberg 3y agoCan you add a link to it?
- mdaniel 3y agoI'm sorry to hear that your searches for that very specific name didn't provide the information you were looking for its show hn: https://news.ycombinator.com/item?id=34970917 https://news.ycombinator.com/item?id=34970917 tfl: https://www.crul.com/ https://www.crul.com/
- sunshadow 3y agoCrul looks nice, though, you cannot imagine how many startups that I've seen failed doing a very similar thing as Crul. Wouldn't rely on it. The problem is complex: Humans generating messy pages
- docyes 3y agoThank you for the positive acknowledgment and insightful observation. As one of the creators of Crul, I fully understand the challenges inherent in this intricate business and software domain. Our initial emphasis on the browser abstraction layer, predating APIs such as SOAP, REST, GraphQL, etc., serves as a data driver and stateless cluster for interpreting DOM nodes. While we initially lacked programmatic extensibility for custom browser control, as you rightly pointed out, addressing complex edge cases often requires such a feature. Looking ahead, we are exploring the possibility of opening up the core, starting with "Krull," the browser cluster. We welcome feedback to gauge interest in this development.
- slig 3y agoThanks for sharing! Just a small nit: the links at the bottom of this page are broken [1]. [1]: https://github.com/philippta/flyscrape/blob/master/docs/readme.md#configuration https://github.com/philippta/flyscrape/blob/master/docs/read...
- philippta 3y agoThanks for spotting, will get this fixed quickly.
- fyzix 3y agoWhat happens if 'find()' returns a list and you call '.text()'. Intuition tells me it should fail but maybe it implicitly gets the text from the first item if it exists. Either way, I think you create a separate method 'find_all()' that returns a list to make the API easier to reason about.
- moehm 3y agoInteresting. Can you compare it to colly? [0] Last time I looked it was the most popular choice for scraping in Go and I have some projects using it. Is it similar? Does it have more/less features or is it more suited for a different use case? (Which one?) [0] https://github.com/gocolly/colly https://github.com/gocolly/colly
- philippta 3y agoColly is a great scraping library if you are a Go developer. Flyscrape on the other hand is a ready-made CLI tool that aims to be easy to use even for someone who is a little familiar with JavaScript. It just happens to be written in Go, but that should not matter to the end user. It does not have full feature parity with Colly but most use cases should be covered.
- 1vuio0pswjnm7 3y ago"default = 100 [requests per second]" How many new TCP connections per second. Is this a "scraper" or a "crawler". It appears to accept a "starting URL" and to follow links. Opening many TCP connections is arguably still a reason why website operators try to prevent crawling (except from Googlebot IPs). As for scraping, it can be done with a single TCP connection. Perhaps "developers" instead opt to use many TCP connections and then complain when they get blocked.
- oefrha 3y agoI had a look at https://github.com/philippta/flyscrape/blob/master/scrape.go https://github.com/philippta/flyscrape/blob/master/scrape.go. It’s just using the builtin HTTP client to fire off requests, with an identifying user agent. Which means it’s useless for scraping most real world sites you may want to scrape, unfortunately. You’ll get served a JavaScrpt challenge, or not even that (many sites will refuse to serve anything if they see a random user agent like flyscrape/1.0).
- 1vuio0pswjnm7 3y agoWhat are some examples of "most real world sites". What are some examples of sites that are not "most real world sites". Is HN a "real world site". What percentage of sites submitted to HN are "most real world sites". (IME, it's a minority fraction.) Why not just delete or change the user-agent line in scrape.go before compiling. (Personal experience: I have been successfully retrieving information from the www for decades without including a user-agent header. The number of sites I have found that require this header is relatively small. It does not rise to the level of "most".) HN replies usually fail to include even a single example. Similarly, do examples provided for scraper programs and libraries ever include "most real world sites".
- specproc 3y agoI've not looked at the source code, but if GP is correct, then absent JS rendering means there's little added value for me (a dude who scrapes a lot). Real world example, I was looking at scraping unjobs.org for a friend the other day. The need for JS rendering turned the job from 15 minutes of requests and beautifulsoup into a full-blown session with selenium, geckodriver etc. I'm not saying the linked framework isn't nice, I've not looked at it, but tools for simple scraping are plenty and easy to use. There's a lot more that a new framework needs to do to distinguish itself. I'd love something that makes JS rendering, proxy-rotation and catchpa solving easier, in a nice package I can deploy myself.
- 1vuio0pswjnm7 3y agoThank you for providing an example. As expected, retrieving the jobs listings from unjobs.org is trivial. Below is a simple demonstration using only common UNIX utilties and minimising the number of TCP conections. No browser. No Javascript. No Selenium. No Geckodriver. No proxies. Step 1 requires 40 TCP connections and completes in under a minute. Step 2 requires one TCP connection and completes in 10min. (NB. The connection minimisation used here, what RFCs used to call "web etiquette", is not possible using a popular headless graphical browser.) 1.htm is 1.5M, 2.htm is 40M # step 1 x="some user-agent string" # e.g., https://raw.githubusercontent.com/51Degrees/Device-Detection/master/data/20000%20User%20Agents.csv n=1;while true;do test $n -le 40||break # unjobs.org only shows listings 1-1000 case $n in 9|17|25|33)sleep 10;esac # need a delay after every 8 requests echo "GET /new/$n HTTP/1.0@Host: unjobs.org@User-Agent: $x@" \ |tr @@ '\r\n' \ |openssl s_client -connect unjobs.org:443 -ign_eof -servername unjobs.org n=$((n+1)); done > 1.htm # step 2 n=1000; grep -Eo -m$n /vacancies/1[0-9]\{12\} 1.htm \ |tail -$n \ |sed "\$!s>.*>GET & HTTP/1.1@Host: $x@Connection: keep-alive@>; \$s>.*>GET & HTTP/1.1@Host: $x@Connection: close@>" \ |tr @@ '\r\n' \ |openssl s_client -connect unjobs.org:443 -ign_eof > 2.htm If provided with an example of what the formatted output should loook like, I will demonstrate how to do it quickly and easily, without Python. Quite sure the text processing methods I use to extract data from HTML are faster than Python.
- krick 3y agoThis looks like something I could use. Maybe not revolutional, but I do that from time to time, and even if only for organizational purposes it seems to make sense to store that stuff as a bucnh of configuration files for some external tool, rather than a bunch of python-scripts that I implement somewhat differently every time. Right now I'm just wrapping my head around how this works, and didn't try it hands-on yet, but I struggle to evaluate from the existing documentation, how useful this actually is. All examples in the repository right now are ultimately one-page scrappers, which, honestly, would be quite useless to me. Pretty much every scraper I write has at least 2-3 logical layers. Like, consider your HN-example, but you want to include top-10 comments for each post. Is it even possible? Well, I guess for HN you could just get by using allowedURLs and treating default function as a parser for the comment-page, but this isn't generic enough. Consider some internet shop. That would be (1) product category tree, sometimes much easier to hard-code, rather than scrape it every time; hard-coding often is generative (e.g. example.com/X/A-B-C, where X is a string from the list, A, B and C are padded numbers, each with a different range) (2) you go into each category, retrieve either a sub-category list (possibly, js-rendered, multiple pages) or product list (same applies) (3) open each product url, do the actual parsing (name, price, specification, etc). Each of json-object from (3) often has to include some minimal parsed data from level (2) (like category name) More advanced, but also way to popular to imagine a generic web-scraper without it: in addition to some json-metadata you download pictures, or pdf-files, etc. (Sometimes you don't even need metadata.) Maybe just text files, but the result is several GBs, and isn't suitable to be handled as a single json-object, but rather a file/directory tree. Is any of this possible with this tool? Also, regardless of being it useful for my cases, some minor comments: 1. Links in docs/readme.md#configuration don't work (but the .md files for them actually exist). 2. I would suggest making "url" in the configuration either a list, or string|list. I suppose, that pretty much doesn't change the logic, but would make a lot of basic use-cases much easier to implement.