Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
stummjr
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
1.
▲
by
stummjr
6y ago
This repo may be helpful to you: https://github.com/satwikkansal/wtfpython It’s a collection of snippets to explain behaviors that may be considered unexpected.
2.
▲
by
stummjr
7y ago
I disagree with this article in so many ways. I have a CS degree and I work with many people who don’t, and they are just as good as (or even better than) me at all the points raised by the article. They meet deadlines, they are incredibly
3.
▲
by
stummjr
9y ago
Ngrok is just awesome! A huge shout out to the developers!
4.
▲
How to Run Regular Python Scripts in Scrapy Cloud
(blog.scrapinghub.com)
3 points
by
stummjr
10y ago
|
0 comments
5.
▲
by
stummjr
10y ago
Crawl-delay is not in the standard robots.txt protocol, and according to Wikipedia, some bots have different interpretations for this value. That's why maybe many websites don't even bother defining the rate limits in robots.txt.
6.
▲
by
stummjr
10y ago
That's kind of what Scrapy's AUTO_THROTTLE middleware does.
7.
▲
by
stummjr
10y ago
Yeah, but that's not just because of web scraping. Plagiarism has been an issue for centuries.
8.
▲
by
stummjr
10y ago
Scrapy is asynchronous, but it provides many settings that you can use to avoid DDoS a website, such as limiting the amount of simultaneous requests for each domain or IP address. And yes, crawling politely requires a bit of effort from bot
9.
▲
How to Crawl the Web Politely with Scrapy
(blog.scrapinghub.com)
139 points
by
stummjr
10y ago
|
41 comments
10.
▲
Introducing Scrapy Cloud with Python 3 Support
(blog.scrapinghub.com)
9 points
by
stummjr
10y ago
|
0 comments
11.
▲
by
stummjr
10y ago
Scrapinghub is 100% remote from day zero. Nowadays there are ~140 people spread around the world, covering almost all timezones.
12.
▲
by
stummjr
10y ago
Hey! I work for Scrapinghub. Feel free to ask any questions.
13.
▲
How to scrape infinite scrolling pages with Scrapy
(blog.scrapinghub.com)
1 points
by
stummjr
10y ago
|
0 comments
14.
▲
Scrapy Tips from the Pros: how to debug your spiders
(blog.scrapinghub.com)
4 points
by
stummjr
10y ago
|
0 comments
15.
▲
by
stummjr
10y ago
Hey, Valdir from Scrapinghub here! Feel free to ask any questions you might have about the platform.
16.
▲
Scrapy Cloud 2.0: run your web crawlers in the cloud
(blog.scrapinghub.com)
8 points
by
stummjr
10y ago
|
0 comments
17.
▲
by
stummjr
10y ago
This is such an inspiring story! I guess a lot of people today go the opposite way, renting or buying places where they know there is a decent internet access. I'm guilty about it. :)
18.
▲
A (not So) Short Story on Getting Decent Internet Access
(blog.scrapinghub.com)
7 points
by
stummjr
10y ago
|
2 comments
19.
▲
by
stummjr
10y ago
Hey, I'm the author. Feel free to ask any questions.
20.
▲
Grok Your Data with the New MonkeyLearn Addon
(blog.scrapinghub.com)
7 points
by
stummjr
10y ago
|
0 comments
21.
▲
How to Become a Whistleblower: From Panama Papers to Open Data
(blog.scrapinghub.com)
9 points
by
stummjr
10y ago
|
0 comments
22.
▲
Web Scraping to Create Open Data
(blog.scrapinghub.com)
176 points
by
stummjr
11y ago
|
55 comments
23.
▲
Scrapy Tips from the Pros, March 2016 Edition
(blog.scrapinghub.com)
3 points
by
stummjr
11y ago
|
0 comments
24.
▲
by
stummjr
11y ago
Hey, of course. We are glad that you are interested in testing Kumo. Please email us (help at scrapinghub . com) your user id, organization ids and the project ids you want to migrate to Kumo. Then we'll get back to you, giving early a
25.
▲
by
stummjr
11y ago
Hey, I'm the author of this post. Feel free to ask any questions or to suggest topics for the next month's post on the "Scrapy tips from the pros" series. :)
26.
▲
Portia: Open-Source Alternative to Kimono Labs
(blog.scrapinghub.com)
53 points
by
stummjr
11y ago
|
1 comments
27.
▲
by
stummjr
11y ago
All sorts of careers, for example: - developers who want to develop some data-based product (a travel agency website, who finds the best deals from airline companies); - lawyers can use it to structure the data from Judgments and Laws (so t
28.
▲
by
stummjr
11y ago
Maybe you should give a try at Portia ( http://scrapinghub.com/portia/ ). It does exactly what you mean. You may also be interested in this library: https://github.com/scrapy/scrapely
29.
▲
by
stummjr
11y ago
Spider Contracts can help you: http://doc.scrapy.org/en/latest/topics/contracts.html
30.
▲
by
stummjr
11y ago
Scrapy can do the job, for sure. We use it to crawl more than 2 billion pages a month. :)
More ›