Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Jake232
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
91.
▲
by
Jake232
12y ago
Here's a comment on the stack overflow answer I linked to in the post: Very performant: (1) Bytecode decompiling is very fast. (2) Since each query has corresponding code object, this code object can be used as a cache key. Because of
92.
▲
BeatScape – BeatDetektor.js + Audio API + Processing.js + CubicVR.js
(cjcliffe.github.io)
1 points
by
Jake232
12y ago
|
0 comments
93.
▲
by
Jake232
12y ago
I'm pretty proficient in a ton of other languages, so I'm hoping I have a pretty good head start and will be able to follow the book, but I'll read Thinking in C++ in the meantime until my order arrives. Thanks for the sugges
94.
▲
by
Jake232
12y ago
Just purchased this. I'm hoping it can make up for a year of not attending my c++ lectures, to learn enough for my exam in 1.5weeks.
95.
▲
by
Jake232
12y ago
That makes sense, thanks.
96.
▲
by
Jake232
12y ago
I've always wondered. How do these guys hit google search so much without hitting limitations? Google is pretty aggressive at banning bots, and I can't imagine Google have given a competitor API access or something like that. Prox
97.
▲
Redis as JSON document store
(matt.sh)
1 points
by
Jake232
12y ago
|
0 comments
98.
▲
by
Jake232
12y ago
I feel this could be easily monetized, if you don't get a sale you should look into it for sure. iTunes have an affiliate system, and they have an API endpoint for search. You could simply have some links to purchase the latest album o
99.
▲
by
Jake232
12y ago
Would also love an answer to this.
100.
▲
by
Jake232
12y ago
Most significantly: "Your data isn't even in the picture. We are simply not interested in any of it." - Pretty sure Facebook are interested in it.
101.
▲
by
Jake232
12y ago
From the Github page: Third, Goji is not magic. One of my favorite existing frameworks is Martini, but I rejected it in favor of building Goji because I thought it was too magical. Goji's web package does not use reflection at all, whi
102.
▲
by
Jake232
12y ago
Or if you're a python enthusiast, then shameless self link: http://jakeaustwick.me/python-web-scraping-resource/
103.
▲
by
Jake232
12y ago
How so? I'm from the UK and most of the people I know would say MS are a reputable brand for sure.
104.
▲
by
Jake232
12y ago
Although this may seem like a massive mistake, that's because we all associate Microsoft as a negative brand, I think you're forgetting that it's associated with negativity in tech circles only. 95%+ of consumers see Micros
105.
▲
by
Jake232
12y ago
The submitted link works for me.
106.
▲
by
Jake232
12y ago
Somebody asked this on the blog post, here was the reply from Linode: If you take the upgrade, you inherit the new plan specs, vcpus and all. We’ve greatly reduced the contention on these new machines compared to our old structure, and in t
107.
▲
by
Jake232
12y ago
The increased CPU on some of the plans compared to DO is what is encouraging me to change. When doing web scraping, it's easy to max the CPU once you are doing 20 paths on a page.
108.
▲
by
Jake232
12y ago
I write my own custom scrapers, I prefer the flexibility and feel safer that the service isn't going to disappear any minute. If anybody is interested, I wrote a detailed article on scraping not so long back that was well received here
109.
▲
by
Jake232
12y ago
I've always wondered why github never had tags.
110.
▲
Heartbleed affects clients too
(jakeaustwick.me)
3 points
by
Jake232
12y ago
|
0 comments
111.
▲
by
Jake232
13y ago
I see no harm in doing so. I have a ton of books in digital formats, just never thought about buying a physical copy before. The only reason I did so this time is because I have some travelling upcoming, so it's ideal to read on the pl
112.
▲
by
Jake232
13y ago
First ever physical copy of a programming book I've bought, lets's hope it's worthwhile.
113.
▲
by
Jake232
13y ago
Sidekiq wasn't around when they made resque.
114.
▲
by
Jake232
13y ago
If you're running on a server with 1Gbps, then yes - it can be an issue. Another issue (I'm going to add a section on it) is that you can peg the CPU at 100% very easily. Parsing HTML / Running xPaths uses a lot of CPU, so if
115.
▲
by
Jake232
13y ago
Hi. I mentioned selenium in the section below that, but I'll drop a note to it in the AJAX section too!
116.
▲
by
Jake232
13y ago
Thanks for the mechanize link, I'll add a note about it. Mechanize is awfully slow though, if you need to crawl quickly it's not asynchronous. I wouldn't want to use it for general crawling. I guess you could patch the stdlib
117.
▲
by
Jake232
13y ago
Thanks, good to know regarding the proxy. There was a couple of other little things that just didn't work the way I wanted though (I honestly don't remember them now). I've built private libraries on top of requests now that
118.
▲
by
Jake232
13y ago
Hey, thanks for the feedback! I'm planning on adding more to the article in the near future, this was just a start. I plan on it being a resource with almost everything in, so people can bookmark it for future use. I really haven'
119.
▲
by
Jake232
13y ago
Hey, I've actually built something similar to this myself, I plan on writing an article in the future with something along these lines. Yours look pretty polished though, good job!
120.
▲
by
Jake232
13y ago
I've wrote a lot of python crawlers, and have never really experienced an issue with lxml. It doesn't really ever seem to choke, and I'm sure I must have come across some pretty funky markup in my time. lxml has always hand
More ›