4 ms·
Just curious what technologies you are using? Also how quickly will this update on the 1st when a new thread is posted...again - just curious. Great job! Mu
by SanjayUttam 14y ago
Just curious what technologies you are using? Also how quickly will this update on the 1st when a new thread is posted...again - just curious.
Great job! Much better than ctrl-f-ing my way through postings [just like you were doing].
- also_on_sunday 14y agoThanks. To answer your first question: I'm using a custom scraper to get the results. I request the job pages through a random http proxy every so often, parse them, and update the cached json of all the comments. Getting a reliable, free, machine-friendly list of http proxies is harder than I thought. It was automatic, but the proxy provider I was using went down (doh!) so now I manually pull one from a list and trigger the scrape. On the server I'm using a tiny express app (node.js) + redis to store hidden posts per use based on a cookie. On the browser just a little bit of twitter boostrap + jquery + some custom UI helpers to generate the page. The searching and hide/show is done within the browser itself. I try to involve the server as little as possible. To answer your second question: It will update as soon as I kick off another scrape. I hope to re-enable the automation on this soon so it will only lag 10-15 minutes behind the actual page. Right now it could be quite a while (up to 12 hours). Yeah! ctrl-f was tough. Especially as the comments spanned multiple pages and weren't sortable by date.
- typpo 14y agoWhy is it necessary to use a proxy?
- also_on_sunday 14y agohacker news black listed the ip address of my server after scraping 1 post's comments about once every fifteen minutes for a day. The best way I could think to get around this was an http proxy.
- etcet 14y agoYou can use this link http://news.ycombinator.com/unban?ip=<your http://news.ycombinator.com/unban?ip=<your ip> to get your IP unbanned. I think it only works once though.
- shurane 14y agoVery awesome. Got any more projects up your sleeve? Is it better to have an xml feed and load that instead of a JSON file? I wrote a scraper that dumped university courses at my school to a JSON file. I would load that JSON file and load the elements to a table. Mine was pretty dang slow. But your site holds up fine. Here's my site if you're at all interested: http://lo.leet.la/jola/bootstrap/docs/sunny.html# http://lo.leet.la/jola/bootstrap/docs/sunny.html# Relvant JS code: http://lo.leet.la/jola/bootstrap/docs/assets/js/sunny.js http://lo.leet.la/jola/bootstrap/docs/assets/js/sunny.js
- SanjayUttam 14y agoThanks for the info !