4 ms·
anyone knows of something equivalent to Dapper out there? I really wish there was since I need this for a project. Thanks!
by danielnicollet 16y ago
anyone knows of something equivalent to Dapper out there? I really wish there was since I need this for a project. Thanks!
- yurylifshits 16y agohttp://webnumbr.com http://webnumbr.com is like dapper for numbers
- mnutt 16y agoIf you're pretty technically inclined and know your way around FireBug/Webkit Inspector, YQL (Yahoo Query Language) is very convenient. It lets you use css selectors to grab data and returns it in JSON or XML. It's great for quick and dirty hacks, but the big question is how long Yahoo will allow it to stick around.
- gtani 16y agothere's (sort of) related (for-pay) web apps as well: http://www.metafy.com/index.html http://www.metafy.com/index.html http://www.diffbot.com/howitworks http://www.diffbot.com/howitworks http://www.aignes.com/ http://www.aignes.com/ http://sharedcopy.com/public/andthensome http://sharedcopy.com/public/andthensome http://webcache.googleusercontent.com/search?q=cache:L5dhj2wQR4gJ:grid.orch8.net/tools/+site:http://grid.orch8.net/tools/&cd=1&hl=en&ct=clnk&gl=us&client=firefox-a http://webcache.googleusercontent.com/search?q=cache:L5dhj2w...
- danielnicollet 16y agoThanks for all those great replies. I will now spend some time reviewing!!
- hartror 16y agoPythoneers have BeautifulSoup, it is fast simple and can deal with real world html. I have used it for site scraping with great success.
- Luyt 16y agolxml.html does a pretty good job too, and offers elementtree and xpath querying. http://codespeak.net/lxml/lxmlhtml.html http://codespeak.net/lxml/lxmlhtml.html Recently I used Beautiful Soup in a very simple program to scrape playlists from Soma.fm: http://www.michielovertoom.com/hobby/somafm-playlists/ http://www.michielovertoom.com/hobby/somafm-playlists/
- andrewljohnson 16y agoFetch, but it's super expensive: http://www.fetch.com/ http://www.fetch.com/
- SudarshanP 16y agohttp://needlebase.com/ http://needlebase.com/ belongs to ITA software being acquired by google... You can find some public datasets at https://pub.needlebase.com/ https://pub.needlebase.com/