3 ms·
My team (5 others) and I built an application that would help you revise or learn a subject by testing you about it, where that subject could be anything, from
by codefined 9y ago
My team (5 others) and I built an application that would help you revise or learn a subject by testing you about it, where that subject could be anything, from mathematics to Shakespeare.
It worked by scraping the top 1000 Bing results and then scraping several levels deep from each of the results to generate a "map" of knowledge. This in turn could be used to ask the user questions, generated fill-in-the-blanks, matchups & whatever else you could think up.
It worked after the five-day hackathon, which was truly surprising, but unfortunately, it used too much computing power so we never released it to the world.
- rprameshwor 9y agoI'm interested to know more about the tools/techniques you used to scrap the sites and clean the results. I am working on scraping contents from a couple of sites and it seems to be a pain, probably because i don't have lot of experience around it.
- wasi0013 9y agoIf you are familiar with python then try scrapy[1]. You can also scrape websites using beautifulsoup4[2] [1] https://scrapy.org/ https://scrapy.org/ [2] https://pypi.python.org/pypi/beautifulsoup4 https://pypi.python.org/pypi/beautifulsoup4
- codefined 9y agoWe actually used Node.JS' request module, combined with some NLP (using natural) in order to pick out the main content. This worked pretty well, but for our purposes we didn't need it to be perfect because anything like headers would be removed when we processed the content (not being full sentences).
- hello_newman 9y agoThat sounds like an awesome tool! If you don't mind me asking, is it publicly available to view anywhere? What language is it in? To second another comment, I am also working on scraping but for articles. In this particular example I am trying to figure out how to get the scraper to start on the first headline all the way down to the bottom, and not touch anything else. Any tips for how to get it to find the first heading on a page?
- assafmo 9y agoI would think it depends on the site... maybe some special html characteristics around the first headline? I usually use lynx, which strips down the html and only outputs the text. News site tend to have the same structure for every article, so after passing the html to lynx, the headline for a news site article will probably be around line 5-6.
- laurieg 9y agoHow did you generate the questions automatically? This is something I have tried to do with natural languages in the past but getting good, error free, unambiguous questions out of an algorithm seems to be really difficult.