4 ms·
I don't know about this scraper, but one thing that mine does (http://www.copymethat.com http://www.copymethat.com) is to "read" the complete page looking for c
by tinebak 9y ago
I don't know about this scraper, but one thing that mine does (http://www.copymethat.com http://www.copymethat.com) is to "read" the complete page looking for certain word combinations that indicate ingredients or steps. It also considers styles and location on the page and looks for keywords that tend to start or end a recipe. It then picks what it considers to be the strongest recipe on the page. This means that it can pick up some weird things if the page doesn't actually contain a recipe; It really wants to find one!
- oelmekki 9y agoInteresting approach, thanks for mentioning it. I guess it means you have a lot of unsuccessful results? Do you try iterate several times on the same page to find different possible sources for a given info and rank them, or is it something more like "if we're not confident enough, forget about that info"?
- tinebak 9y agoThere are hardly any unsuccessful results. (Assuming that the page actually contains a recipe.) People have copied recipes from more than 70,000 websites into their recipe boxes. Of course, I can't check that all the millions of recipes have been accurately copied, but we do check a lot and also get terrific feedback. The parser first goes through all lines/sentences on the page and gives them a rank based on whether it seems to be an ingredient or step. Then it looks at groupings (several steps together) and then the placement of the ingredients compared to the steps.