35 ms·
Scraping Recipe Websites
- jangstrom 6y agoThis is pretty interesting. I wonder how the recipe parsers from MyFitnessPal or Pinterest compare to this. Sometimes I think they do pretty good, but often they do miss the mark. My guess is on Pinterest they only treat something as a Recipe if it contains the metadata mentioned in the article, and do the easy parse if so. MFP seems to try something a bit more advanced, but I've never been super-impressed with its parsing abilities.
- gklitt 6y agoPleasantly surprised to learn that most recipe sites include structured metadata. Makes sense given the combination of a relatively straightforward schema, and SEO incentive from Google.
- stx 6y agoThis could also be useful for websites that do not print well. I have run into a few occasions where adds and other website elements printed with the actual recipe. The result was a small recipe divided on several pages mostly covered with other content. There were pictures and text formatting that I could not copy out. Often for stuff like that I just pull the HTML and edit until it prints well but I would rather have an easier way.
- kmbfjr 6y agoBut how will I read about "Dakota", an avid yoga enthusiast who just happens to be a mom, who enjoys making healthy and savory meals for her family while blogging? Seriously, I hope this spells an end to the Google ranking imposed nonsense that makes the simple act of searching for a recipe so insufferable.
- tmountain 6y agoIt's a running joke in our house. I start off wanting to make some mashed potatoes, and time and time again, I have to suffer through someone's life story--the camping trip in North Dakota when Susan's husband first discovered his love of homemade sour cream--etc. Makes me wonder if a super barebones recipe site that literally just has recipes and absolutely no fluff would be something people would gravitate towards.
- pjc50 6y agoHaving thought about this and my own use of cookbooks as loose inspiration rather than actually to follow in detail, I have come to the conclusion that the "fluff" is what most of the recipe-reading public want. In particular, the fluff has value even if you never cook the recipe! Which saves a lot of time and inconvenience on your part while still giving the same warm fluffy feelings. Actual "I have these things and want to cook something" could practically be automated. The big exception is baking, where precision ratios and time matters a lot. (I'm also reminded of various stories of people trying to trace the origins of much loved family recipes and then discovering that rather than being an authentic traditional Calabrian whatever, their grandmother copied it off the back of a tin. I'm fairly sure my own grandmother's cookies recipe is from Tate&Lyle)
- dvtrn 6y agoI have come to the conclusion that the "fluff" is what most of the recipe-reading public want. I've often heard the prevailing reason why this happens on the web is because of SEO and 'bounce rates'. More time spent on the site improves ranking, so the actual recipe is pushed below the fold so users have to scroll down thereby adding more time on the site. Have often wondered if any SEO wonks with the inside baseball can actually validate this?
- anticsapp 6y agoI can't validate this but have been told this by multiple authors and SEO professionals. So it's anecdotal. One blogger apologized to me and told me she was embarrassed doing it but it's an industry practice. She explained it's because other recipe sites use automated scraping tools and republish their recipes in an effort to outrank original authors. The personal fluff helps slow them down. Also, a lot of recipes are bullshit, they manually steal them from elsewhere and change a few variables to evade copyright claims. Although, I guess changing a few variables is how cuisines evolve.
- kevinmchugh 6y ago
- sseagull 6y agoMy favorite: talking about how this recipe helped you cope with the September 11 attacks (although this intro is shorter than a lot that I have seen). https://cooking.nytimes.com/recipes/1017089-maple-shortbread-bars https://cooking.nytimes.com/recipes/1017089-maple-shortbread...
- zerocrates 6y agoIt ran in the paper 8 days after 9/11... seems pretty reasonable to me. In much the same vein as the rush by many to baking now.
- mattkrause 6y agoThe "9/11 content" is also literally four sentences long. One sentence also introduces the recipe contributor, while another describes the original source she adapted it from. It's not particularly gratuitous, especially given the date and location.
- crankylinuxuser 6y agoMy understanding of why there's a life story copy at the top of recipes is that a recipe is not copyrightable, but a story is. And, an AI can generate a story one time relatively easy. https://lizerbramlaw.com/2015/04/07/copyright-protect-recipes/ https://lizerbramlaw.com/2015/04/07/copyright-protect-recipe...
- tylerrobinson 6y agoIs there evidence that the "annoying recipe sites" in question include algorithmic stories, or are you speaking hypothetically?
- crankylinuxuser 6y agoI don't have any direct proof that recipe sites are doing this. I only look at just how many "stories" there are, and how many recipes. There just looks like too much writing. And it too is also very formulaic. And it's also 'if I was to do a recipe site, I'd use a text generator'
- logfromblammo 6y agoSounds cheaper than paying a freelancer $0.025 per word, too. I suppose if you were just starting a blogspam recipe site, you could initially pay freelancers for the first articles, and use them as your training corpus. But since this sort of templated recipe site is already evil, just scrape all the other recipe sites and use their articles as your corpus. "When we vacationed in { madlib( international_city ) }, we ate a { r.Name }, and it was delicious. When we returned home, we tried to recreate the recipe, but it was never quite right. After months of trying, it's finally perfect." "As a child, I remember my grandmother making { r.Name } and eating it with all my cousins. It was her secret recipe. She never told anyone. When she died, we were all very sad because we thought the recipe was gone forever. But I found this in her old { madlib( noun ) }, and now I'm the most popular cousin at the reunion." Repeat ad nauseam.
- interestica 6y agoI've always thought of copyrighted music as "sound recipes"
- tclancy 6y agoNow we need an bot that parses the comments and applies AI to do . . . something with the people who say the substituted half the ingredients for what they had and left the others out to reply and tell them what recipe they actually made and that they can stuff their 1 star review.
- SkyBelow 6y agoDoes it pose an actual problem? When I search for recipes, I type in the food I want + recipe and then open the top 5 or so links. I quick scan for a list of ingredients. If I don't easily spot on in a few seconds I move on. I'll do this until I have a couple different lists of ingredients for making the item. This ends up taking less than a minute or two. That just isn't a significant portion of time compared to how long I'll spend comparing different recipes to find a common theme to follow. Maybe it is because I never follow a single recipe but instead combine the common themes from a couple that the whole life story before the recipe shtick isn't something that bothers me.
- onion2k 6y agoThis ends up taking less than a minute or two. For you to open 5 websites, dismiss the cookie permission request on each one, dismiss the notifications request, scroll down to find the ingredients list, dismiss the scrolling activated newsletter signup, read the ingredients list, and click the 'next page' button to see the instructions to work out what you can tweak, all in an average of 24s per site is very impressive.
- Freak_NL 6y agoOr just get thrown out completely because you are from the EU.
- SkyBelow 6y agoI just tried it. Went to the top seven sites for chicken tikka masala. Exited one for not loading, had to mute another tab with an annoying video, but got to the recipe in 6 of them in under two minutes. No popups (though I do have an ad blocker that may have prevented them). >read the ingredients list, and click the 'next page' button to see the instructions to work out what you can tweak That wasn't included. I was talking about scanning to verify there was an ingredient list. I pointed that out when I said: >That just isn't a significant portion of time compared to how long I'll spend comparing different recipes to find a common theme to follow.
- 0xdeadbeefbabe 6y ago
- zwieback 6y agoDakota also owns many beautiful bowls and whisks handmade from sustainable materials, featured in 20 photos before the recipe hidden behind another link I hate blogger recipes, luckily there are enough cooking sites that are well curated.
- SamBam 6y agoIt's definitely grown worse now, but I think that this originated from recipe sites that people actually used to follow, because the blogs were interesting and we got to know the writers, and what's changed is more that we're jumping to the first Google hit and we expect them just to grant us the information we wanted. There is a difference between opening up a recipe site, like a favorite blog, or the New York Times (which does the same kind of spiel before its recipes), just to read and find out what interesting thing they have posted, vs doing a search for "pasta carbonara," clicking on the first link, and having to read a life-story. I never mind opening up the recipe section of the New York Times and reading about what's so interesting about this recipe, and memorable times it was served. That's because I trust the article to be vaguely interesting, and reading it is a form of entertainment. There's a reason why no newspaper's recipe section has ever simply been: "Pasta Carbonara: 1 lb pasta. 2 oz Pancetta. 5 egg yolks. Cheese. Combine as directed below." So I feel like the in-vogue hatred of these recipe site styles is more a reflection of how expectations on consuming and searching for recipes has changed, more than significant changes in how recipes have always worked.
- reaperducer 6y agoI think it's also about different types of recipe collections. There are cookbooks that are all recipes. This seems to be what the HN crowd is looking for when they search the internet. There are cookbooks where each recipe is accompanied by a little story. These seem to sell well, judging by the number of them that appear on bookstore shelves. And then there are cookbooks where all of the anecdotes are in the front of the book and the recipes are in the back. These are the ones I like because I can easily find what I'm looking for, but can still read the background about a recipe, if I choose. It's not right in front of me causing the actual recipe steps to continue on another page. I think recipes and the internet don't mix, unless you're just looking up ingredients while shopping. It's one of the areas where a fat, old cookbook is always better, in my experience. 30 years from now, nobody is going to cherish grandma's dog-eared and tattered old iPad full of recipes.
- mdszy 6y agoThere's also cookbooks that are almost more like a textbook, with technical cooking information followed by recipes (almost like textbook information followed by exercises).
- wastedhours 6y agoIt's gone to extremes now, but to be honest - if someone wants to blog and post a recipe at the same time, that's their prerogative? Some people actually enjoy reading those things too. There's a place for straight recipe sites, and a place for personal word-vomit blogs with a recipe at the bottom. The web would be a sad place if you were only allowed to write your recipes in LaTeX.
- Guest0918231 6y agoYes, they can do whatever they desire. The issue is when people are required to write life stories and superfluous content not because that's the direction they want to go, and not because it's appealing to their target audience, but because it's the only way they can rank well on Google. On the user side if you only want straightforward recipes, you're out of luck, because they're never going to be near the top of the search results. If you are one of the people looking to read these recipe background stories, even you get a mediocre experience, because the majority of the content you're reading was written primarily for Google (how many times can I reference salmon, fish, atlantic, norwegian, protein, healthy, omega-3 and smoked in this recipe story to hit all of the important keywords?).
- ghaff 6y ago>The issue is when people are required to write life stories and superfluous content not because that's the direction they want to go, and not because it's appealing to their target audience, but because it's the only way they can rank well on Google. Citation needed. When I do a search on, say, Wiener Schnitzel (which I made last night) it looks to me as if the top searches are pretty to the point. Honestly, I tend to want some context for a recipe other than just a list of ingredients and some directions.
- wastedhours 6y agoSame for me too - especially in the Google recipe snippets at the top of the search. I find those tend to have more straight recipes, whereas the normal search list has more of the story style. And for context, I do like reading about the history of a recipe, why it was created etc... not all the stories float my boat, but literally all it takes is a scroll.
- jurip 6y agoIf a recipe is too hard to find, just move on. If there's a few paragraphs of things you don't care about before the recipe, press page down. Maybe let Dakota write what she wants on her own site. One reason I haven't seen mentioned here for that personal content: it's also a question of building context and trust. If I go to Bon Appetit for a recipe, I know they've tested it a few times and it should be more or less OK even if I don't recognize the author. If I go to a barebones anonymous just-the-recipes site, I have no faith that it ever worked and if it did, it wasn't just a fluke and it was written down right. Having some detail around a recipe from a previously unknown source allows me to build a connection to a persona in my head, genuine or otherwise. If a recipe doesn't work I'll know to avoid the site in the future. If it does work I can remember that connection and come back to the site again with some more confidence.
- m_ke 6y agoAre there any legal issues with scraping recipe sites in a commercial app like that? I'm assuming ingredients and directions are "facts" so can't be copyrighted, but what about the pictures?
- thinkloop 6y agoScraping is LEGAL, all search engines scrape to some degree for example, there is a fair use component, so you can't "scrape" 100% of a site and stick it on your domain, but you can still scrape more than zero. In general it is leaning more acceptable than less.
- m_ke 6y agoYeah I understand that part, my question is about showing the scraped data to your users. https://www.yummly.com/ https://www.yummly.com/ used to have a paid API for recipe search and currently still lets users search their index. Did they have to go and get permission from each site that they index or is it fair use?
- thinkloop 6y agoIt looks a shade more detailed than google's recipe cards, they link back to the original source for the instructions, I would bet they didn't get permission, and that they count as a fair-use search engine. The law isn't (can't be) perfectly prescriptive here, there's some line that you have to sue about to know if it has been crossed.
- dilliwal 6y agoIf the robots.txt file has no restrictions parsing and scraping is fine. Of course not all scrappers respect robots.txt but they should But as an internet’s citizen better to always reference the source
- MatthewWilkes 6y agoWhile a recipe isn't protected by copyright in the US (and many other countries, including the UK), the wording of the recipe could well be an original literary work, the layout of the page could attract a copyright (as it does in cookbooks) and you're right that the images would be protected. All that said, if the import is being used for personal use only and not being edited, then it's little different to printing it out and putting it in a binder. I don't know much about US fair-use laws, but in the UK it would seem that reproducing a recipe in an app for your own use would qualify as fair dealing thanks to being personal study. That only applies if the imports are specific to the person importing them, of course. If they're shared or published, then it's a different story. Also, if you're importing more than one recipe, so it's a significant amount of the published work, then that'd be an issue too. You can't import a whole cookbook and claim it's personal study, but one recipe out of dozens is probably fine.
- papadoc 6y agoIs this post illegal as it contains information on how to commit crimes such as copyright infringement?
- deleted 6y ago[deleted]
- jfk13 6y agoIs the instruction manual for a photocopier illegal?
- throwaway55554 6y agoCopyright does not protect recipes. It does more likely than not protect the images being scraped, though.
- SkyBelow 6y agoEven if the action was a crime, discussion about an illegal action, including instructions on how to do it, are not illegal. It might be illegal in very exact situations where you are giving information directly to a person you believe will use that information to commit a crime, but sharing the information in general as a topic of discussion is protected (at least in the US, laws do differ in other countries but I find the US law as the best one to go by).
- ganstyles 6y agoOf course not, as others have said. But, did you create this account just to post this comment?
- jedieaston 6y agoPaprika 3 (I use the iOS version, but I believe the Mac version has the same function) has a fantastic web scraper for recipes. I've had to correct maybe 1-2 errors across 100 recipes I've brought in from a bunch of different sites. It's super helpful to look through them in a standardized way (and you can sort by ingredient/category) to figure out what to make.
- zimpenfish 6y agoTried this out and I have to say I'm impressed on the first recipe. Scraped it correctly (albeit from the BBC which has a reasonably sane layout) and, since I've only got 75g of dessicated coconut instead of the 85g required, I wanted to scale it by 75/85 ... which worked. I just typed in 75/85 and it worked. Amazing.
- karatestomp 6y agoI think most recipes are published using a microformat that makes this pretty easy, and that's why Paprika (I use it too!) so rarely screws up.
- reaperducer 6y agoYep. And if Google detects that your page contains a recipe and the microdata isn't perfect, or doesn't include all the things that Google wants so it can show your recipe to people without them clicking through to your site, Google sends you an e-mail through Webmaster Tools telling you to fix it, with the implied threat that your page won't be listed if you don't allow Google to use your work for free.
- rrrrrrrrrrrryan 6y agoI'm a little torn with this. On the one hand, it's messed up that Google forces these companies to basically hand over their data, as you put it, but on the other hand, if they don't push companies to do things like this, single-visit webpages like lyrics and recipes inevitably become ad-infested, SEO-driven trash. Maybe if they used a carrot in addition to the stick, it'd feel less sleezy, but I'm not sure what exactly that would look like.
- SeanDav 6y agoAnother tool for difficult-to-scrape sites is OCR. There are a few decent free/opensource options available: https://source.opennews.org/articles/so-many-ocr-options/ https://source.opennews.org/articles/so-many-ocr-options/
- thinkloop 6y agoAny recommendations for a js lib that does all the "easy" scraping (microdata, og tags, jsonld, etc)?
- choward 6y agoI thought that's what this blog post was going to be about but it's just an ad for their app. I just need the scraping functionality.
- hundchenkatze 6y agoWhile they do end with pushing their product, I think they did a good job of outlining how they scrape the recipes. They inform the reader about json+ld, microdata, and how to scrape the sites that don't use those. They even link to a JS lib that handles the parsing for you. I think calling it "just an ad" is inaccurate. > There are libraries like https://github.com/digitalbazaar/jsonld.js/ https://github.com/digitalbazaar/jsonld.js/ to parse JSON-LD + Microdata for you.
- greglindahl 6y agoIn Python I use the "extruct" package from the scrapy people. It's not very good with syntax errors in the markup.
- hundchenkatze 6y agoThe article recommended one. > There are libraries like https://github.com/digitalbazaar/jsonld.js/ https://github.com/digitalbazaar/jsonld.js/ to parse JSON-LD + Microdata for you.
- throwaway55554 6y agoThere is markup specifically for recipes. I wonder why it isn't more often used. EDIT: Yes, the article mentions it, but doesn't give a clue why it isn't more prevalent.
- ateamtexas 6y agohttps://ateam-texas.com/the-ultimate-guide-to-software-development-outsourcing-in-2020/ https://ateam-texas.com/the-ultimate-guide-to-software-devel...
- welanes 6y agoNeat write-up, and thanks for putting me on to jsonld.js - looks useful. I'm building https://simplescraper.io https://simplescraper.io and we're trying to create heuristics to update CSS selectors whenever a website changes. People become unhappy when a scrape task that ran smoothly on Monday suddenly returns nothing on Tuesday so while it's a tough nut to crack it's super important. We use a combination of XPath, historical data and data type (the value may change but the type and length often remain the same or similar) to narrow down the options. Of course there's more sophisticated methods using Machine learning etc. but it's fun to try different approaches to solve this problem.
- saadalem 6y agoNext level : Shazam for cooking shows.
- sunsetMurk 6y agoThat'd be sweet for YouTube videos! I watch so many food/cooking YouTube videos.
- saadalem 6y agoThat's the point : Similar to Shazam for music, you would build a mobile app that would use voice recognition to identify the cooking show and give you the recipe. Additionally you could find a restaurant that makes something similar and offer to have it delivered for a fee.
- WrtCdEvrydy 6y agoHere's the question... why is it so difficult to do this in Android? Seriously, AndroidDriver for Selenium was last updated 2013... and importing it throws an HttpClient error now. Update that client and you get a class duplication hell that is impossible to exit. All I needed was to interact with 2-3 fields on a webpage but it's been eight hours and now I hate my life.
- MatthewWilkes 6y agoI believe that the Webdriver/Chromedriver approach is the current recommended way of doing this.
- openthc 6y agoCheckout BrowserStack -- it's dead easy -- and even if you're not using their platform, their docs are good for showing the Selenium/Driver usage.
- selecsosi 6y agoI highly useful tool in my household for dealing with the SEO/tracking scourge that recipe blogs have become is https://www.paprikaapp.com/ https://www.paprikaapp.com/. Hoping someday to have some spare time to integrate this with https://grocy.info/ https://grocy.info/ and have a pipeline for recipe -> preparation automation.
- taude 6y agoBig fan of this app, and I love it so I don't have to keep revisiting the sites. This is one of the few apps, I've purchased multiple times: for my iOS eco-system (I typically cook with my iPad), my android phone (so I can add recipes on the go), and my partner's devices so she can add recipes to our list. We've developed a little workflow where we put all the recipes we want to try into an "incoming" category, and then move them to one of our custom categories when we make it and decide it's worth keeping. This is a reaction to becoming recipe hoarders when using a site like Pintarest for something similar. The iOs has a really subtle, but nice feature when you're cooking with the app. It prevents the iPad from going to sleep and locking which you have messy fingers.
- snuxoll 6y agoIt also has a timer built in that, unlike the default iOS Clock app, allows multiple timers to be run simultaneously. Just tap on any time in the recipe to start it as well. Paprika is well worth the couple bucks.
- vitiell0 6y agoYou might want to checkout an app I built called Cooklist. It has the features of Paprika + Grocy + Instacart + Pinterest all in one. https://cooklist.co https://cooklist.co
- SpyKiIIer 6y agoworks in Canada?
- 6y ago
- partiallypro 6y agoBe careful with this, some recipes are subject to copyright law. I think you can list ingredients of a recipe with no problem, but once you get to exact measurements and prep it somehow switches over to falling under copyright law. There used to be a bunch of open sourced recipe repos/databases...but almost all of them are gone.
- CrazyStat 6y ago> I think you can list ingredients of a recipe with no problem, but once you get to exact measurements and prep it somehow switches over to falling under copyright law. In the US, recipes are not protected by copyright, including precise measurements and instructions for how to combine them. If you write sufficiently creative instructions those may be protected by copyright, but I can get around that by simply not replicating those instructions. I can give exactly the same list of ingredients and a less-creative set of instructions for the same dish without infringing your copyright. Copyright.gov has a circular which explains this with a couple examples [1]. [1] https://www.copyright.gov/circs/circ33.pdf https://www.copyright.gov/circs/circ33.pdf
- cronopios 6y agoCan you provide links to those open source recipe resources? Aren't there dumps of the bygone ones anywhere?
- biggestdummy 6y agoAmong other things, I am a cookbook author, so I know a fair bit about this. Ingredient amounts are not subject to copyright protection. Any prose - intros, descriptions, instructions are covered by copyright the same way that any other book would be. So, yes, this kind of activity is likely in violation of copyright. Let me also say that I find it a bit insulting that people who make a living creating IP (software) would be happy to disrespect the IP of these recipe authors. By taking the recipes from outside of the revenue source (a book, a banner ad, a cookie, whatever), you are stealing from the author and the publisher. I make 25 cents on every book that is sold. So I don't actually care if you steal from me. The money is tiny. But it is a bit insulting when people - in my own living room, reading my copy of my book - decide that they want to take a picture of a recipe instead of buying a copy. It devalues the hundreds of hours and thousands of dollars that I sunk into the creation, cooking, and photography of the book. So here's the moral of the story. Waiters should leave good tips because they know that waiters depend on tips. And IP creators should know better than to steal IP.
- aodj 6y agoReally nice! I often copy and paste recipes into text files I have locally so this is a great alternative. One feature request (if I may be so bold): it would be great to offer an imperial<->metric convertor. This is predominantly one of the reasons I keep copies of recipes I find and use.
- monksy 6y agoThis is pretty awesome. I'm currently working on a data pipeline to demonstrate recipe scraping with kafka streams. This is going to be a big help in part of it.
- zwieback 6y agoCool, now the next interesting step would be to categorize recipes, maybe some kind of clustering algorithm, to see how similar they are and whether they have a common ancestor. When I look at a recipe and notice some unusual proportions I usually check against Joy of Cooking or some other standard book. I've noticed that often everything old is new again.
- vadansky 6y agoI've been using Tasty. Quick videos showing all the steps and the how it's supposed to look like along the way. That's the only way I can accept recipes anymore.
- julianlam 6y agoI'm surprised nobody has mentioned "Recipe Filter" https://addons.mozilla.org/en-CA/firefox/addon/recipe-filter/ https://addons.mozilla.org/en-CA/firefox/addon/recipe-filter... Cuts the fluff and puts the recipe front and center. I wouldn't be able to find recipes online without this.
- Cactus2018 6y agoIn 2011, Google released "Google Recipe Search". With filtering based on ingredients, cook time, and calories. https://www.wired.com/2011/02/google-recipe-semantic/ https://www.wired.com/2011/02/google-recipe-semantic/ https://latimesblogs.latimes.com/technology/2011/02/google-debuts-recipe-view-search-function-for-cooks.html https://latimesblogs.latimes.com/technology/2011/02/google-d...
- memset 6y agoInteresting! I wrote https://plainoldrecipe.com https://plainoldrecipe.com (open source!) to solve this, an inadvertently discovered many of the metadata tags described here. The irony is that the content is required for SEO purposes, but once you’ve landed on the page you don’t want to see it. I wonder if there would be a way to write SEO that only the google bot sees and hide it from humans...
- phito 6y agoYour header says "plan old recipe"
- yepthatsreality 6y agoWhich is just dripping with irony? serendipity?
- hawski 6y agoThere is a way to present different things for the google bot and humans, but it can and should result in Google ban. You're probably aware and I'm a bit too verbatim.
- peterwwillis 6y agoSo far the best way I've found to search for recipes is to search in a foreign language. Translate what you're looking for, then search and translate back to English. There are still recipe blogs, but 5 instead of 5,000, and usually an authentic dish, not what Michelle The Stir Fry Queen From Michigan thinks constitutes a "Moroccan" dish because it has cinnamon and tomatoes. Would love to see someone put together a search engine that excludes recipe blogs and penalizes SEO.
- sum2000 6y agoNeat! I am interested in developing REST API around it to support more functionality, wanna collaborate?
- Mela1998 6y agoI wish I had this when I first started cooking! I love this concept, but wouldn't this also harm the creator's traffic???
- mark_l_watson 6y agoHey Ben, thanks for that write up! You may not have time for this, but your article and the intersection of food/recipes and computer science would make a good book, at least I would read it. I wrote [1] about 12 years ago in Clojure because for health reasons I had to track my intake of vitamin K, then decided to track all nutrients in the USDA nutrition database. I am working on a semantic web product (with another semantic product in planning) but maybe the end of this year will get to rewriting my food web app in Common Lisp and as a macOS app. I am adding a link to your article and these comments here to my notes for that project. Useful stuff. [1] http://cookingspace.com http://cookingspace.com
- GrantSolar 6y agoI've been working on something similar for the past couple of days, but the trouble comes with wanting static types. There are a few projects out there that offer either a microdata parser, or types derived from schema.org but nothing that combines the two as yet
- jeffrogers 6y agoI’ve been working on this and will have a recipe-specific solution up in a couple weeks. See https://rcpe.io https://rcpe.io
- hamilyon2 6y agoSo, Google actually encourages open semantical web? That is news
- greglindahl 6y agoIt's been the case for quite a while. However, it's mostly useful if you have a huge web crawl, elsewise discoverability is a bit poor.
- dsilver 6y agohttps://www.eater.com/2020/3/31/21201374/why-are-free-online-recipes-so-long-stop-shaming-food-bloggers https://www.eater.com/2020/3/31/21201374/why-are-free-online...
- russellbeattie 6y agoHa! I should totally know better, but for a second, I mixed up .com and .net and thought, "Ben has a blog? And posted about recipes? Did he pull them into his 8 bit computer or something??" I didn't realize my mistake until I clicked.
- logfromblammo 6y agoThe simple truth is that the core recipes are fact-based and non-copyrightable, and the 1000-word blogspam recipe header is both copyrightable and garners better search result rankings. So the business model is to take facts from the public domain, wrap it in bullshit prose, and then SEO the bullshit to have higher ranking than the naked source facts, for more unique visitors and ad revenue. Making comments about "providing recipes for free" are exactly as useful as comments about "providing phone numbers for free" or "providing mailing addresses for free" or "providing the original text of 'Little Women' for free" or "providing the steps of the long division algorithm for free". Obfuscating the public domain is not a valuable service. Automatically removing the obfuscation is valuable. A "Project Gutenberg" style repository of recipes would be recurringly donation-worthy.
- linsomniac 6y agoA surprisingly good UX for recipes is Google Home. Ask it for a recipe, and it will ask if you want directions or ingredients. If you ask for ingredients, it will say them one by one, and pause between them until you ask it for the next one. My son has used it to great effect to make pancakes.
- deleted 6y ago[deleted]
- ohhaimarc 6y agoThis case is a perfect 'recipe' for reinforcement learning. Let me know if you want help here.
- benawad 6y agoI considered going down the ML route, but didn't know where to start. I'd love to hear how you would approach it.
- brendanmcd 6y agoGoogling for recipes drove me to install ad-blocker. Have to say I never considered how google created recipe card featured snippets -- cool stuff!
- qrv3w 6y agoThis is great! Its a wonderful write-up. I've also made something almost identical - a Go library for recipes scrapers for ingredients [1] and instructions [2]. Instead of the LCA method here, in my version I try to find the longest sequence of highest scoring HTML tags and those are "ingredients" or "instructions". It works very well (although I think this one works better). Like the article mentioned, I found that the heuristics for finding HTML elements with ingredients turn out to be surprisingly simple - they usually include just a number, a measurement, and a food! This simple heuristic worked better than other sophisticated things I tried. [1]: https://github.com/schollz/ingredients https://github.com/schollz/ingredients [2]: https://github.com/schollz/instructions https://github.com/schollz/instructions
- imgabe 6y agoThis is great. I made a similar product at No Nonsense Recipes https://nononsense.recipes https://nononsense.recipes because I was also tired of dealing with all the dreck on recipe sites. I did scrape some recipes to seed the site with but haven't integrated it as a feature yet. I did ignore the photos though, since while recipes are not subject to copyright, photos are.
- franciscop 6y agoThis is pretty interesting, I wonder if this meta could be reused for tutorials of any kind (and not only of food, a.k.a. recipes). A tutorial normally has some requisites, and then step by step guide of how to achieve it, and then the final result.
- triyambakam 6y agoHey Ben, if you read this, thanks for your helpful and entertaining youtube videos!
- IncRnd 6y agoAlmost every comment on this page is helpful and from people's direct experiences. Wonderful :) Thank you everyone for all of this information!
- chirau 6y agoIs there any recipe tool out there that can do at least one of the following: 1) Scale the quantity of ingredients and cooking time as number of people to be served increases? 2) Tell me what dishes I can make with the ingredients I have?
- nolroz 6y agoI enjoyed the recipe scaling abilities of Gourmet: https://thinkle.github.io/gourmet/ https://thinkle.github.io/gourmet/ You can also filter and search by ingredient, but that might be somewhat simplistic depending on what you had in mind.
- wantacker 6y agoOff-topic, but I just wanted to mention that Ben's been one of my favorite 'teachers' in YouTube. He has some quality content on React and JS stuff. For those wanting to learn React (including some advanced stuff), check out his channel! And no he didn't pay me to post this here. Hey thanks Ben - I know a bit of React and have used it on a few projects thanks (also) to you.
- fulldecent2 6y agoI saw all the terrible SEOd recipe websites and my first thought was: I should make a better recipe website that is simpler and is better SEOd. --- FIRST EXAMPLE: How to cook chicken on a skillet Step 1 -- get this much chicken [picture] Step 2 -- cook on skillet for 5 minutes OPTIONAL -- here are seasonings you may add [pictures] RELATED: - How to cook a lot of chicken on a skillet [LINK] - How to fry chicken breast [LINK] --- But then I didn't understand how any of these websites are making money so I didn't do it.
- RhodesianHunter 6y agoThe reason all of these websites are so terrible with the long winded intro-stories is precisely because they do better with SEO.
- fulldecent2 6y agoOnly if the page is low quality. Leaving a low quality page after 60 seconds is way better than leaving a low quality page after 5 seconds.
- kevindong 6y agoI personally just find recipes, make it as written from the website, and then (if I actually like it), I'll convert it to be sane for actually following and output into Apple Notes. What I mean by that is most recipes call for using wwwaaayyy more intermediary bowls/plates than actually required (e.g. if spices, chopped veggies, and minced garlic are going into the pot at the same time, there's no point in using three bowls) or list ingredients out of order of how you'd actually use them.
- nicbou 6y agoI just started transcribing every recipe I make. Even if you can extract all the essential information from a recipe site, some changes are needed: - I need to convert recipes to metric. I am neither equipped nor inclined to cook in freedom units. - A "can" or a "packet" is not a standard unit of measurement. - Package sizes vary between countries. I often adjust recipes to avoid wasting food. - I cook by mass, not volume. I convert the units them round them. - Instructions are sometimes too verbose. I make them easier to follow while my hands are busy. - I will make my own changes and I must write them down somewhere. Besides, sites go down and links break. Food.com broke many of my bookmarks a few years ago. Other sites went dark. My recipes are plain text. They are editable, searchable, editable, and available offline.
- tincholio 6y agoI wish I had the willpower to do this consistently...
- ben_utzer 6y agoI did something similar a while ago. I still have somewhere a DB with half a million recipes somewhere. I didn't continue it because I got stuck with the client side and I didn't find anyone interested in helping me.