8 ms·
Google, Bing & Yahoo Unite To Standardize Structured Search Data
- 51Cards 15y agoThis is somewhat off topic and a completely non-technical response but am I the only one who gets a small dose of 'warm fuzzies' when large competitors join forces on something like this? Everything we read is always competition, lawsuits, accusations... and to me it gets tiring. Occasionally seeing something like this warms the cockles of my cold jaded IT heart. Nuff said, let the technical merit debate begin.
- camiller 15y agoForget the technical merit, I want to discuss your point. I'm pretty sure one of the signs of the apocalypse starts with the phrase "Google, Bing and Yahoo unite...". When I see that phrase my spidey sense tingles in a decidedly non 'warm fuzzy' fashion! One of the three will probably try to sabotage this. Likely, one will contribute something to the standard without informing the others that it is patent encumbered, then submarine them. Others my feel free to weigh in on my paranoia.
- fragsworth 15y agoI can't imagine "tagging with metadata" to be very patentable. It's too ubiquitous. Then again, it wouldn't really surprise me.
- andrewcross 15y agoI agree with camiller! The technical part of the article isn't the interesting point. It's the fact that the big boys were able to join together and create something for good! Let's hope "Don't Be Evil" is infectious and no one goes after each other.
- sullivandanny 15y agoSitemaps was the last big thing they united on. It's been very helpful to have. Years before that, they agreed on meta description tags and robots.txt blocking. The latter helped keep many web sites from being overwhelmed by requests, yet enabled search engines to grow. I think it's a good thing.
- contextfree 15y agoMaybe you'll feel better, and 51Cards will feel worse, if you think of it as "companies who have invested in search engines unite to give themselves more competitive advantage over companies who have not invested in search engines, such as Apple" =)
- baguasquirrel 15y agoThey are agreeing on a protocol, they aren't saying that they're going to share their mined and ranked data. No way in Hell is Google going to agree to that. My bet is that Google is worried about its relevance in a world where an increasing amount of the data is not accessible to its crawlers. This is not just FB and Twitter, but Yelp and GroupOn traffic as well. That's where Yahoo and MSFT come in. By agreeing to standardize search APIs, they can hope to net more of the developer mindshare. Now, instead of this just being "lets program for Twitter or FB's or Google", this becomes "lets program for one of the social networks, or for search". Google doesn't lose the monetarily valuable human-driven search results, and YHOO and MSFT get to take a risk that they will be able to monetize potentially valuable developer-driven search along with Google. The thing that's funny about this is that no one seems to have figured out how to monetize APIs yet. If anyone is going to figure it out, it may well be Twitter because it is life or death for them. But Buzz goes to show how much Twitter envy there is going on at Google. They're concerned that if there is a way to monetize that traffic, then they will be shut out. Better to take a larger share of the pie with YHOO and MSFT than for traditionally-crawled to get shut out of the party.
- baguasquirrel 15y agoThey are scared of Facebook and Twitter.
- andrenotgiant 15y agoI am curious: Given all of the trouble with scraped content on the web, why would data publishers want to make their data more machine-readable?
- wccrawford 15y agoGiven all what trouble? You mean those few print publishers who don't understand the internet and are very vocal about it? There has been very little actual trouble. Most business roll with new technology and use it to get an edge on their competitors that don't get it. You've been listening to people who refuse to innovate and thinking they are the majority, or somehow a backbone of our society. They aren't either.
- rbarooah 15y agoMaybe there is a refusal to innovate, but maybe we now have a structural problem wherein we simply don't need as many people to work on these things as a result of automation, and even information work jobs are going to diminish. How do you tell these two scenarios apart?
- felix0702 15y agoI bet Yelp is probably thinking the same thing. June 2, 2011 - "Google Places Now Borrowing Yelp Reviews Without Attribution In iPhone App" http://news.ycombinator.com/item?id=2611985 http://news.ycombinator.com/item?id=2611985
- clistctrl 15y agoIn my opinion yelp has no (exclusive) right to the user contributed data contained in its site. Google should not have to provide attribution to "yelp" if they scrape my yelp review, it should provide attribution to me.
- dhimes 15y agoWithout having read them, it seems like that would depend on the TOS of Yelp.
- doorty 15y agoIs anybody seeing how much extra work this is going to put on the front end developer. We already have to support multiple browsers, multiple display sizes, ARIA roles, and more. I don't mind the extra work for the handicap (ARIA), but looking up and adding all this new meta information merely for search engines is going be a pain. In the end, this extra work makes it easier for the search engine by requiring the developer to do more work. I would prefer the reverse, otherwise maybe I need to get into the search business. They are creating a standard that will make it incredibly easy for future search engines.
- dhimes 15y agoI think that in particular the web page schema should be, at least optionally, read from the natural html. In other words, you should be able to give <div id="WPFooter"> and the search engine should know that this is the footer. In fact, the same could be asked of the rest of the itemtypes so that we can link them with CSS and don't have to repeat ourselves.
- bergie 15y agoMicrodata, why? Compared to RDFa (or even Microformats), I don't see a single compelling reason to use microdata. At least RDFa has a well documented way to define and extend your own schemas...
- andymurd 15y agoBecause most web pages are not valid HTML, let alone valid XHTML and you really need XHTML to make the most of those schemas. In other words "RDFa is hard, let's use Wordpress".
- brudgers 15y agoOn one hand I think a common semantic framework is inevitable. On the other it just looks like the search providers further encouraging the manipulation of search engine algorithms - another round in the SEO arms race in which consumer oriented interests are increasingly drowning out non-commercial speech on the web - e.g. Googling "weather" in the US returns forecasts from commercial sites but not NOAA's local forcast.
- equark 15y agoSo far the schema listed don't look very subtle. Extracting this type of information from unstructured text should really be something that machine learning algorithms can do pretty well.