5 ms·
At least searching and ranking a mere 75,000 documents to find good employees shouldn't be a problem for them.
by cing 16y ago
At least searching and ranking a mere 75,000 documents to find good employees shouldn't be a problem for them.
- colonelxc 16y agoThough, unless the documents have links to other documents, their main algorithm (PageRank) isn't going to be very effective.
- bane 16y agoThat's an interesting idea actually. Assuming that the people you work with represent a graph (say something like Linkedin) and everybody's resumes were online, a traversal of that graph might yield other good employees.
- mayank 16y agoI believe PageRank isn't as important as it was 10 years ago. Besides, you don't need explicit links to create links. Off the top of my head, the following could be used to link resumes: keywords, references, educational institutions, co-authors on papers, and former employers. All these can be extracted fairly well from a document as structured as a resume, and then you can go wild with link analysis. At the very least, it would allow you to filter the chaff, and enough care could be given to a dataset as small as 75,000 documents to take care of preprocessing, some manual curation, etc. and yield some decent results. If you wanted to, you could even compute PageRank scores in R on your laptop for a dataset that small (and that would probably be the only computationally intensive part of it, after POS tagging and ML model fitting). As a bonus, the effort would go a long way to help your recruiting in the future.
- JabavuAdams 16y agoIt's unlikely to tell you who's a jerk, or who has a cocaine habit.
- rudiger 16y agoPageRank hasn't been their main algorithm for years. PageRank remains prominent because Google was so open to talk about their ranking algorithm in the early years (to be fair, Messrs Brin and Page were PhD students at the time). They won't make that mistake again.
- yuhong 16y agoYea, had the AGPL existed in 1998 it might have been different, but it didn't exist back then, so....
- mahmud 16y agoNo, not Page Rank, you can do basic retrieval with tf-idf alone.
- mkramlich 16y agoFunny and I see your point. However, the catch is that treating them as merely documents rather than human beings and also thinking you could whip up a codable algorithm to rank them would be exactly the sort of mistake I'd most expect Google to make. To give just the easiest and most recent example of how that approach breaks: the increasing deluge of content farms and SEO spam in Google results. Many folks consider that non-optimal but Google defends it because their data "tells them" it's optimal. Now picture that approach with people. Rank things by some measurement and only things that measure well by that metric will rank high, shutting you out of great opportunities. The false negative, etc. Related anti-pattern is the concept of Local Maxima.
- spacemanaki 16y ago> but Google defends it because their data "tells them" it's optimal. C'mon, I'm not one to usually come rushing to Google's aid, but from Matt Cutts' comments on HN and recent blog posts on the official blog is seems to me that they are working hard on this problem, which would imply they don't think the status quo is optimal.
- mkramlich 16y agoI understand. However, I've seen public comment by Google folks that basically conveyed the idea that it was optimal -- by certain standards. Granted, they have many people and opinions can differ and vary over time. Secondly, it should be possible to do something as simple as filter out obvious content farm sites in under a week tops. Or give all users a quick and easy to do it. Perhaps have a default blacklist of known content farmers, and then give users the option to remove things from the blacklist when doing their own searches. You might be dealing with a huge company if something as simple as this takes 20k employees more than a week, tops, to put into production. And I'm being generous with the week estimate, because it could be done even faster than that under more agile conditions. And a 'default blacklist of known content farmers' is not even the only solution which could be implemented quickly, you could also do something like add another variable to the PageRank algorithm which has a down weighing effect on any page under a domain known to be a content farmer/spammer. Supposedly, they are already doing this (IIRC) but again, the basic idea could be implemented and put into production quickly, if it hasn't already.