4 ms·
Though, unless the documents have links to other documents, their main algorithm (PageRank) isn't going to be very effective.
by colonelxc 16y ago
Though, unless the documents have links to other documents, their main algorithm (PageRank) isn't going to be very effective.
- bane 16y agoThat's an interesting idea actually. Assuming that the people you work with represent a graph (say something like Linkedin) and everybody's resumes were online, a traversal of that graph might yield other good employees.
- mayank 16y agoI believe PageRank isn't as important as it was 10 years ago. Besides, you don't need explicit links to create links. Off the top of my head, the following could be used to link resumes: keywords, references, educational institutions, co-authors on papers, and former employers. All these can be extracted fairly well from a document as structured as a resume, and then you can go wild with link analysis. At the very least, it would allow you to filter the chaff, and enough care could be given to a dataset as small as 75,000 documents to take care of preprocessing, some manual curation, etc. and yield some decent results. If you wanted to, you could even compute PageRank scores in R on your laptop for a dataset that small (and that would probably be the only computationally intensive part of it, after POS tagging and ML model fitting). As a bonus, the effort would go a long way to help your recruiting in the future.
- JabavuAdams 16y agoIt's unlikely to tell you who's a jerk, or who has a cocaine habit.
- rudiger 16y agoPageRank hasn't been their main algorithm for years. PageRank remains prominent because Google was so open to talk about their ranking algorithm in the early years (to be fair, Messrs Brin and Page were PhD students at the time). They won't make that mistake again.
- yuhong 16y agoYea, had the AGPL existed in 1998 it might have been different, but it didn't exist back then, so....
- mahmud 16y agoNo, not Page Rank, you can do basic retrieval with tf-idf alone.