5 ms·
One way of doing this: 1. Get some million tweets from twitter API (each tweet is exactly like what the article describes as a 'document') 2. After ignoring c
by blr_hack 16y ago
One way of doing this:
1. Get some million tweets from twitter API (each tweet is exactly like what the article describes as a 'document')
2. After ignoring common words (e.g. 'a', 'an', 'the' ... ) in a tweet , start assigning some kind of relevance rank to all the other word pairs found in the tweet
e.g. its likely that 'Cricket' and 'Sachin' (Or 'NBA' and <top NBA player) will both increase each other's relevance rank, WRT to each-other.
3. Process all the tweets like this, while maintaining the output of step 2 in a most suitable data structure. You would also need to start dropping word-pairs (to avoid having the problem of storing Million-C-2 words! ) based on some logic/heuristic.
4. If we have a good logic for having reasonably not-big storage (by avoiding the million-c-2 explosion), then what we have is at the end of processing: A simple look up of a million words, where each word has its 'top' max_allowed(k) semantic words.
PS: The most complex piece in the approach is to come out with a solution for dropping word-pairs (in step 3)
- spencertipping 16y agoI have a preliminary implementation of this, geared for map/reduce-style parallelism in Perl, up at http://github.com/spencertipping/metaoptimize-challenge http://github.com/spencertipping/metaoptimize-challenge (in the fcm directory). It may be a start to solving step 3 -- using reduce-by-two on each step and dropping low-relevance words (I think this is valid, though I'm not 100% certain).
- blr_hack 16y agoHi! Thanks for reading my suggested approach. I can understand how Map-reduce can be used to process the things faster (by processing them in parallel and later using reduce to aggregate the results??) But can't imagine how reduce (of map-reduce) can help in dropping of word-pairs (i.e. in Step 3)
- spencertipping 16y agoI think it's an imperfect solution, though it might still be solid. It depends on your reduction strategy (basically breaking associativity). If you do a fold-left, then you'll accumulate such high relevance that you'll quickly start discarding every word in the right-hand data set. On the other hand, if you binary-split the folding process you're more likely to be OK. I used a binary split and due to the sparseness of the data (and the uninsightfulness of my algorithm) didn't run into too many spurious drop cases. But the implementation I posted is very basic and shows my lack of background in NLP-related pursuits -- I'm lucky it did anything useful at all :)