5 ms·
Small experiment of visualization of wikipedia articles as a graph using d3.js. Articles with more traffic are bigger. I computed the semantic similarity using
by lucamartinetti 15y ago
Small experiment of visualization of wikipedia articles as a graph using d3.js.
Articles with more traffic are bigger.
I computed the semantic similarity using LSI with python (gensim)
You have to scroll down/right a bit!
http://similarityapi.appspot.com/graph/?title=blade%20runner http://similarityapi.appspot.com/graph/?title=blade%20runner
There is also a JSON api:
http://similarityapi.appspot.com/api/v1/?limit=100&title=blade%20runner http://similarityapi.appspot.com/api/v1/?limit=100&title...
All feedback is appreciated:
@lucamartinetti
luca@luca.io
- viscanti 15y agoThe JSON api should degrade gracefully if results aren't found. I.E. There should be a JSON message explaining that that item doesn't exist.
- lucamartinetti 15y agoRight! It could use some input checking / normalization too. It expects the title parameter to be lower case now.
- rplnt 15y agoOption to select language version could be a good feature (defaulting to en as now).
- 3pt14159 15y agoI've had much, much better results with LDA than LSI. Give that a shot if you have a chance, you'll be blown away. Stop word ratios are important, and make the max number of tokens 500,000.
- Radim 15y agohow much data did you use for the semantic analysis?
- lucamartinetti 15y agoThe whole text of all articles from wikipedia english (then filtered those with more the 1k views last month)