3 ms·
It's a very tough problem. Queries such as "square", "blue", "fast", etc... will yield very poor results. PLSI tends to perform very well on more specific quer
by eserorg 18y ago
It's a very tough problem. Queries such as "square", "blue", "fast", etc... will yield very poor results.
PLSI tends to perform very well on more specific queries, such as "Paul Graham", "silicon graphics", etc...
The problem with PLSI is that it is extremely computationally expensive -- which is why most internet-scale search engines don't use it.
Our innovation was figuring out some tricks that have allowed us to improve performance dramatically. However, there is obviously still room for improvement.
Our goal is to satisfy 80% of the queries with decent results -- and to leave the other 20% (square, etc...) to someone else.
The interesting thing about PLSI is that it's able to rank documents from the text alone -- ignoring the link structure and other metadata.
Therefore, we're thinking our algorithm will make the most sense in situations where there is lots of textual data without web-like link metadata.
The two scenarios that come to mind where people need to text-mine documents outside the metadata-rich web are:
(1) windows file shares on corporate intranets
(2) large volumes of legal documents inside law firms
Text-mining wikipedia is a proof-of-concept at this point