4 ms·
The engines we build nowadays are still mostly based on indexed collections of metadata, just as the old card catalogs were. The innovation was in being able to
by Millennium 9y ago
The engines we build nowadays are still mostly based on indexed collections of metadata, just as the old card catalogs were. The innovation was in being able to create new kinds of cards, and then compile catalogs of them automatically from the books (or whatever data sources one cared to use). This is a very useful thing to be able to do, but I'm not sure it can really be called a new kind of discovery engine, just a faster one.
- greglindahl 9y agoThe entire point of my discovery project was to NOT just use new kinds of card catalog cards, but to instead do discovery on individual sentences. I think that new kinds of card catalog cards could also be pretty novel and interesting, for example I'd love to build a histogram of all of the dates seen in a book and include that in the card catalog card.
- Millennium 9y agoIt's a nifty technique, but I'm not sure it refutes my statement. You still have agents crawling the book to build an index of metadata by (in your experiment's case) subjects and years mentioned. You throw away the entries for subjects and years you're not searching for, because they're not needed for your specific problem, but that's an optimization. If you didn't do that, your agents would assemble what amounts to a fine-grained (and correspondingly large) card catalog of sentences, indexed by subject and year, which you could then search.
- philjohn 9y agoHave a look at the latest breed of discovery interfaces that do full text search on their corpus, EDS Discovery, Ex Libris Primo, Serial Solutions Summon. For academic work they are invaluable, letting students and researchers find snippets of related research.