2 ms·
Very interesting! >For most documents, we rely on Apache Tika to transform the original document into a canonical HTML representation, which then gets parsed i
by PokemonNoGo 8y ago
Very interesting!
>For most documents, we rely on Apache Tika to transform the original document into a canonical HTML representation, which then gets parsed in order to extract a list of “tokens” (i.e. words) and their “attributes” (i.e. formatting, position, etc…).
How good is really Apache Tike at this? I've messed about but its hard to find solutions that cover the base cases.
What are the recommendations for covering lets say PDF, OpenXML, and ODF?