2 ms·
Apache Tika is a text extraction toolkit. They store a wide selection of file types for their parser tests: https://github.com/apache/tika/tree/master/tika-par
by doykle 10y ago
Apache Tika is a text extraction toolkit. They store a wide selection of file types for their parser tests:
https://github.com/apache/tika/tree/master/tika-parsers/src/test/resources/test-documents https://github.com/apache/tika/tree/master/tika-parsers/src/...
For large sets of some common media types take a look at the govdocs1 corpus:
http://digitalcorpora.org/corp/files/govdocs1/by_type/ http://digitalcorpora.org/corp/files/govdocs1/by_type/
For an odd example, sometimes a google search will turn something up.:
https://www.google.com/?q=filetype:xlsx https://www.google.com/?q=filetype:xlsx
- brudgers 10y agoThe test files for the parser are available at: https://github.com/apache/tika/tree/master/tika-parsers/src/test/resources/test-documents https://github.com/apache/tika/tree/master/tika-parsers/src/...