3 ms·
This is a pretty small python module at the moment. To fetch metadata, it downloads a 230mb zip from PG and parses it for a few categories. Project Gutenberg h
by sethish 12y ago
This is a pretty small python module at the moment. To fetch metadata, it downloads a 230mb zip from PG and parses it for a few categories. Project Gutenberg has some metadata about Subject, but the info is inconsistent, but there is sometimes a Library of Congress code in the metadata.
Not all books have an html version. Most PG books are plaintext, and _some_ have a separate html variant. A handful are written in a markup language that can become html or plaintext.