4 ms·
Your use-case is not what I built the library for (natural language processing, not text consumption), but let's see what we can do... You can download HTML E-
by c-w 12y ago
Your use-case is not what I built the library for (natural language processing, not text consumption), but let's see what we can do...
You can download HTML E-Books using the following command:
python -m gutenberg.download -vvv --filetypes=html --limit=5mb ./ebooks
This will download 5mb of zipped E-Books for which there exists an HTML version to the ./ebooks directory.
It seems as though the legal disclaimers and copyright notices in the HTML files are all within <pre> tags so we can easily clean-up the files with a small shell script:
EBOOK_DIR="./ebooks"
find "${EBOOK_DIR}" -name *.zip -type f -exec unzip -d "${EBOOK_DIR}" {} \;
find "${EBOOK_DIR}" -name *.html -type f -exec sed -i '/<[pP][rR][eE]>/,/<\/[pP][rR][eE]>/d' {} \;
This will probably not work for all E-Books, but it'll give you something to work with. Note that removing the copyright notices may or may not be against the Project Gutenberg terms of service.
Downloading E-Books via genre, author, etc. is not currently supported but is something that I wanted to implement - so watch this space.