3 ms·
icco Added a mobi version (Thanks!) and I put in an updated version that had a few bugs. There are definitely spacing issues etc but you can pretty easily unde
by dy 15y ago
icco Added a mobi version (Thanks!) and I put in an updated version that had a few bugs. There are definitely spacing issues etc but you can pretty easily understand the original meaning. Here's the main loop of code that extracts the text, if anyone has a better idea let me know:
source = open(link['href']).read
text = Readability::Document.new(source, :tags => %w[div p br font]).content
xhtml_text = Sanitize.clean(text, :elements => ['a', 'div', 'pre', 'br', 'font', 'p', 'img', 'table', 'tr', 'td'], :attributes => {:all => ['class', 'id', 'src', 'href']})
xhtml_text = Mustache.render(xhtml_template, :title => link.text, :content => xhtml_text)
where ePub formats expect something like
xhtml_template = <<XHTML
<?xml version='1.0' encoding='utf-8'?>
<html xmlns="http://www.w3.org/1999/xhtml>
<head>
<title>{{title}}</title>
</head>
<body>
<h2>{{title}}</h2>
{{{content}}}
</body>
</html>
XHTML
- dpapathanasiou 15y agoInstead of using this xhtml_template, you should pass the content through tidy with the -asxhtml switch on. This will produce a valid xhtml file that passes epub validation (http://code.google.com/p/epubcheck/ http://code.google.com/p/epubcheck/).
- dy 15y agoAwesome - exactly the information I was looking for, will set it up later tonight and post both a new ePub and the source.