Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sethish
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
15 ms
·
91.
▲
by
sethish
12y ago
I forked Project Gutenberg to github[0]. I've been doing some some work on newsdiffs[1]. I'm working on a python library for dealing with Library of Congress Subject Headings[2]. And I'm trying to get access to and release hu
92.
▲
by
sethish
12y ago
Oh yes. GITenberg is not a replacement for DP in the slightest. New books aren't likely to come out of GITenberg as it currently exists. That is what distributed proofreaders is for. At some point, I would like to investigate DP tool
93.
▲
by
sethish
12y ago
Yep. The RDF/XML data. I have a mirror of it on github: https://github.com/sethwoodworth/PG_rdf_metadata I would love to have a complete python parser for the metadata. I strongly recommend collaborating with the
94.
▲
by
sethish
12y ago
This is a fantastic overview of this part of the publishing space. GITenberg has a mailing list, and needs to start collecting breakdowns like this. I would love it if you would join us: https://groups.google.com/forum
95.
▲
by
sethish
12y ago
GITenberg has a mailing list: https://groups.google.com/forum/#!forum/gitenberg-project And there has been some conversation about the project on the OKFN-humanities list: https://lists.okfn.org/ma
96.
▲
by
sethish
12y ago
If folks are interested in contributing, the mailing list is here: https://groups.google.com/forum/#!forum/gitenberg-project
97.
▲
by
sethish
12y ago
There are a number of transcription errors in many books. There is an example PR https://github.com/GITenberg/Chess-Strategy_5614/pull/1
98.
▲
by
sethish
12y ago
Gah! This website was thrown up quickly. I wasn't intending to post to HN until after I had fixed up the website. Someone beat me too it :-S Pull requests welcome: https://github.com/GITenberg/gitenberg.github.com
99.
▲
by
sethish
12y ago
The main thing that GITenberg could provide authors is a toolset and workflow for using git to store books, and some kind of toolchain to turn them into epub or print-ready pdf. That toolchain is something I effectively have to build for GI
100.
▲
by
sethish
12y ago
Because parsing the original metadata from Project Gutenberg is time consuming to write. I wasn't going to submit it to HN until I had an index/search api, but someone beat me to it.
101.
▲
by
sethish
12y ago
It's all checked into git now, and I have an API to fetch the repo names via github's api. Migrating from github to self-hosted would now be easier than doing it from scratch. Github does have issues with Unicode repo names, so i
102.
▲
by
sethish
12y ago
GITenberg is more of a reference to Project Gutenberg at this point than Johannes. But no, I didn't know that at the time :-(
103.
▲
by
sethish
12y ago
I would love to help standardize this workflow, but it has to be a community process and discussion. I'm meeting with the inkling folks sometime next week to collect more information. re: Metadata. I'm slowly working on that probl
104.
▲
by
sethish
12y ago
I can't think of a VCS with a better online tool than git and github. With editing books, it is entirely possible to use only the github editor, which effectively abstracts the git command line interface.
105.
▲
by
sethish
12y ago
In retrospect, I didn't need to make 80k+ commits with my own account.
106.
▲
by
sethish
12y ago
Idea first, pun after. But thanks!
107.
▲
by
sethish
12y ago
PGDP does a series of passes on a book. They don't continually update it. If a book makes it through their process with a typo, or more common, made it through the transcription process 20 years ago, there aren't good systems in
108.
▲
by
sethish
12y ago
This is a pretty small python module at the moment. To fetch metadata, it downloads a 230mb zip from PG and parses it for a few categories. Project Gutenberg has some metadata about Subject, but the info is inconsistent, but there is somet
109.
▲
by
sethish
12y ago
To better differentiate from Project Gutenberg, GIT + Gutenberg. Case is ambigious on github, so I can change it later without breaking anyone's URLs.
110.
▲
by
sethish
12y ago
I think the previous version of the metadata included a path to the ftp server. Splitting the book id (4443 -> 4/4/4/4443) works for _most_ books, but there were somewhere between 800 and 3000 books organized in a differe
111.
▲
by
sethish
12y ago
The github repos are intended to collect issues and received pull requests. Project Gutenberg doesn't have a public bugtracker, nor do they use version control.
112.
▲
by
sethish
12y ago
This is fantastic! I just made a github repo for each Gutenber book: https://github.com/GITenberg This will be very helpful, the XML/RDF files are a hassle.
113.
▲
by
sethish
12y ago
Maintained by Tron guy! <3<3<3
114.
▲
by
sethish
13y ago
Slippery slope argument? Really?
115.
▲
by
sethish
13y ago
Justine Tunney is not the founder of the Occupy movement.
116.
▲
by
sethish
13y ago
This licensing terms of this article is not compatible with the GPL or the Debian Free Software Guidelines.
117.
▲
by
sethish
13y ago
Chromium is packaged in Ubuntu and Debian. It's not packaged in Fedora because Chromium forked a half-dozen libraries [instead of pushing changes upstream]( http://ostatic.com/blog/making-projects-easier-to-package
118.
▲
by
sethish
13y ago
Congratulations, you are a bigot.
119.
▲
by
sethish
13y ago
I disagree. Companies should allow and embrace cultures other than those of their founders. Not doing so is ignoring the issue. Yes, schools need to reform and a majority of the blame lies there well before SF companies have a chance.
120.
▲
by
sethish
13y ago
Chromium, the open-source-only version of Chrome is also a viable option.
More ›