3 ms·
I'm familiar with the Synopticon, which would be fun to structure. I didn’t do OCR myself, except for the topic index and to fill in a few gaps. I started from
by ahaspel 5mo ago
I'm familiar with the Synopticon, which would be fun to structure.
I didn’t do OCR myself, except for the topic index and to fill in a few gaps. I started from existing Wikisource text and then built a pipeline around that: cleaning (headers, hyphenation, etc.), detecting article boundaries, reconstructing sections, and linking things back to the original page images. Most of the effort went into rendering the complex layouts, and handling the cross-linking, not the initial ingestion.
Glad to go into more detail if you’re interested, but that’s the gist of it.
- peterldowns 5mo agoAh ok thanks very much!
- xnobodyx 5mo agowould love to hear more details. are you familiar with the semantic lab at pratt's work - https://semlab.io/projects https://semlab.io/projects (also see https://tools.semlab.io https://tools.semlab.io )?