2 ms·
Since we're on the subject, I'd like to take the time to mention UNIHAN, and also an effort I've created to make historical / regional variants of CJK character
by git-pull 8y ago
Since we're on the subject, I'd like to take the time to mention UNIHAN, and also an effort I've created to make historical / regional variants of CJK characters more easily accessible.
UNIHAN is Unicode's Han Unification effort. It handles the variant issue - but actually also goes a step further, citing information from paper books, including stroke information, definitions, and even pronunciations [1].
I've compiled a overview of UNIHAN at https://unihan-etl.git-pull.com/en/latest/unihan.html https://unihan-etl.git-pull.com/en/latest/unihan.html.
unihan-etl is a project I've created that allows extracting the contents' of UNIHAN's database: https://unihan-etl.git-pull.com https://unihan-etl.git-pull.com. It can be used as a Python library, or a self-serve export of the database to a tabular or structured format.
In addition, there is something I've worked on to make this data also available in SQLAlchemy / DB form: https://unihan-db.git-pull.com https://unihan-db.git-pull.com
And also using it as a basis for a spiritual successor to cjklib [2]: https://cihai.git-pull.com https://cihai.git-pull.com
[1] https://www.unicode.org/reports/tr38/ https://www.unicode.org/reports/tr38/
[2] https://github.com/cburgmer/cjklib https://github.com/cburgmer/cjklib