4 ms·
A few years ago I downloaded several hundreds of megabytes of Japanese subtitles, split into 3 categories: live action/drama, anime and foreign film/tv I’ve li
by olsgaarddk 7y ago
A few years ago I downloaded several hundreds of megabytes of Japanese subtitles, split into 3 categories: live action/drama, anime and foreign film/tv
I’ve listed them in a google sheets together with a few other corpora
https://docs.google.com/spreadsheets/d/1yb5dq4ahdwc_g0aQTL3YM6i2mKiZ2m-AvhpwygbZD4A https://docs.google.com/spreadsheets/d/1yb5dq4ahdwc_g0aQTL3Y...
Choose the jimaku tab for subtitles to see how big the variation between corpus can be.
According to other comments here, it appears that OP list is based on a newspaper corpus from 1993.
- echelon 7y agoThe source links appear to no longer work. Do you know where we can download Japanese subtitles? I would love to attempt to segment a bunch of Japanese subtitles into words and then do frequency analysis. My interest is in increasing my listening ability, so I want to put the most frequently spoken words into SRS/Anki, and perhaps even break it down by anime. Alternatively, has anyone already done this?
- olsgaarddk 7y agoThat was my initial goal, but I had a lot of trouble with vanilla MeCab not understanding a lot of the text. But this was before neologd, so i think it would work better now. I don’t have the source code on me, but I scraped it from a website that publishes subtitles. The scraping was easy, the cleaning not, and I believe this spreadsheet is generated from my first attempt at cleaning. A lot of sources in Japanese nlp and linguistics have a bad habit of changing url often, so it bitrots easily. Sorry.