3 ms·
just curious, how do you mark a word as common? I often used jisho.org or classic.jisho.org (from wwwjdic) and from what i remember, a lot of those words marked
by deeteecee 11y ago
just curious, how do you mark a word as common? I often used jisho.org or classic.jisho.org (from wwwjdic) and from what i remember, a lot of those words marked as "common" were mostly definitely not.
- chrisvasselli 11y agoYeah, I found the commonalities in WWWJDic/JMDict to be pretty problematic, so I came up with my own. I use a combination of corpuses including newspapers, novels, literature/poetry, and spoken language. Some of these have been hand-parsed by humans, and others I parse using the same parsing as clippings. I think the result is pretty good!
- glandium 11y agoCommonality is a tricky thing. There are words that most native won't actually know, words they know but rarely use, words that used to be trendy but aren't anymore, etc. It's hard to know from "common", "uncommon", and "rare" in what kind of bucket a word or expression would enter, while "native won't actually know" is a very important distinction to make. Moreover, it's not clear which of "uncommon" and "rare" ranks above the other and below "common", both words being synonyms. To give an example, since I gave a try to your app (best I've found so far, by the way, but unfortunately doesn't match my own needs): 自問自答 is marked rare. 出勤 is marked uncommon. Now, looking at this and some other words, my guess is that uncommon is above rare. Fine. 自問自答 might be rare, but there's no Japanese I know (excluding small kids) that wouldn't understand it. (And, in fact, it stuck in my mind because I keep hearing it). So in your classification, there is no room to tell apart those words that most japanese won't know.
- chrisvasselli 11y agoThanks, it's useful to hear that the difference between "uncommon" and "rare" was unclear to you. I'll think about how to make that better. I definitely would love to have a distinction in the app for "this is a word that any educated Japanese person would know". I haven't found the dataset yet that can let me build that, unfortunately. I checked, and in the case of 自問自答, it actually doesn't appear even once in any of the corpuses I'm using. So it seems like we need some other source of data.
- glandium 11y agoThe data used for commonality in JMDict is a little dated (1998) and biased (exclusively based on newspapers, which tend to have specific vocabulary), which I guess is why you mentioned they were problematic. However, there are more recent data sets available on the Monash ftp archive. http://ftp.monash.edu.au/pub/nihongo/ http://ftp.monash.edu.au/pub/nihongo/ . For example, there's one dataset from 2008 using blog entries from goo.ne.jp, and another with novels. They could be good additions to your corpus (fwiw, for 自問自答 there are 18038 occurrences indicated in the dataset for goo.ne.jp and 68 in the dataset from novels). Certainly, doing some similar work with current data from the net would be useful too. I wish there was some regular scraping done, so that we could always use fresh data. Hell, I wish Google, Bing or any other search engine were just giving out such word frequencies from their spider bots data (and not just for japanese).
- chrisvasselli 11y agoHmm, I actually use that novels corpus. Sounds like you may have found a bug in Nihongo. I'll look into it. Thanks!