3 ms·
A lot of foreign words are frequently met in English texts, especially in narrow domains like philosophy, biology and medicine, etc. And seing that: http://en
by pretoriusB 14y ago
A lot of foreign words are frequently met in English texts, especially in narrow domains like philosophy, biology and medicine, etc.
And seing that:
http://en.wikipedia.org/wiki/Deutsche_Forschungsgemeinschaft http://en.wikipedia.org/wiki/Deutsche_Forschungsgemeinschaft
is a research institute, one would expect tons of mentions for this word in _english_ scientific papers.
But that is an outlier, and given the vastness of the corpus those would have been sorted out.
That said, there would have been an easy way to filter non English books out of the way automatically: do a statistical analysis on each book (e.g on letter frequency) and reject the ones that stray too much from the norm -- or send them to a secondary filtering stage, e.g by word presence or a verification by a human. Done carefully that filtering would not harm the actual results at all (e.g by presuposing a specific letter frequency, because it would only reject extreme outliers that would indeed by non-english works).
- arrrg 14y agoIt’s not a research institute. It gives grants to researchers (it’s the biggest organization in Germany giving research grants), it’s a foundation. As such it is also often mentioned in scientific papers. (“This study was funded in part by a grant from the …”)