4 ms·
Taking it in a slightly different direction – because in my work chunking is nearly never necessary (too little text or too much memory) whereas Unicode charact
by Mlller 4y ago
Taking it in a slightly different direction – because in my work chunking is nearly never necessary (too little text or too much memory) whereas Unicode characters and punctuation nearly always is – I would tend to something like:
from collections import Counter
from re import finditer
word_counts = Counter()
with open(PATH, encoding = 'utf-8') as file:
doc = file.read()
word_counts.update(match.group().lower() for match in finditer(r'\w+', doc))
In Python, this could be a fairly performant way of taking unicode into account, because it doesnʼt use Pythonʼs for-loop, and regular expressions are rather optimized compared with writing low-level-style Python. (Maybe it should use casefold instead of lower.)