3 ms·
Right out of the bag lots of wasted lines word_bag: dict[str, int] = dict() # Multiset for line in raw_corpus: words = line.split()
by SethTro 3y ago
Right out of the bag lots of wasted lines
word_bag: dict[str, int] = dict() # Multiset
for line in raw_corpus:
words = line.split()
for word in words:
if word in word_bag:
word_bag[word] += 1
else:
word_bag[word] = 1
keys_to_drop = []
for k, v in word_bag.items():
if v < min_frequency:
keys_to_drop.append(k)
for k in keys_to_drop:
del word_bag[k]
pprint(word_bag)
print(len(word_bag))
vocabulary = word_bag.keys()
return set(vocabulary)
Can be a python one/two liner
# word_bag = Counter() # defaultdict(int)
word_bag = Counter(word for line in raw_corpus for word in line.split())
return set(word for word, count in word_bag.items() if count >= min_frequency)
- CamperBob2 3y agoSometimes Python's conciseness works against its value as a didactic language.
- Y_Y 3y agoThe original code looked like someone had learned old-school C++ and just shoehorned that into python. This phenomenon is all over physics. The fixed code isn't just short, it's idiomatic and clear (YMMV) and hence much easier to understand.
- nine_k 3y agoReasonably modern C++ would allow you to define a Counter class, and to define maps and filters, if the stdlib versions don't work for you for whatever reason. Fortran, on the other hand,...
- NathanFulton 3y agoI collect mini implementations of ML things for teaching purposes. In this case, the longer-form version is a better artifact. Pithy readable implementations of core ideas have a lot of value. I don't see much value in code golfing besides having fun :)
- nine_k 3y agoI politely disagree. Both pieces of code have a teaching value, but different. The iterative code is busy and long, but allows to track exactly how the pretty trivial calculation is happening, down to elementary(-ish) operations. The comprehension-based code is more declarative; it succinctly shows what is happening, in almost plain English, without the minute details cluttering out the purpose of the code. For anyone who is not a Python beginner, but is an ML beginner, the shorter version is much more approachable, as it puts the subject matter more front-and-center. (Imagine that every matrix multiplication would be written using explicit loops, instead of one "multiply" operation. Would it clarify linear algebra for you, or the other way around?)
- NathanFulton 3y ago> For anyone who is not a Python beginner, but is an ML beginner, the shorter version is much more approachable, as it puts the subject matter more front-and-center. It certainly depends on the audience. Interestingly, I had the opposite conclusion about Python beginners in my head before reaching this line! I think it's more about the learner's prior background. Lately, I've mostly been helping friends who do a lot of scientific computing get started in ML. For that audience, the "nested loops" presentation is typically much easier to grok. > (Imagine that every matrix multiplication would be written using explicit loops, instead of one "multiply" operation. Would it clarify linear algebra for you, or the other way around?) Obviously "every" would be terrible. But there's a real question here if we flip "every" to "first". For a work-a-day mathematician who doesn't write code often, certainly not! For a work-a-day programmer who didn't take or doesn't remember linear algebra, the loopy version is probably worth showing once before moving on. On a related note: I sometimes find folds easier to understand than loops. Other times find loops easier to understand than folds. I'm not particularly sure why. Probably having both versions stashed away and either exercising judgement based on the learner at hand -- or just showing both -- is the best option.
- BeetleB 3y agoI sympathize with your position, but in this case the two liner is significantly more readable. I had no problem digesting it. I then looked at the longer version and it was a much higher cognitive load to digest that. I have to look at multiple for loops to realize they're merely counting the words in a corpus. The shorter version lets me see it immediately.