7 ms·
def get_corrected_stats(ids): corrected_counts = {} i = 0 while i < len(ids) - 1: current_pair = (ids[i], ids[i+1]) # If the pair i
by gzer0 3y ago
def get_corrected_stats(ids):
corrected_counts = {}
i = 0
while i < len(ids) - 1:
current_pair = (ids[i], ids[i+1])
# If the pair is new or if the current element is not a repeat of the previous one (or it's the first element)
if current_pair not in corrected_counts:
corrected_counts[current_pair] = 1
else:
if i == 0 or ids[i] != ids[i-1]:
corrected_counts[current_pair] += 1
# If the current and next element are the same, skip to the next unique element
if i < len(ids) - 2 and ids[i] == ids[i+1]:
while i < len(ids) - 1 and ids[i] == ids[i+1]:
i += 1
else:
i += 1
I have a basic grasp of this topic. Can you explain why the proposed solution is (or isn’t) effective in this scenario?
- thesz 3y agoI do not understand why your solution skips to the next unique elements. Consider sequence [1,1,1,1,1,1]. We should count there 3 (three) (1,1) pairs. Karpathy's (and mine's) algorithm would count 5 pairs, yours' will count 1. One should just skip one or two elements counting pairs, one element when pair's elements are different, two when they are same. This may break vectorization, though. I also should note that error is here, but no-one hits it. My experiments show that (A,A) pairs gets counted as most frequent ahead of their natural turn about half a dozen symbols earlier. I also never saw that symbol B, assigned for pair (A,A), is used in any most frequent pair before an actual turn for (A,A). One need to construct special string for that to happen.
- theGnuMe 3y agoIt sounds like that sequence is being k-merized. In the case of byte-pair that should be a look-ahead of 1 and then a shift up of 1. Maybe I am missing something.