10 ms·
I'm pretty sure a simple while loop is better than either: size_t strlen(const char * s) { size_t index = 0; while(s[index] != 0) { index += 1;
by pbsd 10y ago
I'm pretty sure a simple while loop is better than either:
size_t strlen(const char * s) {
size_t index = 0;
while(s[index] != 0) {
index += 1;
}
return index;
}
- lmm 10y agoThat's substantially harder to follow than the code in the article - you have to jump around to see where index changes and how that interacts with the loop condition.
- matthewaveryusa 10y agoJust out of curiosity how would you rank strlen{1,2,3} in terms of difficulty to read? size_t strlen1(const char * s) { const char* it =s ; while(*it != '\0') { ++it; } return it - s; } size_t ptr_distance(const void *begin, const void* end) { return end - begin; } size_t strlen2(const char * s) { const char* it =s ; while(*it != '\0') { ++it; } return ptr_distance(s,it); } int is_null(const char* p) { return *p == '\0'; } const char* find_first(const char* s, int (*predicate)(const char*)) { while(1) { if(predicate(s)) { return s; } ++s; } } size_t strlen3(const char * s) { return ptr_distance(s,find_first(s,is_null)); }
- kbenson 10y agoAs someone whos' written almost no C in over a decade, and read minimal C in that time, I would say strlen1 is clearest because it's so simple and succinct, mainly because of how simple the underlying algorithm is. strlen3 is easiest to read, as you can make fairly valid assumptions about the helper functions, and easily confirm them if needed, and it very clearly breaks down the algorithm into its conceptual components. I think it would also be clearest (and thus winner) if the algorithm was slightly more involved, as the gains in clarity would outweigh the gains of having everything defined in just a few lines. strlen2 is worst, because the helper function gains you very little for the cognitive overload of having the implementation being somewhere else. You are just as well served in this case by a trailing comment, IMO.
- ThatGeoGuy 10y agoI agree with this comment more or less, however I would like to point out that strlen{1,2,3} are all wrong, at least in the capacity that they take pointer differences and return a size_t. In truth, `ptr_distance` and `return it - s;` should both be of type ptrdiff_t, which is not necessarily the same as size_t. In some ways this is why using an explicit index of type size_t that can be incremented is better, because you can avoid some type casting.
- lmm 10y agoI would say strlen3 is easiest to read (assuming we're in a codebase that makes widespread use of find_first, ptr_distance and is_null), then strlen1, then strlen2. I think thinking in terms of ptr distances is still needlessly confusing. It's hard to express what we need directly in C (even the Rust version is using nonlocal return which is not the best thing for readability - we could emulate it with goto but that has its own issues) where we don't have generics, anonymous functions, multiple return or tail call elimination. I mean what I'd write in my head (or another language) translates into C as something like: typedef struct { boolean continue; size_t value; } result_size_t; int indexed_cata_size_t( const char *s, void (*operate)(const char, size_t, result_size_t*)) { const char *it = s; size_t idx = 0; // May get initialization syntax slightly wrong but you get the idea result_size_t result; result.continue = true; while(result.continue) operate(it++, idx++, &result); return result -> value; } void index_of_first_null_u(const char c, size_t idx, result_size_t *result) { if(c == '\0') { result -> value = idx; result -> continue = false; } } size_t strlen4(const char *s) { return indexed_cata_size_t(s, index_of_first_null_u); } but I'm not sure that ends up any clearer.
- burfog 10y agoNice work, maybe. I really don't know if this is a joke. Wow. FWIW, strlen in C is 5 to 7 lines of code depending on how you format your code. See my other comment. For those who are not experienced C programmers, yes the shorter code is more readable.
- lmm 10y agoresult_size_t and indexed_cata_size_t are generic things that would only need to be written once / in the standard library. (perhaps as macros so that they could be generic i.e. not tied to size_t). So the thing to compare is index_of_first_null_u and strlen4 vs alternate implementation, not the whole set of code that I wrote. And yeah, it's still not nice. In my preferred language it would be something like: def strlen(s: String) = s.cata[size_t] { case Cons(head, tail) => if('\0' == head) 0 else 1 + tail } (yes, a cons list is not the same as an array and this distinction is important at this level - but the logic should look similar) I do think there's value in having a clear separation between what kind of looping you're doing and the per-entry logic (index_of_first_null_u). Is it worth the overhead of doing that in C? Probably not for native C programmers, but if I was writing C to be maintained by functional programmers (or me) then I probably would do it.
- deleted 10y ago[deleted]