8 ms·
I stopped writing C because string manipulation sucked.
by drb91 7y ago
I stopped writing C because string manipulation sucked.
- aap_ 7y agoI never really got that complaint, I'd like to see some examples of what people consider so ugly about C strings.
- dgellow 7y agoWhy is this comment being down voted? Not everybody here is familiar with C string manipulation, if you down vote or complain at least give more detail that "it sucks". @aap_, I asked a similar question some time ago and got some answers, you can check the thread here: https://news.ycombinator.com/item?id=19302581 https://news.ycombinator.com/item?id=19302581 The direct answer I got was: > I'm guessing because an off-by-one or an extra skip might mean you miss the end of the string and go off into la-la land feeding whatever garbage happens to be in memory to your parser? That would mostly be a C issue (as it has no string abstraction at all).
- coldtea 7y ago>Why is this comment being down voted? Not everybody here is familiar with C string manipulation, if you down vote or complain at least give more detail that "it sucks". Well, if someone is not familiar, why do they read a subthread on the matter? Shouldn't they better start with a tutorial on C/C strings? Even if people on this thread gave arguments, how would they (not familiar with C and C strings) would evaluate them? They could be totally bogus.
- flohofwoe 7y agoI agree, but C (the language) doesn't even have the concept of a 'string'. It's just the convention how some C standard library functions interpret an array of bytes with a zero at the end. At least in C it's quite obvious that strings are not trivial if you want both an intuitive way to work with strings, and high performance. The C++ std::string type is neither intuitive to work with, nor does it allow to write high-performance code. For string processing it's really better to use another language with different trade-offs.
- pjmlp 7y agoC++ std::string is better and more secure and anything that C ever produced. As for string processing in general, I do agree that other languages are better suited.
- aap_ 7y agoMore often that not I find myself missing C-type strings in other languages. Being able to just walk through the characters and manipulating them is something I found rather ugly in python for instance. The NUL character is in my experience not so terrible, you typically have null pointers at the end of a linked list or whatever as well and nobody complains about that. Now I have to admit I had a bug recently that took me longer to fix then I would like to admit because I wasn't walking a string right, but usually I have very little trouble with them.
- adrianN 7y agoWhat's a character? UTF-8 makes that a bit difficult to answer. If you want arrays of ASCII bytes, you can have those in most programming languages.
- lugg 7y ago1 to 4 bytes. How does utf8 make that difficult to answer? When was the last time you iterated over a string of unicode points and said you know what would be handy right now? If these code points were split up into arbitrary and unusable bytes of memory.
- adrianN 7y agoWell, ä is a character in German. You can either write it as LATIN SMALL LETTER A WITH DIAERESIS, or you can use COMBINING DIAERESIS and a. When you iterate over the German word Mädchen as Unicode code points you might be confused. Other languages do much crazier things.
- 7y ago
- pjmlp 7y ago1 - string contents and actual length are handled in separate variables without correlation 2 - no enforcement that a null terminator actually exists in the string 3 - C brags about performance and is probably the slowest language to compute string length 4 - manipulating strings requires very carefull handling of buffers, usually forcing everyone to use the heap as easier way out
- owl57 7y ago3 - C brags about performance and is probably the slowest language to compute string length I have a feeling that of all axes of performace C cares the most about memory overhead. Then the obvious idea is to have it at exatly one byte per "simple" string, and you get to pick the class of programs that can't get away with that default string type: • One-byte terminator: complicated text-handling application with a lot of (longer than a couple of pointers on average) string slices. • One-byte length: anything that needs strings longer than 255 chars. And then of these two solutions you pick the obviously more general one. What could possibly go wrong?
- pjmlp 7y agoAnything that needs strings longer than 255 chars already had a solution in existing systems programming languages back when C was born. Character arrays with open length, bound checked. Naturally it requires better compiler support than C authors were willing to implement.
- coldtea 7y ago>Naturally it requires better compiler support than C authors were willing to implement. Which is still the case with many things in Go, a language of close origin to C (though this time not about strings).
- pjmlp 7y agoIt is interesting how in both cases they disregarded what was being made around them.
- coldtea 7y agoIf this a joke? Length not known, so prone to overflows at anytime, atrocious standard library, ... (and let's not even go into the Unicode situation).
- mhh__ 7y agoYou basically have to write your own high level string library just to approximate the features of just about any other language...
- drb91 7y agoI am finding hard to imagine not seeing the difficulty here, so instead I’m just gonna point out simple operations like stripping whitespace, splitting strings on a character pattern, changing case, dealing with character encodings, regex matches all require manually iterating and mutating or copying strings and in the case of regexes require compiling and auditing various libraries. The abstractions other standard libraries have used, such as rust, make it much easier to simply express the string operations as high level operations and spend your time elsewhere while retaining relatively high levels of performance. Often, string processing is not in the inner loop and does not benefit from things like combining multiple string operations into a single pass, traditionally a thing that might make c perform better all other things being equal.
- ChrisRR 7y agoI agree with that. Recently I started reading "Writing an interpreter in Go" and thought I'd follow along using C. From the first chapter, the Go code starts using strings as a short-cut to represent tokens . In other languages this is trivial because strings are very easy to create, resize, change, etc. Using C, this became an issue though, as using strings became a roadblock where I started having to implement different solutions rather than focusing my attention on the contents of the book.