4 ms·
Yes, that is a file with zero lines that ends with an "incomplete line". Processing of such files by standard line-oriented utilities is undefined in the opengr
by jepler 3y ago
Yes, that is a file with zero lines that ends with an "incomplete line". Processing of such files by standard line-oriented utilities is undefined in the opengroup spec. So, for instance, the effect of "grep"ping such a file is not defined. Heck, even "cat"ting such a file gives non-ideal results, such as colliding with the regular shell prompt. For this reason, a lot of software projects I work on check and correct this condition whenever creating a commit.
https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1_chap03.html#tag_03_403 https://pubs.opengroup.org/onlinepubs/9699919799/basedefs/V1... ("text file")
- rovr138 3y ago> Yes, that is a file with zero lines that ends with an "incomplete line". It's a file with zero complete lines. But it has 1 line, that's incomplete, right? The file starts empty. Anything in it starts "a line". So it's 1 incomplete line. I hate weird states.
- xyzzy_plugh 3y agoNo, it is valid for a file to have content but no lines. Semantically many libraries treat that as a line because while \n<EOF> means "the end of the last line" having just <EOF> adds additional complexity the user has to handle to read the remaining input. But by the book it's not "a line". If I said "ten buckets of water" does that mean ten full buckets? Or does a bucket with a drop in it count as "a bucket of water?" If I asked for ten buckets of water and you brought me nine and one half-full, is that acceptable? What about ten half-full buckets? A line ends in a newline. A file with no newlines in it has no lines.
- joshjje 3y agoThats beyond ridiculous. Most languages when you are reading a line from a file, and it doesn't have a \n terminator, its going to give you that line, not say, oops, this isn't a line sorry.
- LK5ZJwMwgBbHuVI 3y agoThat's a relatively recent invention compared to tools like `wc` (or your favorite `sh` for that matter). See also: https://perldoc.perl.org/functions/chop https://perldoc.perl.org/functions/chop wherein the norm was "just cut off the last character of the line, it will always be a newline"
- squeaky-clean 3y agoMost languages but not all. I've even been bit by this recently in cron. Assuming that EOF is identical to \\nEOF will end up causing trouble for you one day, because it's not actually identical.
- int_19h 3y agoI don't think you can meaningfully generalize to "most languages" here. To give an example, two extremely popular languages are C and Python. Both have a standard library function to read a line from a text stream - fgets() for C, readline() for Python. In both cases, the behavior is to read up to and including the newline character, but also to stop if EOF is encountered before then. Which means that the return value is different for terminated vs unterminated final lines in both languages - in particular, if there's no \n before EOF, the value returned is not a line (as it does not end with a newline), and you have to explicitly write your code to accommodate that.
- nativeit 3y agoI get this is largely a semantic debate, but find it a little ironic so many programmers seem put off with the idea of a line count that starts at “0”.
- akdev1l 3y agoNo, a line is defined as a sequence of characters (bytes?) with a line terminator at the end. Technically as per posix a file as you describe is actually a binary file without any lines. Basically just random binary data that happens to kind of look like a line.
- mort96 3y agoIt's a file with 0 lines and some trailing garbage.
- DougBTX 3y agoAnother way to look at it is that concatenating files should sum the line count. Concatenating two empty files produces an empty file, so 0 + 0 = 0. If “incomplete lines” are not counted as lines, then the maths still works out. If they counted as lines, it would end up as 1 + 1 = 1.
- coryrc 3y agoPedantically, if it doesn't end with a newline, it's considered a binary file and not a text file. Binary files don't have lines. In practice, most utilities expecting text files will still operate on it.
- PaulDavisThe1st 3y agoNo file has lines. "Lines" are a convention established by (or not) software reading a data stream.
- coryrc 3y agoAckshully
- wtetzner 3y agoThat's a weird way to look at it. Binary files might not have "lines", but there's no reason they couldn't include a byte with value 10 (the ASCII value for \n). Software reading that file wouldn't know the difference, right? Also, why couldn't you have a text file without any lines?
- coryrc 3y agoAll I'm addressing is GP's comment: It's a file with zero complete lines. But it has 1 line, that's incomplete, right? Because the Unix definition of text file requires the file to end with a newline. "Lines" only exist in the context of text files. If there's no terminating newline, it's (pedantically) not a text file and so has no lines. Now, in practice, if you open() that file in text mode, it doesn't TMK return an error if the terminating newline isn't present, but it's undefined behaviour. And if you do have a terminating newline, then you have at least one line :).
- pxc 3y agoHere's another way to think about this: This isn't a weird state. It's a language problem. An 'incomplete line' isn't a type of line, it's an unfortunate name for a thing that is not a line. Just like how the 'wor' is an incomplete word (the word 'word'), but 'wor' is, of course, not a word. Same thing for formalisms like equations in algebra or formulas in propositional logic— we have the phrase 'well-formed formula', and we might describe some sequences of terms as 'incomplete formulas' or perhaps 'ill-formed formulas', but those phrases don't describe anything that meets the formal system's definition of 'formula' at all— they are not formulas. 'Ill-formed formula' is not a compositional phrase where 'ill-formed' describes a feature of a 'formula'. It's a bit of convenient language for what we can intuitively or metaphorically recognize as a formula-ish thing.
- rerdavies 3y agoThe opengroup spec says no such thing.
- simonh 3y ago3.206 Line A sequence of zero or more non- <newline> characters plus a terminating <newline> character. See also ‘3.403 Text File’ for the definition of a text file. No new line characters, no lines. No lines, not a text file.
- wtetzner 3y ago> No lines, not a text file. That seems like a broken (maybe just bad?) definition/specification to me. A blob of JSON in a file isn't "text" if there's no newline character trailing it?
- simonh 3y agoThere are other definitions of a text file than the opengroup spec, particularly for specific OS platforms. I’m not sure what convention JSON follows. As a spec it’s fine. It defines a text file in such a way that you can easily write code to process such a file deterministicaly.