4 ms·
I can see wanting to stick to plain ASCII for identifiers, but what of embedded strings? Localization might prefer to be data-driven, but what of unit tests su
by MaulingMonkey 4y ago
I can see wanting to stick to plain ASCII for identifiers, but what of embedded strings? Localization might prefer to be data-driven, but what of unit tests surrounding text management - including locale-specific sorting tests, glyph rendering, etc.? I don't think you're against making the source code UTF-8 per se, just "UTF-8" identifiers. Which seems quite valid. In rust you might write:
#![forbid(non_ascii_idents)]
pub fn fiancée() {}
And get an appropriate error:
error: identifier contains non-ASCII characters
--> src/lib.rs:3:8
|
3 | pub fn fiancée() {}
| ^^^^^^^
|
note: the lint level is defined here
--> src/lib.rs:1:11
|
1 | #![forbid(non_ascii_idents)]
| ^^^^^^^^^^^^^^^^
Despite the source code indeed being UTF-8.
- geokon 4y agoWhy would you not just put test input in a separate test-input file that you pipe in? That way you can test multiple different encodings and all sorts of edge cases While I'm also team ASCII, it's possible I'm not appreciating some edge case here.
- MaulingMonkey 4y ago> Why would you not just put test input in a separate test-input file that you pipe in? You just added a dependency on File I/O (can't easily run the test on filesystemless embedded/wasm targets), deserialization, the current working directory if using relative paths, filesystem layout if using absolute paths - we lose type checking as part of our compile step, so deserialization might fail, we lose intellisense... > That way you can test multiple different encodings and all sorts of edge cases That can be done in code too. Let's take something concrete like https://github.com/openjdk/jdk/blob/master/test/jdk/java/util/Locale/bcp47u/CurrencyTests.java https://github.com/openjdk/jdk/blob/master/test/jdk/java/uti... : {JPY, Locale.forLanguageTag("ja-JP-u-rg-uszzzz"), "\uffe5"}, {JPY, Locale.forLanguageTag("en-US-u-rg-jpzzzz"), "\u00a5"}, {JPY, Locale.forLanguageTag("ko-KR-u-rg-jpzzzz"), "JP\u00a5"}, Well, a mixture of half and fullwidth yen symbols (¥¥). A bit awkward to eyeball as they've been escaped via \u#### codes - dodging the whole "what encoding are our source files" problem - but given the graphical variants of that glyph, I could see keeping the escaped versions. Now, the array this is helping construct could certainly be deserialized from, say, a JSON or XML test-input file. But what would we gain, exactly, besides more boilerplate and context switching to wade through? > While I'm also team ASCII, it's possible I'm not appreciating some edge case here. To be clear I'm not saying you can't use files, and sometimes files are more appropriate and convenient than loading everything into code despite the caveats I mentioned... but I'm not seeing much of a boon for exiling test data to files here.
- geokon 4y agothanks. Interesting examples - and thanks for actually finding in-code examples :) I didn't realize OpenJDK targets systems without filesystems. That's cool. I thought the Java Smart Card days are behind and the Java folks only care about server applications now. However having a requirement that all tests be inline seems a bit onerous. Do they inline images as well for testing BufferedImage and company? I can see how an external file makes it a bit more of an "integration test" that would in effect be testing several moving pieces. I don't agree it would create more boiler plate or make things any less clear though. It makes it much easier to test many different complex and large inputs and to feed in new tests without needing to recompile. It's also easy to version control and introduce new tests with pathological cases. It presents a clear separation of test inputs from code - but maybe I can understand the counterargument. It does look neater to have everything together
- xmcqdpt2 4y agoYeah the resources feature in Java https://www.jetbrains.com/help/idea/resource-files.html https://www.jetbrains.com/help/idea/resource-files.html makes it easy to embed arbitrary data with your code (or in this case, your tests, which are usually compiled separately) so no, you can easily test against files without access to a filesystem (beyond the access required to read the jars). I still think having test cases with strings in the code is often a lot clearer personally. (Unrelated: I wish more PLs made it easy to embed files within executables and access them with an OS-independent filesystem like API, it's often very useful.)
- MaulingMonkey 4y ago> I didn't realize OpenJDK targets systems without filesystems. To be fair I'm not sure it does, really, per se. I know embedded Java is/was a thing, though, and it wouldn't suprise me if someone somewhere tortured their own personal fork into running a test suite on embedded stuff - if only for legacy support testing purpouses. And poking around in that directory did make it clear some tests use files containing test vectors. > I don't agree it would create more boiler plate To get concrete again, I consider this boilerplate: https://github.com/openjdk/jdk/blob/9fc518ff8cadbbb731a016d8f53ed065b3561a7c/test/jdk/java/util/Locale/Bug4184873Test.java#L105-L113 https://github.com/openjdk/jdk/blob/9fc518ff8cadbbb731a016d8... https://github.com/openjdk/jdk/blob/9fc518ff8cadbbb731a016d8f53ed065b3561a7c/test/jdk/java/util/Locale/Bug4184873Test.java#L125-L133 https://github.com/openjdk/jdk/blob/9fc518ff8cadbbb731a016d8... And this is merely (de)serializing pure unstructured strings without any kind of data format or failiable schema beyond charset encoding, making this a poster child for exiling test data to files (probably part of the reason why it was exiled to files!) > It makes it much easier to test many different complex and large inputs and to feed in new tests without needing to recompile. For bulk plain text I'd agree. For codebases with slow incremental builds, structured data might also benefit from being exiled for iteration speed. Or there can be benefits to skipping having a full developer environment. But if incremental builds are fast (a worth goal), and if full developer environment is reasonably assumed (dev-focused unit/integration testing), compiler assistance with structured data is often more convenient, and has better error reporting for syntax errors etc. than what you'll get from many/most simple and straightforward uses of deserializers.
- dataflow 4y ago> what of embedded strings Escape sequences?
- Freak_NL 4y agoA great way to reduce the legibility of the code; not very useful beyond that.
- singron 4y agoIf you are specifically trying to test Unicode support, then you might want to try multiple formulations for the same glyph. E.g. you can make é either directly with 00e9 or with a combining mark 0065 0301. You can't tell the difference in a string literal, but escape codes will make it clear.