27 ms·
The Norway Problem
- pintxo 6y ago> “While the website went down and we were losing money we chased down a number of loose ends until finally finding the root cause.” And that's why you have a staging environment. Or you debug in production, whatever you prefer.
- eitland 6y agoEverybody has a testing environment. Some people are lucky enough enough to have a totally separate environment to run production in. https://mobile.twitter.com/stahnma/status/634849376343429120 https://mobile.twitter.com/stahnma/status/634849376343429120
- atoav 6y agoI'd go further and say this is why you write tests. Creating tests that cover a lot (or all) possible inputs is sometimes not that hard and really pays off if you manage to catch a very common error like the Norway thing. Even better if you catch something that would have been a nightmare to fix in production. I say this because two days ago I wrote a test that used all country codes as input. It took 15 minutes to write that test. During the whole testing session I found at least 5 mistakes of which 3 would have been quite dramatic.
- simion314 6y ago>I say this because two days ago I wrote a test that used all country codes as input. It took 15 minutes to write that test. During the whole testing session I found at least 5 mistakes of which 3 would have been quite dramatic. And how many minutes to test all city/state/region/street/person names ? It can also happen that you test s will become outdated, like when url standard changed and more characters codes were allowed.
- atoav 6y agoFor something like URLs I'd use the hypothesis Python module and rely on their implementation of URLs (and if that changes the test will fail for newly formated URLs), for everything "custom" I would extract problematic test cases and include them as examples. Testing doesn't take too long on my machine (maybe 10 seconds), but even if it would, it would be totally acceptable as I run it pre commit only.
- groundCode 6y agoBugs make it to production no matter how careful you are. What matters is how you deal with incidents as an organisation, not that you should never release a bug.
- mrighele 6y agoOr you just return to the previos (and working) version of the website while you fix the issue. At least if you a good old monolith; if you have 10s of microservices it may be more complicated
- pietroppeter 6y agoI like strict yaml but I have used it very little. Anyone who uses it more that can give more feedback?
- grenoire 6y agoI was helping out a friend of mine in the risk department of a Big 4; he was parsing CSV data from a client's portfolio. Once he started parsing it, he was getting random NaNs (pandas' nan type, to be more accurate). I couldn't get access to the original dataset but the column gave it away. Namibia's 2-letter ISO country code is NA—which happens to be in pandas' default list of NaN equivalent strings. It was a headache and a half...
- grenoire 6y agoVerbatim from the docs, on read-csv: na_valuesscalar, str, list-like, or dict, default None Additional strings to recognize as NA/NaN. If dict passed, specific per-column NA values. By default the following values are interpreted as NaN: ‘’, ‘#N/A’, ‘#N/A N/A’, ‘#NA’, ‘-1.#IND’, ‘-1.#QNAN’, ‘-NaN’, ‘-nan’, ‘1.#IND’, ‘1.#QNAN’, ‘<NA>’, ‘N/A’, ‘NA’, ‘NULL’, ‘NaN’, ‘n/a’, ‘nan’, ‘null’. You fix it by using `keep_default_na=False`, by the way.
- mseepgood 6y agoA Ms True also broke Apple's iCloud: https://twitter.com/RachelTrue/status/1365461618977476610 https://twitter.com/RachelTrue/status/1365461618977476610
- grenoire 6y agoThat looks like an interesting hard-coded check, I wonder what it intended to fix.
- fanf2 6y agoThere’s some analysis in this twitter thread: https://twitter.com/badedgecases/status/1368362392573317120 https://twitter.com/badedgecases/status/1368362392573317120 tl;dr: there are a bunch of fields of various types that arrive as strings, and they get coerced but without paying attention to which field should have which type
- yakshaving_jgt 6y ago> it’s equally true that extremely strict type systems require a lot more upfront and the law of diminishing returns applies to type strictness - a cogent answer to the question “why is so little software written in haskell?“ I was with the article up until that point. I don't agree that diminishing returns with regards to type strictness applies linearly. Term-level Haskell is not massively harder than writing most equivalent code in JavaScript — in fact I'd say it's easier and you reap greater benefit. Perhaps it's a different story when you go all-in on type-level programming, but I'm not sure that's what the author was getting at. This smells of the Middle Ground logical fallacy to me. Or of course the comment was tongue-in-cheek and I'm overreacting.
- choeger 6y agoThat law of diminishing returns might actually apply, I am not 100% sure. But more powerful type systems allow for the more complex composition of more complex interfaces in a safe manner. Think of higher-level modules and data structures. Or dependent types and input handling. Or linear types and resource handling.
- samvher 6y agoI agree. I would say that Erlang goes ~80% of the way compared to Haskell's type system and the last 20% really matter, to the point that in many cases I find myself not really using Erlang's (optional) type system at all. Better type coverage and more descriptive types allow the compiler to infer more and I'd say this is the opposite of diminishing returns.
- 7952 6y agoI had to rewrite some JavaScript code in Postgres recently that measured the overlap between different elevation ranges. In JS I had to write it myself and deal with the edge cases and bugs. In Postgres I just use the range type and some operators. It was brilliant in comparison. The tiny effort of learning it was worth it. The list of data types I use all the time is bigger than just string, numbers and booleans. Serialisation formats should support them. Particularly as there are often text format standards that already exist for a lot of them. Give me wkt geometry and iso formatted dates. It's not that difficult and totally with it.
- kstenerud 6y agoThe worst tragedy of this is the security implications of subtly different parsers. As your application surface increases, you're likely to mix languages (and thus different parsers), which means that the same input data will produce different output data depending on whether your parser replaces, truncates, ignores, or otherwise attempts to automatically "fix up" the data. A carefully crafted document could exploit this to trick your data storage layer into storing truncated data that elevates privileges or sets zero cost, while your access control layer that ignores or replaces the data is perfectly happy to let the bad document pass by. And here's something else to keep you up at night: Just think of how many unintentional land mines lurk in your serialized data, waiting to blow up spectacularly (or even worse, silently) as soon as you attempt to change implementation technologies! This is why I've been so anal about consistent decoder behavior in Concise Encoding https://github.com/kstenerud/concise-encoding/blob/master/ce-structure.md#security-and-limits https://github.com/kstenerud/concise-encoding/blob/master/ce... https://concise-encoding.org/ https://concise-encoding.org/
- yellowapple 6y agoThis is exactly why configuration/serialization formats should make as few assumptions about value types as possible. Once parsing's done, everything should be a string (or possibly a symbol/atom, if the program ingesting such a file supports those), and it should be up to the application to convert values to the types it expects. This is Tcl's approach, and it's about as sensible as it gets. ...which is why it pains me to admit that in my own project for a Tcl-like scripting/config language[1] I missed the float v. string issue, so it'll currently "cleverly" return different types for 1.2 (float) v. 1.2.3 (atom). Coincidentally, I started work on a "stringy" alternative interpreter that hews closer to Tcl's philosophy (to fix a separate issue - namely, to avoid dynamically generating atoms, and therefore avoid crashing the Erlang VM when given potentially-adversarial input), so I'm gonna fix that case for at least the "stringy" mode (by emitting strings instead of numbers, too), knocking out two birds with one stone for the upcoming 0.3.0 release :) ---- [1]: https://otpcl.github.io https://otpcl.github.io, for those curious
- dkersten 6y agoIt’s reasons like this that I want my configuration languages to be explicit and unambiguous. This is why I use JSON or if I want a human friendly format, TOML. Strings are always “quoted” and numbers are always unquoted 1.2, it can never accidentally parse one as the other. The convenience of omitting quotes is just not worth the potential for ambiguity or edge cases to me.
- deleted 6y ago[deleted]
- progval 6y ago> Once parsing's done, everything should be a string Or give a schema to the parser, defining what type is expected in each field.
- kenshoen 6y agoYes, that looks like a right way to handle this problem without ignoring YAML spec. Define what to parse upfront.
- eitland 6y agoThere exists a couple of mainstream languages that are full of these sorts of interesting behavior, one of them is supposedly cool and productive and the other is supposedly ugly and evil.
- drno123 6y agoPython vs JavaScript?
- speedgoose 6y agoPython vs PHP also.
- pansa2 6y ago> full of these sorts of interesting behavior I don’t think that applies to Python - it’s quite strongly (although not statically) typed. I agree that it does apply to JavaScript and PHP.
- eitland 6y agoJavascript and PHP is correct.
- exyi 6y agoI think this applies to Python pretty well. Although certainly not as bad as PHP, most JS traps also exist in Python (falsy values, optional glitchy semicolons, function scoped variables, mutable closure). There is many JS specific traps like this and also other Python specific ones (like static fields are also instance fields, Python versions and library dependency hell). However I find it easier to avoid them in JS than in Python with TypeScript, avoiding classes, ...
- lunfard00 6y agoand yet I don't see anyone complain about bash which is arguably far worse than those 2. When things get hard on bash, you will start to see python scripts on CI and whole thing is complete unreadable mess
- RcouF1uZ4gsC 6y agoYAML seems like a really neat idea, but over time, I have I have come to regard it as being too complicated for me to use for configuration. My personal favorite is TOML, but I would even prefer plain JSON over YAML The last thing I want at 2 AM when trying to look figure out if an outage is due to a configuration change is having to think if each line of my configuration is doing the thing I want. YAML prizes making data look nicely formatted over simplicity or precision. That for me, is not a tradeoff, I am willing to make.
- Arnavion 6y agoThey all have their downsides. JSON: - no comments, unless you fake them with fake properties, unless your configuration has a schema that doesn't allow extra fake properties - no trailing commas; makes editing more annoying - no raw strings YAML: - the automatic type coercion - the many ways to encode strings ( https://yaml-multiline.info/ https://yaml-multiline.info/ ) - the roulette wheel of whether this particular parser is anal about two-space indentation or accepts anything as long as it's used consistently - the roulette wheel of whether this particular parser supports uncommon features like anchors TOML: - runtime footguns in automated serialization ( https://news.ycombinator.com/item?id=24853386 https://news.ycombinator.com/item?id=24853386 ) - hard to represent deeply-nested structures, unless you switch to inline tables which are like JSON but just different enough to be annoying
- perlgeek 6y agoFor hand-writing I love jsonnet, which produces JSON, is much more convenient to write, and has some templating, functions etc. https://jsonnet.org/ https://jsonnet.org/ You wouldn't serialize data structures to jsonnet though, you'd just generate JSON.
- anticristi 6y agoThis makes me sad. It's 2021 and we still haven't figure out how to serialize configuration in a format that is easy-to-edit and predictable.
- 6y ago
- groundCode 6y agoI've hit this exact same problem loading YAML in Ruby. Luckily caught it before it hit prod, but still, it made me go argh for a while.
- LeonM 6y agoI was bitten with this issue some time ago. The Stripe library has constants for which type of VAT number is supplied. One of those constants is 'NO_VAT'... Needless to say, this caused me some grey hairs
- schoen 6y agoRecent related HN discussion: https://news.ycombinator.com/item?id=26365365 https://news.ycombinator.com/item?id=26365365
- azernik 6y agoYAML had a worse example, once. For the ease of entering time units YAML 1.1 parsed any set of two digits, separated by colons, as a number in sexagesimal (base 60). So 1:11:00 would parse to the integer 4260, as in 1 hour and 11 minutes equals 4260 seconds. Now try plugging MAC addresses into that parser. The most annoying part is that the MAC addresses would only be mis-parsed if there were no hex digits in the string. Like the bug in this post, it could only be reproduced with specific values. Generally, if you're doing implicit typing, you need to keep the number of cases as low as possible, and preferably error out in case of ambiguity.
- m463 6y agoslightly related, on my microwave 99 > 100, even 61 > 100
- bonzini 6y agoWhy does your microwave compare numbers?
- mckirk 6y agoHow else would you prove it's turing complete and can run Doom?
- pdpi 6y agoNot the OP, but I have the same problem. For some reason that escapes me, pressing the “10 sec” button 7 times produces 00 70 instead of 01 10. If you then press the “1 min” button you get 01 70
- nerdponx 6y agoMost microwaves (in the USA) do this, at least in my experience. They treat the ":" like a sum of two sexagesimal numbers, rather than a sexagesimal digit separator.
- sokoloff 6y ago
- keeperofdakeys 6y agoThis is part of more general problem, they had to rename a gene to stop excel auto-completing it into a date. https://www.theverge.com/2020/8/6/21355674/human-genes-rename-microsoft-excel-misreading-dates https://www.theverge.com/2020/8/6/21355674/human-genes-renam... Edit: Apparently Excel has its own Norway Problem ... https://answers.microsoft.com/en-us/msoffice/forum/msoffice_excel-mso_win10-mso_o365b/norway-not-showing-on-maps-in-excel/a0a9d886-598b-46e1-a2f9-35794960b89a https://answers.microsoft.com/en-us/msoffice/forum/msoffice_...
- masklinn 6y ago> This is part of more general problem The more general problem basically being sentinel values (which these sorts of inferences can be treated as) in stringly-typed contexts: if everything is a string and you match some of those for special consideration, you will eventually match them in a context where that's wholly incorrect, and break something.
- chrisdone 6y agoThat’s a shrewd observation. Static types help with this somewhat. E.g. in Inflex, if I import some CSV and the string “00.10” as 0.1, then later when you try to do work on it like x == “00.10” You’ll get a type error that x is a decimal and the string literal is a string. So then you know you have to reimport it in the right way. So the type system told you that an assumption was violated. This won’t always happen, though. E.g. sort by this field will happily do a decimal sort instead of the string 00.10. The best approach is to ask the user at import time “here is my guess, feel free to correct me”. Excel/Inflex have this opportunity, but YAML doesn’t. That is, aside from explicit schemas. Mostly, we don’t have a schema.
- dalbasal 6y agoIf we're talking about general problems, then I don't think we can be satisfied with "sometimes it's a problem with types and sometimes it's a UI bug." That's not general.
- 6y ago
- sdfhbdf 6y agoWhat I am most baffled by with Yaml is the fact that it’s a superset of JSON. Whenever an input accepts YAML you can actually pass in JSON there and it’ll be valid It really surprised me when I found out and I use JSON Whenever possible since then since it’s much stricter https://en.m.wikipedia.org/wiki/JSON#YAML https://en.m.wikipedia.org/wiki/JSON#YAML
- dragonwriter 6y ago> Whenever an input accepts YAML you can actually pass in JSON there and it’ll be valid Strictly speaking, this is only true of YAML 1.2, not YAML 1.0-1.1 (the article here addresses YAML 1.1 behavior, the headline example od which was removed ib YAML 1.2 twelve years ago), though it calla YAML 1.1 “YAML 2.0”, which doesn’t actually exists. Of course, there are lots of features, like custom types, that JSON doesn’t support, but you can still use YAML’s JSON-style syntax instead of actual JSON, for them.
- alephu5 6y agoYes this is usually the best way. If you need some features for code reuse there are several preprocessors. I personally use Dhall to configure everything and then convert it to JSON for my application to consume. It is a lot more powerful than YAML and has a very safety-oriented type system.
- norrius 6y ago> Whenever an input accepts YAML you can actually pass in JSON there and it’ll be valid ...unless your parser strictly implements YAML 1.1, in which case you should be careful to add whitespace around commas (and a few other minor things). This is a valid JSON that some YAML parsers will have problems with: {"foo":"bar","\/":10e1} The very first result Google gives me for "yaml parser" is https://yaml-online-parser.appspot.com https://yaml-online-parser.appspot.com, which breaks on the backslash-forward slash sequence.
- paxys 6y agoI will never understand why YAML didn't just require quoted strings. Did the creator not anticipate how many problems the ambiguity would cause?
- mattmanser 6y agoNever's a strong word, seems quite easy to understand why to me. You've got ease of use reasons, historical reasons like the mis-guided Robustness principle, etc. And these sort of things happen time and time again. And although officially JSON requires quoted strings, almost none of the parsers actually enforce that, and so you will find a huge amount of JSON out there that is not actually compliant with the official spec. Just like browsers have huge hacks in them to handle misformed HTML.
- Safety1stClyde 6y ago> And although officially JSON requires quoted strings, almost none of the parsers actually enforce that What programming language? I'm not familiar with those parsers, the ones I know of very much do enforce quoted strings. > you will find a huge amount of JSON out there that is not actually compliant with the official spec The parsers I use all follow the current JSON RFC specification, and I've never encountered any JSON from APIs which they reject. > Just like browsers have huge hacks in them to handle misformed HTML. Web browsers do deal with a variety of things, not so much JSON parsers in my experience.
- Macha 6y agoI think the point is that they accept more than the spec dictates - do your JSON parsers accept e.g. the vs code config file (JSON with comments) or JSON with unquoted keys?
- jokethrowaway 6y agoCan you name a JSON parser which accept comments or unquoted keys? I've never seen one
- enaaem 6y agoRelated xkcd https://xkcd.com/327/ https://xkcd.com/327/
- tpowell 6y agoLittle Bobby Tables. I came here to post this.
- joshxyz 6y agoThis is why i love JSON. It's only string, number, boolean, arrays, objects/dictionaries, unless you write custom serializer and deserializers..
- lokedhs 6y agoExcept that its numbers are underspecified and cannot be used safely outside of a certain range. The spec explicitly states that the precision of numbers is not defined, meaning that N and N+1 may be the same number, and its behaviour would depend on the parser you're using. The number one rule when creating a serialisation format should be that serialisation and deserialisation is predictable. It's quite remarkable that two of the most popular formats doesn't do this. I'm actually surprised we haven't seen any major security issues caused by this.
- BlueTemplar 6y agoHow well does JSON deals with hexadecimal numbers ?
- deleted 6y ago[deleted]
- WalterBright 6y ago> The most tragic aspect of this bug, howevere, is that it is intended behavior according to the YAML 2.0 specification. This is one of those great ideas that sadly one needs experience to realize are really bad ideas. Every new generation of programmers has to relearn it. Other bad ideas that resurface constantly: 1. implicit declaration of variables 2. don't really need a ; as a statement terminator 3. assert should not abort because one can recover from assert failures
- atleta 6y agoI agree with the general observation, but the need for ";" ? Quite a few languages (over a few generations) have been doing fine without the semicolon. Just to mention two: python and haskell. (Yes, python has the semicolon but you'll only ever use it to put multiple statements on a single line.)
- yakshaving_jgt 6y ago> Yes, python has the semicolon but you'll only ever use it to put multiple statements on a single line. This is also true of Haskell btw.
- Cu3PO42 6y agoHaskell has the semicolon for the same reason!
- lelanthran 6y ago> I agree with the general observation, but the need for ";" ? Quite a few languages (over a few generations) have been doing fine without the semicolon. Just to mention two: python and haskell. (Yes, python has the semicolon but you'll only ever use it to put multiple statements on a single line.) But then it's inconsistent and has unnecessary complexity because now there's one (or more) exceptions to the rules to remember: when the ';' is needed. And of course if you get it wrong you'll only discover it at runtime. "Consistent applications of a general rule" is preferable to "An easier general rule but with exceptions to the rule".
- lifeisstillgood 6y agoHang on. The strict model seems off. In the first model entering GB 9.3 gets you a string and a number. But the second gets you two strings? Both are wrong in my opinion. "GB" 9.3 is the correct approach Explicit beats implicit every time.
- lenkite 6y agoWhy does YAML accept unquoted strings ? Be Strict. Be Safe.
- Hendrikto 6y ago> The real fix requires explicitly disregaring the spec Or… just quote your strings.
- dragonwriter 6y agoOr, “use an appropriate schema”. Or, for several of the specific problems identified in the source article, use YAML 1.2 (2009) instead of YAML 1.1 (2005), which the article misidentifies as “YAML 2.0” and acts as if it is the current spec.
- abujazar 6y agoNorwegian here. I’d say the problem is YAML, not Norway :D
- deleted 6y ago[deleted]
- bmn__ 6y agoThe problem is insufficiently analysed by the article author and the commenters in this thread so far. It is very superficial. The recent thread "Can’t use iCloud with “true” as the last name" https://news.ycombinator.com/item?id=26364993 https://news.ycombinator.com/item?id=26364993 went deeper. Let me take up its relevant particulars into this thread. The article author hitchdev does not say it outright, but it is heavily implied that the YAML file was edited by hand. This is the immediate cause of the problem. The indirect root of the problem is that the spec authors chose a plain text serialisation format and thus created an affordance http://enwp.org/Affordance#As_perceived_action_possibilities http://enwp.org/Affordance#As_perceived_action_possibilities to be edited by hand. This turns out the be unsafe/source of bugs because YAML end-users are not capable of correctly applying the serialisation rules considering the edge cases detailed in the article because humans are creatures of habit, applying analogy and common sense, making assumptions and then sometimes go wrong, whereas a piece of software will not make the Norway, Null etc. mistakes. hitchdev even writes that quoting the string is "a fix for sure, but kind of a hack", but that's a grave misunderstanding. Quoting the string here is actually applying the serialisation rules correctly. The tangential at the end of the article about typing is also orthogonal/irrelevant. YAML is strictly/strongly/unambiguously typed, and so is the mentioned variant Strict YAML. The difference is that Strict YAML has serialisation rules that are more amenable to or aligning with the human factors of habit etc. and thus work better in practice. My personal recommendation is to never edit YAML by hand and always use a serialiser. This is less convenient, but safe. In closing, I would like the reader of this comment to make an effort to distinguish between "what is" and "what ought to be" in their head, otherwise the ideas here will be very muddled.
- dragonwriter 6y ago> The problem is insufficiently analysed by the article author The article author also misidentifies the version of the YAML spec (calling it 2.0, which doesn’t exist; the behavior is from YAML 1.1, and this class of problems motivated a bunch of changes in YAML 1.2, which has been out since 2009.) But the article author isn’t trying to analyze the problem, he’s trying to rationalize why what is notionally a YAML-processing library just ignores the spec.
- Aeolun 6y ago
- dragonwriter 6y agoits weird that this is a 2019 article misrepresenting behavior in the YAML 1.1 spec (2005) most of which reverted in the YAML 1.2 spec (2009) as being part of a nonexistent YAML 2.0 spec and justifying a library that purports to handle “YAML” ignoring the spec.
- atombender 6y agoYou're right, but it's worth noting that much of the world is still on YAML 1.1, for whatever reason, so in practice, these are actual problems that will be encountered in the real world. For example, Ruby's standard library only supports YAML 1.1. It relies on libyaml, which is not yet compliant with 1.2. Meanwhile, Python's popular PyYAML library only supports 1.1, and asks users to migrate to a newer fork called ruamel.yaml for 1.2 support.
- dragonwriter 6y ago> You're right, but it's worth noting that much of the world is still on YAML 1.1 This is an article justifying use of (and justifying design decisions of) a particular Python quasi-YAML parsing library. If you are in a position to select a non-YAML-1.1-compliant parsing library for Python, or to take the articles advice on design of a YAML(-ish) parsing library, you are, necessarily, not stuck with YAML 1.1. > for whatever reason Articles like this spreading misinformation about the current state of standard YAML are part of the reason. LibYAML lagging support is another since so much of the ecosystem depends on libYAML (though, while the documentation situation is terrible, it looks like maybe libYAML has some level of 1.2 support since 0.23.) > For example, Ruby's standard library only supports YAML 1.1. It relies on libyaml, [...] Python's popular PyYAML library only supports 1.1 Which, also, is dependent on libYAML. > and asks users to migrate to a newer fork called ruamel.yaml for 1.2 support. Which makes a lot more sense than migrating to a library thar supports neither 1.1 nor 1.2, but a nonstandard variant that addresses some of the same issues resolved years ago in 1.2, especially when a library supporting 1.2 is available for the same language.
- jansan 6y agoNorway is one of the luckiest countries in the world. They have a vast amount of resources, can produce their electrical energy entirely from hydropower, have a great democracy, a government they can trust, a beautiful landscape and great people. I must say that I feel a little bit of relief to see that they have problems that nobody else has, besides insanely expensive alcohol that is only sold in "wine monopoly" stores that are more heavily guarded than banks.
- makkesk8 6y agoWe too had this problem, we solved it using the 3 letter country code instead.
- throwaway4good 6y agoJust use json.
- ancarda 6y agoYou can catch this with yamllint (https://github.com/adrienverge/yamllint https://github.com/adrienverge/yamllint): % cat countries.yml --- countries: - US - GB - NO - FR % yamllint countries.yml countries.yml 5:4 warning truthy value should be one of [false, true] (truthy)
- jmartrican 6y agoIt seems like we need to treat yaml like json and quote all strings. Would that help resolve these issues? Just trying to figure out a rule I can implement to prevent these issues.
- dudeinjapan 6y agoYou say Norway, I say Yesway.
- dan-robertson 6y agoOther reasons to not want types happening during parse time: - “modified” numbers, e.g. $50, 35%, 1.2345568896347853246863477 - Dates. If your language tries to convert a date to Unix time or Julian Day, you can have problems with time zones or distant or historical dates. - strings vs symbols. The person writing config shouldn’t have to care about this distinction. - Automatic deduplication for fields of objects can be a problem.
- ravanave 6y agoBtw, the reason Haskell isn’t used more isn’t type system per se, as all types can be inferred at the compilation time. People would sometimes use this feature even to see if GHCi guesses the type correctly (by correctly I mean exactly how the user wants, technically it’s correct always) first time and save them some time writing it either with an extension or just copy&paste from the interpreter window. When it gets hairy is that most programming languages have low entrance barrier. To write Haskell effectively you’ve got to unlearn a lot of rooted bad habits and you get to dive into the “mathematical” aspect of the language. Not only you got monads, but there’s plethora of other types you need to get comfortably onboard with and the whole branch of mathematics talking about types (you don’t need to even know that such a field as category theory exists to use it). However, since most people just want to write X, or just want hire a dev team at price they can afford, Haskell rarely is the first choice language.
- Kliment 6y agoIn my opinion the mathematical concepts and abstractions are not the issue with Haskell. The issue is that it's a pain to use in practice because of: 1. Really annoying to do any kind of i/o 2. Extremely poor interoperability with non-Haskell code 3. (opinion) Unpleasant, inconsistent, hairy syntax
- deleted 6y ago[deleted]
- andrewclunn 6y agooh no, we want this value to be parsed as a string, so we need to put quotes around it. the humanity!
- DennisP 6y agoSomething that used to plague me is that I had database processes importing Excel docs from clients, and if the first few rows in a column were numbers, SQLServer assumed that all the values must be numbers. Then it would run into cells containing other strings, and instead of revising its assumption, it would just import them as null. Since clients often didn't have great data hygiene, it was a problem. I finally solved it by exporting to csv, and using third-party software that handled its own import and did it correctly.
- dpratt71 6y agoI don't understand why Haskell gets brought up in the middle of an otherwise interesting and useful article. This sort of thing cannot happen in Haskell. And while Haskell is not universally admired, I can't recall seeing Haskell's flavor of type inference being a reason why someone claimed to dislike Haskell.
- Waterluvian 6y agoI have never gotten far into a project and thought, "my config files are too verbose. I wish there were clever shorthands." Does Yaml have any sort of strict mode? I imagine I could find a linter that disallows implicit strings.
- exyi 6y agoNot YAML by itself, but there are libraries that parse a YAML-like format that is typed. For example this one: https://hitchdev.com/strictyaml/ https://hitchdev.com/strictyaml/. Technically, it is not compatible with the YAML spec.
- thrower123 6y agoI've never seen anything that used YAML that I didn't want to douse with gasoline, nuke from orbit, and then salt the ground where it once stood. I cry and rage and rend my clothes when I stumble upon some new thing that makes me have to use it.
- mcv 6y agoIf it ignores part of the spec, I don't think "strictyaml" is the correct name here. Instead, if it interprets everything as string, perhaps "stringyaml" would have been more accurate, though I'm sure that's not as good PR. I'm reminded of the discussion we had a few days ago about environment variables; one problem there is that env variables are always strings, and sometimes you do want different types in your config. But clearly having the system automatically interpret whether it's a string or something else is a major source of bugs. Maybe having an explicit definition of which field should be which type would help, but then you end up with the heavy-handed XML with its XSD schema. Or you just use JSON, which is light-weight, easy to read, but unambiguous about its types. I guess there's a good reason it's so popular. Maybe other systems like yaml and environment variables should only ever be used for strings, and not for anything else, and I suppose replacing regular yaml with 'strictyaml' could play a role there. Or cause unending confusion, because it does violate the spec.
- povik 6y ago“saneyaml” would not make for bad PR
- marcinzm 6y ago>If it ignores part of the spec, I don't think "strictyaml" is the correct name here. The article didn't fully explain it but strictyaml requires a typed schema or defaults to string (or list or dict) if one is not provided. So it strictly follows the provided schema.
- mcv 6y agoThat makes a big difference indeed. It wasn't clear to me from the article, but string yaml + optional schema sounds like a useful combination.
- msiemens 6y ago> JSON, which is [...] unambiguous about its types With the one exception that with floatig point values the precision is not specified in the JSON spec and thus is implementation defined[1] which may lead to its own issues and corner cases. It for sure is better than YAML's 'NO' problem, but depending on your needs JSON may have issues as well [1]: https://stackoverflow.com/questions/35709595/why-would-you-use-a-string-in-json-to-represent-a-decimal-number https://stackoverflow.com/questions/35709595/why-would-you-u...
- bsenftner 6y agoThis is such a core issue with a tool like YAML, how the hell did this program get so popular? Are there that many developers willy nilly using tools that fail in critical, silent ways, and the horde of no-nothings follows them?
- IHLayman 6y agoReminds me of the multiple YAML bugs that have plagued Kubernetes such as https://github.com/kubernetes/kubernetes/issues/82296 https://github.com/kubernetes/kubernetes/issues/82296 It is interesting how the standard of any language seems to diverge due to just the implementation from different parsers.
- paulintrognon 6y agoI am sometimes annoyed by the fact you have to put double quotes around string properties in JSON. It would be so much lighter to use JS syntax..! Then I read articles like this one. Thank you JSON for not trying to be smart.
- earthboundkid 6y agoIt's fashionable to hate XML because it was used in a lot of places it was a bad fit in the 00s, but at least it's a pretty good document language. YAML though is always a bad fit. If you want machine readable config, use JSON; human readable, use TOML. When does YAML ever fit?
- dangoor 6y agoCue also solves this problem. The "no" example is right on the front page: https://cuelang.org https://cuelang.org I used it for configuration of a Go program recently and found it pleasant to work with. I hope the language is declared stable soon, because it's a good model.
- jgalt212 6y agoanother gotcha: 2020-03-25 -> datetime.date(2020, 3, 25), not "2020-03-25"
- lolinder 6y agoThis comment was buried in a thread, but I'm bringing it out because it's very relevant to the conversation: https://news.ycombinator.com/item?id=26679728 https://news.ycombinator.com/item?id=26679728 > the article refers to YAML 2.O, a nonexistent spec, and to PyYAML, a real parser which supports only YAML 1.1. > Both the unquoted-YES/NO-as-boolean and sexagesimal literals were removed in YAML 1.2.
- BlueTemplar 6y agoYeah, I'd bet that YAML "two point oh" (rather than "two point zero") doesn't exist ! :p
- suttree 6y agoReminds me that the reasoning behind austerity came from an Excel calculation that didn't include all the relevant rows :~/ https://www.theguardian.com/politics/2013/apr/18/uncovered-error-george-osborne-austerity https://www.theguardian.com/politics/2013/apr/18/uncovered-e... https://www.bbc.co.uk/news/magazine-22223190 https://www.bbc.co.uk/news/magazine-22223190 https://theconversation.com/the-reinhart-rogoff-error-or-how-not-to-excel-at-economics-13646 https://theconversation.com/the-reinhart-rogoff-error-or-how...
- BlueTemplar 6y agoWell, the main issue here seems to be that their work has not been peer reviewed ?
- blunte 6y agoIf you want no misunderstandings, be explicit. This applies to YAML and life in general. There's an annoying but fairly accurate saying about assumptions that applies. If you want something to be a specific type, you better have an explicit way of indicating that. If you say quotes will always indicate a string, great. Of course we know it's not that simple, since there are character sets to consider. The safest answer is to do something like XML with DTDs. But that imposes a LOT of overhead. Naturally we hate that, so we make some "convention over configuration" choices. But eventually, we hit a point where the invisible magic bites us. This is one case where tests would catch the problem, if those tests are thorough enough - explicitly testing every possibility or better yet, generative testing.
- korijn 6y agoOr just opening your browser and trying out norwegian on a QA environment.
- jasode 6y agoThat author's blog post sent me down a rabbit hole of insanity with YAML and the PyYAML parser idiosyncrasies. First, he mentions "YAML 2.0" but there's no such reference about "2.0" from yaml.org or Google/Bing searches. Yaml.org and wikipedia says yaml is at 1.2. Apparently the other commenters in this thread clarified that the older "YAML 1.1" is what the author is referring to. Ok, if we look at the official YAML 1.1 spec[1], it has this excerpt for implicit bool conversions: y|Y|yes|Yes|YES|n|N|no|No|NO |true|True|TRUE|false|False|FALSE |on|On|ON|off|Off|OFF But the pyyaml code excerpts[2][3] from resolver.py has this: u'tag:yaml.org,2002:bool', re.compile(ur'''^(?:yes|Yes|YES|n|N|no|No|NO |true|True|TRUE|false|False|FALSE |on|On|ON|off|Off|OFF)$''', re.X), The programmer omitted the single character options of 'y' and 'Y' but it still has 'n' and 'N' ?!? The lack of symmetry makes the parser inconsistent. And btw for trivia... PyYAML also converts strings with leading zeros to numbers like MS Excel: https://stackoverflow.com/questions/54820256/how-to-read-load-yaml-parameters-with-leading-zeros-as-a-string https://stackoverflow.com/questions/54820256/how-to-read-loa... [1] https://yaml.org/type/bool.html https://yaml.org/type/bool.html [2] 2020 latest: https://github.com/yaml/pyyaml/blob/ee37f4653c08fc07aecff69cfd92848e6b1a540e/lib/yaml/resolver.py#L172 https://github.com/yaml/pyyaml/blob/ee37f4653c08fc07aecff69c... [3] 2006 original : https://github.com/yaml/pyyaml/blob/4c570faa8bc4608609f0e531452f9941e4a272b3/lib/yaml/resolver.py#L96 https://github.com/yaml/pyyaml/blob/4c570faa8bc4608609f0e531...
- kleer001 6y agoAh, tool centric (not centered in the person) time-offset (the ones making the original mistake doesn't see it) paredolia. Good 'el TCTOP...
- zomglings 6y agoHave had a similar issue when adding git revisions to YAML documents. The problem is that if a YAML parser sees a string like this: "0123e04" It interprets it as a number: 123 * 10^4 Our hacky solution was to prefix the revision hashes like sha-0123e04, but still this was quite annoying. After that experience, I have stopped using YAML for any of my own configuration. Have started preferring putting my configurations in code. And when I don't want that, have found JSON good enough for my purposes.
- deleted 6y ago[deleted]
- BlueTemplar 6y agoBut hashes are numbers, not strings, so is it really only YAML at fault ?
- zomglings 6y agoHashes are NOT numbers in base 10 scientific notation, which is how the hash that I showed you would be interpreted by YAML. The point is that this behavior is sporadic. It doesn't apply consistently across all git hashes, which is the real problem. It is easy to be caught unawares by this behavior.
- BlueTemplar 6y agoYes, I got that, but why have you declared a hash, which is a number, though a different kind of number than a base 10 scientific notation, as a string ?
- zomglings 6y agoBecause we were not using any numerical properties of the hash. We were not adding it to other hashes, seeing if it was greater than or less than other hashes, etc. Literally the only thing we were doing was passing it between shell commands, helm charts, Kubernetes deployments and then back (if we needed to debug). It sounds like you have a more attractive alternative in this case than to treat hashes as strings. Would love to hear it.
- atombender 6y agoThe world desperately needs a replacement for YAML. TOML is fine for configuration, but not an adequate solution for representing arbitrary data. JSON is a fine data exchange format, but is not particularly human-friendly, and is especially poor for editable content: Lacks comments, multi-line strings, is far too strict about unimportant syntax, etc. Jsonnet (a derivative of Google's internal configuration language) is very good, but has failed to reach widespread adoption. Cue is a newer Jsonnet-inspired language that ticks a lot of boxes for me (strict, schema support, human-readable, compact), but has not seen wide adoption. Protobuf has a JSON-like text format that's friendlier, but I don't think it's widely adopted, and as I recall, it inherits a lot of Protobufisms. Dhall is interesting, but a bit too complex to replace YAML. Starlark is a neat language, but has the same problem as Dhall. It's essentially a stripped-down Python. Amazon Ion [1] is neat, but I've not seen any adoption outside of AWS. NestedText [2] looks promising, but it's just a Python library. StrictYAML [3] is a nice attempt at cleaning up YAML. But we need a new language with wide adoption across many popular languages, and this is Python only. Any others? [1] https://amzn.github.io/ion-docs/ https://amzn.github.io/ion-docs/ [2] https://nestedtext.org/ https://nestedtext.org/ [3] https://github.com/crdoconnor/strictyaml/ https://github.com/crdoconnor/strictyaml/
- geraldbauer 6y agoYou might look at JSON Next variants (if you remember - "classic" JSON is a subset of YAML), see https://github.com/json-next/awesome-json-next https://github.com/json-next/awesome-json-next My own little JSON Next entry / format is called JSON 1.1 or JSONX, that is, JSON with eXtensions, see https://json-next.github.io https://json-next.github.io
- orthoxerox 6y agoThe list is missing http://www.relaxedjson.org/ http://www.relaxedjson.org/ Also, there's no explanation what <..-..> and <..+..> do.
- IshKebab 6y agoJSON5 is the best option currently. A fair number of tools in the JS ecosystem support it.
- mrzool 6y agoWow, YAML has definitely some pretty quirky edges.
- DoofusOfDeath 6y agoFunny coincidence. Around 2000, I worked for a company that coined the term "Norway problem" for a different software problem. Their product used an MVCC database (I think ObjectStore). One of their customers in Norway had a problem where updates to the database seemed to not show up. IIRC the problem was a bug in this company's software that caused MVCC to show an older version of the database content than expected.
- NaturalPhallacy 6y agoThis is why implicit typing is an invitation to errors.
- namelosw 6y agoI prefer JSON over YAML because I spend more time confused and burned by the problems caused by it. I understand that people don't like directly use JSON because it's not very friendly: no comments, no multi-line string, etc. A great alternative IMHO is cson[0]. It's like JSON to JavaScript but for CoffeeScript (though nobody talks about it nowadays). It has indentation-based syntax, comments, and multiline string which usually don't need to escape. The advantage is it's close enough to JSON which is the canonical format that everybody can agree on nowadays. For YAML and TOML there are too many visual part-aways from JSON. Or just create a JSON variant that enables comments and the backtick multiline string from JavaScript. [0] https://github.com/bevry/cson https://github.com/bevry/cson
- knorker 6y agoNO problem.
- korijn 6y agoEdit: downvoters, thanks! I realize this is not an easily agreeable opinion ("let's all chant 'death to YAML!'") but it's really easy to avoid losing money on something like this. Just do proper testing. Aren't you setting yourself up for surprises if you write file formats such as TOML and YAML without reading the documentation, learning and experimenting first? How about unit testing? Or verifying the type in your config parser? Have you tried opening your site in the norway config on your development or testing environment? Or even in production? It all seems very basic and not at all blog post or even HN worthy. I'm going to assume the authors still haven't learned their lesson and are going to experience many more surprises in the future working with plain text file formats.
- teddyh 6y agoThis is another good argument against weak types in general. Strong types are better, and explicit is better than implict.
- vanshg 6y agoWhy not just use Python itself for storing configurations? You can be explicit about the data type and no need to parse anything
- jmull 6y ago> “While the website went down and we were losing money we chased down a number of loose ends until finally finding the root cause.” Hopefully not a real story. If you’re trying out new configurations in production and have no mechanism to rollback problematic changes, you’ve got bigger problems than YAML. To me, though, YAML, including “StrictYAML” doesn’t solve any problems JSON, perhaps w/comments, already solves.
- gitowiec 6y agoI don't like YAML because when I need to write configuration in it I waste time to remember what is the syntax. I have much better understanding of JSON, because I use it almost on daily basis.
- Meniteos4 6y agoThey decided to go against the YAML standard and therefore are no longer a YAML parser. The actual answer to this problem would have been to use a better storage format. Perhaps JSON5? or TOML?
- runeks 6y agoWhy not just enclose your strings in quotes and be done with it? As far as I can see, this has nothing to do with typing and everything to do with syntax (of literals). If strings were required to be quoted this problem wouldn’t appear. This is the reason no programming language has this issue — regardless of type system (JS/Python/Java/Haskell). If you want a string here you need quotes. Haskell could even be regarded as what the author calls “implicitly typed” — since types are derived from literals — and I’ve never heard a Haskeller complain about this issue.
- lopespm 6y ago> Christopher Null has a name that is notorious for breaking software code - airlines, banks, every bug caused by a programmer who didn’t know a type from their elbow has hit him. This one made chuckle, and TIL that Null is a real life surname.
- dwheeler 6y agoThe YAML specification eliminated this problem in 2009, with the release of version 1.2. That spec also eliminated some other problematic problems. The real problem is that YAML parsers in wide use have not been updated to the spec that was released TWELVE years ago. So who's going to help the common YAML parser developers update their implementations to support version 1.2? I think that would be a big help. Maybe the Norwegian government can chip in some money & time to get them updated, that would probably quietly eliminate a number of problems.
- dwheeler 6y agoI'm replying-to-myself, because I think this text from YAML 1.2 (explaining its changes) are key: > The primary objective of this revision is to bring YAML into compliance with JSON as an official subset. YAML 1.2 is compatible with 1.1 for most practical applications - this is a minor revision. An expected source of incompatibility with prior versions of YAML, especially the syck implementation, is the change in implicit typing rules. We have removed unique implicit typing rules and have updated these rules to align them with JSON's productions. In this version of YAML, boolean values may be serialized as “true” or “false”; the empty scalar as “null”. Unquoted numeric values are a superset of JSON's numeric production. Other changes in the specification were the removal of the Unicode line breaks and production bug fixes. We also define 3 built-in implicit typing rule sets: untyped, strict JSON, and a more flexible YAML rule set that extends JSON typing." Since "no" is not the same as false, the Norway problem disappears. It's safer to always quote single-word strings like 'this', just like you always have to quote all strings in JSON.
- mleonhard 6y agoDon't use ambiguous formats. Use TOML. https://toml.io/ https://toml.io/