13 ms·
What even is a JSON number?
- billpg 3y ago"ID numbers start from 2^53 and are allocated sequentially including odd numbers that are not compatible with "double" types. Please ensure you are reading this value as a 64-bit integer."
- paulddraper 3y ago> I-JSON messages SHOULD NOT include numbers that express greater magnitude or precision than an IEEE 754 double precision number provides I'm confused by this. What is the precision of 0.1, relative to IEEE 754? If I read it correctly, that statement is saying: json_number_precision(json_number) <= ieee_754_precision ^ How do I calculate these values?
- bterlson 3y agoI think the spec just means, assume IEEE 754. In the case of 0.1, which cannot be represented exactly, software should assume that `0.1` will be represented as `0.100000000000000005551115123126`. Depending on `0.1` being parsed as the exact value `0.1` is not widely interoperable.
- ebolyen 3y agoRelatedly, what about integers like 9007199254740995. Is that a legal integer since it rounds to 9007199254740996? It does seem unclear what it means to exceed precision (given rounding is such an expected part of the way we use these numbers). Magnitude feels easier as at least you definitely run out of bits in the exponent.
- hugh-avherald 3y agoI think the spec is saying that it is the message that should not express greater magnitude or precision, not 'the number'. So including the string "0.1" in a message is fine because v = 0.1 implies 0.05 < v < 0.15, but including 0.100000000000000000000000000000000000 would not be.
- EdSchouten 3y agoI think the description for Go is inaccurate/incomplete. You can call this function to instruct the decoder to leave numbers in unparsed string form: https://pkg.go.dev/encoding/json#Decoder.UseNumber https://pkg.go.dev/encoding/json#Decoder.UseNumber That allows you to capture/forward numbers without any loss of precision.
- bterlson 3y agoI have added this note, thanks! In the blog I am mostly trying to show the behavior you get using the (maybe defacto) stdlib with its default configuration, but this is useful data to call out.
- js2 3y agoIf you're going to extend Go the courtesy of customizing the parser, oughtn't you do the same for Python (and all the languages)? To wit, Python's json module has `parse_float` and `parse_int` hooks: https://docs.python.org/3/library/json.html#encoders-and-decoders https://docs.python.org/3/library/json.html#encoders-and-dec... Example: >>> json.loads('{"int":12345,"float":123.45}', parse_int=str, parse_float=str) {'int': '12345', 'float': '123.45'} FWIW, when I've cared about interop and controlled the schema, I've specified JSON strings for numbers, along with the range, precision, and representation. This is no worse (nor better) than using RFC 3339 for dates.
- bterlson 3y agoI'm just a JS guy trying to understand the world around me and documenting what I find, not trying to be discourteous (or even courteous). I'll add the note about Python, thanks for calling it out. FWIW JS does not have a similar capability so I can't add a note there.
- js2 3y agoFair enough! Thank you for the writeup.
- 3y ago
- timvdalen 3y agoA little off topic, but fun to see that someone else has adopted that magical CSS theme! (https://css.winterveil.net/ https://css.winterveil.net/)
- deleted 3y ago[deleted]
- ape4 3y agoSince JSON is so widely used it should be modified to support more types - Mongo DB's Extended JSON supports all the BSON (Binary) types: Array Binary Date Decimal128 Document Double Int32 Int64 MaxKey MinKey ObjectId Regular Expression Timestamp https://www.mongodb.com/docs/manual/reference/mongodb-extended-json/ https://www.mongodb.com/docs/manual/reference/mongodb-extend...
- bterlson 3y agoJS is likely to get a hook to be able to handle serialization/deserialization of such values without swapping out the entire implementation[1]. Native support for these types, without additional code or configuration, would likely break the Internet badly, so is unlikely to happen unfortunately. 1: https://github.com/tc39/proposal-json-parse-with-source https://github.com/tc39/proposal-json-parse-with-source
- apantel 3y agoMuch more valuable than any such extension would be a way to annotate types and byte lengths of keys and values so that parsers could work more efficiently. I’ve spent a lot of time making a fast JSON parser in Java and the thing that makes it so hard is you don’t know how many bytes anything is, or what type. It’s hard to do better than naive byte-by-byte parsing.
- your_fin 3y agoIf you control the underlying data, I must reccomend Amazon Ion! Its text format is a strict superset of JSON, but they also maintain binary format that will round-trip data and is designed for efficient scanning. There's even prefixed annotations if you want them :) It also specs proper decimal values, mitigating the issues presented in the OP. https://amazon-ion.github.io/ion-docs/ https://amazon-ion.github.io/ion-docs/
- zilti 3y agoOr maaaybe use XML for such cases
- 3y ago
- frizlab 3y agoIt’s missing Swift tests, but otherwise it’s a great post.
- bterlson 3y agoIf you would like to contribute Swift tests, I would be happy to take it! You can send a PR into this document, updating the data tables and adding a code sample at the end: https://github.com/bterlson/blog/blob/main/content/blog/what-is-a-json-number.md https://github.com/bterlson/blog/blob/main/content/blog/what.... No need to test openapi-tools swift codegen unless you really want to!
- frizlab 3y agoI’m having a lot on my plate currently, but I’m adding this to my TODO list!
- whalesalad 3y agoI tend to end up encoding everything as an integer (multiply by 1000, 10000 etc) and then turn it back into a float/decimal on decode. For instance if I am building a system dealing with dollar amounts I will store cent amounts everywhere, communicate cent amounts over the wire, etc. then treat it as a presentation concern to render it as a dollar amount.
- eddd-ddde 3y agoThis is great as long as you always make clear which value is pre post encoding. I remember one of my first production bugs was giving users 100 times the credit they actually bought. Oops.
- Supermancho 3y agoI have tried to encode all non-trivial numbers as strings. If it's too big (or small), or if it's a float, I'll have to change my JSON schema. Bake the need to decode numbers into the transforms for consistency.
- jerf 3y agoIt's worth bearing in mind when you do that that the largest integer that is "generally safe" in JSON is 2^53-1, so if you scale by a factor of 10000 you're taking 13-14 more bits off that maximum. That leaves you about 2^40, or about a trillion, before you may start losing precision or seeing systems disagree about the decoded values. Whether that's a problem depends on your domain.
- nedt 3y agoSo basically you use fixpoint numbers. Especially for currency that’s a very good idea anyway, because of rounding errors, even more so in IEEE 754
- bterlson 3y agoPedantically, IEEE 754 defines decimal floating point formats (like decimal128) which are appropriate for representing currency. Representing currency in non-integer values in any of the binary floating point formats is indeed a recipe for disaster though.
- erik_seaberg 3y agoIt's weird that any parser that loses digits is tolerated. A parser that forces strings into uppercase US-ASCII never would be.
- msm_ 3y agoThat's true for every floating point number in every programming language you have ever used, though. $ python3 Python 3.10.13 (main, Aug 24 2023, 12:59:26) [GCC 12.2.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> 100000.000000000017 100000.00000000001
- Izkata 3y agoThis is why Decimal exists: Python 3.8.10 (default, Nov 22 2023, 10:22:35) [GCC 9.4.0] on linux Type "help", "copyright", "credits" or "license" for more information. >>> from decimal import Decimal >>> Decimal('100000.000000000017') Decimal('100000.000000000017') For example: >>> import json >>> json.loads('{"a": 100000.000000000017}') {'a': 100000.00000000001} >>> json.loads('{"a": 100000.000000000017}', parse_float=Decimal) {'a': Decimal('100000.000000000017')}
- cerved 3y agobut serializing/deserializing decimal using the json module is futile
- yawaramin 3y agoWhy is it futile? It can be serialized/deserialized perfectly through its string representation.
- dhosek 3y agoAnd not every programming language offers a Decimal type and on most of those, there’s usually a performance penalty associated with it not to mention issues of interoperability and developer knowledge of its existence. For financial calculations, usually using integers with an implicit decimal offset (e.g., US currency amounts being expressed in cents rather than dollars), while other contexts will often determine that the inherent inaccuracy of IEEE floating types is a non-issue. The biggest potential problem lies in treating values that act kind of like numbers and look like numbers as numbers, e.g., Dewey Decimal classification numbers or the topic in a Library of Congress classification.¹ ⸻ 1. This is a bit on my mind lately as I discovered that LibraryThing’s sort by LoC classification seems to be broken so I exported my library (discovering that they export as ISO8859-1 with no option for UTF-8) and wrote a custom sorter for LOC classification codes for use in finally arranging the books on my shelves after my move last year.
- kccqzy 3y agoI'll add that for Haskell, the library everyone uses for JSON parses numbers into Scientific types with almost unlimited size and precision. I say almost unlimited because they use a decimal coefficient-and-exponent representation where the exponent is a 64-bit integer. The documentation is quite paranoid that if you are dealing with untrusted inputs, you could parse two JSON numbers from the untrusted source fine and then performing an addition on them could cause your memory to fill up. Exciting new DoS vector. Of course in practice people end up parsing them into custom types with 64-bit integers, so this is only a problem if you are manipulating JSON directly which is very rare in Haskell.
- akubera 3y agoI was attempting to solve this very problem in the Rust BigDecimal crate this weekend. Is it better to just let it crash with an out of memory error, or have a compile-time constant limit (I was thinking ~8 billion digits) and panic if any operation would exceed that limit with a more specific error-message (does that mean it's no longer arbitrary-precision?). Or keep some kind of overflow-state/nan, but then the complexity is shifted into checking for NaNs, which I've been trying to avoid. Sounds like Haskell made the right call: put warnings in the docs and steer the user in the right direction. Keeps implementation simple and users in control. To the point of the article, serde_json support is improving in the next version of BigDecimal, so you'll be able to decorate your BigDecimal fields and it'll parse numeric fields from the JSON source, rather than json -> f64 -> BigDecimal. #[derive(Serialize, Deserialize)] pub struct MyStruct { #[serde(with = "bigdecimal::serde::json_num")] value: BigDecimal, } Whether or not this is a good idea is debatable[^], but it's certainly something people have been asking for. [^] Is every part of your system, or your users' systems, going to parse with full precision?
- AaronFriel 3y agoI'd strongly recommend against this default - it's a major blocker for using the Haskell library with web APIs as it transforms JSON RPC into into readily available denial of service attacks. 8 billion digits (~100 bits?) is far more than should be used. Would it possible to use const generics to expose a `BigDecimal<N>` or `BigDecimal<MinExp, MaxExp, Precision>` type with bounded precision for serde, and disallow this unsafe `BigDecimal` entirely? If not, I expect BigDecimal will be flagged in a CVE in the near future for causing a denial of service.
- egwor 3y agoI think the thing folk miss is when there’s an error like divide by zero, or the calculation would return NaN. I feel like this is the main gap/concern with using JSON and it seems to be rarely discussed.
- olejorgenb 3y agoAgreed, this can be a pain. Python by default serialize and de-serialize the `NaN` literal, making you pay some cleanup cost once you need to interopt with other systems. (same for `Inf`) Say what you want about NaN, but IEEE 754 is the facto way of dealing with floating points in computers and even if NaNs and Infs are a bit "fringe" it's unfortunate that the most popular serialization format can not represent these.
- int_19h 3y agoThere are so many things that are poorly thought out or underspecified in JSON, it's amazing that it got so widely adopted for interop. No wonder that it became a perpetual source of serialization bugs.
- lifthrasiir 3y agoEspecially annoying given that they could have been easily adopted. Infinity could've been encoded as `1/0` (among most other possibilities). NaN could've been encoded as `0/0` (again, among most other possibilities). JSON doesn't allow all possible JavaScript literals anyway, so these encodings might have been worked if they were somehow standardized.
- nigeltao 3y agoWhen I wrote my jsonptr tool a few years ago, I noticed that some JSON libraries (in both C++ and Rust) don't even do "parse a string of decimal digits as a float64" properly. I don't mean that in the "0.3 isn't exactly representable; have 0.30000000000000004 instead" sense. I mean that rapidjson (C++) parsed the string "0.99999999999999999" as the number 1.0000000000000003. Apart from just looking weird, it's a different float64 bit-pattern: 0x3FF0000000000000 vs 0x3FF0000000000001. Similarly, serde-json (Rust) parsed "122.416294033786585" as 122.4162940337866. This isn't as obvious a difference, but the bit-patterns differ by one: 0x405E9AA48FBB2888 vs 0x405E9AA48FBB2889. Serde-json does have an "float_roundtrip" feature flag, but it's opt-in, not enabled by default. For details, look for "rapidjson issue #1773" and "serde_json issue #707" at https://nigeltao.github.io/blog/2020/jsonptr.html https://nigeltao.github.io/blog/2020/jsonptr.html
- 01HNNWZ0MV43FF 3y agoOh wow. So serde_json doesn't roundtrip floats by default, it uses some imprecise faster algorithm https://github.com/serde-rs/json/issues/707 https://github.com/serde-rs/json/issues/707 Good thing there's msgpack I guess.
- lofenfew 3y agothis requires multiple precision to do properly and isn't useful most of the time. its odd to describe this as "not properly". you might say "with exact rounding", but that makes it clearer that this isn't that useful a feature, especially since we usually expect floats to be inexact in the first place.
- int_19h 3y agoWith JSON, there's essentially no such thing as "properly" when it comes to parsing numbers, since the spec doesn't limit the ability of the implementation to constrain width and precision. It only says that float64 is common and therefore "good interoperability can be achieved by implementations that expect no more precision or range than these provide", but note the complete absence of any guarantees in that wording. The only sane thing with JSON is to avoid numbers altogether and just use decimal-encoded strings. This forces the person parsing it on the other end to at least look up the actual limits defined by your schema.
- ctrw 3y agoI still get a laugh of ecma 404. The first time I looked it up I refreshed the page a large number of times before I realized it wasn't an error.
- hinkley 3y agoOne of the first Ajax projects I worked on was multi tenant, and someone decided to solve the industrial espionage problem by using random 64 bit identifiers for all records in the system. You have about a .1% chance of generating an ID that gets truncated in JavaScript, which is just enough that you might make it past MVP before anyone figures out it’s broken, and that’s exactly what happened to us. So we had to go through all the code adding quotes to all the ID fields. That was a giant pain in my ass.
- golergka 3y agoI've been burned by a similar issue too. Lesson here is never to use numbers for things you are not planning to do math on. Ids should always be strings.
- deleted 3y ago[deleted]
- greggyb 3y agoUntil you want faster joins, in which case, comparisons of integers tend to be much faster on hardware I am aware of than string comparisons.
- golergka 3y agoWe're talking about deserialising JSONs in the application server here, nobody stops you from treating ids as numbers on the database side of things. But also, this sounds like a premature optimisation. Most applications will never reach a level where their performance is actually impacted by string comparison, and when you reach that stage, you're likely have already thrown out a lot of other common sense stuff like db normalisation to get there, and we shouldn't judge "regular people" advice because it doesn't usually apply to you anyway. Out of curiosity, have you ever seen an application that was meaningfully impacted by this? How gigantic was it? ---- Scratch that. I've actually thought about it some more, and now I'm not 100% sure it's premature, I have to investigate further to be sure. Question still stands though.
- bevekspldnw 3y agoTry to get a DECIMAL value out of a Postgres database into a JSON API response and you’ll learn all this and more in the most painful way possible!
- Someone 3y agoOther values one could test for: - “+1” (not a valid number, according to ECMA-404 and RFC-8259) - “+0” (also not a valid number, but trickier than “+1” because IEEE floats have “+0” and “-0”) - “070” (not a valid number, but may get parsed as octal 56) - “1.” (not a valid number in json) - “.1” (not a valid number in json) - “0E-0” (a valid number in json) There probably are others.
- vbezhenar 3y agoMy opinion is that a safe approach is to use either 52-bit integer number or 64-bit floating number to keep JavaScript compatibility. JavaScript is too important and at the same time, the errors are too terrific (JS will silently round to the nearest 52-bit integer number which could lead to various exploits) to skip on that. If you need anything else, just use strings.
- yawaramin 3y agoLong story short: don't use JSON numbers to represent money or monetary rates. Always use decimals encoded as string. It's surprising how many APIs fall short of this basic bar.
- jillesvangurp 3y agoDepends on the language. On the JVM you are fine. With Javascript, doing math on big numbers is probably going to end in tears unless you know what you are doing. Either way, have some tests for this and make sure your code is doing what you expect. Encoding numbers as string because you are using a language and parser that can't deal with numbers properly (even 64 bit doubles), is a bit of a hack. Basically the rest of the world giving up because Javascript can't get its shit together is not a great plan.
- wruza 3y agoAccounting for the lowest common denominator that has a huge share in it is always a great plan. Every trading platform out there uses "+-ddd.ddd" format, even binary-born protocols completely unrelated to js used it since forever.
- incorrecthorse 3y agoNo. Use integers to store the smallest money decimal, and store the currency name alongside.
- sjducb 3y agoWhat happens if you’re sure that four decimal places is the smallest, then suddenly a partner system starts sending you 6 decimal places?
- Macha 3y agoPrecision should be part of the spec for integrations. With the integer multiple of minimal unit, that makes it clear in the API what it is. e.g. it doesn't make sense to support billing in sub-currency unit amounts just by allowing it in your API definition, as you're going to need to batch that until you get a billable amount which is larger than the fee for issuing a bill. Even for something like $100,000.1234, the bank doesn't let you do a transfer for 0.34c. For cases where sub-currency unit billing is a thing, it should be agreed what the minimal unit is (e.g. advertising has largely standardised on millicents)
- datagreed 3y agoThat font >.<
- p0w3n3d 3y agoIt does something to my eyes
- Tabular-Iceberg 3y agoDoes anyone have any idea why Crockford decided that at least one digit is required after the decimal point, as opposed to JavaScript which has zero or more?
- mulmen 3y agoThere are no numbers in JSON. There are only strings.
- Cloudef 3y agoFirst thing that I check before using a JSON parser library is that if it lets me to get the number as a string and let me do my own conversion. Libraries that try to treat the number as double or bring in a large bigint/decimal library gets usually pass from me.
- speedgoose 3y agoIf you need a specific exotic JSON parser to parse the numbers you have correctly, I would argue that you should serialise them as strings and not as numbers. That's what prometheus is doing for example. https://prometheus.io/docs/prometheus/latest/querying/api/ https://prometheus.io/docs/prometheus/latest/querying/api/
- Cloudef 3y agoThat only works if you are the one who serializes the json in the first place.
- p0w3n3d 3y agoI think that good decision regarding numbers in API (as it was made in my project) is to put meaningful decimal numbers into string and let them be handled by exact decimal calculation framework, e.g. BigDecimal in Java etc.
- j16sdiz 3y agoThe font choice for inline text is so distracting This "Averia Serif Libre" is unreadable for me.
- yau8edq12i 3y agoJSON is a notation. It's syntax. The semantics are left up to the implementation. The question has no answer.
- eternityforest 2y agoI like to think of floating point values as noisy analog voltages, with the extra propery that they can store small integers perfectly, and they can be copied within code but not round trip serialized and deserializer without noise. They're not really noisy, but if an application would work with some random noise added, it will probably work with floats, and if it wouldn't work with noise added, it's probably easier to just not use floats and expext people to reason about IEEE details, while risking subtle bugs if different float representations get mixed. Of course I'm not doing a lot of high performance algorithms, I would imagine in some applications you really do need to reason about floats.