9 ms·
Concise Encoding: A secure data format for a modern world
- kstenerud 4y agoHey guys, sorry for the slow updates on this. Ukrainian issues have swallowed up most of my time recently, and my OSS projects have suffered as a result. Rest assured that this project IS an ongoing concern, and is the foundation of many more technologies I intend to bring to bear over the coming years.
- samatman 4y agoI'm glad to see this surface again, I lost track of it for a couple years and surfaced it about six months ago. Are you open to feedback in the spirit of collaboration? I've considered opening an issue or two on the repo, but it always feels a bit presumptuous to do that when one person is the author of the entire project. I like the way you think about these things and I find CE quite promising.
- kstenerud 4y agoI'm absolutely open to feedback, criticism, collaboration of any kind! The whole point of a format is to be usable, and while I have a number of opinions on the subject, I'm no superman. I may not always agree with other peoples' opinions, but even when I don't, the whole point is the conversation it sparks; that's the nexus of greatness. One of the main reasons why I haven't released version 1 yet is because it's been so much of a one-man-show that I'm under no illusions of being immune to tunnel vision. In the immortal words of Johnny 5: I need input!
- munro 4y agoLooks really cool--I love the overview, we've definitely grown out of JSON, most of the time I sadly end up using pickle when using Python. Too many question marks in my head to be honest. The more I think about it, I believe code still may be the the best way to describe really complex data. What is interesting to me is the holistic view of all the data types we most commonly use--but then I don't really know how this differs from Thrift/Avro/etc--and then further, why haven't we as a community moved to one of those?
- manholio 4y agoThis is pretty cool, definitely would have come handy 15 - 20 years ago when people like Cohen and Nakamoto were inventing Bittorrent and Bitcoin, rolling their own binary serialization atrocities and unleashing them on humanity.
- FridgeSeal 4y agoYou mean you don’t like swapping between endianness 15 million times when you’re processing Bitcoin????
- BoppreH 4y agoThings I liked: - Versioning. - Time zone identifier instead of just a fixed offset (which are ambiguous for future events). - Native encoding of binary values. - Graph notation with support for labels. - Comments! - Trying to escape lookalike characters, even though I think that's a lost cause. Things I'm not so keen about: - NUL character in strings being platform and settings dependent. - Line break not being forced to a consistent value. - The most complicated number encoding scheme I've ever seen (e.g. 0xa,3fb8p+42). - Entity references are a footgun for anyone writing a depth-first or breadth-first algorithm. - Arrays-vs-list feels like it doesn't belong in encoding formats.
- profquail 4y agoThe complicated number encoding scheme you mentioned is a hexfloat: C has them too. Hexfloat can be really useful when you need precise/exact floating-point constants for numerical methods. Without them, you end up having to do more-complicated hacks to preserve exact constant values when code gets compiled, or you have to live with compilers (sometimes) subtly altering constants. I wish more languages supported hexfloats.
- BoppreH 4y agoThanks, I had never seen that before. But Concise Encoding did complicate them further by accepting commas as decimal separator.
- unwind 4y agoJust to clarify a bit what "C has them" means; the standard library string-to-float parsing function strtof() [1] support this format. You can also use them as floating-point literals in your code. [1]: https://linux.die.net/man/3/strtof https://linux.die.net/man/3/strtof
- formerly_proven 4y agoHexfloat is also vastly easier to parse from/to IEEE 752 format than decimal floats.
- lifthrasiir 4y ago
- baq 4y agoIMHO schema support should be part of the initial release. https://json-schema.org/ https://json-schema.org/ being an afterthought to the original JSON was a mistake.
- kstenerud 4y agoYup, that's the idea. I've been looking at CUE for this https://cuelang.org/ https://cuelang.org/
- barnabee 4y agoWould be nice to support ISO8601 dates and times rather than (or in addition to) defining a similar custom format IMO
- deleted 4y ago[deleted]
- davibu 4y agoMany data set are stored like that (Json separated by RC): {"name":"name1", "phone":"+769989823"} {"name":"name2", "phone":"+769244563"} {"name":"name3", "phone":"+769989295"} ... There should be a way to do the same thing with concise.
- lifthrasiir 4y agoNot Concise, but CBOR does have a proposal for such data [1]. This is implemented via CBOR's tagging system where you can attach a tag to otherwise mundane data to hint otherwise, which is in my opinion one of the biggest idea behind CBOR and why CBOR is not really a ripoff of MessagePack. [1] https://github.com/kriszyp/cbor-records https://github.com/kriszyp/cbor-records
- sgammon 4y agothanks i hate it
- sgammon 4y agowe do not need special types for UUIDs. why are we doing this to ourselves
- sgammon 4y ago> Times are different from the carefree days that brought us XML and JSON. no they are not lol