8 ms·
A fast EDN (Extensible Data Notation) reader written in C11 with SIMD boost
- medv 10mo agoA very impressinve implementation with SIMD and WASM!
- Jeaye 10mo agoThis is superb. Thank you for making it and licensing it MIT. I think this is a contender to replace the lexer within jank. I'll do some benchmarking next year and we'll see!
- drob518 10mo agoOooo that’d be nice.
- delaguardo 10mo agoWow, that is a greate news!) Thanks for looking at it from this perspective! There are some benchmarks already available in the project - https://github.com/DotFox/edn.c/blob/main/bench/bench_integration.c https://github.com/DotFox/edn.c/blob/main/bench/bench_integr... you can run it locally with `make bench bench-clj bench-wasm` Let me know if I can do anything to help you with support in jank.
- Jeaye 10mo agoIt looks like the key missing part which would be needed for a lexer is source information (bare minimum: byte offset and size). I don't think edn.c can be used as a lexer without that, since error reporting requires accurate source information. As a side note, I'm curious how much AI was used in the creation of edn.c. These days, I like to get a measure of that for every library I use.
- delaguardo 10mo agoIt should be easy to add source info for every token, some of them already keep both (size and offset) I can create a branch for that. > I'm curious how much AI was used in the creation of edn.c A fair amount. This is my first big public project written in pure C. I did consult LLM about best practices for code organisation, memory management, difference in SIMD instructions between platforms, etc. All the things Clojure developer typically don't think about (luxury of a hosted language). Ultimately, the goal was to learn some part of C programming, working reader is a side effect of that. > These days, I like to get a measure of that for every library I use. Btw, I'm curious, what kind of measuring you are looking for?
- deleted 10mo ago[deleted]
- HexDecOctBin 10mo agoCan the metadata feature be used to ergonomically emulate HTML attributes? It's not clear from the docs, and the spec doesn't seem to document the feature at all.
- nerdponx 10mo agoI'm not sure how the metadata syntax works, but you might not need it because you can do this: (html (head (title "Hello!")) (body (div (p "This is an example of a hyperlink: " (a "Example" :href "https://example.org/")))))
- delaguardo 10mo agoI think you can use metadata to model html attributes but in clojure people are using plain vector for that. https://github.com/weavejester/hiccup https://github.com/weavejester/hiccup tl;dr first element of the vector is a tag, second is a map of attributes test are children nodes: [:h1 {:font-size "2em" :font-weight bold} "General Kenobi, you are a bold one"]
- geokon 10mo agoAlso Hickory。。 https://github.com/clj-commons/hickory https://github.com/clj-commons/hickory i feel hiccup indexed based magic is a common design pattern you see in early Clojure and its less common now (id say hiccup is the exception that lives on)
- huahaiy 10mo agoVery nice. Is there a plan to have an EDN writer in C as well?
- delaguardo 10mo agoYes, plan is there but didn't have time yet. Most likely will be available next week
- huahaiy 10mo agoWonderful. Looking forward to it.
- zzo38computer 10mo agoI think it would be better to not use Unicode (so that you can use any character set), and to use "0o" instead of "0" prefix for octal numbers. Also, EDN seems to lack a proper format for binary data. I think ASN.1 (and ASN.1X which is I added a few additional types such as key/value list and TRON string) is better. (I also made up a text-based ASN.1 format called TER which is intended to be converted to the binary DER format. It is also intended that extensions and subsets of TER can be made for specific applications if needed.) (I also wrote a DER decoder/encoder library in C, and programs that use that library, to convert TER to DER and to convert JSON to DER.) ASN.1 (and ASN.1X) has many similar types than EDN, and a comparison can be made: - Null (called "nil" in EDN) and booleans are available in ASN.1. - Strings in ASN.1 are fortunately not limited to Unicode; you can also use ISO 2022, as well as octet strings and bit strings. However, there is no "single character" type. - ASN.1 does have a Enumerated type, although the enumeration is made as numbers rather than as names. The EDN "keywords" type seems to be intended for enumerations. - The integer and floating point types in ASN.1 are already arbitrary precision. If a reader requires a limited precision (e.g. 64-bits), it is easy to detect if it is out of range and result in an error condition. - ASN.1 does not have a separate "list" and "vector" type, but does have a "set" type and a "sequence" type. A key/value list ("map") type is a nonstandard type in ASN.1X, but standard ASN.1 does not have a key/value list type. - ASN.1 does have tagging, although its working is difference from EDN. ASN.1 does already have a date/time type though, so this extension is not needed. Extensions are possible by application types and private types, as well as by other methods such as External, Embedded PDV, and the nonstandard - The rational number type (in edn.c but the main EDN specification does not seems to mention it), is not a standard type in ASN.1 but ASN.1X does have such a type. (Some people complain that ASN.1 is complicated; this is not wrong, but you will only need to implement the parts that you will use (which is simpler when using DER rather than BER; I think BER is not very good and DER is much better), which ends up making it simpler while also capable of doing the things that would be desirable.) (But, EDN does solve some of the problems with JSON, such as comments and a proper integer type.)
- delaguardo 10mo ago> EDN seems to lack a proper format for binary data The best part of EDN that it is extendable :) #binary/base64 "SGVsbG8sIHp6bzM4Y29tcHV0ZXIhIEhvdyBhcmUgeW91IGRvaW5nPw==" This is a tagged literal that can be read by provided (if provided) custom reader during reading of the document. The result could be any type you want.
- pgoudjo 10mo ago[dead]
- Hammershaft 10mo agoI'm grateful for this! Love seeing EDN find its way into new places.
- delaguardo 10mo agoAnd also with experimental things that might eventually find its way to Clojure :) Check out digit delimiters and text block features: {:pi 3.141_592_653_589 :c 299_792_458 :hex 0xDE_AD_BE_EF :tb """ #!/usr/bin/env bb (require '[babashka.http-client :as http]) (defn get-url [url] (println "Downloading url:" url) (http/get url)) """}
- exceptione 10mo agoInteresting, I had to look up what EDN is. Important to note that EDN doesn't have a concept of a schema like JSON Schema. This is a `map`, which bears semblence with a Json object. The following might look like an incorrect paylood, but will actually parse as valid EDN: {:a 1, "foo" :bar, [1 2 3] four} // Note that keys and values can be elements of any type. // The use of commas above is optional, as they are parsed as whitespace. If one wants to exchange complex data structures, Aterm is also an option: https://homepages.cwi.nl/~daybuild/daily-books/technology/aterm-guide/aterm-guide.html https://homepages.cwi.nl/~daybuild/daily-books/technology/at... Some projects in Haskell use Aterms, as it is suitable for exchanging Sum and Product types.
- delaguardo 10mo agoOne of a key design principles in EDN is to be exclusively data exchange format. Which is true even for JSON where json-schema is something that sits on top of JSON itself. Same goes to EDN - in Clojure there is clojure.spec that adds schema like notation, validation rules and conformation. https://clojure.org/about/spec https://clojure.org/about/spec , something like this could be implemented in other languages as well.
- fulafel 10mo agoJSON doesn't have schemas either, JSON Schema is just a separate schema spec that happens to build on JSON, but you might be using for example Zod instead of that. Similarly systems that consume EDN can have various schema systems. For example spec or malli in the Clojure world. (Or you could be using Zod with EDN, etc).
- ndr 10mo agoOne clever difference is that you can use namespaced keywords for instance no more doubt of what :username might mean, you could use :com.ycombinator.news/username as key. If you worked in exclusively static typed languages it takes a while to grasp how convenient this is when you're mixing data from different sources.
- eliasdejong 10mo ago[dead]
- delaguardo 10mo agoThanks for the link! Yes, EDN is a textual format intended to be human-readable. There is also a format called Transit used to serialise EDN elements. Unlike raw EDN, Transit is designed purely for program-to-program communication and drops human readability in favor of performance. It can encode data into either binary (MessagePack) or text (JSON), but in both cases, it preserves all EDN data types and originates from the Clojure language. https://github.com/cognitect/transit-format https://github.com/cognitect/transit-format
- zzo38computer 10mo ago> EDN also has no builtin 'raw bytes' type. That was my complaint too. > I am working on a format consisting of serialized B-tree. It is essentially a dictionary, but serialized I had wanted something a bit similar; a serialized B-tree (or a similar structure) but with only a 'raw bytes' type, for keys and values (I will use DER for the values; I have my own library to work with DER already), and the ability to easily find all records whose key matches a specified prefix.
- sevensor 10mo agoI don’t wish to pick on this post, it looks quite well done. However, in general, I have some doubts about data formats with typed primitives. JSON, TOML, ASN.1, what have you. There’s very little you can do with the data unless you apply a schema, so why decode before then? The schema tells you what type you need anyway, so why add syntax complexity if you have to double check the result of parsing?
- delaguardo 10mo agoI can do a lot without applying schema at all. For that I only need handful of types defined in EDN specification and Clojure programming language.
- sevensor 10mo agoSuppose you have the EDN text ( { :name "Fred" :age 35 } { :name 37 :age "Wilma" } ) There's a semantic error here; the name and age fields have been swapped in the second element of the list. At some point, somebody has to check whether :name is a string and :age is a number. If your application is going to do that anyway, why do syntax typing? You might as well just try to construct a number from "Wilma" at the point where you know you need a number. Obviously I have an opinion here, but I'm putting it out there in the hope of being contradicted. The whole world seems to run on JSON, and I'm struggling to understand how syntax typing helps with JSON document validation rather than needlessly complicating the syntax.
- delaguardo 10mo agoWhat do you mean under "syntax typing" and complications in the syntax? > The whole world seems to run on JSON That is true, and I don't like that :) From my perspective JSON syntax is too "light" and that translates to many complications typically in the form of convention: {"id": {"__MW__type": "LONG NUMBER", "value": "9999999999999999999999999"}}.
- 10mo ago
- deleted 10mo ago[deleted]
- rurban 10mo agoStopped reading at this insanity https://github.com/DotFox/edn.c?tab=readme-ov-file#control-characters-in-identifiers https://github.com/DotFox/edn.c?tab=readme-ov-file#control-c... Call it symbol, if it's not identifiable
- delaguardo 10mo agoFrom the spec: about symbols - "Symbols are used to represent identifiers" about keywords - "Keywords are identifiers that typically designate themselves." more than that: in this reader implementation nil, true, false are also valid identifiers with special handling to turn them into nil and bool. I encourage you to reread original specification and then continue after "this insanity" :)