6 ms·
This looks absolutely amazing. One thing I do find interesting (and a wish were different) is that only programming languages are supported, rather than data f
by emacsen 5y ago
This looks absolutely amazing.
One thing I do find interesting (and a wish were different) is that only programming languages are supported, rather than data formats as well.
For example, two JSON documents may be valid but formatted slightly differently, or a common task for me is comparing two YAML files.
Comparing config files that have a well defined syntax and or can be abstracted into a tree (JSON, YAML, TOML, etc.) would be absolutely lovely, even and including (if possible) Markdown and its ilk.
- mark_and_sweep 5y agoJSON is supported. HTML and XML are missing, too.
- emacsen 5y agoYou're right. I missed JSON. Sadly YAML, TOML and the others I mentioned are not there (yet?)
- softwarebeware 5y agoThere’s always room for contributions!
- simonw 5y agoI would naively expect that this problem is easiest to solve for languages like JSON that have an unambiguous way to be pretty printed.
- chockchocschoir 5y agoIndeed. One could just do `diff $(jq . $fileOne) $(jq . $fileTwo)` and you'll end up with a "nice enough" diff even if $fileOne and $fileTwo were very differently formatted.
- lstamour 5y agoThe problem is when a file also needs to be normalized - e.g. object keys in a different order, YAML syntax expansion. It can be very useful to indicate when a JSON file is identical to another JSON file but some of the properties or array items are out of order and that requires more in-depth knowledge of the data format. Let's not mention that you could UTF-8 encode characters or write out the same character using backslash notation, numeric or boolean data that might be wrapped in a string in one file but not in another, etc. There can still be a lot of modelling and interpretation to consider when comparing data files rather than code files.
- chockchocschoir 5y agoI'm not too familiar with YAML, so can't answer to that. But re JSON: > object keys in a different order They can't be "in a different order" as JSON keys are not ordered. They can be whatever order, and would still be considered the same. > array items are out of order Then it's different, as JSON arrays are ordered. ["a", "b"] is not the same as ["b", "a"] while {a: 1, b: 1} and {b: 1, a: 1} is the same. > you could UTF-8 encode characters or write out the same character using backslash notation, numeric or boolean data that might be wrapped in a string in one file but not in another Then again, they are different. If the data inside is different, it's different. I understand that logically, they are the same, but not syntax-wise, which is why I included the "differently formatted" "disclaimer", it wouldn't obviously understand that "one" and "1" is the same, but then again, should you? Depends on use case I'd say, hard to generalize.
- deleted 5y ago[deleted]
- stormbrew 5y ago> They can't be "in a different order" as JSON keys are not ordered. They can be whatever order, and would still be considered the same. This is what GP is saying, I'm pretty sure. Object member order is non-semantic in json, so in order to do a semantic diff (one that understands structure), you need to canonicalize the order of the two sides. Simply diffing the output of jq doesn't do that, because (afaik) jq doesn't alter the order. Basically, if you want this to come up the same: {"a":"b","c":"d"} {"c":"d","a":"b"} you need more than just `diff $(jq) $(jq)`. Can argue about whether a tool like difftastic should do that, I guess, but I would personally lean towards that it should be smart enough to see this because it's precisely the sort of thing that both humans and line-based diff can be awful at seeing.
- deleted 5y ago[deleted]
- Wilfred 5y agohttps://github.com/andreyvit/json-diff https://github.com/andreyvit/json-diff works really well for JSON diffing in my experience. It's more simplistic than difftastic though: it considers `1` and `[1]` to have nothing in common.
- paxys 5y agoThis isn't going to add anything to existing diff tools for JSON or YAML though. Those formats barely have any syntax highlighting or complex structures.
- Wilfred 5y agoJSON and CSS are supported today, and I'm interested in adding more structured text formats. If a format has a tree-sitter parser, it can be added to difftastic. The TOML tree-sitter parser looks good, but there isn't a mature markdown parser for tree-sitter. There are other markdown parsers available, so in principle difftastic could support markdown that way. The display logic might need a little tuning for prose-heavy formats like markdown though. I'm not happy with how difftastic handles block comments yet either. I'm not sure about formats that contain more prose, such as markdown or HTML.
- zmix 5y agoI think supporting XML would be something, a lot of people would appreciate. That XML is difficult to diff comes up again and again... However, one would need to decide, whether one wants to compare by syntax or by meaning. Latter one may be preferrable, but would require the XML to be canonicalized on both sides, first.
- linsomniac 5y agoI would love a great XML diff tool, and after seeing the demo of this I was sad to see XML not in there. Would pay for.
- d0gsg0w00f 5y agoThis is kind of like the problem of programmatically analyzing AWS IAM roles and policies to understand impact of changes. Very difficult to do in JSON format but worth tons of money to CISOs if it can be solved.
- alxmrs 5y agoSimilarly, I would love it if Pandoc’s AST were supported. Or, if this could be extended to compare any documents taking formatting into account, or document-to-document conversions.
- tomatowurst 5y agosame, I don't know how many times I do a diff and wish there was a smarter solution that could take account formatting and whitespaces. This is it. Wish git diff would incorporate this, would be a real treat.