13 ms·
TXR: A Programming Language for Convenient Data Munging
- danso 12y agoAs I come to see more of my data-related work be consumed by data munging/cleaning work, I'm convinced that a language/framework devoted to data munging is at least as important as those devoted to data visualization.
- slackstation 12y agoIt looks ugly and akward to type. It's doesn't seem like it would be a pleasure in which to write programs.
- AnkhMorporkian 12y agoIt's a very ugly language, I don't think anyone is going to disagree there. That being said, it has some intriguing features that I'm not going to dismiss. I work with COBOL on a daily basis, so I'm not going to say no to a new language just because it's ugly. There seems to be a lot of utility here.
- hawkw 12y agoPurely out of curiosity: what is it that you do that forces you to work with COBOL on a daily basis?
- e_modad 12y agoLikewise! I'm curious too. Why is COBOL your main language? Bank legacy servers?
- AnkhMorporkian 12y agoI mostly do legacy code conversion for the banking sector. It's all contract work, so it varies, but 95% of the time that's my deal.
- mhd 12y agoYes, all those at-signs make it look like a combination of lisp and perl, which probably won't excite too many people. But I'd say that data munging is inherently ugly. I don't really see myself using this as the next tool to write clever algorithms that will stand the test of time, but if you offer me this as a stand-in for the usual shell-script/awk/sed/perl/printf/regexp mess you need for ad-hoc file transformations, I'm suddenly listening.
- kazinator 12y agoThe at sign in the TXR pattern language is that way because TXR can match reams of literal text. This is hard to show in small examples, so small examples become dense with the notation. Just like, say, tiny examples of HTML become a dense soup of tags. Note that TXR Lisp doesn't have the at signs. You can write a pure TXR Lisp program by wrapping the whole file with @(do ... ). TXR looks a lot better with syntax highlighting; unfortunately, this only exists for Vim. On the other hand, the syntax highlighting definition file for Vim is quite good.
- kazinator 12y agoOh it can be. Typically, if I need to do some text transformation or extraction, I start by getting sample data and renaming it to a .txr suffix. Then just generalize that data into the TXR pattern that matches it and gets out what is needed. As an example, I was doing some kernel work and needed patches to conform to the kernel's "checkpatch.pl" script. Unfortunately, this thing outputs diagnostics in a way that Vim's quickfix doesn't understand; I wanted to be able to navigate among the numerous sources of errors in the editor. First I looked at the checkpatch.pl script hoping that of course they would have the diagnostic output in one place, right? Nope: formatting of messages is scattered throughout the script by cut-and-paste coding. TXR to the rescue: Sample output: WARNING: line over 80 characters #279: FILE: arch/arm/common/knllog.c:1519: +static void knllog_dump_backtrace_entry(unsigned long where, unsigned long from WARNING: line over 80 characters #321: FILE: arch/arm/include/asm/unwind.h:50: +extern void unwind_backtrace_callback(struct pt_regs *regs, struct task_struct WARNING: line over 80 characters #322: FILE: arch/arm/include/asm/unwind.h:51: + void dump_backtrace_entry_fn(unsigned long where, WARNING: line over 80 characters #323: FILE: arch/arm/include/asm/unwind.h:52: + unsigned long from, Quick and easy TXR to the rescue: @(repeat) @type: @message #@code: FILE: @path:@lineno: @(output) @path:@lineno:@type (#@code):@message @(end) @(end) Result (redirected into errors.err, loads with vim -q): arch/arm/common/knllog.c:1519:WARNING (#279):line over 80 characters arch/arm/include/asm/unwind.h:50:WARNING (#321):line over 80 characters arch/arm/include/asm/unwind.h:51:WARNING (#322):line over 80 characters arch/arm/include/asm/unwind.h:52:WARNING (#323):line over 80 characters arch/arm/include/asm/unwind.h:53:WARNING (#324):line over 80 characters arch/arm/kernel/unwind.c:352:ERROR (#337):inline keyword should sit between storage class and type The nice thing is that we know what the above does when we revisit it six months later.
- klibertp 12y agoTXR has some really cool features and seems very well suited to the domain. If you're going to dismiss perfectly good tool just because it "looks ugly" then you're just being unprofessional. And if it's "awkward to type" you can always write a transpiler if you really need it, or more likely a couple of macros/snippets for your editor.
- _delirium 12y agoMy first reaction was also that it didn't look very clean. But after some admittedly cursory comparison of how you'd do something in TXR to a few existing scripts I have (some Perl, some awk, and some chaining Unix utilities), it doesn't look terrible, and maybe even good. I should emphasize this is based on like 30 minutes of looking at it though, not serious knowledge of how TXR works. One part that seems nice is that it handles multi-line constructs in a way that isn't horrible. Perl and awk have a big complexity jump once you go past one-line records, and most of the traditional Unix utilities just don't handle them at all (stuff like cut/join/sort only works on single-line, delimited records). Since constructs like Perl's while(<INPUT>) stop automatically doing the Right Thing once you get multi-line records, the usual next stop is that you're manually maintaining a state machine.
- spullara 12y agoReminds of a trick I do with mustache.java. Templates can not only be used to generate output, but because of the declarative nature of the mustache language they can be used to parse output back into data that in combination with the template would generate that output. Makes for pretty intuitive parsers. In my case all text that isn't templating declarations are regexes.
- vdm 12y agohttps://github.com/spullara/mustache.java/search?p=1&q=invert&utf8=%E2%9C%93 https://github.com/spullara/mustache.java/search?p=1&q=inver...
- rout39574 12y agoI wish their page included something along the lines of "Why do I care?" Maybe a few examples of "data munging" tasks which the authors view as poor fits for [language X] and how their stuff solves the problem better. Maybe something like "why is our language better than regexps in whatever language environment you already know?"
- kazinator 12y agoThere is page with a navigation frame giving Rosetta Code examples, syntax colored, with back links to Rosetta: http://www.nongnu.org/txr/rosetta-solutions.html http://www.nongnu.org/txr/rosetta-solutions.html TXR has regexps is you need them. The regex engine is geared in a different direction from mainstream regex: it doesn't have anchoring, register capture or Perl features like lookbehind assertions. On the other hand it has intersection and negation (without backtracking). TXR translations of Clojure, Common Lisp and Racket solutions to the same problem: http://www.nongnu.org/txr/rosetta-solutions-main.html#Self-referential%20sequence http://www.nongnu.org/txr/rosetta-solutions-main.html#Self-r...
- rout39574 12y agoI saw those; what I miss is "This is why I think this new way is better". If it's supposed to be obvious by inspection, well... I guess I'm too unenlightened.
- ezequiel-garzon 12y agoI'd say its multi-line approach makes it quite unique when compared to, say, sed or awk.
- deleted 12y ago[deleted]
- stdbrouw 12y agoHmm, this looks more like parsing than munging to me, but then I guess "munging" is not exactly scientific terminology. My own take on easy data transformations, if you'll allow me the plug: https://github.com/stdbrouw/refract https://github.com/stdbrouw/refract
- bane 12y agoCool ideas. I really like that it has support for grammars. What's the performance like compared to Perl on similar tasks?
- hyp0 12y agoAt first I thought this was TXL, for source code transformation http://www.txl.ca/ http://www.txl.ca/
- nieve 12y agoTXR looks rather like the CRM114 language that's been used to implement some rather amazingly accurate text classifiers (some better than most people on their own mail), though a bit less bizarre and I think more accessible: http://crm114.sourceforge.net/docs/INTRO.txt http://crm114.sourceforge.net/docs/INTRO.txt CRM114 too treats pattern matching as the fundamental construct and has blazing performance for it and certain kinds of number crunching (it has to), but I don't think it's nearly as useful for the average hacker trying to munge a couple of text files. Still, worth a look both to users and possibly to language implementors. I'm definit
- deleted 12y ago[deleted]
- aurelius 12y agoKaz Kylheku is one of the kooks from comp.lang.lisp where lisp is the One True Language. The funny thing is that TXR is written in C! Kaz: How come you didn't write TXR in lisp?
- kazinator 12y agoBecause I'm also one of the kooks from comp.lang.c where C is the One True Language. But seriously, TXR is built on its own Lisp: an infrastructure which provides the managed environment and data representations which also support the TXR Lisp dialect. This is no different from any Lisp implementation based on a C kernel, like CLISP, GNU Emacs, ... If you do it from scratch, you lose a lot: you don't have a mature, optimized dynamic language implementation. But, by the same token, you can experiment in ways that you normally wouldn't. You get to dictate things like, oh, what is a cons cell. I have lazy conses that look like ordinary conses: they satisfy consp, and work with car, cdr, rplaca and rplacd. You can invent new evaluation rules. I came up with a way to have Lisp-1 and Lisp-2 in a single dialect, seamlessly, with the conveniences of both. I have Python-like array access. I made traditional Lisp list operations work with vectors and strings: you can mapcar through a string and so on. Sequences and hashes are functions. For instance orf is a combinator that combines functions analogously to the Lisp or operator. If hash1 and hash2 are hash tables, you can do something like [orf hash1 hash2 func] to create a function which takes one one-argument that will look that argument in hash1; then if that returns nil, it will try hash2, and if that returns nil, it will pass the key to func and return whatever that returns. Or ["abc" 1] returns the character #\b. [mapcar "abc" '(2 0 1)] yields "cab": the numeric indices are mapped through "abc", as if it were a index to character function. Fun things like this are good reasons to experiment with your dialect. I believe TXR is a great companion if you're a Lisper working in ... one of those other environments. Ah, one more thing. Well, two, or maybe three. Part of why I used C was to create a project whose tidy, clean internals stand in stark contrast to some of popular written-in-C scripting languages. You know, to sock it to them! See, there is a hidden agenda: the call of "I can do this better". If you use C, then a more direct comparison is possible. Secondly, people widely understand C. Give them a cleanly written project in C, and maybe they will hack on it, and from there understand something about Lisp too. C means low dependencies from the point of view of packaging: easy porting with just basic shell environment with make and a C compiler. Cross-compiling for ARM or whatever is a piece of cake. Easy work for package maintainers, ...