13 ms·
Ohm – A library and language for building parsers, interpreters, compilers, etc.
- tobr 6y agoSpeaking of - what’s the status of HARC? Is it defunct?
- jagger27 6y agoDefunct enough to let their TLS cert expire.
- azeirah 6y agoYep, HARC is no more. I don't recall the exact history but iirc SAP withdrew its funding and HARC basically ceased to exist. Now, ohm survives as an open-source project, Bret Victor continues work with Dynamicland and Vi Hart is currently employed at Microsoft Research.
- corysama 6y agoThis is a follow-up to a major component of the http://vpri.org/writings.php http://vpri.org/writings.php project that created an self-contained office suite, OS and compiler suite in something like 100-200k lines of code without external dependencies.
- hobo_mark 6y agoDo you have a link to the project? I'm failing to find it on that page.
- beagle3 6y agoNot op, and can’t google now but the project was called STEPS, they did a down-to-metal os including network and GUI (and mote) in 20k lines. Don’t remember anything about office suite. Related names I remember are Alan Kay, Dan Amelang, Alessandro Wirth and Ian Piumarta.
- renox 6y agoThe 'Word' equivalent was called Frank but AFAIK nobody has been able to reproduce what was demonstrated.. Quite painfully ironic for a software research project that they didn't use properly a VCS..
- elgertam 6y agoThey did use VCS, actually, but a lot of them used SVN and each person in the STEPS project was hosting their own code. Most of those servers have gone dark now, though you can find random ports over to GitHub (rarely with the version history). As far as I can tell, Dan Amelang and Alex Warth were the only two who used git or moved their code over to git.
- hobo_mark 6y agoThank you, funnily enough this lead me back to the orange website: "STEPS Toward the Reinvention of Programming, 2012 Final Report Submitted to the National Science Foundation (NSF) October 2012" https://news.ycombinator.com/item?id=11686325 https://news.ycombinator.com/item?id=11686325
- elgertam 6y agoThe biggest artifact from STEPS was Frank, which was at the time bootstrapped using Squeak Smalltalk and included the work from Ian Piumarta (IDST/Maru, which was a fully bootstrapped LISP down to the metal), Dan Amelang (Nile, the graphics language, and Gezira, the 2.5D graphics library implemented in Nile, which both depended on Maru), Alex Warth (OMeta, which had some sort of relationship to Ian's work on Maru), Yoshiki Ohshima (a lot of the experimental things from Alan's demos of Frank were made by Yoshiki) and then several other names. I got close to getting Frank working, but honestly, I'm not sure it's worth it at this point. A lot of the work is 10-15 years old, and the last time I dove in, I ran into issues running 32-bit binaries. The individual components are more interesting and could be packaged together in some other way. Since it was a research project, STEPS never quite achieved a cohesive, unified experience, but they proved that the individual components could be substantially minimized and the cost of developing them amortized over a large project like a full GUI environment. Nile and some of the applications of Maru, like a minimal but functioning TCP/IP stack that can be compiled to bare metal by virtue of being made in Maru, still fascinate me. Work on Maru is ongoing, albeit run by a community (with some input from Ian), Nile has been somewhat reborn of late, Ohm is again under active development as the successor to OMeta and Alan is still around. (Source: Dan is a friend and colleague, and I've met a few of the STEPS/VPRI people that way.)
- e12e 6y agoSee: https://en.m.wikipedia.org/wiki/Ometa https://en.m.wikipedia.org/wiki/Ometa (including reference section) Or go to: http://www.vpri.org/writings.php http://www.vpri.org/writings.php If I recall correctly you want: "STEPS Toward the Reinvention of Programming, 2012 Final Report Submitted to the National Science Foundation (NSF) October 2012" (and earlier reports) Discussed on hn: https://news.ycombinator.com/item?id=11686325 https://news.ycombinator.com/item?id=11686325 And: https://news.ycombinator.com/item?id=585360 https://news.ycombinator.com/item?id=585360 Notable for implementing tcp/ip by parsing the rfc. "A Tiny TCP/IP Using Non-deterministic Parsing Principal Researcher: Ian Piumarta For many reasons this has been on our list as a prime target for extreme reduction. (...) See Appendix E for a more complete explanation of how this “Tiny TCP” was realized in well under 200 lines of code, including the definitions of the languages for decoding header format and for controlling the flow of packets." (...) "Appendix E: Extended Example: A Tiny TCP/IP Done as a Parser (by Ian Piumarta) Elevating syntax to a 'first-class citizen' of the programmer's toolset suggests some unusually expres- sive alternatives to complex, repetitive, opaque and/or error-prone code. Network protocols are a per- fect example of the clumsiness of traditional programming languages obfuscating the simplicity of the protocols and the internal structure of the packets they exchange. We thought it would be instructive to see just how transparent we could make a simple TCP/IP implementation. Our first task is to describe the format of network packets. Perfectly good descriptions already exist in the various IETF Requests For Comments (RFCs) in the form of "ASCII-art diagrams". This form was probably chosen because the structure of a packet is immediately obvious just from glancing at the pictogram. For example: +-------------+-------------+-------------------------+----------+----------------------------------------+ | 00 01 02 03 | 04 05 06 07 | 08 09 10 11 12 13 14 15 | 16 17 18 | 19 20 21 22 23 24 25 26 27 28 29 30 31 | +-------------+-------------+-------------------------+----------+----------------------------------------+ | version | headerSize | typeOfService | length | +-------------+-------------+-------------------------+----------+----------------------------------------+ | identification | flags | offset | +---------------------------+-------------------------+----------+----------------------------------------+ | timeToLive | protocol | checksum | +---------------------------+-------------------------+---------------------------------------------------+ | sourceAddress | +---------------------------------------------------------------------------------------------------------+ | destinationAddress | +---------------------------------------------------------------------------------------------------------+ If we teach our programming language to recognize pictograms as definitions of accessors for bit fields within structures, our program is the clearest of its own meaning. The following expression cre- ates an IS grammar that describes ASCII art diagrams."
- infinite8s 6y agoThey were trying for 10k lines of code(I think I saw Alan Kay mention online that they got to about 20k lines).
- tovej 6y agoCompiler compilers are great, I love writing DSLs for my projects. I usually use yacc/lex, or write my own compiler (typically in go these days). However (and this is just me talking), I don't see the point in a javascript-based compiler. Surely any file format/DSL/programming language you write will be parsed server-side?
- TheRealPomax 6y agoIf your ecosystem is JS, having a JS based compiler is pretty convenient. As long as it's just "slower by some constant", rather than by a runtime order, the fact that it's not as fast as yacc/bison etc. is pretty much irrelevant, so being able to keep everything JS is quite powerful for people new to the idea having started their programming career using JS, as well as seasoned devs working in large JS codebases. (and you can always decide that you need more speed - if you have a grammar defined, it's almost trivial to feed it to some other parser-generator)
- deleted 6y ago[deleted]
- peterhunt 6y agoThere’s definitely a use for js based parsing for tooling that runs in the browser (autocomplete, documentation browsing etc). Integration with the Monaco editor is a common use case.
- branneman 6y agoIn that case, way I ask why you are not a Racket user? Sounds like it'll save you a ton of time and keep your implementations high level.
- chrisseaton 6y ago> I don't see the point in a javascript-based compiler JavaScript is a full programming language. Why wouldn't it be a fine choice to write a compiler in? People have a funny idea that compilers are more complex software or are somehow something low-level? In reality they're conceptually simple - as long as your language lets you write a function from one array of bytes to another array of bytes, then you can write a compiler in it. And for practicalities beyond that you just need basic records or objects or some other kind of structure, and you can have a pleasant experience writing a compiler. > Surely any file format/DSL/programming language you write will be parsed server-side? JavaScript can be used user-side, or anywhere else. It's just a regular programming language.
- TheRealPomax 6y agoIt'd be cool if the online editor dispensed with the need to "write the grammar" entirely. A node based parser-generator in addition to Ohm being yet another grammar based parser-generator would be pretty great.
- ampdepolymerase 6y agoEven better would be to generate parser from examples. See the Microsoft Research Excel Flash Fill paper.
- branneman 6y agoWhen should one use Ohm over Racket?
- coldtea 6y agoWhen they want a library and toolkit for building parsers and languages, rather than a general programming language based on Scheme.
- branneman 6y ago... but racket basically exists to create parsers and languages. It happens to also be a general programming language. But so is JS nowadays with Node.
- dunefox 6y agoSo, I guess you don't know why OP specifically asked about Racket: https://www.cs.utah.edu/plt/dagstuhl19/ https://www.cs.utah.edu/plt/dagstuhl19/ https://beautifulracket.com/stacker/why-make-languages.html https://beautifulracket.com/stacker/why-make-languages.html
- coldtea 6y agoNah, I know about Racket's DSL support and touting itself as friengly to language writing, but it's still not the same as a dedicated parsing toolkit, the same way I wouldn't consider a Lisp with reader macros equivalent either...
- hardwaregeek 6y agoI've used PEGs in the past. They're nice since they combine the mental model of LL grammars with the automation of LALR parser generators. However, it is quite easy to accidentally write rules where you never parse the second rule due to the ordering priority for rules. For instance: ident ::= name | name ("." name)+ Because with PEGs, the parser tries the first rule, then the second, and because whenever the second rule matches, the first one will also match, we will never parse the second rule. That's kinda annoying. Of course with PEG tools you could probably solve this by computing the first sets for both rules and noticing that they're the same. Hopefully that's what this tool does.
- sleavey 6y agoThis is what's called left-recursion, and there's indeed a way to deal with it in PEG parsers: https://github.com/PhilippeSigaud/Pegged/wiki/Left-Recursion https://github.com/PhilippeSigaud/Pegged/wiki/Left-Recursion.
- pjmlp 6y agoLove it, this is great for teaching purposes.
- f430 6y agoIf I want to modify GraphQL to support custom syntax, would Ohm work? Or does a solution exist already for my needs?
- gklitt 6y agoOhm’s key selling point for me is the visual editor environment, which shows how the parser is executing on various sample inputs as you modify the grammar. It makes writing parsers fun rather than tedious. One of the best applications of “live programming” I’ve seen. https://ohmlang.github.io/editor/ https://ohmlang.github.io/editor/
- Waterluvian 6y agoA lot of regex testers do this and I can't imagine writing a regex or a parser without.
- anon_tor_12345 6y ago>a parser without can you show me a parser generator that produces this kind of visualization?
- 4silvertooth 6y agoantlr has many that produces such visualization, vscode has some great plugins for it.
- anon_tor_12345 6y agogreat thanks!
- BeKindAndLearn 6y agoNot a syntactic parser generator, but [chevrotain](https://chevrotain.io/playground/ https://chevrotain.io/playground/) can generate flow diagrams for your ruleset. IIRC not just in the playground, but in general.
- thesz 6y agoI used to debug parsing process for VHDL grammar (which is ambiguous on lexem level) with parsing combinators and Haskell REPL. Whenever my "compiler" found a syntax error in test suite, I was able to load part of source around error and investigate where my parser's error or omission is by running parser of smaller and smaller part of grammar on smaller and smaller parts of input. It was 12 years ago. And yes, it is fun. ;)
- crazypython 6y agoThis title is misleading. It's a library and language for building parsers. Full stop. Parsing toolkit, as they say themselves.
- exdsq 6y agoThe title copies the second sentence of their readme: > You can use it to parse custom file formats or quickly build parsers, interpreters, and compilers for programming languages.
- UncleMeat 6y agoI guess it depends on what it means to somebody to build a compiler. Something like yacc says "compiler compiler" in the name but really it is a parser generator. The hard part of industrial compilers is the optimization.
- rkagerer 6y agoOHM is also the acronym for Open Hardware Monitor, a great open-source project for monitoring computer temperatures, fan speeds, voltages, etc: https://openhardwaremonitor.org/ https://openhardwaremonitor.org/
- dw-im-here 6y agoI'd rather put my hand in boiling water than develop a compiler in a dynamic weak typed language.
- deleted 6y ago[deleted]
- pwdisswordfish6 6y agoWrite a compiler in a strongly typed language, and then remove all the type annotations. This may come as a shock, but this is what a compiler (or any codebase) could look like when developed in a weakly typed language.
- mintplant 6y agoThat doesn't work if you're using types for anything beyond correctness-checking. Type-driven dispatch, for example, which tends to be used heavily in big compiler and interpreter projects. And tagged unions (or algebraic datatypes), a natural fit for representing ASTs, become more unwieldy without type-directed features like pattern matching.
- pwdisswordfish6 6y agoSounds like a double standard and possibly moving the goalposts. There are strongly typed languages that don't have those features, and compiler codebases that don't use that kind of architecture. Do they get a pass or not?
- mintplant 6y agoI'm directly responding to the claim that you can "write a compiler in a strongly typed language, and then remove all the type annotations", and that's what compiler architecture looks like in a "weakly typed language". Compiler projects built in languages with richer typing can and do use it for purposes beyond correctness checking, and the idea that you can simply erase all the types and expect the code to work the same is a misconception.
- j0e1 6y agoThis is an example of a library we built using Ohm: https://github.com/Bridgeconn/usfm-grammar https://github.com/Bridgeconn/usfm-grammar [1] It works great for our use-case though I have been eyeing tree-sitter[2] for its ability to do partial parses. [1] USFM: https://ubsicap.github.io/usfm/ https://ubsicap.github.io/usfm/ [2] https://tree-sitter.github.io/tree-sitter/ https://tree-sitter.github.io/tree-sitter/
- joshmarinacci 6y agoI'm so happy to see this on HN. I've used Ohm for several projects. If you want a tutorial for building a simple programming language using Ohm, check out this series I put on GitHub. https://github.com/joshmarinacci/meowlang https://github.com/joshmarinacci/meowlang
- fjfaase 6y agoI recently wrote a similar parser, maybe less fancy, for a workshop on parsing. It does display the the abstract syntax tree with d3.js and also has a build evaluator for a limited set of language constructs. https://fransfaase.github.io/ParserWorkshop/Online_inter_parser.html https://fransfaase.github.io/ParserWorkshop/Online_inter_par... It is based on a parser I implemented in C++.
- recursivedoubts 6y agoAlways fun to find the first commit: https://github.com/harc/ohm/commit/4611bf63c5ecb90d782112d68f73a2277f87ca7d https://github.com/harc/ohm/commit/4611bf63c5ecb90d782112d68... 2014 Neat tool. I write parsers by hand though. More fun, and you can be a lot sleazier.
- codr7 6y agoI recently created a library for the other part of an interpreter. https://github.com/codr7/liblg https://github.com/codr7/liblg https://github.com/codr7/liblgpp https://github.com/codr7/liblgpp
- scroot 6y agoWe are using Ohmjs on a project at work and it is fantastic. I'm hoping one day that Ohmjs and Ohm/s (Squeak) can be compatible again -- would love to have the Smalltalk version of our interpreter and environment we built using this
- jweissman 6y agoI’ve built a number of toy language projects with Ohm and it’s really wonderful. Just a joy to use the visual tooling also. All around really beautiful machinery
- Sparkyte 6y agoOhm just makes me want to watch Nausicaa of the Valley of the Wind.
- PaulHoule 6y agoEach PEG generator promises a revolution but only burns a car. I was disappointed with how they do operator precedence; they use the usual trick to make a PEG do operator precedence which looks cool when you apply it to two levels of precedence but if tried to implement C or Python in it it gets unwieldy. Most of your AST winds up being nodes that exist just to force precedence in your grammar, working with the AST is a mess. For all the horrors of the Bell C compilers, having an explicit numeric precedence for operators was a feature in yacc that newer parser gens often don't have. I worked out the math and it is totally possible to add a stage that adds the nodes to a PEG to make numeric precedence work and also delete the fake nodes from the parsed AST. Unparsing I'm not so sure of, since if someone wrote int a = (b + c); how badly you want to keep the parens is up to you; a system like that MUST have an unparse-parse identity in terms of 'value of the expression', but for sw eng automation you want to keep the text of the source code as stable as you can.
- glrsbstrd 6y agoU realise ohm is used for maesuring ressitance in electiricty dont u?