3 ms·
Don't worry about offending me with that comment. I have a pretty strong belief in why I'm coding this way, so I'm glad to have the opportunity to work with peo
by arcfide 10y ago
Don't worry about offending me with that comment. I have a pretty strong belief in why I'm coding this way, so I'm glad to have the opportunity to work with people like you who find the code scary and disgusting and see if I can't either change your mind or change the code to be better.
Firstly, about this being a personal project, it's actually a bit more than that. It serves as a research platform for a research agenda around the usability of programming languages and the HCI and pedagogy of computation, yes. However, it's also a commercially funded compiler that is commercially licensed and distributed. The compiler is still in early stages, so it's large a boutique offering at this point, but that's scheduled to change this year or maybe the next. And yes, the development team that works with me on this compiler has read the code, and while they are not as fluent in it as I am, they understand how to work with it and we can talk about the compiler and work through issues in the compiler that comes up.
Indeed, the fact that the compiler is so easy to track through at a macro level has allowed us to avoid needing extra documentation throughout, because when we have a question about some level of architecture design, we can usually pull up a page of the compiler and work through it without needing any other documentation.
And, the point of the post above was that small code is a useful metric for pushing for simplicity. There is a difference between obfuscation and small code, but my code is not obfuscated to those who need to work with it. It is obfuscatory for anyone who expects to read it like a normal program.
At the heart of this is the meaning of readability. You've implicitly defined readability as being a state for code bases that allows anyone else to read and understand the code. That's a high bar. If I write a standard proof of the uncountability of the real numbers, it's a rather high bar to say that everyone should be able to read that proof.
Also, if you look at the way that the Clang codebase is engineered, for instance, if you take any one snippet of 50 lines or so of code, it's all nice, neat, and readable. But when it comes to understanding the entire compiler as a whole, the codebase is completely unreadable. It requires external documentation to understand almost any part of that code at a macro level.
But Clang uses best industry practices and is, on the whole, what most people would consider very cleanly written code. And yet, it is essentially impenetrable from a macro level without other documentation.
Instead, I'd submit that readability is something that we should consider valuable for those who have the relevant pre-requisite understandings of key ideas and concepts when looking at a new code base.
Part of the problem is that this code base is introducing new, research level ideas into the coding space. There is a fundamental difference between it and other compilers in the approach that it is taking, and thus, you can't just look for the same patterns.
I've already touched on malleability elsewhere. SICP has a classic quote about the importance of malleability in code (amoeba vs. pyramid programming).
I'll give a simple description of the architecture. If you understand everything said in the following sentence, then you'll have no trouble understanding how the code is written, and if you don't, then learning these sub-domains of programming skillsets will go a long way in helping to clarify the design.
It's a three part dfns->C++ offline batch compiler overloading the standard Quad-Fix system function in APL interpreters to compile whole, closed namespace scripts built of pure functional dfns on the Dyalog 15.0 primitive vocabulary sans-guards through a PEG parser to a core compiler written in a Nanopass compiler architecture over a linearized Quad-XML style matrix AST representation where each pass is written as a data-flow, data-parallel function train leading to a single dispatch code generator with a runtime library header prepended to each output file containing implementations of each implemented APL primitive.
Those would be the basic set of techniques and skills that are being put to use in the compiler. If you already understand Nanopass, PEG parsers, the Quad-XML tree linearization format, function trains, and so forth, then the structure and format and design of the compiler is obvious and easy to work with after about 5 minutes of orientation. If you don't have that background, then understanding that part of the compiler is rather a difficult one. In addition to this, there are new techniques being used and applied to solve problems in this compiler itself, and those are being documented through the papers that I'm publishing on these techniques:
https://github.com/arcfide/Co-dfns#publications https://github.com/arcfide/Co-dfns#publications
Most people don't have a strong data-parallel, array style programming background, which makes the micro-level code the hardest part to understand for them. However, if you are experienced in that background, then working with the compiler passes is not difficult, provided that you take the time to understand the core idioms in play.
So, in summary, I'd say that you're right that the code looks horrendous, because your heuristics are designed for code that is completely, almost assuredly, fundamentally different than this code. However, like I said, come to the live session and see me explicate the architecture of the compiler. I'll explain a lot of the ideas I mention in the above sentence enough to allow you to walk through the code easily. If you still think it's scary, okay. I'd appreciate some ways to make it easier to work with it on a day to day basis.
- libeclipse 10y agoThanks for the reply, and let me just say that you take criticism really well. I still disagree with the premise that the code is clean, beautiful or readable, but I could concede that this may be due to me not having an in-depth understanding of it. You speak about a small code base being a metric for a simple code base, and while this is true sometimes, it starts to fail as a metric when the code is more obfuscated than concise. This is the kind of code I'd see in a webshell, not a compiler. I'll try and make it to your live session if I can; I look forward to it. P.S. I don't know if this was intended or not, but you really came across as trying to make everything seem more complicated than it actually is in the "sentence" where you went over the architecture. Dumping a load of complex-sounding words in a giant sentence just sounds as if you're trying to justify your design decisions by further obscuring understanding of the project. I'm fairly sure this isn't your intention since you are hosting a live session, but it's just how it looks. :P
- arcfide 10y agoI look forward to convincing you of the simplicity of the code base. :-) The sentence was a bit of a tongue-in-cheek sort of rhetoric. In particular, if you look up most of those words in the relevant domains, they're all standard practice ideas that are well understood around their various parts. Importanntly, none of the words or anything said in the sentence is really complicated, and if you were familiar with all of those ideas, then it would be an easy sentence. However, I wrote it in a way so that it appears to be obtuse and a bit ridiculous on the face of it. Basically, attempting a bit to mirror the code itself. The sentence very concisely and neatly describes the architecture of the compiler, but only if you know what you're reading. One of the biggest issues with reading the compiler is that most people will have a "part" of the picture based on their backgrounds. If you have written compilers before, the overall compiler design and the strategies at a macro-level used for it will make perfect sense, but you'll balk at the data-parallel programming style. If you are an APL programmer you'll be very familiar with the basic tricks being used in the compiler and you'll easily be able to see at a micro level what's happening with the code, but not having the background in Programming Language Design (such as you might receive at Indiana University's PL course path), the overall design and the intent of the whole system won't be intuitive to you. This means that likely for anyone new coming into the code, there will be parts of the code base which feel very foreign to them. Of course, I don't know of a way of doing something new without making people learn a thing or two to understand it. Fortunately I'm not the only one in the world that has experience in all of these areas. And even better, the code is straightforward enough that once you do learn the basic skills, it's easy to work with. But there are precious few APL/Array oriented implementers out there, and probably less than a handful of compiler writers out there that specialize in this area of languages. I hope that will change and that I can convince people that this approach is actually a very neat and easy one to work with. But that's not easy to do.