17 ms·
Smaller Code, Better Code
- nattaylor 10y ago>for every one of those 750 lines, I've had to examine, rework, and reject around 5400 lines of code. I guess there's no such thing as "good enough" with a compiler? Those are staggering numbers to me. Kudos to the author.
- rakoo 10y agoReminds me of that good ol' folklore: http://www.folklore.org/StoryView.py?story=Negative_2000_Lines_Of_Code.txt http://www.folklore.org/StoryView.py?story=Negative_2000_Lin...
- ythn 10y agoThe kind of managers who use lines of code as a performance metric are the kinds of managers I avoid like the plague. It usually indicates that they don't understand what I am working on and are likely to reward code firefighters more than code surgeons.
- EGreg 10y agoWhat do you think is a good measure of productivity? In my opinion it is the number of requested features shipped minus the number of bugs introduced. Weighted by the importance of each, as collectively decided on by everyone or client.
- arcfide 10y agoFeatures is the wrong metric. There is no good quantitative measure, because what matters at the end is the degree of effectiveness at addressing a human need, which is always fuzzy at the heart of it, even though we may try to abstract that fuzziness into something with which we can work.
- ScottBurson 10y agoI think that's about as close as you can get to a usable metric. Of course, it still has problems. The metric probably has to be calculated long before the number of serious bugs is even known. On the other hand, good developers anticipate needs; there may be features implemented that the users haven't even realized they wanted yet. And of course, it's hard to come up with numerical weights for the features and bugs. But the worst problem with this metric is that it doesn't count the maintainability of the code. That's an even harder thing to measure, of course.
- curun1r 10y agoNot that I believe that you can boil down productivity to a single metric, but I've often found one of the best proxies for productivity is the number of automated tests added or modified. It's one that can be gamed, but absent that, it comes pretty close to capturing the amount of work that a programmer has completed since more complex features require more automated testing. It's also pretty easy to pull from CI build results, along with how often a developer was breaking and/or fixing the build. Combine that with a code review process that will surface commits with excessive or insufficient tests and it was one of a couple ways I validated my feel for the work of my direct reports when it came time to fill out reviews.
- tigershark 10y agoWhy the number of times that I break the build should matter in any case? I think that I break the build several times per week because when I test the application using F5 in visual studio it doesn't do a full build and so the tests may not compile. Furthermore, even if they compile they may be broken. This is the primary purpose of the build system, to run the tests and to fail when something is wrong so I don't have to spend time running tests on my machine. Once they fail on teamcity I fix them. And I will be very surprised if anyone of my colleagues complains about it given that only I use my feature branch.
- curun1r 10y agoBreaking your feature branch isn't breaking the build anymore than breaking the build on your local laptop is. "Breaking the build" means breaking master or a release branch. Those are the ones that inconvenience others and limit response time to critical production issues.
- tonyedgecombe 10y agoWhat do you think is a good measure of productivity? Profit? The trouble is it is so far removed from the day to day work we do it's almost impossible to draw any direct conclusions. So instead we start trying to use proxies like function points, bugs or lines of code. In the most dysfunctional organisations worse metrics get used like time keeping a seat warm or political skills.
- arcfide 10y agoThere is a development method called "Direct Development" which has arisen as a term to describe the organic programming model that many profitable APL endeavors have followed. It helps to eliminate the issue of metrics by eliminating the divide between the programming/IT unit and the users of the software. In companies that are using Direct Development, the metric is "are they giving me what I need?" And the way they accomplish this is by pair programming with the users of the software themselves. That's right, the users of the code actually participate directly in the development of the code. They write the code in such a way that the code itself becomes the specification analogue and is in fact the shared knowledge between the coders and the users. Not the comments, but the code itself. The users read the code, work with the programmers, and they make changes on the fly (with appropriate QA). If your users have the confidence that they can walk up and get any feature they need implemented with you, that's about the best metric of success I can think of.
- ScottBurson 10y agoThat's a great story. I can't resist quoting Dijkstra: If we wish to count lines of code, we should not regard them as "lines produced", but as "lines spent" -- the current conventional wisdom is so foolish as to book that count on the wrong side of the ledger.
- nickpsecurity 10y ago"the current conventional wisdom is so foolish as to book that count on the wrong side of the ledger." Never seen that. Thanks for sharing as it might be a great way to drive point home to business types. The amount of code certainly turns [for developers] from asset into a liability as it's size and usage grows. Hmm. Maybe such a presentation could always consider the amount of code a liability or neutral asset that produced benefits on the other side or reduced them. A well-stated connection between the two might justify reducing technical debt.
- arcfide 10y agoI am very inspired by Dijkstra's high level ideas on programming. Importantly, one of the fundamental assumptions of Dijkstra was that you could actually understand your code base and reason about it. The creation of excessive abstraction may create a degree of robustness that protects against programmer's who don't understand the code base, under the assumption that no one will, but at the cost of eventually ensuring that no one will be able to understand the entirety of the code base or even reason at the macro and micro levels efficiently at the same time for a large part of the code base.
- franciscop 10y agoThe numbers are nothing like this, but I had a really similar experience to the author when doing Umbrella JS. With exceptions, but I've tried to keep every function down to few lines of code by doing heavy code reuse: // src/addclass/addclass.js // Add class(es) to the matched nodes u.prototype.addClass = function () { return this.eacharg(arguments, function (el, name) { el.classList.add(name); }); }; While they don't do exactly the same (Umbrella JS is more flexible but jQuery supports IE9), compare that to jQuery's addClass(): addClass: function( value ) { var classes, elem, cur, curValue, clazz, j, finalValue, i = 0; if ( jQuery.isFunction( value ) ) { return this.each( function( j ) { jQuery( this ).addClass( value.call( this, j, getClass( this ) ) ); } ); } if ( typeof value === "string" && value ) { classes = value.match( rnothtmlwhite ) || []; while ( ( elem = this[ i++ ] ) ) { curValue = getClass( elem ); cur = elem.nodeType === 1 && ( " " + stripAndCollapse( curValue ) + " " ); if ( cur ) { j = 0; while ( ( clazz = classes[ j++ ] ) ) { if ( cur.indexOf( " " + clazz + " " ) < 0 ) { cur += clazz + " "; } } // Only assign if different to avoid unneeded rendering. finalValue = stripAndCollapse( cur ); if ( curValue !== finalValue ) { elem.setAttribute( "class", finalValue ); } } } } return this; },
- burgerdev 10y agoAt first I was wondering how he managed to write a compiler in 750 loc. Then I noticed it's for APL, which I would call terse: Y0←{⊃,/((⍳≢⊃n⍵)((⊣sts¨(⊃l),¨∘⊃s),'}',nl,⊣ste¨(⊃n)var¨∘⊃r)⍵),'}',nl} See also https://en.wikipedia.org/wiki/APL_(programming_language)#Examples https://en.wikipedia.org/wiki/APL_(programming_language)#Exa...
- zzzcpan 10y agoHe also replaces long names with short ones, so it's more like an obsession. First commit I clicked on was replacing "penv" with "p" just to make it shorter.
- coldtea 10y agoThat wouldn't affect line count.
- devmunchies 10y agoIt would if you have a max length rule for lines and some need to broken up.
- hobarrera 10y agoIt does affect deleted/added lines. You can quickly have several dozed deletions and additions renaming a single variable.
- arcfide 10y agoThere's a specific reason I made that switch, which for a long time had appeared to be a silly change. Eventually I realized that "penv" as a name was so different from the rest of the naming conventions that it was causing cognitive dissonance in my programming that was taking me out of the flow and making it more difficult to work with the code. Move to the name "p" did shorten the code, but more importantly, brought more consistency, predictability, and regularity into the code base. It is a case of synergizing simplicity and brevity and how they work together.
- dude01 10y agoWoah! From the article: "added roughly 4,062,847 lines of code to the code base, and deleted roughly 3,753,677".
- zzzcpan 10y agoThis is not a good thing though, meaning the language and abstractions are not expressive and not reusable enough. Self-hosting compilers, like the author's, feel wrong to me because of that, meta DSLs for compilers should serve as much better abstractions and save a lot of work.
- arcfide 10y agoExcept that your meta DSL probably isn't able to solve the problem that this compiler is solving, which is putting an entire compiler natively onto the GPU in a way that the code is actually maintainable in a "native GPU" version, rather than requiring translation from some other state. This compiler has gone through many core paradigm shifts in an attempt to find an appropriate way to express a solution to the problems that it encountered. Each iteration revealed some new insight into how to solve the problem, but inevitably lead to a need to rethink the system. Now, the system is so expressive and capable that reusability isn't even an issue. At this point reusability is about as useful in the compiler as having a new word to represent the word "the". Why? Why not just write the? Anything else you could write is likely to create confounding layers of indirection and distance between definition and use in the code that will actually obscure clarity. Instead, I take the intentional approach to make the code as "disposable" as possible. Why change a compiler pass that is two lines long when you can just rewrite it from scratch in less time? By leveraging a different aesthetic, architecture, and language, I'm able to have more expressivity by removing unnecessary abstraction and making it as easy as possible to re-engineer the whole thing at the drop of a hat. This means that I never have to "live with" code bloat or some design decision that's annoying me. The cost to re-engineer is so low that I have almost no technical debt. If an architecture fails to scale, replace it and move on, without any loss of productivity, and a net gain since the code gets easier and easier to work with on each iteration.
- finin 10y agoI've found the when teaching, I sometimes work on an example program too much, producing what I think is elegant and compact code, but that the students find hard to understand. I suspect that the same may be true when I am collaborating with others on a program. There can be value in writing code in a straightforward, easy to comprehend style.
- Silhouette 10y agoMy experience has been similar. A good coding style for teaching, when the reader doesn't yet recognise the building blocks of a language or their idiomatic usage, is often very different to a good coding style for professional use, when the reader can be assumed to understand the concepts and idioms already.
- akkartik 10y agoThis affects more than just beginning programmers. Many best practices we teach programmers today help insiders manage a project but hinder understanding in newcomers to the project (even if they're familiar with the language, libraries, etc.). In a strange new project straight-line code is usually easier to follow than lots of indirection and abstractions. Comments are of limited value because most comments explain local features, but fail to put them in a global context. Build systems that automate a lot of work in our specialized industrial-strength setup turn out to be brittle on someone's laptop when running for the first time.
- Silhouette 10y agoYou raise interesting points, though I don't think this one is obviously true: In a strange new project straight-line code is usually easier to follow than lots of indirection and abstractions. I would argue that indirection and abstraction can always be harmful to code readability if the amount of complexity they hide is less than the added complexity from using them. For example, if your abstractions are leaky and you use them to break a long algorithm down into a hierarchy of very short functions, a reader probably still often needs to look through the implementation to figure out what is really happening, but now they have to follow several levels of indirection to find that information. I would also argue that if you choose levels of abstraction that really do hide a lot of complexity most of the time, this can be helpful for beginners learning a new system as well. For example, this can happen when the abstractions in question have intuitive meanings, such as representing real world objects or other recognisable concepts from the problem domain you're working with. It can also happen when the abstractions represent common patterns of behaviour in some reasonably clean and concise form, such as navigating a data structure in a particular way. In short, while I agree with you that straight-line code can be clearer than lots of indirection and abstractions, I think that is often because of the poor choices of the latter rather than the experience level of the reader with that particular project. If they were very new, they'd have to learn what the key concepts and common patterns of behaviour were anyway, and once they do understand those ideas, good code using them should be easier to follow than code written using more primitive concepts.
- natch 10y agoFrom the project: ... rth,←' A zs;A rs=scl(r.v(0));rr##mf(zs,rs,p);if(c==1){z.v=zs.v;R;}\',nl rth,←' array v=array(z.s,zs.v.type());v(0)=zs.v(0);\',nl rth,←' DO(c-1,rs.v=r.v(i+1);rr##mf(zs,rs,p);v(i+1)=zs.v(0))z.v=v;)\',nl rth,←' DL(zz,if(rr##scl){rr##df(z,l,r,p);R;}\',nl ... No. And commit messages like "Hopefully that does it." No again.
- RodgerTheGreat 10y agoAre you going to articulate your objection to that code or just sneer at it unconstructively?
- dang 10y agoSnarky dismissals are not ok on Hacker News, especially not when they're advocating an entirely conventional and dare I say middlebrow position. When faced with something unconventional, the reaction we're hoping for from HN users is first to pause—and then to reflect. If after pausing and reflecting you want to argue that the conventional position is right, you'll be able to do that thoughtfully and with some sense of nuance.
- foxhill 10y agoin defence of the parent's snarkiness, this code is disgusting. imagine being presented with this and tasked with maintaining this. or adding a language feature. i'm certain the author could do it without much effort, but this code is as short as to be obfuscated - i have had more understanding from ioccc entries than this. code exists as a common language for humans to understand and collaborate. this code is nightmare-ish.
- dang 10y agoThere are quite a few assumptions in your comment that you could investigate if you wanted to. That might be more interesting than just being disgusted.
- 10y ago
- jfoutz 10y agoAs pointed out in paip, clarity and concision are at odds. It takes good taste to balance the two.
- BurningFrog 10y agoI think it also takes empathy. In the sense that you can imagine how the code would read to someone else, who was new to it.
- edblarney 10y agoSmaller is better, but that does not mean 'fancy pants super dense cryptic code'. I think 'simpler' would be a better term than 'smaller'. Also - every line of code has cost. A lot of cost. Maintenance of code and complexity is not only expensive, but it adds to the maintenance of other code. So less code to solve the problem is almost always better.
- arcfide 10y agoAt the heart you are absolutely right. We're after simplicity and clarity. However, I have found that "small" really does make a difference, especially if you push yourself to be small on the macro, rather than micro level. If I just chose "simple," it is too easy to believe that it's "simple enough." If I force myself to maintain poetry like small-ness, then I'm not just able to get by with "simple enough" but have to seek macro levels of simplification that we can often fail to see when the code is so large that all we look at is the single, local view of a single function. By forcing myself to ever greater degrees of ascetical code sizes, small, cute micro hacks in a given function don't work. At that point the "fancy pants" hacks fail, and I am forced to create macro simplifications that obviate the need for whole classes of programming techniques. So, yes, we want simple, but it's about how we can push ourselves and our minds to get there.
- edblarney 10y agoYup, I agree on the 'smaller architecture' bit. One more point: I find that there are a lot of very common things that we, as developers, have not 'standardized' on - but if we did, it would be beneficial. The underscore/lodash JS libraries are great examples of this. They are not just a bunch of 'helper functions' - they are really a series of new 'functional keywords' that in a way represent a new paradigm in software: we all get used to these 'mini patterns' and call them the same thing, and when used in code they can make things a lost simpler. Map, reduce, find, each, pull, filter etc. etc. - at first glance it would seem compulsive to jam all these into some code - but once the developers are familiar with them ... guess what - they become almost part of the programming language itself. So I think this is a pretty good example of a 'meta' way to facilitate simplicity: agree on names for very common patterns, and abstract them away with tools or linguistic constructs.
- n0mad01 10y agothats roughly 1369 loc added per commit or 1855 loc per day.
- arcfide 10y agoAs the author of this code in question, I'd like to make the offer to the Hacker News community and anyone at large. I'll do a live screen cast demonstration for interested persons and walk you through the entire compiler in 30 minutes to 1 hour. In the end you won't have a complete understanding of the compiler, but if you have reasonable prior programming experience, I claim that you will have a better, more full, and complete understanding of the compiler than if you had spent the same amount of time learning most other compiler designs. At that point, you would be able to continue your own self-study and would be able to start making contributions to the compiler rather quickly. This is an offer to demystify the code to people so that they have an opportunity to see how it really does make the whole compiler simpler and easier to work with. If people express interest, I'll run such a live session and let people judge for themselves what they think of the code and my approach to "simplicity" after they've been introduced personally to the code base.
- jpt4 10y agoI would observe such a live session.
- dang 10y agoThat's a great idea. If you'd be interested in doing this semi-officially on HN (maybe something along the lines of an AMA) please email hn@ycombinator.com and let's co-ordinate it!
- arcfide 10y agoDone.
- camelspade 10y agoI would like to see this as well, sounds very interesting
- chetanbhasin 10y agoI'd be down for such a session. Sounds like a great idea!
- 10y ago
- jcoffland 10y agoIt's interesting to note that the author has written more lines here in this thread than are contained in the compiler in question. The English language is not nearly as concise as APL.
- skybrian 10y agoIt seems like there is a missing explanation of the language this compiler compiles and why someone would want to use it? (Searches on "dfns" and "co-dfns" don't find much.)
- known 10y agoAKA https://en.wikipedia.org/wiki/Pareto_principle https://en.wikipedia.org/wiki/Pareto_principle
- arcfide 10y agoThe live session is up and running now. You can find more information about the stream and ask your questions at the following post: https://news.ycombinator.com/item?id=13638086 https://news.ycombinator.com/item?id=13638086
- fourier 10y agoHere is the link: https://www.youtube.com/watch?v=gcUWTa16Jc0 https://www.youtube.com/watch?v=gcUWTa16Jc0 and proper q/a thread: https://news.ycombinator.com/item?id=13638086 https://news.ycombinator.com/item?id=13638086