9 ms·
Norris Numbers – Walls you hit in program size (2014)
- montrose 9y ago"Absolutely refuse to add any feature or line of code unless you need it right now, and need it badly." As I've gotten older I've seen the value of this rule. Though it sounds merely negative, it can produce effects that seem little short of genius.
- reificator 9y agoExcepting those features which allow you to cull large chunks of code. Those are my favorite features, rare as they are.
- phkahler 9y agoThose are not exceptions. We all need those now and need them badly ;-)
- kelihlodversson 9y agoIn my experience this is often possible by removing duplicated functionality into a shared implementation. That is by exchanging multiple features of linear complexity into a single implementation. Due to being shared, that implementation will tend to have geometric complexity. That means the duplication has to be highly significant and variations need to be few for it to pay off. ... And you'll better hope you have a good set of tests to verify everything is still working. In most cases it's still a good idea, but it's still a case where you need to argue that you really need it, so it's no exeption at all.
- chewbacha 9y agoSounds like the perennial redundancy vs dependency struggle [0]. I do like the explanation in terms of linear vs geometric complexity though. [0] https://yosefk.com/blog/redundancy-vs-dependencies-which-is-worse.html https://yosefk.com/blog/redundancy-vs-dependencies-which-is-...
- reificator 9y agoThat's one way to do it. Another way is when the requirements change and you no longer need most of a feature.
- phkahler 9y agoPeople often try to implement things with a hypothetical future use in mind. It may even be a realistic future, but it is never the less a future that does not need to be dealt with right now. Adding code for things that aren't actually in play means adding something you can't actually test. It means people may abuse what you set up later (and introduce bugs) if the actual future turns out different than you planned. Taken to the extreme, this planning for the future turns one into an architecture astronaut. https://www.joelonsoftware.com/2001/04/21/dont-let-architecture-astronauts-scare-you/ https://www.joelonsoftware.com/2001/04/21/dont-let-architect...
- Cthulhu_ 9y agoIt's a more fleshed out version of YAGNI: https://en.wikipedia.org/wiki/You_aren%27t_gonna_need_it; https://en.wikipedia.org/wiki/You_aren%27t_gonna_need_it; that page has a number of other similar acronyms; if it ain't broke, don't fix it, KISS, feature creep; all names for the same problem. Mind you, software developers aren't the only ones that cause this; a product owner / lead / whoever calls the shots is also very likely to keep bolting features onto a product. Either they have to keep that impulse down, or they too need a system to manage it. Also they shouldn't be afraid of culling existing features; what usually seems to happen is that more and more stuff is bolted on, until it's too much and they just rewrite the whole thing from scratch, leading to a clean, lean new application - often missing some features that older users like. Either they do it, or a competitor does.
- Joeri 9y agoMy experience is that it is very hard to cull features, even with a sympathetic product manager. Sales wants to have as many boxes ticked as possible, so they fight tooth and nail against feature removal, all the while forgetting that simplicity is a feature too.
- dang 9y agoDiscussed at the time: https://news.ycombinator.com/item?id=8072730 https://news.ycombinator.com/item?id=8072730 and in 2015: https://news.ycombinator.com/item?id=10191540 https://news.ycombinator.com/item?id=10191540
- fredley 9y agoDoes this depend on language? 20k lines of Perl is a different beast from 20k lines of Python, or 20k lines of verbose Java, I would expect. If not, does it suggest we should be aiming to use more expressive languages, that can get more done in less lines?
- evincarofautumn 9y agoJust from my own experience, I’d say yes, with qualification—rather like age is only an approximate indicator of maturity, line count isn’t a meaningful metric per se, but a proxy for more important, less tangible things like complexity, navigability, fitting the program in mind, and so on. In Haskell, for example, I can write software that does more things, using less code, with higher confidence that it’s correct, and much higher ability to change it, than I ever could in C++, my previous top language. That’s partly the language, partly the fact that it matches my style of thinking, and partly the culture around it—writing reusable code, heavily encapsulating state, and making as much of your program as possible “correct by construction” so that invalid states and state transitions are inexpressible—or at least harder to express, such that the easy path is also the right one. A 500-line Haskell module tends to feel very large to me, on par with 1–2000 lines or so of C++. I don’t necessarily mean to just evangelise Haskell or functional programming here. Any language & ecosystem with an expressive orthogonal core, which matches your problem space and not only matches but improves your style of thinking, can give you similar benefits if you’re deliberate about taking advantage of it.
- userbinator 9y ago...and 20k lines of APL is... probably not something any APL programmer has ever reached for any single application.
- ryanmarsh 9y ago20k lines of Perl is a different beast from 20k lines of Python, or 20k lines of verbose Java Very good point. We’re all familiar with the “I can do ${lambda} in ${n} lines with ${lang} language”. It seems to take me 10x more lines to do something in Java vs. Python, and 10x lines to do something in Python vs. Perl. Except that the Java is laborious to read, the Python is nicely readable, and the Perl will never be understood by a programmer. I had great fun writing Perl many years ago but I used to joke that Perl was a “write only” language.
- maximexx 9y agoThis can be true for a beginner, for a developer who never hit that wall before or for a developer without enough talent to learn how to do it right after all. But once you start using the right methodologies, making things highly modular and as much as possible independent from each other, you can go quite a distance before new "walls" arise. Sometimes I take more time to think of a proper and scalable name for variables and methods than the time it actually took to write the implementation. While refactoring I might do another round of thinking about naming and make sure there are no or an absolute minimum of possible side effects for the implementation, etc.. I'm coding almost 30 years now. I'm still learning and make my mistakes of course, but most of the time they occur because I rush for some reason. Writing good code takes time, especially to rethink what you're doing, refactoring, making the right adjustments so it completes the codebase. Before I complete the beta release of a codebase I've made thousands and thousands of decisions, where only 1 wrong decision can cause a terrible amount of trouble later on. For me it's a creative process. It goes in waves. I cannot always be a top performer, I've accepted that. When I recognise I was in a low during some implementation I might do a total rewrite of it or apply some serious refactoring(and force myself to take the time for that). I still experience coding to be much harder to get right than I ever expected it to be. For me a codebase is a highly complex system of maybe hundreds of files with API's working together, not just a bunch of algorithms.
- jacquesm 9y agoIf he had that insight in 2011 then he was pretty late to the party. By then the 1K to 2K line limit for 'beginner programmers' was very well established. To go beyond that you need a few tricks of the trade, mostly structured programming and avoiding global scope will get you to 10K and up. After 50K or so you will need even more tricks: version control will become a must, naming conventions matter and you're probably well into modularization territory. Some actual high level documentation would also not be a bad idea. By the time you hit 100's of thousands of lines you are probably looking at a life-time's worth of achievement by a single programmer, or more likely you are looking at teamwork. And so you'll need yet more infrastructure around your project to keep it afloat. You can see quite a bit of this at work when you browse through open source repositories at github, those are roughly the points at which plenty of projects and up being abandoned, as often as not through lack of insight in how to organize a larger project as through the demotivation that comes from lack of adoption.
- sasa_buklijas 9y agoagree, but it also depends on programing language. I do not see how 1K of C and Python lines are same.
- jacquesm 9y agoThat's true, and for languages even more dense the limits might be lower but it's 'order of magnitude' correct in all cases.
- fapjacks 9y ago... And then there's embedded programming, which flips this whole thing on its head, and a programmer's skill seems to be measured in how much smaller (more efficient) you can make your code while also keeping and improving clarity, security, maintainability, etc.
- jacquesm 9y agoThat's one reason why I love embedded stuff. I see all phone software as embedded stuff by the way, but the lessons do not seem to apply at all!
- theSage 9y agoAre there projects we can undertake to intentionally hit those walls and measure ourselves?
- kybernetikos 9y agoLike the equivalent of 'memory-hard' problems - 'lines-of-code' hard problems. I suppose that's close to what https://en.wikipedia.org/wiki/Kolmogorov_complexity https://en.wikipedia.org/wiki/Kolmogorov_complexity is.
- theSage 9y agoAh no. I meant what task can we undertake to see this in action? For example, writing your own compiler is bound to make you hit the 1k wall. If you're able to write it chances are you have passed the 1k wall some time in your past. That kind of measurement.
- kybernetikos 9y agoPart of the problem is that people who are good at dealing with large codebases tend to write fewer lines of code if they can at all get away with it.
- sifoo 9y agoI would say writing your own non-trivial interpreter is a pretty good metric from my own experience. That will take you into 10k land, and the problem is complex enough to lure beginners into creating a tangled mess. The reason I'm fairly sure, is that I just did; and I had to use most tricks I've learned to get this far: https://github.com/basic-gongfu/cixl https://github.com/basic-gongfu/cixl
- deleted 9y ago[deleted]
- Cthulhu_ 9y agoI think it's only something you really encounter in either your day to day work, or some OS projects but those have, I think, a very different day to day structure.
- g5095 9y agoshard your problem space.
- mrweasel 9y agoIndeed, micro services can be a pain in the butt to debug, due to communication, but they can help you to view the world as a collection of smaller easily understood programs, rather than one huge monolith.
- dublin 9y agoNot exactly a new idea: https://en.wikipedia.org/wiki/Unix_philosophy https://en.wikipedia.org/wiki/Unix_philosophy
- watmough 9y agoI hadn't come across this before, and sure enough, looking at the current personal project I'm working on, it's composed of 3 main files of C++, a Windows program, a parser [1] and a custom OpenGL control [2], each is 500 - 600 lines. For me, I use Stepwise Refinement [3] to get to this point, but to get further, I have to start breaking out a more abstract approach. I found this definition of Structured Programming that puts it very nicely [4]. There's also a perhaps self-imposed wall, where you might trial a solution and implement a prototype, then use that to realize that it no point in pressing on without some serious redesign of a key component. For my example above, the rendering part is old-style OpenGL, which works pretty nicely at 60 fps, but I'm holding off doing more until I can slot in and benchmark a better approach using vertex buffers and shaders, with the goals of enabling me to shape and scale the renderings in an abstracted coordinate system, and scale to rendering hundreds of files instead of just one. [1] https://twitter.com/watmough/status/962470455037841409 https://twitter.com/watmough/status/962470455037841409 [2] https://twitter.com/watmough/status/965007110391128064 https://twitter.com/watmough/status/965007110391128064 [3] http://www.informatik.uni-bremen.de/gdpa/def/def_s/STEPWISE_REFINEMENT.htm http://www.informatik.uni-bremen.de/gdpa/def/def_s/STEPWISE_... [4] https://www.encyclopedia.com/computing/dictionaries-thesauruses-pictures-and-press-releases/structured-programming https://www.encyclopedia.com/computing/dictionaries-thesauru...
- deleted 9y ago[deleted]
- sifoo 9y agoIt's also worth mentioning the value of staying below 20 kloc, and dividing into separate projects; especially if you don't have a team to back you up. Besides YAGNI, these are the most important heuristics I've picked up over the years: Don't get too attached to ideas, they may well prove sub optimal with more experience. Be ruthless in ripping out what doesn't pull it's weight any more, no matter how much effort already went into it. First, make it do the right thing; until then nothing else matters. Skip as many corners as you possibly can get away with, don't get stuck in the loop of checking boxes for the sake of checking boxes.
- gmoes 9y agoI feel that there is some naiveté in this perspective, although the OP does touch on it somewhat. A novice most likely would write their code in a very monolithic fashion. That same approach fails significantly with larger code bases. As a seasoned developer I have come to realize that one of the most important things to be a good developer is organizational skills. Unfortunately it seems that ways to organize code bases, including things like naming, mutable state, modularization, cohesion/coupling, etc., are not as well developed or understood in general in our industry as they should be. While understanding and knowing the right algorithms is important. I sometimes wonder if our emphasis on the knowing algorithms off the top of your head interviewing process contributes to putting the emphasis on the wrong things in software development.
- wellpast 9y agoAs a thought experiment imagine an org that hired for organizational/abstraction skills and landed a single worker good at algorithm design. (This is a distilled/simplistic example but bear to the point.) 10-20% of time in my career building business/productivity application systems have I needed to design & impl a complex algo. Okay so my fictional team abstracts away that need and my algo guy colors in the deets. (In most cases this is possible; in rarer cases the performance of the algo needs to cross-cut.) But the point is what you’re saying—one skill is pluggable, the other is not. Kind of making this less interesting (or more?) is that learning algorithm dev tools/skills is far, far easier than learning the architectural/organizational skills. So no wonder the current prevailing hiring emphasis.
- ksk 9y ago>Unfortunately it seems that ways to organize code bases, including things like naming, mutable state, modularization, cohesion/coupling, etc., are not as well developed or understood in general in our industry as they should be. True, but I think any attempt to standardize or to make it another engineering discipline would mean the end of high programmer salaries. Just follow guidelines and standard procedure in a book and you'll end up with a reasonable solution that is reasonably good and reasonably reliable at a reasonable cost. Most businesses would jump at that... And I'd argue that would be a good thing for all the sectors where programming is important, but no good programmer wants to actually join because its not sexy. (coal mining, medical equipment, etc) > I sometimes wonder if our emphasis on the knowing algorithms off the top of your head interviewing process contributes to putting the emphasis on the wrong things in software development. Yes but no business actually cares about creative solutions, unless algorithms are core to their business (and even then, other human factors outweigh finding the optimal, bestest, fastest solution). They simply want to use computers to solve a business problem. They want a runner to run from A to B, not an Olympic sprinter who is going to break a world record. Do you know a reasonably decent sort algorithm? Good, just use it. Profiling? Optimizing? That's for Olympic sprinters. Design patterns? Blindly apply a GoF design pattern that approximates your problem, etc etc.
- aaavl2821 9y agoI'm a novice just building my first 2k+ program and the thing that seems crazy to me is that based on what you learn in intro courses, thinks like namespaces and organization and modularization not only aren't emphasized, but seem like unimportant distractions from the core task of programming Based on all the blogs and HN comments and comments here, that couldn't be further from the truth. Are these just things you learn from experience / mentors?
- DougN7 9y agoYou learn them when you need them. They're just extra overhead on smaller projects so they're ignored... until the project grows and they can't be ignored any longer.
- phunehehe0 9y agoI suppose the courses were encouraging to start with something, anything, instead of trying to get everything perfect on first try. And then perhaps the "and then learn these next" part gets lost...
- user5994461 9y agoDepends what language you work is. Java for instance is very good at this and everything is always organized, whereas C has no notion of namespace or package.
- epicide 9y agoLearn to crawl before you walk. Learn to walk before you run. As far as where to learn these things: there are some books that talk about more broad topics than just how to use a particular tool/language. Those are more supplemental to experience/mentors, though.
- dahart 9y ago> thinks like namespaces and organization and modularization not only aren't emphasized, but seem like unimportant distractions from the core task of programming There's an analogy I've heard, and I can't remember where. It might be a Feynman quote I'm recalling... It'd be crazy to start design work on an airplane, if you're hungry and want to go to the corner store a block away for chips. It'd also be crazy to try to walk to China, since it's a long way, and also you might drown. The approach you take to solving the problem of getting from here to there depends completely on where here and there are, and what else you want to carry with you. I don't use namespaces and build modules for an Arduino project, or a simple command line tool. But for large projects, I literally can't live without namespacing and modularization. The core task of programming constitutes very different activities at different scales, as different as comparing walking to building jet engines.
- jorgec 9y agoI'm shocked to know how many businesses don't practice SRP. My current project (that started three weeks ago) has 250.000 lines of code but its clear and easy to find any bug, and I am not doing magic or something so special.
- Paul_S 9y agoA bit over dramatic. I remember Unreal 3 was close to 2 mil and that's the normal region for games (and has been for a decade) and trust me, gamedev studios are not staffed by seasoned programmers.
- epicide 9y ago> cannot debug or modify it without herculean effort. This could explain release dates getting pushed back :)
- avinium 9y agoIs that just scripting, or does that include the engine? If it's just game scripting, then I could imagine it's possible to squeeze 2 mil lines out of junior devs. I assume there are very few things that can go drastically "wrong".
- Paul_S 9y ago"Scripting" vs coding is not a clear cut case also a lot of studios (back in the day that is, since now the unreal model is different and everyone has access to the source) had source access so games would usually make changes to the engine. 2 mil is just the base engine. There's of course loads of middleware and plugins and integration for it all, probably doubling the size and of course the "scripting" which adds complexity like anything else and it relies on support for things in the main codebase. Customising the engine is not a matter of tweaking variables but extending the base classes. And for game logic you might be writing completely separate systems with their own architecture. Repos were big. Really big. Especially since we're talking about people who will check in FMVs into source control. Ah... what a horrible world.
- psyc 9y agoInteresting. My main project is getting near the 20k range. I haven't hit a wall, but it is getting noticeably harder to stay fluent in every part of the program. When I switch from one major system to another, there is a few days ramp-up. Still happy with readability and complexity.
- photojosh 9y agoSame. There are some parts of the code that I haven't touched in a few years, and when I do there's a significant effort to refamiliarise myself... and a whole lot of "what stupid idiot wrote this", oh yeah, that was me three years ago. :) I just finished the Python 2 -> 3 upgrade on it, that was fun.
- cortesoft 9y agoI avoid these walls entirely by never using newlines in my code.
- HumanDrivenDev 9y agoThat's webscale af
- agumonkey 9y agowhat about paulg onlisp and leveraging nested macros ? what about kay vpri efforts to reduce a full system to 100K (OMeta)