5 ms·
Any human line-by-line/application-by-application analysis is (for this particular discussion) out of the scope. The size of the thing and the way we thought w
by user1241320 12y ago
Any human line-by-line/application-by-application analysis is (for this particular discussion) out of the scope.
The size of the thing and the way we thought we were going to work is quite different.
For instance, suppose we produce AST for all the routines/pieces of logic/you_name_it we wanted to then find similar patterns or clusters that would give us hint to then work on a "pareto-like" way.
As already stated it's not ONE project, it's an old (but still running), poorly-documented codebase produce in decades around this big firm we work for.
- jerven 12y agoDon't try to figure out how the code does what does yet. Figure out what systems exists inside it: 1. What kind of modules? 2. Which servers/hardware? 3. Which databases/datastores? 4. What systems talk to what? 5. What test systems exist or existed? 6. Which api/frameworks where used? 7. Who is currently working on them/maintaining it? 8. Is anyone left who used to? 9. Why is a rewrite on the table? 10. Is there any way you can work on smaller pieces at a time? 11. What are the pain points of the current users (will tell you what area to focus on)? 12. Can you document what comes in and out? In my experience with such large code bases, there is never one way to do things. i.e. I once worked on a smaller system with 4 ways to talk to the same database. On one with 100 million lines I would expect even more ways to rome ;) If you do want to go down the static analysis path, start with existing tools before trying to build your own. If needed get external help for this. A 100 Million lines of code is not so bizarre. The project I work on is currently about 300,000 lines and a project some 300 times larger is quite imaginable for me.
- jacquesm 12y agoCode complexity does not increase linearly. 100 M lines is stupendous.
- jerven 12y agoNot really, my project is 3 FTE for 8 years. Double it to 16 years and then multiple the number of developers by 100. Consider a large enterprise having 300 developers in multiple teams, I am not at all surprised that they can manage to write 100 million lines or so. Also I think this is really a system of systems, and in my experience probably has large parts developed by the lowest bidding firms. Which when software is developed for 10 years or more means more than one way to do the same thing. Also having to deal with lots off ancient systems and working around weird bugs probably fixed years ago. You know things like bugs in Java 1.2 on HP-UX and stuff like that, or errors in Oracle 7i etc... Plus functional duplication because team A did not know subteam C2 build the same thing... Editing my comment instead of replying to the excellent comment by @jacquesm as hn does not allow me to reply to the reply. Actually completing the transfer of a codebase like that is unlikely to a new team without much much more of a handover. But some high/middle frustrated with the current system manager asking a team to start rebuilding before it gets shut down a few months/years later is very possible. An other plausible option is a corporate take over... But then I would expect a very experienced team to work on it who do not need to ask HN for this kind of thing. I personally have been in a situation where code moved between companies and no documentation or old developers where available. Not as large as this only a 1 or 2 million lines of code/xml. But I am no longer surprised by the stupid acts that large corporations can perform. And this must be systems not just one. You can't build a single jar file of 50+ millions of lines of code than could have loaded in a JVM around 1.3.1 on even high end hardware for the time.
- jacquesm 12y agoHow realistic does it sound to you that a codebase that size would be transferred to a new team without any of the old team, outside of a hack of a bank or a reversal of some outsourcing decision or something like that? Typically the value of such a codebase is determined by the quality of the team maintaining it and the degree to which it is documented. Complexity of software constructs is not linear and enough books have been written about simply multiplying and dividing manyears and lines of code that I don't think we need to hash that out all over again. See 'the mythical man-month' and many similar books and articles.
- smm2000 12y ago
- radicalbyte 12y agoWhen you go above the 1m-or-so LOC it gets much easier to move to more data-driven designs. Of course the OP could be including all test cases, data etc in his LOC; in which case you could easily reach the hundred-million LOC mark...