8 ms·
The case for continuous documentation
- Aeolun 5y agoSomeone is promoting swimm.io in a sort of sideways way? I sort of agree with the author, but the only solution presented is this unknown product.
- ycombiswimm 5y agoI couldn't find any other product that solves this basic problem, happy to hear about them if they exist.
- ericholscher 5y agohttps://docs.readthedocs.io/en/stable/ https://docs.readthedocs.io/en/stable/
- yosefk 5y agoIf you essentially have one live version of your program, like you do when it runs on your servers, coupled documentation is probably a good idea. If you have multiple active versions, like you often do with software released to run on someone else's machines, coupled documentation "works," but has a big downside. Namely, it prompts people to "refactor mercilessly"/change everything all the time, together with the documentation. When you need to maintain multiple versions at a time, having a single version of the document explaining the differences between all the live versions can somewhat curb the enthusiasm for gratuitous changes (since whoever does the changes must also maintain the increasingly long and ugly description in the single document describing all the versions.) And someone needing to work with all those versions has these differences nicely laid out and those areas not having differences also clearly visible. Whereas with multiple versions of the document you need to "diff" these versions if you want to build a mental model of what changed. Sadly (for those agreeing with this), I presume that the above is a minority opinion.
- cryptica 5y agoI think this argument makes perfect sense. Personally I don't like in-line documentation. It's mostly popular among people who depend on bulky proprietary IDEs. I hate all these approaches which try to trick people into using proprietary tech. I enjoy reading a nice documentation website maintained by the open source organization; it also gives me a touch-point with the organization which created the library. I also agree with your nuanced argument concerning incentives. Arguments related to incentives are almost always discarded by managers but they are very important. I do think decoupling the documentation encourages people to think about documentation more carefully as a distinct and important activity. I find that in-line documentation tends to be neglected; as a developer, when you're in the middle of coding an important feature which requires your full attention, you don't want to be distracted with updating the comments all the time because it breaks your train of thought. Usually developers tell themselves that they will do it later and they often forget. Comments are often neglected in the PR review process too. There is no way around it, you need to set aside some time to write or update the documentation as a distinct activity. There is a time for coding and there is a time for explaining.
- iib 5y agoCan you not just not write the documentation, even if it resides in the same repository, and then later make a commit to update it, as a separate task?
- cryptica 5y agoYes, but why take up space in the actual source code and repo? It increases space usage both in terms of disk space (which means more download time) for the library and requires more scrolling when reading the code. Also, if the library is a sub-dependency which the developer doesn't interact with directly, why should they download the documentation for it? They will never read those comments in the code anyway.
- michael1999 5y agoI see the value in internalizing the cost of breaking changes. But rather than just suffer a doc burden to discourage change, why not fix it with something like Stripe’s version conversions?
- simonw 5y agoI'm adamant that the documentation for a project should live in the same repository as the code itself. This is crucial for a number of reasons: 1. If the docs are in the same repo, a commit that changes the code can update the relevant documentation (in addition to the tests) as part of the same unit of work 2. This means it can be enforced during code review: if a developer forgets to update the docs they can be reminded before they land their PR 3. This also provides a version history for the documentation which is synchronized with the code history. This is really useful when looking at history and trying to figure out what changed when. 4. This also works great with branches, PRs and releases. New features can have their documentation developed alongside the code in a branch, which makes it easier to understand a proposed change. If your software is deployed in multiple places as multiple versions (or even just staging vs production) you have a way to view the correct documentation for each individual deployment. 5. Added together, all of this builds trust. A common problem I've seen with internal documentation is that no-one trusts it to be up-to-date. Making it part of the regular code development lifecycle can fix this. 6. If you do this, you can write automated tests that enforce aspects of your documentation! I call these documentation unit tests, and wrote about them here: https://simonwillison.net/2018/Jul/28/documentation-unit-tests/ https://simonwillison.net/2018/Jul/28/documentation-unit-tes... - even something as simple as a test that fails if a new API endpoint isn't mentioned in a markdown file using simple string matching can ensure no-one forgets about the docs when they add a new feature.
- Noumenon72 5y agoHow do you read the in-repo documentation? Search for all files named readme.md? I have never learned about a library from documentation scattered about the repo. There's the readme at the root, and everything else is on a web page, which is a better way to organize and browse documentation.
- simonw 5y agoI use documentation systems that publish the documentation from the repo to a website. Most of my projects use Sphinx and reStructuredText for this, but I recently tried MyST (Markdown for Sphinx) and I like that a lot. Some examples: - https://docs.datasette.io https://docs.datasette.io serves documentation from https://github.com/simonw/datasette/tree/main/docs https://github.com/simonw/datasette/tree/main/docs - which has documentation unit tests here: https://github.com/simonw/datasette/blob/main/tests/test_docs.py https://github.com/simonw/datasette/blob/main/tests/test_doc... - https://sqlite-utils.datasette.io/ https://sqlite-utils.datasette.io/ serves from https://github.com/simonw/sqlite-utils/tree/main/docs https://github.com/simonw/sqlite-utils/tree/main/docs - unit tests here: https://github.com/simonw/sqlite-utils/blob/main/tests/test_docs.py https://github.com/simonw/sqlite-utils/blob/main/tests/test_... - https://django-sql-dashboard.datasette.io/ https://django-sql-dashboard.datasette.io/ serves from markdown in https://github.com/simonw/django-sql-dashboard/tree/main/docs https://github.com/simonw/django-sql-dashboard/tree/main/doc... - I don't have documentation unit tests for that yet Those three are all hosted on https://www.readthedocs.org https://www.readthedocs.org but I've also used this trick on web app projects that host their own documentation deployed as part of the build process.
- CraigJPerry 5y agoIf it's technical docs for developers, you'll get more bang for your buck by making executable documentation first - tests, deployment automation, build automation. Make it so that to do 1 logical action then there's only 1 step needed. How do i build this? Run the build command. How do i test this? Run the test command. How do i run only the unit tests? Run the unit test command. How do i start this locally? Run the start-local command. How do these components interact? Run the contract-test command. ... That fixes "reference" type docs better than any reference doc but there's still a place for technical guides around a code base but short screen recordings voiced over by an experienced dev on the project navigating their IDE will beat any written guide on any metric (time to write, usefulness etc.) If it's customer facing docs, treat them as code and host them inside the application in some way. There's few things worse than reading the wrong version of a doc.
- simonw 5y agoGitHub have a pattern for this called "scripts to rule them all" - https://github.com/github/scripts-to-rule-them-all https://github.com/github/scripts-to-rule-them-all - I've not fully adopted it yet but I probably should, it looks very well thought-out.
- zie 5y agoWe just abuse Make, so make test, make bootstrap, etc.
- LadyCailin 5y agoI believe this to be one of the success stories of my programming language, MethodScript[0]. Early on I made the strange decision (in the sense that I’ve never seen it elsewhere) to make the documentation for each api element be part of the code itself. The documentation generator is part of the code as well, so every single build of the software is capable of generating bespoke documentation for that exact version. The website simply hosts the newest version, but you can always generate your own locally. Of course I also enforce that contributors must add/modify documentation at the same time as the code, but that’s easy, because if you modify most of the code, the documentation is also right next to it. [0] https://methodscript.com https://methodscript.com
- blowski 5y agoAn example of how to do this is the documentation for the PHP framework Symfony. Code examples from the documentation are run in the CI server. If a pull request breaks a code example, then that example must also be fixed as part of the pull request. That's a fantastic feature for a popular open-source framework, as it means the documentation remains up to date. I'm not sure at what point it's worth the effort for an internal project, though. If you have a cultural problem with incorrect, out-of-date or missing documentation this could make things worse. I'd look for the root cause of that first (training, motivation?), before trying to enforce it with technology.
- lucb1e 5y agoTesting your examples from documentation is actually a great idea. As a security tester, I can't even count the number of times I've gotten API documentation with omitted info like perhaps-trivial-to-them-but-blocker-to-me how to actually authenticate against the API. If they actually had it running somewhere, that sort of thing can't be missing. (Also, the number of times I've asked for API docs in an API-only test and they go "umm, let me task someone to write that real quick"... like, what did you think I was going to work with, balloons and thin air?) When appropriate, I'll definitely be recommending clients to include their documentation examples in testing. But they will probably ignore it like the rest of our non-high-risk advice (today's 'low' findings are tomorrow's stepping stones for ransomware).
- MauranKilom 5y agoI'm a fan of having code samples in the documentation, and making sure (at e.g. build/test time) that those samples actually work. Given the headline, I thought the article would talk about this, but it's more of a general "why and how you should keep your documentation up to date".
- ycombiswimm 5y agoThe product actually makes sure code samples stay up-to-date when the code changes.
- Larryreverse 5y agoI just saw their demo. Fresh approach.
- pram 5y agoDocumentation is great and all, but no one ever talks about when theres too much. Maybe because it's rare? At big companies I've seen "architects" churn out page after page of diagrams, design docs, runbooks, checklists, descriptions, etc. There is so much information that it becomes practically useless in aggregate, because no one is reasonably going to read it all. I'm not going to pretend that I have the patience or the attention span to have the gnostic mysteries of our Kubernetes infrastructure revealed to me.
- zelphirkalt 5y agoIn such cases I wish for a "cookbook" style of additional documentation, that I can search, to find examples for doing, what I want to do.
- qayxc 5y agoNothing a good search engine and proper requirements can fix. In theory it should work like this: I) "User should be able to do X" <- requirement II) "X can be achieved by performing steps A, B, and C" <- description of the implementation (high-level, user-perspective) III) "A works by using components 1 and 2" <- technical documentation (design-level, architecture perspective) etc. I) generates your index (what can I even do with the software) II) generates the documentation (how can I do it) III) and below is for technical use only (extending, modifying, porting) Stuff like rationales for design decisions can be structured in the same layered way. I don't know how something like this can be extracted after the fact, but no matter the development model (waterfall/agile), a structure like this should arise naturally anyway and the absolute amount of documentation isn't a problem. Lack of proper structure, however, is.
- 0xbadcafebee 5y agoThere is no such thing as too much documentation. There is out of date documentation, inaccessible documentation, unindexed documentation, poor documentation, redundant documentation, etc... But what you described is amazingly valuable. Just because it's not valuable to you, right now, doesn't mean there's too much of it.
- ycombiswimm 5y agoI completely agree that documentation should be part of the CI/CD and that it should be part of the code.
- Larryreverse 5y agoAre you the guy who wrote this? Would love to interview you for my podcast.
- systematical 5y agoThe only way for documentation to be part of CI (IMO) is for missing documentation to cause failed builds. There are a few ways I could think of to enforce this but is there anything off the shelf that does this?
- simonw 5y agoI've been doing this for nearly three years now - it works really well. It's not particularly sophisticated - just some tests which introspect the code and then use dumb pattern matching against the documentation text to check that different concepts from the code are mentioned at least once in the docs: https://simonwillison.net/2018/Jul/28/documentation-unit-tests/ https://simonwillison.net/2018/Jul/28/documentation-unit-tes...
- papito 5y agoIn the least - your repo should be the main gateway to a proper WIKI. The problem with decoupled documentation is that it's the proverbial tree in a forest - no one knows it's there when it "drops". Docs are like code - the less your write of it, the less you have to maintain. Documentation should be treated as inherently evil. The only worse thing than no documentation is documentation that is not maintained and out of date. There is nothing more infuriating than following the docs only to find out from someone later that it was antiquated. Why is it there? There are common sense rules. Why would you have the docs on how to set up a fresh checkout NOT live with the checkout? How would I know it lives somewhere else? How would anyone update those steps if those steps were not code?
- euroderf 5y agoYes, "Wrong documentation is worse than no documentation."
- cryptica 5y agoI don't agree with this at all. There are many projects such as Node.js which have excellent, up-to-date documentation on their websites (for all past versions too). This is good for the Node.js project because it forces developers visit the website which gives the open source project an opportunity to connect with their developers, to potentially monetize and stay independent. On the other hand, in-code documentation is hard to follow because it's scattered all over the source code, relies on special IDEs (more corporate lock-in) and developers often forget to update the documentation anyway (even more easily than they would forget to update the website). Not to mention that it takes up a LOT of space and requires more scrolling; IMO this has a negative impact on the readability of the code. Well written code is simple enough that it doesn't need much in-line documentation. I don't know why, but these days, when it comes to software development, I find that I disagree with 90% of all the top links that make it to the top of the HN front page. A lot of the practices which are being advocated are inefficient, bureaucratic and they seem to align with corporate interests as opposed to developer interests. The agenda seems to be about making developers more reliant on proprietary tools, IDEs, subscription SaaS services - All at the expense of free software principles. There is also an agenda around making developers more reliant on teams and less independent in the software development process. I remember coming across some outrageous claims such as "Good full stack developers don't exist". Also there is a push towards monorepos and other corporate structures which limit the degree of possible decentralization and autonomy of different projects and their dependencies. The shift towards static typing is also part of the trend towards centralization, de-modularization and high inter-dependency with proprietary tools and services. It's kind of ironic that tight coupling used to be considered one of the main signs of low-quality code but this concept is barely mentioned these days and the agenda is to promote it without saying outright what is going on.
- cryptica 5y agoEver wondered why the most popular package managers which have the most modules are all for dynamically typed languages? e.g. npm, Ruby Gems, pip... It's because dynamically typed languages are more modular since they have less rigid interfaces. With statically typed languages, there is a possibility that the type system of library Y might not correspond very elegantly with the type system of your own project X. Static typing require stronger coupling between the project and its libraries; it's typical that projects written by different teams will follow completely different typing conventions and names (for many different reasons); this adds friction. A very common one is when a library was written before some new Type/Interface was introduced as part of the core language and the library had invented its own abstraction which does the same thing... So the interface exposed by the library became redundant. Statically typed libraries require a lot more maintenance and this may also explain why companies are increasingly pushing for a monorepo structure which facilitate this constant maintenance which would have been unnecessary with a dynamically typed language.
- 0xbadcafebee 5y agoThere are still unsolved problems with documentation that I'd like to find solutions for. Everything can be made into code, but at a certain point it's just so complicated to do that you're spending more time and money automating your docs than your applications. So until we have solutions for all that, you will have to maintain some docs manually. For those manual docs, how do you keep them fresh? I've thought of automatically sending an email to warn that in 30 days the document would be deleted unless someone updated it, but even if people agreed to such a system, they could just update some punctuation and it would remain stale. Even blank pages, people seem to want to keep around rather than fix. How do you navigate your docs? Search engines actually suck for the most part. Search is a hard problem to solve, and a home rolled search will usually net terrible results. On the other hand, most people don't have the time to maintain a governance structure for their docs, much less an enforcement mechanism, so the docs invariably become terribly organized. People also seem to need training to learn how to write good docs. I know there are some trendy pages being passed around about some kind of "golden framework for docs" but they don't explain how to write them either. I know how to write docs, but I feel like I'd need to write a whole book to get it across. One thing I found really useful was Atlassian's newer Confluence page templates, which come well organized and primed with examples of how to write the docs. As a philosophy, I really, really love GitLab's Handbook First model. Their handbook is incredibly detailed and covers pretty much their whole organization, and is fairly easy to update. I feel like this one one of the magical missing links in getting more documentation for the important things that aren't code.
- deeblering4 5y agoOh cool, so books and websites in general are “bad” now.
- dllthomas 5y agoNot, by this metric, if they live in the same repo as the code. When they don't, they have the same problem as any strongly coupled systems maintained across multiple repos, or you are paying the cost of keeping the two uncoupled.
- navotgil 5y agoThis is why projects with micro service architecture are best saved in a mono repo
- dllthomas 5y agoI think the monorepo question is complicated, but this is generally a strong element in the pros column. One exception to that is when you want decoupling for other reasons, the pressures of multiple repos can help motivate it in the day-to-day.
- nwmcsween 5y agoLiterate programming is probably the best documentation
- navotgil 5y agoThat does not solve describing flows or patterns in the code
- remoquete 5y agoThe article seems to skim over the complexities of docs workflows and the role of docs. In fact, docs are nowhere to be found throughout the article. What are those “docs” that the OP is talking about? Perhaps it’d be fair to rename that article “The Case for Continuous READMEs”.
- ericholscher 5y agoThis is also the goal of Read the Docs. We even use the same wording in our docs: https://docs.readthedocs.io/en/stable/ https://docs.readthedocs.io/en/stable/ The Python ecosystem has been building docs this way for over 10 years now, and it works great.