16 ms·
Show HN: I built an AI that turns GitHub codebases into easy tutorials
https://the-pocket.github.io/Tutorial-Codebase-Knowledge/ https://the-pocket.github.io/Tutorial-Codebase-Knowledge/
- MoonieSzzS 1y ago[dead]
- badmonster 1y agodo you have plans to expand this to include more advanced topics like architecture-level reasoning, refactoring patterns, or onboarding workflows for large-scale repositories?
- zh2408 1y agoYes! This is an initial prototype. Good to see the interest, and I'm considering digging deeper by creating more tailored tutorials for different types of projects. E.g., if we know it's web dev, we could generate tutorials based more on request flows, API endpoints, database interactions, etc. If we know it's a more long-term maintained projects, we can focus on identifying refactoring patterns.
- kristopolous 1y agoHave you ever seen komment.ai? Is so did you have any issues with the limitation of the product? I haven't used it, but it looks like it's in the same space and I've been curious about it for a while. I've tried my own homebrew solutions, creating embedding databases by having something like aider or simonw's llm make an ingests json from every function, then using it as a rag in qdrant to do an architecture document, then using that to do contextual inline function commenting and make a doxygen then using all of that once again as an mcp with playwright to hook that up through roo. It's a weird pipeline and it's been ok, not great but ok. I'm looking into perplexica as part of the chain, mostly as a negation tool
- zh2408 1y agoNo, I haven't, but I will check it out! One thing to note is that the tutorial generation depends largely on Gemini 2.5 Pro. Its code understanding ability is very good, combined with its large 1M context window for a holistic understanding of the code. This leads to very satisfactory tutorial results. However, Gemini 2.5 Pro was released just late last month. Since Komment.ai launched earlier this year, I don't think models at that time could generate results of that quality.
- kristopolous 1y agoI've been using llama 4 Maverick through openrouter. Gemini was my go to but I switched basically the day it came out to try it out. I haven't switched back. At least for my use cases it's been meeting my expectations. I haven't tried Microsoft's new 1.58 bit model but it may be a great swap out for sentencellm, the legendary all-MiniLM-L6-v2. I found that if I'm unfamiliar with the knowledge domain I'm mostly using AI but then as I dive in the ratio of AI to human changes to the point where it's AI at 0 and it's all human. Basically AI wins at day 1 but isn't any better at day 50. If this can change then it's the next step
- zh2408 1y agoYeah, I'd recommend trying Gemini 2.5 Pro. I know early Gemini weren't great, but the recent one is really impressive in terms of coding ability. This project is kind of designed around the recent breakthrough.
- kristopolous 1y agoI've used it, I used to be a huge booster! Give llama 4 maverick a try, really.
- ryao 1y agoI would find this more interesting if it made tutorials out if the Linux, LLVM, OpenZFS and FreeBSD codebases.
- wordofx 1y agoI would find this comment more interesting if it didn’t dismiss the project just because you didn’t find it valuable.
- zh2408 1y agoThe Linux repository has ~50M tokens, which goes beyond the 1M token limit for Gemini 2.5 Pro. I think there are two paths forward: (1) decompose the repository into smaller parts (e.g., kernel, shell, file system, etc.), or (2) wait for larger-context models with a 50M+ input limit.
- rtolsma 1y agoYou can use the AST for some languages to identify modular components that are smaller and can fit into the 1M window
- achierius 1y agoSome huge percentage of that is just drivers. The kernel is likely what would be of interest to someone in this regard; moreover, much of that is architecture specific. IIRC the x86 kernel is <1M lines, though probably not <1M tokens.
- 1y ago
- Retr0id 1y agoThe overview diagrams it creates are pretty interesting, but the tone/style of the AI-generated text is insufferable to me - e.g. https://the-pocket.github.io/Tutorial-Codebase-Knowledge/Requests/01_functional_api.html#whats-the-functional-api https://the-pocket.github.io/Tutorial-Codebase-Knowledge/Req...
- zh2408 1y agoHaha. The project is fully open-sourced, so you can tune the prompt for the tone/style you prefer: https://github.com/The-Pocket/Tutorial-Codebase-Knowledge/blob/08b2cade4fff9bb210b8e77e2a88cf5a30332f56/nodes.py#L547 https://github.com/The-Pocket/Tutorial-Codebase-Knowledge/bl...
- deleted 1y ago[deleted]
- vivzkestrel 1y agomind explaining what exactly was insufferable here?
- Retr0id 1y agoIf you don't feel the same way from reading it, I'm not sure it can be explained.
- stevedonovan 1y agoI agree, it's hopelessly over-cheerful and tries to be cute. The pizza metaphor fell flat for me as well
- fn-mote 1y agoI guess you already know what a “Functional API” is and feel patronized. Also possibly you dislike the “cute analogy” factor. I think this could be solved with an “assume the reader knows …” part of the prompt. Definitely looks like ELI5 writing there, but many technical documents assume too much knowledge (especially implicit knowledge of the context) so even though I’m not a fan of this section either, I’m not so quick to dismiss it as having no value.
- chairhairair 1y agoA company (mutable ai) was acquired by Google last year for essentially doing this but outputting a wiki instead of a tutorial.
- zh2408 1y agoTheir site seems to be down. I can't find their results.
- codetrotter 1y agoWere they acquired? Or did they give up and the CEO found work at Google? https://news.ycombinator.com/item?id=42542512 https://news.ycombinator.com/item?id=42542512 The latter is what this thread claims ^
- nxobject 1y agoIt sounds like it'd be perfect for Google's NotebookLM portfolio -- at least if they wanted to scale it up.
- chairhairair 1y agoI don’t know the details of the deal, but their YC profile indicates they were acquired.
- cowsandmilk 1y agoyou're going to trust the person who started the thread with no idea what happened to the company and then jumped to conclusions based on LinkedIn?
- deleted 1y ago[deleted]
- kaycebasques 1y agoI meant to write a blog post about mutable.ai but didn't get around to it before the product shut down. I did however archive the wiki that it generated for the project I work on: https://web.archive.org/web/20240815184418/wiki.mutable.ai/google/pigweed https://web.archive.org/web/20240815184418/wiki.mutable.ai/g... (The images aren't working. I believe those were auto-generated class inheritance or dependency diagrams.) * The first paragraph is pretty good. * The second paragraph is incorrect to call pw_rpc the "core" of Pigweed. That implies that you must always use pw_rpc and that all other modules depend on it, which is not true. * The subsequent descriptions of modules all seemed decent, IIRC. * The big issue is that the wiki is just a grab bag summary of different parts of the codebase. It doesn't feel coherent. And it doesn't mention the other 100+ modules that the Pigweed codebase contains. When working on a big codebase, I imagine that tools like mutable.ai and Pocket Flow will need specific instruction on what aspects of the codebase to document.
- manofmanysmiles 1y agoI love it! I effectively achieve similar results by asking Cursor lots of questions! Like at least one other person in the comments mentioned, I would like a slightly different tone. Perhaps good feature would be a "style template", that can be chosen to match your preferred writing style. I may submit a PR though not if it takes a lot of time.
- zh2408 1y agoThanks—would really appreciate your PR!
- TheTaytay 1y agoWoah, this is really neat. My first step for many new libraries is to clone the repo, launch Claude code, and ask it to write good documentation for me. This would save a lot of steps for me!
- randomcatuser 1y agoExactly what I did today! (for Codex!) The output here is actually slightly better! I bet in the next few months we'll be getting dynamic, personalized documentation for every library!! Good times
- CalChris 1y agoDo one for LLVM and I'll definitely look at it.
- throwaway314155 1y agoI suppose I'm just a little bit bothered by your saying you "built an AI" when all the heavy lifting is done by a pretrained LLM. Saying you made an AI-based program or hell, even saying you made an AI agent, would be more genuine than saying you "built an AI" which is such an all-encompassing thing that I don't even know what it means. At the very least it should imply use of some sort of training via gradient descent though.
- j45 1y agoIt is an application of AI which is just software, and applying it to solve a problem or need.
- dahuangf 1y ago[dead]
- lionturtle 1y ago>:( :3
- esjeon 1y agoAt the top are some neat high-level stuffs, but, below that, it quickly turns into code-written-in-human-language. I think it should be possible to extract some more useful usage patterns by poking into related unit tests. How to use should be what matters to most tutorial readers.
- chbkall 1y agoLove this. These are the kind of AI applications we need which aid our learning and discovery.
- android521 1y agoFor anyone doubting AI as pure hype, this is the counter example of its usefulness
- croes 1y agoNobody said AI isn’t useful. The hype is that AI isn’t a tool but the developer.
- hackernewds 1y agoI've seen a lot of developers that are absolute tools. But I've yet to see such a succinct use of AI. Kudos to the author.
- croes 1y agoExactly, kudos to the author because AI didn’t came up with that. But that’s what they sell, that AI could do what the author did with AI. The question is, is it worth to put all that money and energy in AI. MS sacrificed its CO2 goals for email summaries and better autocomplete not to mention all the useless things we do with AI
- relativ575 1y ago> But that’s what they sell, that AI could do what the author did with AI. Can you give an example of what you meant here? The author did use AI. What does "AI coming up with that" mean?
- murkt 1y agoGP commenter complains that it’s not AI that came up with an idea and implemented it, but a human did. In the few years we will see complaints that it’s not AI that built a power station and a datacenter, so it doesn’t count as well.
- 1y ago
- zarkenfrood 1y agoReally nice work and thank you for sharing. These are great demonstrations of the value of LLMs which help to go against the negative view on the impacts to junior engineers. This helps bridge the gap of most projects lacking updated documentation.
- bilalq 1y agoThis is actually really cool. I just tried it out using an AI studio API key and was pretty impressed. One issue I noticed was that the output was a little too much "for dummies". Spending paragraphs to explain what an API is through restaurant analogies is a little unnecessary. And then followed up with more paragraphs on what GraphQL is. Every chapter seems to suffer from this. The generated documentation seems more suited for a slightly technical PM moreso than a software engineer. This can probably be mitigated by refining the prompt. The prompt would also maybe be better if it encouraged variety in diagrams. For somethings, a flow chart would fit better than a sequence diagram (e.g., a durable state machine workflow written using AWS Step Functions).
- hackernewds 1y agoexactly it is. I'd rather impressive but at the same time the audience is always going to be engineers, so perhaps it can be curated to still be technical to a degree? I can't imagine a scenario where I have to explain to the VP my ETL pipeline
- trcf21 1y agoFrom flow.py Ensure the tone is welcoming and easy for a newcomer to understand{tone_note}. - Output only the Markdown content for this chapter. Now, directly provide a super beginner-friendly Markdown output (DON'T need ```markdown``` tags) So just a change here might do the trick if you’re interested. But I wonder how Gemini would manage different levels. From my take (mostly edtech and not in English) it’s really hard to tone the answer properly and not just have a black and white (5 year old vs expert talk) answer. Anyone has advice on that?
- porridgeraisin 1y agoThis has given me decent success: "Write simple, rigorous statements, starting from first principles, and making sure to take things to their logical conclusion. Write in straightforward prose, no bullet points and summaries. Avoid truisms and overly high-level statements. (Optionally) Assume that the reader {now put your original prompt whatever you had e.g 5 yo}" Sometimes I write a few more lines with the same meaning as above, and sometimes less, they all work more or less OK. Randomly I get better results sometimes with small tweaks but nothing to make a pattern out of -- a useless endeavour anyway since these models change in minute ways every release, and in neural nets the blast radius of a small change is huge.
- andrewrn 1y agoThis is brilliant. I would make great use of this.
- wg0 1y agoThat's a game changer for a new Open source contributor's onboarding. Put in postgres or redis codebase, get a good understanding and get going to contribute.
- tgv 1y agoIsn't that overly optimistic? The postgres source code is really complex, and reading a dummy tutorial isn't going to make you a database engine ninja. If a simple tutorial can, imagine what a book on the topic could do.
- wg0 1y agoNo, I am not that optimistic about LLMs. I just think that something is better then nothing. The burden of understanding still is with the engineers. All you would get is some (partially inaccurate at places) good overview of where to look for.
- gregpr07 1y agoI built browser use. Dayum, the results for our lib are really impressive, you didn’t touch outputs at all? One problem we have is maintaining the docs with current codebase (code examples break sometimes). Wonder if I could use parts of Pocket to help with that.
- cehrlich 1y agoAs a maintainer of a different library, I think there’s something here. A revised version of this tool that also gets fed the docs and asked to find inaccuracies could be great. Even if false positives and false negatives are let’s say 20% each, it would still be better than before as final decisions are made by a human.
- zh2408 1y agoThank you! And correct, I didn't modify the outputs. For small changes, you can just feed the commit history and ask an LLM to modify the docs. If there are lots of architecture-level changes, it would be easier to just feed the old docs and rewrite - it usually takes <10 minutes.
- mraza007 1y agoImpressive work. With the rise of AI understanding software will become relatively easy
- stephantul 1y agoThe dspy tutorial is amazing. I think dspy is super difficult to understand conceptually, but the tutorial explained it really well
- mvATM99 1y agoThis is really cool and very practical. definitely will try it out for some projects soon. Can see some finetuning after generation being required, but assuming you know your own codebase that's not an issue anyway.
- ganessh 1y agoDoes it use the docs in the repository or only the code?
- zh2408 1y agoBy default we use both based on regex: DEFAULT_INCLUDE_PATTERNS = { ".py", ".js", ".jsx", ".ts", ".tsx", ".go", ".java", ".pyi", ".pyx", ".c", ".cc", ".cpp", ".h", ".md", ".rst", "Dockerfile", "Makefile", ".yaml", ".yml", } DEFAULT_EXCLUDE_PATTERNS = { "test", "tests/", "docs/", "examples/", "v1/", "dist/", "build/", "experimental/", "deprecated/", "legacy/", ".git/", ".github/", ".next/", ".vscode/", "obj/", "bin/", "node_modules/", ".log" }
- m0rde 1y agoHave you tried giving it tests? Curious if you found they made things worse.
- Tokumei-no-hito 1y agowhy exclude tests and docs by default?
- pknerd 1y agoInteresting..would you like to share some technical details? it did not seem you have used RAG here?
- zh2408 1y agoYeah, RAG is not the best option here. Check out the design doc: https://github.com/The-Pocket/Tutorial-Codebase-Knowledge/blob/main/docs/design.md https://github.com/The-Pocket/Tutorial-Codebase-Knowledge/bl... I also have a YouTube Dev Tutorial. The link is on the repo.
- potamic 1y agoDid you measure how much it cost to run it against your examples? Trying to gauge how much it would cost to run this against my repos.
- pitched 1y agoLooks like there are 4 prompts and the last one can run up to 10 times for the chapter content. You might get two or three tutorials built for yourself inside the free 25/day limit, depending on how many chapters it needs.
- throwaway290 1y agoYou didn't "build an AI". It's more like you wrote a prompt. I wonder why all examples are from projects with great docs already so it doesn't even need to read the actual code.
- afro88 1y ago> You didn't "build an AI". True > It's more like you wrote a prompt. False > I wonder why all examples are from projects with great docs already so it doesn't even need to read the actual code. False. This: https://github.com/browser-use/browser-use/tree/main/browser_use https://github.com/browser-use/browser-use/tree/main/browser... Became this: https://the-pocket.github.io/Tutorial-Codebase-Knowledge/Browser%20Use/ https://the-pocket.github.io/Tutorial-Codebase-Knowledge/Bro...
- quantumHazer 1y agoThe example you made has, in fact, a documentation https://docs.browser-use.com/introduction https://docs.browser-use.com/introduction
- afro88 1y agoYou don't point this tool at the documentation though. You point it at a repo. Granted, this example (and others) have plenty of inline documentation. And, public documentation is likely in the training data for LLMs. But, this is more than just a prompt. The tool generates really nicely structured and readable tutorials that let you understand codebases at a conceptual level easier than reading docstrings and code. Even if it's only useful for public repos with documentation, that's still useful, and flippant dismissals are counterproductive. I am keen to try this with one of my own (private, badly documented) codebases and see how it fares. I've actually found LLMs quite useful at explaining code, so I have high hopes.
- quantumHazer 1y agoI’m not saying that the tool is useless, I was confuting your argument about being a project WITHOUT docs. LLM can write passable docs, but obviously can write better docs of project well documented in training data. And this example is probably in training data as of April 2025
- saberience 1y agoI hate this language: "built an AI", did you train a new model to do this? Or are you in fact calling ChatGPT 4o, or Sonnet 3.7 with some specific prompts? If you trained a model from scratch to do this I would say you "built an AI", but if you're just calling existing models in a loop then you didn't build an AI. You just wrote some prompts and loops and did some RAG. Which isn't building an AI and isn't particularly novel.
- eapriv 1y ago> “I built an AI” > look inside > it’s a ChatGPT wrapper
- anshulbhide 1y agoLove this kind of stuff on HN
- firesteelrain 1y agoCan this work with Codeium Enterprise?
- amelius 1y agoI've said this a few times on HN: why don't we use LLMs to generate documentation? But then came the naysayers ...
- tdy_err 1y agoAlternatively, document your code.
- cushychicken 1y agoId have bought a lot of lunches for myself if I had a dollar for every time I’ve pushed my team for documentation and had it turn into a discussion of “Well, how does that stack up against other priorities?” It’s a pretty foolproof way for smart political operators to get out of a relatively dreary - but high leverage - task. AI doesn’t complain. It just writes it. Makes the whole task a lot faster when a human is a reviewer for correctness instead of an author and reviewer.
- runeks 1y agoUseful documentation explains why the code does what it does. Ie. why is this code there? An LLM can't magically figure out your motivation behind doing something a certain way.
- johnnyyyy 1y agoAre you implying that only the creator of the code can write documentation?
- oblio 1y agoNo, they're saying that LLMs (and really, most other humans) can't really write the best documentation. Frankly, most documentation is useless fluff. LLMs will be able to write a ton of that for sure :-)
- bonzini 1y agoTell it the why of the API and ask it to write individual function and class docs then.
- chyueli 1y agoGreat, I'll try it next time, thanks for sharing
- trash_cat 1y agoThis is literally what I use AI for. Excellent project.
- souhail_dev 1y agothat's amazing, I was looking for that a while ago Thanks
- istjohn 1y agoThis is neat, but I did find an error in the output pretty quickly. (Disregard the mangled indentation) # Use the Session as a context manager with requests.Session() as s: s.get('https://httpbin.org/cookies/set/contextcookie/abc') response = s.get(url) # ??? print("Cookies sent within 'with' block:", response.json()) https://the-pocket.github.io/Tutorial-Codebase-Knowledge/Requests/03_session.html https://the-pocket.github.io/Tutorial-Codebase-Knowledge/Req...
- zh2408 1y agoThis code creates an HTTP session, sets a cookie within that session, makes another request that automatically includes the cookie, and then prints the response showing the cookies that were sent. I may miss the error, but could you elaborate where it is?
- deleted 1y ago[deleted]
- Armazon 1y agoThe url variable is never defined
- foobarbecue 1y agoYes it is, just before the pasted section.
- Armazon 1y agooh you are right, nevermind then
- totally 1y agoIf only the AI could explain the errors that the AI outputs.
- 1y ago
- mooreds 1y agoI had not used gemini before, so spent a fair bit of time yak shaving to get access to the right APIs and set up my Google project. (I have an OpenAPI key but it wasn't clear how to use that service.) I changed it to use this line: api_key=os.getenv("GEMINI_API_KEY", "your-api_key") instead of the default project/location option. and I changed it to use a different model: model = os.getenv("GEMINI_MODEL", "gemini-2.5-pro-preview-03-25") I used the preview model because I got rate limited and the error message suggested it. I used this on a few projects from my employer: - https://github.com/prime-framework/prime-mvc https://github.com/prime-framework/prime-mvc a largish open source MVC java framework my company uses. I'm not overly familiar with this, though I've read a lot of code written in this framework. - https://github.com/FusionAuth/fusionauth-quickstart-ruby-on-rails-web/ https://github.com/FusionAuth/fusionauth-quickstart-ruby-on-... a smaller example application I reviewed and am quite familiar with. - https://github.com/fusionauth/fusionauth-jwt https://github.com/fusionauth/fusionauth-jwt a JWT java library that I've used but not contributed to. Overall thoughts: Lots of exclamation points. Thorough overview, including of some things that were not application specific (rails routing). Great analogies. Seems to lean on them pretty heavily. Didn't see any inaccuracies in the tutorials I reviewed. Pretty amazing overall!
- mooreds 1y agoIf you want to see what output looks like (for smaller projects--the OP shared some for other, more popular projects), I posted a few of the tutorials to my GitHub: https://github.com/mooreds/prime-mvc-tutorial https://github.com/mooreds/prime-mvc-tutorial https://github.com/mooreds/railsquickstart-tutorial https://github.com/mooreds/railsquickstart-tutorial https://github.com/mooreds/fusionauth-jwt-tutorial/ https://github.com/mooreds/fusionauth-jwt-tutorial/ Other than renaming the index.md file to README.md and modifying it slightly, I made no changes. Edit: added note that there are examples in the original link.
- mooreds 1y agoUpdate, billing was delayed, but for 4 tutorials it cost about $5.
- thom 1y agoThis is definitely a cromulent idea, although I’ve realised lately that ChatGPT with search turned on is a great balance of tailoring to my exact use case and avoiding hallucinations.
- dangoodmanUT 1y agoit appears like it's leveraging the docs and learned tokens more than the actual code. For example I don't believe it could achieve that understanding of levelDB without the prior knowledge and extensive material it's probably learned on already
- touristtam 1y agoJust need to find one way to integrate into the deployment pipeline and output some markdown (or other format) to send them to what ever your company is using (or simply a live website), I'd say.
- fforflo 1y agoIf you want to use Ollama to run local models, here’s a simple example: from ollama import chat, ChatResponse def call_llm(prompt, use_cache: bool = True, model="phi4") -> str: response: ChatResponse = chat( model=model, messages=[{ 'role': 'user', 'content': prompt, }] ) return response.message.content
- theptip 1y agoYes! AI for docs is one of the usecases I’m bullish on. There is a nice feedback loop where these docs will help LLMs to understand your code too. You can write a GH action to check if your code change / release changes the docs, so they stay fresh. And run your tutorials to ensure that they remain correct.
- mooreds 1y ago> And run your tutorials to ensure that they remain correct. Do you have examples of LLMs running tutorials you can share?
- orsenthil 1y agoIt will be good to integrate a local web server to fire up and read the doc. I use vscode, markdown preview. And it works too. Cool project.
- las_nish 1y agoNice project. I need to try this
- fforflo 1y agoWith $GEMINI_MODE=gemini-2.0-flash I also got some decent results for libraries like simonw/llm and pgcli. You can tell that because simonw writes quite heavily-documented code an the logic is pretty straightforward, it helps the model a lot! https://github.com/Florents-Tselai/Tutorial-Codebase-Knowledge/tree/more-examples/docs/llm https://github.com/Florents-Tselai/Tutorial-Codebase-Knowled... https://github.com/Florents-Tselai/Tutorial-Codebase-Knowledge/tree/more-examples/docs/pgcli https://github.com/Florents-Tselai/Tutorial-Codebase-Knowled...
- 3abiton 1y agoHow does it perform for undocumented repos?
- polishdude20 1y agoIs there an easy way to have this visit a private repository? I've got a new codebase to learn and it's behind credentials.
- bionhoward 1y ago“I built an AI” Looks inside REST API calls
- lasarkolja 1y agoCan anyone turn nextcloud/server into an easy tutorial
- andybak 1y agoIs there a way to limit the number of exclamation marks in the output? It seems a trifle... overexcited at times.
- mattfrommars 1y agoWTF You built in in one afternoon? I need to figure out these mythical abilities. I've thought about this idea few weeks back but could not figure out how to implement it. Amazing job OP
- kaycebasques 1y agoVery cool, thanks for sharing. I imagine that this will make a lot of my fellow technical writers (even more) nervous about the future of our industry. I think the reality is more along the lines of: * Previously, it was simply infeasible for most codebases to get a decent tutorial for one reason or another. E.g. the codebase is someone's side project and they don't have the time or energy to maintain docs, let alone a tutorial, which is widely regarded as one of the most labor-intensive types of docs. * It's always been hard to persuade businesses to hire more technical writers because it's perenially hard to connect our work to the bottom or top line. * We may actually see more demand for technical writers because it's now more feasible (and expected) for software projects of all types to have decent docs. The key future skill would be knowing how to orchestrate ML tools to produce (and update) docs. (But I'm also under no delusion: it definitely possible for TWs to go the way of the dodo bird and animatronics professionals.) I think I have a very good way to evaluate this "turn GitHub codebases into easy tutorials" tool but it'll take me a few days to write up. I'll post my first impressions to https://technicalwriting.dev https://technicalwriting.dev P.S. there has been a flurry of recent YC startups focused on automating docs. I think it's a tough space. The market is very fragmented. Because docs are such a widespread and common need I imagine that a lot of the best practices will get commoditized and open sourced (exactly like Pocket Flow is doing here)
- kaycebasques 1y agoHere's my write-up: https://technicalwriting.dev/ml/pocketflow/index.html https://technicalwriting.dev/ml/pocketflow/index.html
- teknico 1y agoThank you, very detailed and useful.
- bdg001 1y agoI was using gitdiagram but llms are very bad at generating good error free mermaid code! Thanks buddy! this will be very helpful !!
- gbraad 1y agoInteresting, but gawd awful analogy: "like a takeout order app". It tries to be amicable, which feels uncanny.
- citizenpaul 1y agoThis is really cool. One of the best AI things I've seen in the last two years.
- Too 1y agoHow well does this work on unknown code bases? The tutorial on requests looks uncanny for being generated with no prior context. The use cases and examples it gives are too specific. It is making up terminology, for concepts that are not mentioned once in the repository, like "functional api" and "hooks checkpoints". There must be thousands of tutorials on requests online that every AI was already trained on. How do we know that it is not using them?
- remoquete 1y agoThis is nice and fun for getting some fast indications on an unknown codebase, but, as others said here and elsewhere, it doesn't replace human-made documentation. https://passo.uno/whats-wrong-ai-generated-docs/ https://passo.uno/whats-wrong-ai-generated-docs/
- kaycebasques 1y agoMy bet is that the combination of humans and language models is stronger than humans alone or models alone. In other words there's a virtuous cycle developing where the codebases that embrace machine documentation tools end up getting higher quality docs in the long run. For example, last week I tried out a codebase summary tool. It had some inaccuracies and I knew exactly where it was pulling the incorrect data from. I fixed that data, re-ran the summarization tool, and was satisfied to see a more accurate summary. But yes, it's probably key to keep human technical writers (like myself!) in the loop.
- remoquete 1y agoIndeed. Augmentation is the way forward.
- lastdong 1y agoGreat stuff, I may try it with a local model. I think the core logic for the final output is all in the nodes.py file, so I guess one can try and tweak the prompts, or create a template system.
- imposter 1y ago[dead]
- swashbuck1r 1y agoWhile the doc generator is a useful example app, the really interesting part is how you used Cursor to start a PocketFlow design doc for you, then you fine-tuned the details of the design doc to describe the PocketFlow execution graph and utilities you wanted the design of the doc-generator to follow…and then you used used Cursor to generate all the code for the doc-generator application. This really shows off that the simple node graph, shared storage and utilities patterns you have defined in your PocketFlow framework are useful for helping the AI translate your documented design into (mostly) working code. Impressive project! See design doc https://github.com/The-Pocket/Tutorial-Codebase-Knowledge/blob/main/docs/design.md https://github.com/The-Pocket/Tutorial-Codebase-Knowledge/bl... And video https://m.youtube.com/watch?v=AFY67zOpbSo https://m.youtube.com/watch?v=AFY67zOpbSo
- 1899-12-30 1y agoAs an extension to this general idea: AI generated interactive tutorials for software usage might be a good product. Assuming it was trained on the defined usage paths present in the code, it would be able to guide the user through those usages.
- lummm 1y agoI actually have created something very similar here: https://github.com/Black-Tusk-Data/crushmycode https://github.com/Black-Tusk-Data/crushmycode, although with a greater focus on 'pulling apart' the codebase for onboarding. So many potential applications of the resultant knowledge graph.
- rtcoms 1y agoI would be very interested in knowing how did you build this ?
- deleted 1y ago[deleted]
- nitinram 1y agoThis is super cool! I attempted to use this on a project and kept running into "This model's maximum context length is 200000 tokens. However, your messages resulted in 459974 tokens. Please reduce the length of the messages." I used open ai o4-mini. Is there an easy way to handle this gracefully? Basically if you had thoughts on how to make some tutorials for really large codebases or project directories?
- zh2408 1y agoCould you try to use gemini 2.5 pro? It's free every day for first 25 requests, and can handle 1M input tokens
- iamsaitam 1y agoBLOATED. This project is 100 lines of code, but everything that is non-code related is bloated like a gas giant. All the text and videos are written by an LLM. The author would learn from understanding that QUANTITY isn't QUALITY, toning down the verbiage would benefit greatly what they are trying to communicate. PS: The generated "design documents" are 2k+ lines long. This seems like a great way to exceed quotas.
- axelr340 1y agoWe are also building a tool to understand codebases. Our tool shows the features implemented in a codebase visually, along with their hierarchy, and with traceability to associated code. Here is an example feature map for the Spot robot SDK from Boston Dynamics with 100k lines of code: https://product-map.ai/app/public?url=https://github.com/boston-dynamics/spot-sdk https://product-map.ai/app/public?url=https://github.com/bos...