11 ms·
Show HN: DocsGPT, open-source documentation assistant, fully aware of libraries
Hi, This is a very early preview of a new project, I think it could be very useful. Would love to hear some feedback/comments
- verdverm 4y agoVery cool, I'm glad that you provided some instructions on how to use this on a different set of docs. Would this support Markdown format? Context, I'm using Hugo with Markdown + injected code snippets. Might need to actually crawl the site... Do you think including Slack Q&As in training would help too?
- sadrobin 4y agoYeah, its still very early preview. we are working on making sure parsing works well on different formats. But for now you can make sure it looks well in txt. Would love to see you in our discord. Very soon it will support different formats, so make sure you are updated.
- o_____________o 4y agoWhat are you doing on top of langchain? More details would be nice.
- eternalban 4y agoHow to create the vector store: this gets passed along (i assume) in the longchain context: https://github.com/arc53/docsgpt/blob/main/scripts/ingest_rst.py https://github.com/arc53/docsgpt/blob/main/scripts/ingest_rs... This is the prompt that is used: https://github.com/arc53/docsgpt/blob/main/application/combine_prompt.txt https://github.com/arc53/docsgpt/blob/main/application/combi... And this is where it calls the openai. Looks pretty straightforward. https://github.com/arc53/docsgpt/blob/main/application/app.py https://github.com/arc53/docsgpt/blob/main/application/app.p...
- atonse 4y agoWhere can one go to learn how to build something like this with chatgpt? Are people asking a question under the hood? Like “explain these 30 code files to me”
- swalsh 4y agoYou could probably ask ChatGPT how to get started with GPT, but I think a lot of people are building apps taking advantage of GPT's Fine Tuning ability. https://platform.openai.com/docs/guides/fine-tuning https://platform.openai.com/docs/guides/fine-tuning
- petilon 4y agoIf I take advantage of this for allowing my customers to ask questions about our product documentation, how do I limit questions to questions about my product documentation? I don't want to be paying Open OPI for questions unrelated to my product.
- th3h4mm3r 4y agoThis is really a good question. Could be interesting also which is the best method to provide to GPT a Product documentato and ask question inherents to that and nothing else.
- shagie 4y agoOne idea would be to use a much cheaper (and faster) classifier to come back with a "yes" or "no" if the question asked is about your product documentation. Using Ada or Babbage is about 1% of the cost of Davinci (and Curie is 10% the cost of Davinci). Without any real tuning, this responds quite promptly (and the various tests I've done, correctly): curl https://api.openai.com/v1/completions \ -H "Content-Type: application/json" \ -H "Authorization: Bearer $OPENAI_API_KEY" \ -d '{ "model": "text-ada-001", "prompt": "Identify if the following question is about the Olympic Games. Answer with {yes}, {no}, {maybe}.\n\nWhat category is pole vaulting in?", "temperature": 0.7, "max_tokens": 38, "top_p": 1, "frequency_penalty": 0, "presence_penalty": 0 }'
- xwowsersx 4y agoWhat exactly is it aware of though? I guess it's aware of certain packages in certain languages? I asked how to create an AWS API gateway with terraform. It said: > You can use the terraform-aws-api-gateway module. This module provides a set of Terraform resources for creating and managing an API gateway. It allows you to define the API gateway, its resources, methods, and stages. It also provides support for custom domain names, API keys, and usage plans. I followed up with "Show me the hcl configuration for it." > There is no hcl configuration for it.
- ib33 4y agoYesterday, an undergraduate from Sri Lanka released KnowledgeGPT[1] which allows you to upload your docs and get answers from ChatGPT. It also uses FAISS so I'm wondering if DocsGPT is somehow related or inspired by the former. It also appears the Github library for DocsGPT was created shortly after the release of KnowledgeGPT. 1: https://github.com/mmz-001/knowledge_gpt https://github.com/mmz-001/knowledge_gpt
- JPKab 4y agoSeems like this app is very stripped down streamlit app more dependent on API calls vs the DocsGPT. Doubt they are related.
- ib33 4y agoYou're right. There's an interesting difference in the approach.
- osullip 4y agoI will support and pay for this project. I thought about the same problem. Companies have lots of marketing material and data sheets that are not easily queried but would be very useful to support staff when dealing with customer's questions. I didn't build it, they did. They get my support.
- Taek 4y agoAnything that relies on the ChatGPT API is subject to being insta-killed the moment OpenAI decides its not good to have them around. Productivity tools shouldn't depend on external parties like that.
- osullip 4y agoI agree. But the problem will be solved regardless of GPT. I don't think the whole stack relies on GPT for this solution. Some parts maybe, but the whole of the project is more than that.
- droopyEyelids 4y agoIm curious how this works out over time. It seems like adding a translation layer into conversational speech could make good docs harder to use? And if theyre not good docs to start, what will gpt base its answer on? Could be useful for a question that spans multiple technologies maybe
- jacooper 4y agoMan, imagine If you can run it locally and let it parse all of your documents, so you can ask it questions like when does my passport expires, or what is my Wife's ID number and so on.
- arlcode 4y agoThe future will most likely be "upload it to the cloud and ask the question there". Really sad but but it seems like most people seem to not care about privacy even for the most sensitive parts of their lives. Or even worse they mistrust "the government" but place great trust in "companies which are easily pressured or raided by more than one government". It doesn't make sense at all. (Edit: I don't think a well run government in a functioning democracy is inherently evil but pretending that companies are not collaborating because they are "private entities" is foolish)
- hackernewds 4y agoYou merely need to see the Twitter example to see how risky monopoly-level tech is in the wrong hands. Imagine they have access to all politicians DMs now, ripe for blackmail and kompromat
- amrb 4y ago"Do you want megacorps? Because that's how you get megacorps" -Archer Control is why I want to see more "open source" version of machine learning models be it over data or auditing the output and if I can p2p download a model I know we're in a good place. Take the internet, I could see my path being very different if we had to pay per minute like a phone plan to get information, having open access was a big bootstrap here. Finally of course there are problems with the Linux project, but would servers dictated only by M$oft/0r@cle be a positive outcome, gonna say no.
- aioprisan 4y agoAll you need is to 1. convert documents to embeddings (OpenAI or some other local library for this), 2. store all embeddings in an index like (SaaS like Pinecone or FAISS locally), 3. run query against embeddings. Just a document search is likely good enough, but if not you can use a completion algorithm in a simple LLM of the docs to generate answers.
- bfeynman 4y agoThis is really the quality of projects we're upvoting? This is a python script that calls another API.
- zerop 4y agoIt's not about this project. It's a use-case of ChatGPT someone has explored. ChatGPT is new and we all want to learn about how it works in many new use cases.
- shrimpx 4y agoBtw this project doesn't use ChatGPT, it uses text-davinci-003.
- miadabdi 4y agoCan you provide more info? I set up a telegram bot and connected it to my OpenAI account with API keys and it works, but as I'm aware ChatGPT is not available as api yet, so I'm guessing the repo I got the telegram bot kinda lied that it's using ChatGPT? [the repo](https://github.com/karfly/chatgpt_telegram_bot https://github.com/karfly/chatgpt_telegram_bot)
- spencerchubb 4y agofacebook was just a php script to display some pictures
- disgruntledphd2 4y agonot even pictures until 2007 or 2008
- antiatheist 4y agoI agree, the code is pretty average, inconsistently using quotation marks, looks copy pasted and developer comments trying to understand what they're doing. The concept isn't too novel either, LLM usage in knowledge base querying is nothing new, I know some lawyers looking into it for regulatory compliance. The KnowledgeGPT repo linked by another commenter seems more interesting.
- grensley 4y agoSeeing https://github.com/arc53/docsgpt/blob/main/application/combine_prompt.txt https://github.com/arc53/docsgpt/blob/main/application/combi... is just fascinating to me. It's absolutely where this all is headed, but to see it as source code evokes something in me.
- kfarr 4y agoIt's like the new crud but instead of designing the database tables and columns you're tweaking the prompts.
- deleted 4y ago[deleted]
- sadrobin 4y agoI know its crazy, there are also some other promts that I have not edited from LangChain that acutally run in the background
- _nalply 4y agoI would try to fine-tune GPT such that I don't need to repeat the first part for every query. Since OpenAI is billed by the token (a thousand tokens are about 750 words), it makes sense to fine-tune once then only submit the changing part. Caveat emptor: I didn't try this out yet. No idea whether this would work.
- eyegor 4y agoThis is a cool idea. Right now my goto tends to be https://devdocs.io/ https://devdocs.io/, but the idea of a conversational type of layer is fascinating. It's always a struggle in a new set of docs trying to figure out their phrasing for merge/join/combine or how they describe aggregations for example. A lot of the time when you're looking at documentation, you're trying to look up "how to do x (with y)" but most docs are written in a "common language" and end up describing things in jargon you may not be aware of yet.
- johnywalks 4y agoWhat ChatGPT can do combine knowledge and provide personalised examples. This is usually done manually and it takes hours of research with trial & error.
- mdmglr 4y agoSo just so I understand- this is all based on taking input from the user, injecting it in a template prompt that instructs chatgpt to answer the question based on providing it all the source material? What happened to building your own models to run offline?
- Kiro 4y agoThat's like complaining that a musician doesn't build their own piano. Or actually, it's like asking why they don't build their own piano factory. No sole developer has the skills or resources to build something like GPT. Even if it was open source no user would be able to run it locally anyway.
- shzhdbi09gv8ioi 4y agoNo, you completely miss the point. It's like complaining that a musician cannot practice their instrument without software that requires you to be always-online. > Even if it was open source no user would be able to run it locally anyway. You are just stating this, it does not make it true. Several of us are running GPT-3 workloads locally.
- jw1224 4y ago> Several of us are running GPT-3 workloads locally. Unless you work for OpenAI (and your “local” machine somehow has 1TB+ of VRAM, equivalent to roughly 25 A100s), this cannot possibly be true…
- bdhcuidbebe 4y ago1.7k forks… https://github.com/openai/gpt-3 https://github.com/openai/gpt-3 apart from that, several pre trained corpuses had been around for a while
- jw1224 4y agoThat repo just seems to be a bunch of JSON files and sample outputs. Am I missing something?
- gth158a 4y agoIn the screenshot, the last sentence reads: >"This will return a new DataFrame with all the columns from both tables, and only the rows that match the 'key' column". That is incorrect. The 'how' parameter is 'left' not 'inner'.
- shzhdbi09gv8ioi 4y agoAre you insinuating that ChatGPT gives incorrect answers? Careful about the pitchforks.
- sorokod 4y agoThe first paragraph in the "What is..." section states that the purpose is to provide answers. Nowhere does it say that answers should be correct or accurate by any measure. That this is not a problem for many is worrying.
- bjackman 4y agoAI to help me read docs does seem somewhat handy but I feel like if there's already documentation it's really just gonna save a couple of minutes? I will be much more excited when AI can explain undocumented systems to me. This feels like it can't be far away, and it will be a game changer. I guess for this to be helpful it is gonna need out-of-band info, but maybe just the git log would be a pretty good start. If you could add a mailing list or chat history of developers I imagine things could get more powerful.
- sadrobin 4y agoStep 1. Is documented systems, plain code is next. Thank you for your suggestion, great ideas!
- monkeydust 4y agoWe have plenty of partially documented systems, it's a constant challenge to keep it up to date. I wonder if this could be brought more in-line with production by merging documents with support request logs and even code ?
- bjackman 4y agoI think it would be silly to have AI _write docs_ for us! Documentation is an obsolete concept at that point. You can just ask the AI exactly what you need to know.
- westurner 4y agoFrom https://news.ycombinator.com/item?id=34659668 https://news.ycombinator.com/item?id=34659668 : >> How do the responses compare to auto-summarization in terms of Big E notation and usefulness? > Automatic summarization: https://en.wikipedia.org/wiki/Automatic_summarization https://en.wikipedia.org/wiki/Automatic_summarization > "Automatic summarization" GH topic: https://github.com/topics/automatic-summarization https://github.com/topics/automatic-summarization Though now archived, > Microsoft/nlp-recipes lists current NLP tasks that would be helpful for a docs bot: https://github.com/microsoft/nlp-recipes#content https://github.com/microsoft/nlp-recipes#content NLP Tasks: Text Classification, Named Entity Recognition, Text Summarization, Entailment, Question Answering, Sentence Similarity, Embeddings, Sentiment Analysis, Model Explainability, and Auto-Annotatiom
- westurner 4y agoOn further review, there are more GitHub projects labeled with https://github.com/topics/text-summarization https://github.com/topics/text-summarization than "automatic-summarization"; e.g. awesome-text-summarization: https://github.com/icoxfog417/awesome-text-summarization https://github.com/icoxfog417/awesome-text-summarization and https://github.com/luopeixiang/awesome-text-summarization https://github.com/luopeixiang/awesome-text-summarization , which links to what look like relatively current benchmarks for SOTA performance in text summarization from the gh source repo of https://nlpprogress.com/ https://nlpprogress.com/ : https://github.com/sebastianruder/NLP-progress/blob/master/english/summarization.md https://github.com/sebastianruder/NLP-progress/blob/master/e...