5 ms·
Turning a pile of documents into a searchable useable knowledge base
- linuxrebe1 3mo agoI had an issue. A documents folder with over 12k objects in it. A hodgepodge of folders and sub-folders. That over time had created a mess that no amount of file movement was ever going to make it usable. I wanted: 1) To keep my data local 2) be able to filter out PII and other data 3) Be able to find and delete duplicates 4) Get short synopsis of what a document is 5) Semantic and keyword search 6) All of this kept local to me requiring no internet access and no tokens spent to train someone elses AI. The result I call DocuBrowser and in it's current form is FOSS (GPL-3) licensed for your personal use. The UI is in your browser. The AI models used are held local and are tiny, Available for Linux(RPM,Deb, and tgz) Windows and Mac. Let me know what you think and thanks for taking the time to try it out.
- bobim 3mo agoCould it be extended so it also extracts pictures from pptx and xlsx and run vision to get a description to be added to the text content before indexing?
- linuxrebe1 3mo agoLet me look into this
- clif_mcIrvin 3mo agoHow about jpegs or other scanner images files? We have hundreds of scanned documents that were never pdf wrapped.
- linuxrebe1 3mo agohmmmm :)
- esperent 3mo agoI've been working on something related - extracting tons of data from various formats to allow searching them - and the solution I chose for xlxs and xls files was headless LibreOffice to convert them to CSV. There's also exceljs but I found it didn't work for many old xls files. I didn't find screenshotting of spreadsheets worked well, vision wasn't very accurate on them. I do use it for PDFs though. For docx it's probably fine either way but I went with LibreOffice -> markdown.
- linuxrebe1 3mo agoI went with the python libraries (pydoc and pyxls for example), because it's portable and doesn't require a big download to a users system if they don't already have it installed.
- bobim 3mo agoMy take was on pictures embedded into those documents, I'm not sure screenshotting would help as the text/numeric data is already there. Just saying.
- seb1204 3mo agoSounds similar to https://docs.paperless-ngx.com/ https://docs.paperless-ngx.com/ Key difference I see is that you point it to a folder instead of uploading to a system.
- vsviridov 3mo agoI think paperless devs are working on AI integration, and there are 3rd party solutions. I'm holding out for an official one, so far. It's pretty cool, I've set up a share where the scanner scans, and it automatically picks it up from there and ingests it into the system.
- password4321 3mo agoPersonal use? I need this at work, dragging useful info from tarpits like Teams and GitLab. Also need to search git repos including all branches and history (TIL/xkcd#153'd GitLab's web search can basically only do one branch at a time).
- linuxrebe1 3mo agoI creating DocuRepo as well. though not as fleshed out.
- rukshn 3mo agoBut how’d you access teams when it’s work teams and don’t have api access ?
- password4321 3mo agoMicrosoft Graph API
- gatnoodle 3mo agoThis looks really cool. Can you tell me the minimum specs required to run this? It would nice if you could add it to the readme as well.
- linuxrebe1 3mo agoI've run it on a VM with 4G ram and no GPU. It runs, But I really recommend 8G ram at least. If you have a GPU (like I do) with 4G vRAM that is ideal. Will get this in the readme. Thanks for the suggestion. I really tried to build this to minimal spec.
- fnordian 3mo agoIt’s either restricted to personal use, or it’s GPL-3. How can you have both?
- linuxrebe1 3mo agoBy restricted for personal use I mean it's not networked. It's running on your system only. It's not a networked commercial product able to do SSO etc. It's not an enterprise level product.
- aucisson_masque 3mo agoI'm a huge fan of recall, going to test this out. This looks very interesting.
- rahimnathwani 3mo agoDid you mean Recoll (https://www.recoll.org/ https://www.recoll.org/)?
- aucisson_masque 3mo agoIndeed
- asciimoo 3mo agoWe need projects like this. Automatically classifying the files is smart. I'm working on a similar application called Hister (https://github.com/asciimoo/hister https://github.com/asciimoo/hister). I should borrow some of your ideas. =]
- hankbond 3mo agoI have not set up Hister yet but it's on my list to try out. How would I do something like host it on my Unraid box but have it index/persist my local MacBook browsing history?
- linuxrebe1 3mo agoI just had a wild thought. Combine Hister with my RepoSearch app. Point it at a companies Internal github/gitlab and have a searchable knowledge base of your git repos.
- asciimoo 3mo agoI like the idea. Could you share your RepoSearch app? Also, we have Discord & IRC channels, please join and start brainstorming.
- NKosmatos 3mo agoLooks good, definitely going to try it. Extra thanks for creating something fully local, we need more projects like this one!
- linuxrebe1 3mo agothankyou
- toomuchtodo 3mo agoHow do you feel about supporting an S3 compatible target as a feature request?
- linuxrebe1 3mo agoI'm actually thinking of this for a commercial product feature. However, if you use a tool like Rclone on Windows, Linux or Mac. Mount the s3 bucket and you can then run DocuBrowse as if the s3 bucket were local.
- subhobroto 3mo agoI love your project on many fronts. One, you're using Claude. Two, you used Python - but most importantly, you personally care about it. I will be using this, and I will be making contributions to it as well. > I'm actually thinking of this for a commercial product feature Would you consider writing down which features you would like to make commercial product features and how you would like to price them?
- linuxrebe1 3mo agoConsider it yes, However having experience in this ... not really. For now there is a file called Decisions.md in the repo that is my "notes to self" if you will about where and what I need to do.
- deleted 3mo ago[deleted]
- drizzler 3mo agoI just installed this and, after a few hiccups, got it up and running on my Ubuntu system. Works great, looks great. Thank you for this. Half of my documents are OpenDocument format. Is there any chance you'll be supporting ODF in the future?
- linuxrebe1 3mo agoYes, not supporting it is an oversight I will correct.
- linuxrebe1 3mo agoWill have version 0.9.1 out later today to support ODF formats.
- linuxrebe1 3mo agov0.9.1 is in the repo and packages have been built. It now does all of the ODT formats.
- jphorism 3mo agoNice, what are you hoping to accomplish with this project?
- NamlchakKhandro 3mo agoA resume
- passwordoops 3mo agoCare to elaborate?
- linuxrebe1 3mo ago- Filling a need I personally have. - Learning how to leverage AI for real world use not just to fill up a data center. - Personal knowledge -developing skills Pretty much in that order
- Avery29 3mo agoThe hardest part of these projects is usually not making documents searchable
- karmakaze 3mo agoI learned a solution is to turn the documents into vectors in say PostgreSQL (with pgvector) and do a cosine similarity search with a search vector. Doing a search for embed models on HuggingFace shows nomic-ai/nomic-embed-text-v1.5 and Qwen/Qwen3-Embedding-0.6B. I might have used a larger one like Qwen/Qwen3-Embedding-4B. There's some info for AnythingLLM[0] which supports RAG. AnythingLLM has LanceDB out of the box but also supports others including pgvector. [0] https://docs.anythingllm.com/features/embedding-models https://docs.anythingllm.com/features/embedding-models
- mune2gu-chan 3mo agoNot a fan of pushing every personal document to someone else's cloud. Nice to see a tool that keeps everything on disk instead.
- nickweb 3mo agoHonestly. This with Paperless-NGX might be game changing if both pointed to the same folder.
- kamranjon 3mo agoWanted to share Antfly which I think serves a similar niche: https://antfly.io/ https://antfly.io/ https://github.com/antflydb/antfly https://github.com/antflydb/antfly They’ve put a lot of effort into optimizing the local llm pipelines and I have a lot of faith in the devs working on it.
- appstorelottery 3mo agoAnyone getting a bunch of permission errors when running (e.g. Traceback (most recent call last): File "/Users/tron/Applications/DocuBrowse/venv/lib/python3.9/site-packages/psutil/_psosx.py", line 347, in wrapper return fun(self, args, *kwargs) File "/Users/tron/Applications/DocuBrowse/venv/lib/python3.9/site-packages/psutil/_psosx.py", line 508, in net_connections rawlist = cext.proc_net_connections(self.pid, families, types) PermissionError: [Errno 1] Operation not permitted (originated from proc_pidinfo(PROC_PIDLISTFDS) 1/2)
- appstorelottery 3mo agoLiving in bizarro world of AI. Install open source project, fails, feed into OpenCode w/DeepSeekFlash 4 -> feed error into it get fixed. The kill_port function only catches ImportError from the psutil block, so when psutil is installed but raises AccessDenied (common on macOS), it crashes instead of falling back to lsof. In platform_paths.py - add two lines after line 250: except psutil.Error: pass Fixed. Now when psutil raises AccessDenied (as happens on macOS without elevated privileges), it falls through to the lsof/fuser fallback instead of crashing. Try docubrowser start again.
- appstorelottery 3mo agoDisappointed that it wasn't returning a list of paragraphs from eBooks that semantically match; search only appears to list the publications - not the actual match within the document.
- linuxrebe1 3mo agoNoted the bug.
- linuxrebe1 3mo agoI fixed this in version 0.9.1 (just released) thanks for the bug (seriously)
- linuxrebe1 3mo ago
- Ozzie_osman 3mo agoThis is really cool. Can it play nice with gdrive or Dropbox? For better or worse, that's just where my data lives now but I'd love this layer.
- linuxrebe1 3mo agoI use rclone to "mount" them locally. Then it becomes searchable.
- hunmernop 3mo agoCan you make a dockerfile and docker compose file?
- LawrenceKerr 3mo agoMany such open source projects already (which is fantastic). I lose track of them. Today it happened I needed a simple way to embed & query 1TB+ of documents, and I was looking at open source options. Can anyone tell me what their go-to solution is now? Could this be the one? And what are the key differences vs. other open source RAG tools like kotaemon?