6 ms·
Your File System Is Already A Graph Database
- alxndr 6mo ago> […] the knowledge base isn’t just for research. It’s a context engineering system. You’re building the exact input your LLM needs to do useful work. > […] there’s a real difference between prompting “help me write a design doc for a rate limiting service” and prompting an LLM that has access to your project folder with six months of meeting notes, three prior design docs, the Slack thread where the team debated the approach, and your notes on the existing architecture.
- deleted 6mo ago[deleted]
- WillAdams 6mo agoI've found a similar structure along with a naming convention useful at my day job --- the big thing is the names are such that when copied as a filepath, the filepath and extension deleted, and underscores replaced by tabs, the text may then be pasted into a spreadsheet and summed up or otherwise manipulated. In somewhat of an inversion, I've been getting the initial naming done by an LLM (well, I was, until CoPilot imposed file upload limits and the new VPN blocked access to it) --- for want of that, I just name each scan by Invoice ID, then use a .bat file made by concatenating columns in a spreadsheet to rename them to the initial state ready for entry.
- deleted 6mo ago[deleted]
- rubises 6mo ago[dead]
- embedding-shape 6mo agoI've been playing around with the same, but trying to use local models as my Obsidian vault obviously contain a bunch of private things I'm not willing to share with for-profit companies, but I have yet to find any model that comes close to working out as well as just codex or cc with the small models, even with 96GB of VRAM to play around with. I've started to think about maybe a fine-tuned model is needed, specifically for "journal data retrieval" or something like that, is anyone aware of any existing models for things like this? I'd do it myself, but since I'm unwilling to send larger parts of my data to 3rd parties, I'm struggling collecting actual data I could use for fine-tuning myself, ending up in a bit of a catch 22. For some clients projects I've experimented with the same idea too, with less restrictions, and I guess one valuable experience is that letting LLMs write docs and add them to a "knowledge repository" tends to up with a mess, best success we've had is limiting the LLMs jobs to organizing and moving things around, but never actually add their own written text, seems to slowly degrade their quality as their context fills up with their own text, compared to when they only rely on human-written notes.
- weitendorf 6mo agoThis is exactly what we're working on, is there any application in particular you're interested in the most? > I'm struggling collecting actual data I could use for fine-tuning myself, Journalling or otherwise writing is by far the best way to do this IMO but it doesn't take very much audio to accurately do a voice-clone. The hard thing about journalling is that it can actually be really biased away from the actual "distribution" of you, whether it's more aspirational or emotional or less rigorous/precise with language. What I'm starting to do is save as many of my prompts as possible, because I realized a lot of my professional writing was there and it was actually pretty valuable data (especially paired with outputs and knowledge of what went well and waht didn't) for finetuning on my own workloads. Secondly is assembling/curating a collection of tools and products that I can drop into each new context with LLMs and also use for finetuning them on my own needs. Unlike "knowledge repositories" these both accurately model my actual needs and work and don't require me to do really do anything unnatural. The other thing I'm about to start doing is "natural" in a certain sense but kinda weird, basically recording myself talking to my computer (verbalizing my thoughts more so it can be embedded alongside my actions, which may be much sparser from the computer's perspective) / screen recordings of my session as I work with it. This is something I've had to look into building more specialized tools for, because it creates too much data to save all of it. But basically there are small models, transcoding libraries, and pipelines you can use for audio/temporal/visual segmentation and transcription to compress the data back down into tokens and normal-sized images. This is basically creating a semantic search engine of yourself as you work, kinda weird, but IMO it's just much weirder that your computer can actually talk back and learn about you now. With 96GB you can definitely do it BTW. I successfully finetuned an audio workload on gemma 4 2b yesterday on a 16GB mac mini. With 96GB you could do a lot. > letting LLMs write docs and add them to a "knowledge repository" I think what you actually want them to do is send them to go looking for stuff for you, or actively seeking out "learning" about something like that for their own role/purposes, so they can embed the useful information and better retrieve it when they need it, or produce traces grounded in positive signals (eg having access to this piece of information or tool, or applying this technique or pattern, measurably improves performance at something in-distribution to whatever you have them working on) they can use in fine-tuning themselves.
- exossho 6mo agoI can't remember how many file structures I've already tried... LLMs seem to be a great help here. Also used CC to organize my messy harddrive. Now just need to find a good way to maintain the order...
- freedomben 6mo ago> Also used CC to organize my messy harddrive. Do you still have your prompt by chance, and willing to share it? I took a stab at this and it didn't want to make much change. I think I need to be more specific but am not sure how to do that in a general way
- exossho 6mo agoI don't have the exact prompt anymore, but it was very lean. I first asked to do an assessment: "Review the content of the whole folder structure. I want you to assess it, and suggest a better setup and structure based on its content. Don't change anything yet, just assess" and then worked from there, giving feedback on the proposed folder structure, until I was happy
- itake 6mo agoI'm wonder though: 1. Why does AI need that folder structure? Why not a flat list of files and let the AI agent explore with BM25 / grep, etc. 2. pre-compute compression vs compute at query time. Kaparthy (and you) are recommending pre-compressing and sorting based on hard coded human abstraction opinions that may match how the data might be queried into human-friendly buckets and language. Why not just let the AI calculate this at run time? Many of these use cases have very few files and for a low traffic knowledge store, it probably costs less tokens if you only tokenize the files you need.
- laurowyn 6mo ago> Why does AI need that folder structure? Why not a flat list of files and let the AI agent explore with BM25 / grep, etc. It doesn't. The human creating the files needs it, to make it easier to traverse in future as the file count grows. At 52k files, that's a horrendous list to scroll through to find the thing you're looking for. Meanwhile, an AI can just `find . -type f -exec whatever {} \;` and be able to process it however it needs. Human doesn't need to change the way they work to appease the magic rock in the box under the desk.
- itake 6mo ago> The human creating the files needs it why? The human would just talk to the AI agent. Why would they need to scroll through that many files? I made a similar system with 232k files (1 file might be a slack message, gitlab comment, etc). it does a decent job at answering questions with only keyword search, but I think i can have better results with RAG+BM25.
- laurowyn 6mo agoAnd when the system fails for whatever reason? Just because AI exists doesn't mean we can neglect basic design principles. If we throw everything out the window, why don't we just name every file as a hash of its content? Why bother with ASCII names at all? Fundamentally, it's the human that needs to maintain the system and fix it when it breaks, and that becomes significantly easier if it's designed in a way a human would interact with it. Take the AI away, and you still have a perfectly reasonable data store that a human can continue using.
- stared 6mo agoFilesystem is a tree - a particular, constrained graph. Advanced topics usually require a lot of interconnections. Maybe it is why mind maps never spoke to me. I felt that a tree structure (or even - planar graphs) were not enough to cover any sufficiently complex topic.
- nutjob2 6mo agoIf it has hard or soft links, its a proper graph.
- calgoo 6mo agoThat what i was thinking! Instead of Wiki links, use Symlinks (i guess windows would not like it?)
- zahlman 6mo agoOn Linux at least, hard links can't be made to directories, except for the magic . and .. links. So this only allows for a DAG. Symbolic links can form a graph, and you can process them as needed using readlink etc. to traverse the graph, but they'll still be considered broken if they form a cycle.
- Retr0id 6mo agoConsidered broken by what?
- rleigh 6mo agoHistorically, it made deletion rather difficult with some problematic edge-cases. You could unlink a directory and create an orphan cycle that would never be deleted. Combine that with race conditions on a multi-user systems, plus the indeterminate cost of cycle-detection, and it turns out to be a rather complex problem to solve properly, and banning hard-links is a very simple way to keep the problem tractable, and result in fast, robust and reliable filesystem operations.
- bullen 6mo agoYep, my distributed JSON over HTTP database uses the ext4 binary tree for indexing: http://root.rupy.se http://root.rupy.se It can only handle 3 way multiple cross references by using 2 folders and a file now (meta) and it's very verbose on the disk (needs type=small otherwise inodes run out before disk space)... but it's incredibly fast and practially unstoppable in read uptime! Also the simplicity in using text and the file system sort of guarantees longevity and stability even if most people like the monolithic garbled mess that is relational databases binary table formats...
- appsoftware 6mo agoI created AS Notes (https://www.asnotes.io https://www.asnotes.io) (an extension for VS Code, Antigravity etc) partly because of this use case. It works like Obsidian, being markdown based, with wikilinks, mermaid rendering and task management. In VS Code, we have access to really good Agent harnesses and can navigate our notes and documents in a file system like manner. Further, using AGENTS.md, idea files etc we can instruct the agent how to interact, add to our notes etc. I've found working with my notes like this really useful, and provided I trim anything generated by an AI that's not going to be useful, provides an investment in the information I've gathered as the information is retained in markdown rather than getting lost in multiple chatbot UI s.
- itmitica 6mo agoI can see over engineering when I look at one. And premature optimization. Anyway, why care how the data is stored? You need a catalog. You need an index. You need automation. Helps keeping order and helps with inevitable changes and flips and pivots and whims and trends and moods and backups and restoration and snapshots and history and versioning and moon travels and collaboration and compatibility and long summer evening walks and portability.
- stingraycharles 6mo agoUsing the same logic, a key/value database is also a graph database? Isn’t the biggest benefit of graph databases the indexing and additional query constructs they support, like shortest path finding and whatnot?
- sorokod 6mo agoYes, the author is likely unaware of this. They see markdown files with links, so a graph and the set of those files, so a "database". https://neo4j.com/docs/graph-data-science/current/algorithms/ https://neo4j.com/docs/graph-data-science/current/algorithms...
- esafak 6mo agoHis argument is that the LLM is the query engine. By that logic you can approximate anything since LLMs can.
- lamasery 6mo agoNeo4j looooooves the "if you think about it, everything is graphs!" marketing maneuver. They (their marketing department) were the very first thing I thought of when I read this headline.
- zadikian 6mo ago"Everything is graphs, so let's use a graph DBMS for anything" is a classic blunder
- mzelling 6mo agoIf I understand this right, the difference between the author's suggested approach and simply chatting with an AI agent over your files is hyperlinks: if your files contain links to other relevant files, the agent has an easier time identifying relevant material.
- rcdwealth 6mo ago[dead]
- kenforthewin 6mo agoI keep harping on this, but the question is not "can you use your filesystem as a graph database" - of course you can - but whether this performs better or worse than a vector database approach, especially at scale. The premise of Atomic, the knowledge base project I'm currently working on, is that there is still significant value in vectors, even in an agentic context. https://github.com/kenforthewin/atomic https://github.com/kenforthewin/atomic
- iwontberude 6mo agoOh neat! Whenever I was working on RAG proof of concepts vector databases seemed to generate noisiest outputs that happened to include my information from my chunks but it was unable to draw reasonable contextual associations. I swap RAG out with a web search tool, all of a sudden the quality goes way up. Is RAG ever going to be easier to hold or should lay people like me just stay moving on?
- kenforthewin 6mo agoI think agentic RAG still has its place. a hybrid semantic/keyword search tool in addition to other research tools outperforms the baseline in my experience.
- pyinstallwoes 6mo agoCool project.
- kenforthewin 6mo agothanks for checking it out!
- zadikian 6mo agoOn the other hand, I get why cloud drive users completely disregard file structure and search everything. Two files usually don't have the same name unless you're laying it out programmatically like this. I use dir trees for code ofc, but everything else is flat in my ~/Documents. Deep inside a project dir, feels like some the ease of LLMs is just not having to cd into the correct directory, but you shouldn't need an LLM to do that. I'm gonna try setting up some aliases like "auto cd to wherever foo/main.py is" and see how that goes.
- embedding-shape 6mo ago> I use dir trees for code ofc, but everything else is flat in my ~/Documents. Which is great, but on all major OSes you'd eventually hit performance issues with flat directories like this. Might not be an issue in month one, or even year one, but after 10 years of note taking/journaling that approach will show the issue with large flat directories. So eventually you'd need to shard it somehow, so might as well start categorizing/sorting things from the get go, at least in some broad major categories at least, because doing so once you already have 10K entries in a directory, it sucks big time to do it.
- zadikian 6mo agoIf it's just performance, cd ~/Documents && mkdir old && mv ./* old/ (or today's date instead of old). I actually have that layout on one PC. If real organization is needed, seems like that'd be easier in hindsight than having foresight
- embedding-shape 6mo agoSo then you have one intentionally slow directory ("old/" in this case) and one fast directory? Personally I'd categorize stuff, but you do you, there really isn't any wrong way to do it, if it works it works :)
- zadikian 6mo agoI meant you mv into old before it gets too big. I've never actually seen a dir get slow like this. Only seen that with programmatic things like making 1M json files.
- bhewes 6mo agoWow just strings in files. Are you jumping node to node via pointers index free?
- game_the0ry 6mo agoThere for sure a "second brain" product hiding in plain site for one of the frontier AI companies. Google/Gemini should be all over this right now.
- aleksiy123 6mo agoI’ve been thinking about this in a couple of contexts and pretty much how I’ve come to think about it. Folders give you hierarchical categories. You still want tags for horizontal grouping. And links and references for precise edges. But that gives you a really nice foundation that should get you pretty damn far. I also now am telling the llm to add a summary as the first section of the file is longer.
- visarga 6mo ago[dead]
- estetlinus 6mo agoI am more curious on the note taking. How do you ingest data here? Export from slack via LLM:s? Store it in GitHub? My “knowledge” is spread out on various SaaS (Google, slack, linear, notion, etc). I don’t see how I can centralize my “knowledge” without a lot of manual labour.
- LocalPCGuy 6mo agoUnless you're forced into using certain tool (work, etc), start by standardizing on a single tool. That's one reason a lot of people like Obsidian, but there are plenty of similar tools, or you can just write markdown in your editor of choice. Then set of some sort of sync so you have it everywhere you are (mobile can be a bit tricky for some set-ups) and commit to using that method as much as possible for your notes. You may want to do as described and link to Slack messages (etc), but just remember any external link should be treated as ephemeral. You may not have access to the Slack anymore, for example. That may mean you don't need that note either, or it may mean you lost access to a node on your knowledge graph, you have to determine whether that matters. By starting now, at least everything going forward is captured in a way you can both own and utilize it. Then it may be a bit of a pain and some manual work to get existing notes into your tool of choice, but you can determine what needs to be in there from other tools as you go forward.
- kesor 6mo agoSo you have some folders with markdown files ... which are insanely hard to query without a tool ... impossible to traverse via their relationships ... and you call that a graph database? WHAT?! Clicked the link expecting to see some tool or method that actually allows graph-like queries and traversals on files in a file system, all I found was some rant about someone on the internet being wrong. Waste of time.
- Jayakumark 6mo agoInteresting approach but how do you download Google Docs, XLS and Slack threads etc.. and how is it saved in obsidian, are they all converted to markdown before saving or summarized to extract key topics and saved. What about images ?
- themafia 6mo agoSure. It just fails to be atomic. Which is a property I really like.
- SoftTalker 6mo agoI will always be in awe of people who can remain diligent doing this level of journaling/personal information management. I've got scraps of paper and legal pads and post-it notes and just throw them away after they've been sitting around for a while and I forget what they are about.
- evanjrowley 6mo agoI thought he was gonna talk about inodes: https://en.wikipedia.org/wiki/Inode https://en.wikipedia.org/wiki/Inode Maybe someday someone will expose these in the form of a graph database API (just for fun).
- inflam52 6mo agoOne of the benefits of graph databases is that you can measure strengths of connections and also infer connections that don’t actually exist (edge prediction) among many other path traversal techniques. It’s not always just about the connection itself. Many have these algorithms built in so you don’t have to reinvent it.
- deleted 6mo ago[deleted]
- lvca 6mo ago[dead]