5 ms·
Show HN: Continual Learning with .md
I have a proposal that addresses long-term memory problems for LLMs when new data arrives continuously (cheaply!). The program involves no code, but two Markdown files.
For retrieval, there is a semantic filesystem that makes it easy for LLMs to search using shell commands.
It is currently a scrappy v1, but it works better than anything I have tried.
Curious for any feedback!
- thomasquarre 6mo ago[dead]
- sudb 6mo agoI really like the simplicity of this! What's retrieval performance and speed like?
- wenhan_zhou 6mo agoMinimalism is my design philosophy :-) Good question. Since it is just an LLM reading files, it depends entirely on how fast it can call tools, so it depends on the token/s of the model. Haven't done a formal benchmark, but from the vibes, it feels like a few seconds for GPT-5.4-high per query. There is an implicit "caching" mechanism, so the more you use it, the smoother it will feel.
- namanyayg 6mo agoI've seen a lot of such systems come and go. One of my friends is working on probably the best (VC-funded) memory system right now. The problem always is that when there are too many memories, the context gets overloaded and the AI starts ignoring the system prompt. Definitely not a solved problem, and there need to be benchmarks to evaluate these solutions. Benchmarks themselves can be easily gamed and not universally applicable.
- xwowsersx 6mo agoWhat is the memory system you are referring to? I've been trying Memori with OpenClaw. Haven't had a ton of time to really kick the tires on it, so the jury's still out.
- natpalmer1776 6mo agoThe armchair ML engineer in me says our current context management approach is the issue. With a proper memory management system wired up to it’s own LLM-driven orchestrator, memories should be pulled in and pushed out between prompts, and ideally, in the middle of a “thinking” cycle. You can enhance this to be performant using vector databases and such but the core principle remains the same and is oft repeated by parents across the world: “Clean up your toys before you pull a new one out!” Also since I thought for another 30 seconds, the “too many memories!” Problem imo is the same problem as context management and compaction and requires the same approach: more AI telling AI what AI should be thinking about. De-rank “memories” in the context manager as irrelevant and don’t pass them to the outer context. If a memory is de-ranked often and not used enough it gets purged.
- dummydummy1234 6mo agoMid thinking cycle seems dangerous as it will probably kill caching.
- natpalmer1776 6mo agoThe mid thinking cycle would require significant architecture change to current state of art and imo is a key blocker to AGI
- wenhan_zhou 6mo agoFair concern. ReadMe does support loading memories mid-reasoning! It is simply an agent reading files. Although GPT-5.4 currently likes to explore a lot upfront, and only then responds. But that is more of a model behaviour (adjustable through prompting) rather than an architectural limitation.
- natpalmer1776 6mo agoAh, I mean bi-directional management of context. Add and remove. Basically just the remove bit since we have adding down.
- wenhan_zhou 6mo agoContext bloat is real, but the architecture has the potential to solve it. You need clever naming for the filesystem and exploration policy in AGENTS.md. (not trivial!) The benchmark is definitely the core bottleneck. I don't know any good benchmark for this, probably an open research question in itself.
- alexbike 6mo ago[flagged]
- in-silico 6mo agoThis assumes that the model's behavior and memories are faithful to their english/human language representation, and don't stray into (even subtle) "neuralese".
- SyneRyder 6mo agoHaving run a Markdown memory system with Claude for over a year, I don't think I've seen any evidence of neuralese. That's even with Claude being regularly encouraged to write "reflections" on each session, including automated sessions, and weekly summaries of those reflections. The bigger problem is avoiding what I call the Memento Effect. I won't spoil the movie for anyone, but Memento involves a character who cannot make new memories, so he has to take meticulous notes about everything. But if any of those notes are vague or incorrect, they still get accept as truth when next reviewed. So you really need your Markdown memory to be pristine and mustn't allow it to become polluted.
- wenhan_zhou 6mo agoI think what's missing is a benchmark that measures how well the memories contribute to future interactions.
- verdverm 6mo agoIs there anything (besides plumbing) that prevents both? i.e. when the file is edited, all the representations are updated
- Dahvay 6mo ago[dead]
- wenhan_zhou 6mo agoThe editability is surely an underrated advantage, both for the program itself and the memories it generated. I think in terms of noise, it is less problematic here because not everything is being retrieved. The agent can selectively explore subsets of the tree (plus you can edit the exploration policy by yourself). Since there is no context bloat, it is quite forgivable to just write things down.
- dhruv3006 6mo agoI love how you approached this with markdown ! I guess the markdown approach really has a advantage over others. PS : Something I built on markdown : https://voiden.md/ https://voiden.md/
- wenhan_zhou 6mo agoYep. Markdown is the future :-)
- 0xchamin 6mo agois this based on Karapathy's LLM Wiki idea (link: https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f https://gist.github.com/karpathy/442a6bf555914893e9891c11519...)? . I leveraged Karapathy's wiki idea and built MCPTube- it's a CLI and also an MCP server that turns YouTube videos into a compounding knowledge base. Check it out and let me know what you think (link: https://github.com/0xchamin/mcptube https://github.com/0xchamin/mcptube)
- wenhan_zhou 6mo agoI just read LLM Wiki in more detail. I have heard about it second-hand before this project. The "no-code" idea was inspired by Karpathy. As I have understood it, in LLM Wiki, the human is very much in the loop in what gets written. In ReadMe, the human control is mostly on the policy (prompt) level, and it is done once, the agent then goes full autonomously afterwards. After a quick skim of your project. I have tried an embedding-based knowledge base as well, but it is a bit tricky to make the embedding match a user query. For example, "What happened?" is not at all similar to "Batman defeats Joker." You need to reformulate the query using an LLM, which is tricky given that the query is conditioned on the whole chat history. That's partly why I abandoned embedding-based methods. But given that MCPTube already works on Gemini CLI, I could see it work natively without embeddings. Gemini is capable of reading video files natively. Worth a try?
- jusasiiv 6mo agoSeems interesting. Ill give it a try on my agent, memory is definitely an ongoing issue. How long have you been running this in a continuous state? Also have you tried other LLM's and seen a difference on how well they can use it?
- wenhan_zhou 6mo agoAlthough I have been working on memory before, ReadMe is very fresh. The moment I saw it running, I published it. So, no continuous running nor LLM ablation studies. Treat it as an MVP, would love to hear how your agent performs!
- inveflo 6mo ago[dead]
- _zer0c00l_ 6mo agoYour example is with Codex - OpenAI could implement this easily on their end right? Every prompt of yours was an API call and they have a log, they can easily re-create a quick history of what you did/asked for before?
- wenhan_zhou 6mo agoIn theory, yes. Although the privacy setting says otherwise. But in the end, it doesn't really matter; it is public on GitHub, so anyone can use it.
- esafranchik 6mo agoHave you noticed an relationship between recall and the number of files/memories?
- wenhan_zhou 6mo agoCurrently working on a benchmark!
- Bronzado 6mo ago[dead]
- anatoliikmt 6mo agoUsing the same approach for dev documentation storage: https://ctxlayer.dev/ https://ctxlayer.dev/ Has been working for me for a couple of months already. So far human curation of context is the way to go.
- wenhan_zhou 6mo agoHow does the agent intelligently synthesize information across different files?
- anatoliikmt 6mo agoI found progressive discovery works very well. I have an INDEX.md file in every folder with a table describing what files are in that folder and what they do.
- wenhan_zhou 6mo agoAh, so you are effectively offloading the file exploration mechanism to the INDEX.md in the sub-directories rather than writing a complex prompt?
- anatoliikmt 6mo agoYes, exactly. I made the Context Layer into a skill where I describe the structure. So when I say, "reindex", the agent knows to check and update the index files. Also it knows if it modifies any file, it needs to also update the index. Works really well!