23 ms·
I've been very interested in this recently. I'm pretty sure that every project should have a ./papers directory of annotated papers in it like I do in Qlatt[0].
by ctoth 6mo ago
I've been very interested in this recently. I'm pretty sure that every project should have a ./papers directory of annotated papers in it like I do in Qlatt[0].
Literally every project. If it's something that's been done a million times then that means it has good literature on it? If not, then even more important to find related stuff! And not just crunchy CS stuff like databases or compilers or whatever. Are you creating a UI? There's probably been great UI research you can base off of! Will this game loop be fun in the game you're building? There's probably been research about it!
[0]: https://github.com/ctoth/Qlatt/blob/master/papers/ https://github.com/ctoth/Qlatt/blob/master/papers/
- alex000kim 6mo agoThat directory is huge already! I guess the index.md helps the agent find what it needs, but even the markdown file is very long - this would consume a ton of tokens. Also I wonder who/what decides what papers go in there. In the blog post, the agent is allowed to do its own search.
- ctoth 6mo agoCheck out the Researcher and Process Leads skill in ctoth/research-papers-plugin. I have basically completely automated the literature review.
- pstuart 6mo agoHaving a "indexed global data collection" of the markdown would be a kumbaya moment for AI. There's so much data out there but finite disk space. Maybe torrents or IPFS could work for this?
- ctoth 6mo agoI'm actually sort of working on this! https://github.com/ctoth/propstore https://github.com/ctoth/propstore -- it's like Cyc, but there is no one answer. Plus knowledge bases are literally git repos that you can fork/merge. Research-papers-plugin is the frontend, we extract the knowledge, then we need somewhere to put it :)
- pstuart 6mo agoAwesome! TIL about Cyc, and it's quite intriguing. I'd been thinking about how being able to integrate Prolog or similar tools might be a valuable endeavor (although I've yet to write anything in Prolog myself).
- zzleeper 6mo agoWow this is amazing. Did you write all those MD files by hand, or used an LLM for the simple stuff like extracting abstracts?
- ctoth 6mo agoI used https://github.com/ctoth/research-papers-plugin https://github.com/ctoth/research-papers-plugin to produce the annotations. The thing that's really cool is how they surface the cross-links in the collection, for instance look at https://github.com/ctoth/Qlatt/blob/master/papers/Fant_1988_LFFrequencyDomainInterpretation/notes.md https://github.com/ctoth/Qlatt/blob/master/papers/Fant_1988_... Claude is much faster and better at reading papers than Codex (some of this is nested skill dispatch) but they both work quite incredibly for this. Compile your set of papers, queue it up and hit /ingest-collection and go sleep, and come back to a remarkable knowledge base :)