7 ms·
Personal computing paves the way for personal library science
- walterbell 2y ago> Personal Library Science is the leverage of LLM technology, applied to a personal library. A personal library differs from a impersonal library in the fact that a personal library is an interpretation of a source material. These interpretations include: photographs from different photographers at the same event, or favorite scenes from a movie, or favorite passages from books, parts of songs that bring you to tears, etc. Importantly, these interpretations create unique sets that go on to create unique problems which require unique, idiosyncratic solutions. Would an LLM-driven "Personal Library" require manually annotated textual interpretation of each curated item, or could it derive personal interpretations from user history and the uniqueness of curated items/sets? For those who have been using local, offline LLMs with a manually curated text/image corpus, what have been the most valuable or surprising use cases? Author demo video (2023), https://youtube.com/watch?v=7TgqMRz2r3M https://youtube.com/watch?v=7TgqMRz2r3M & tooling comment (2024), https://news.ycombinator.com/item?id=39789712 https://news.ycombinator.com/item?id=39789712 > Inspired by the commonplace book format, I take highlights from Kindle and embed them in a DB. From there I build (multiple) downstream apps but the central one, Commonplace Bot is a bot that serves as a retrieval and transformer for said highlights. Related: https://en.wikipedia.org/wiki/Lifelog https://en.wikipedia.org/wiki/Lifelog
- dartos 2y ago> Would an LLM-driven "Personal Library" require manually annotated textual interpretation of each curated item No. In something like this you’d probably have the LLM annotate and curate your personal library for you. Potentially by creating and assigning tags or topics based on the content of your library.
- unshavedyak 2y agoYea. I had this discussion not too long ago about this. I'd love to have a combination of a library (Personal Knowledge Management style), data ingestions, and a current world view/state. The PKM is the stored info to write to and query against (both for LLMs and humans). The data ingests are just a pipeline of digital inputs to the system, like chat logs, maybe (transcribed) webcam feeds, files i'm currently editing on desktop, browsing history, etc. The current world view is the interpretation of what i'm doing - to tie all the ingests together and give them context. Eg in isolation browsing some Rust crates might not be that useful. But if i'm also editing Project X on my computer then it's reasonable to assume the searching is related to X. However if it's been 8 hours since any Project X activity, it's less likely related. Same goes for context-less chat logs (as happens frequently in my house) where they are extensions of a voice conversation, etc. All of this stuff is of course insanely privacy invading, so i'd only implement this locally. I also wouldn't even store most of it for fear of data invasion, but using it to fuel a PKM automatically seems pretty sexy. Like browser history, but for your life. This is all just wishful thinking though, LLMs have been moving too fast for me to even bother toying with this. I should note though that i did not intend for LLMs to be "smart". Rather, in a RAG-like fashion (i think is the term), i want to just let LLMs do what they're good at - summarization & autocomplete, and let the world view / PKM store the real data.
- dartos 2y agoFWIW LLMs have been advancing on benchmarks but the practical usage of them (RAG, React, CoT, etc) hasn’t really changed much in the past year.
- _bramses 2y ago> Would an LLM-driven "Personal Library" require manually annotated textual interpretation of each curated item, or could it derive personal interpretations from user history and the uniqueness of curated items/sets? I’ve personally found that tagging is less robust than LLM embeddings (mainly due to dimensionality), but human appended thoughts about a source — also embedded — serve even better as tags. Example: “this is a quote about dinosaurs…” (Old way of doing things) Tags: dinosaurs, jurassic, history Query: “dinosaurs” > results = 1… (New way of doing things) Embedded Quote: [0.182…] User Added Thought: “this dinosaur reminds me of a time i went to six flags with my cousins and…” Embedded User Added Thought: [0.284…] Query: “dinosaurs” > results = 2 (indexes = sources, thoughts) The "thoughts" index can do a second layer cosine similarity search and serve as a tag on its own to fetch similar concepts. Basically a tree search created by similarity from user input/feedback loops.
- skybrian 2y agoI imagine an LLM could work well for doing autocomplete while saving and annotating documents. But it’s not personal unless you edit the result to say what you want to say.
- mistrial9 2y agothere may be an important divergence implied by this essay .. people here ask about using an LLM.. but the essay refers to "different photographs of the same scene from different photographers" or other personal collection items that are related but subjective or not-authoritative There is a rush in public to condense and summarize many authoritative publications to find patterns, or to replace a human expert with automated results.. yet that is fundamentally different than taking multiple incomplete perspectives to add to a human library-owners knowledge and investigations. It is subtle to speak it but not subtle in its implications.. taking "data as facts" and condensing them or reordering them or rewriting an output based on them, using automation, is different than a human mind taking in many inputs for human mind knowledge and enabling new outputs from a human author.
- _bramses 2y ago> There is a rush in public to condense and summarize many authoritative publications to find patterns, or to replace a human expert with automated results.. yet that is fundamentally different than taking multiple incomplete perspectives to add to a human library-owners knowledge and investigations. It is subtle to speak it but not subtle in its implications.. taking "data as facts" and condensing them or reordering them or rewriting an output based on them, using automation, is different than a human mind taking in many inputs for human mind knowledge and enabling new outputs from a human author. You nailed it! Thanks for noticing the divergence!
- walterbell 2y agoThere's lots of interesting work that came out of BCL in 1960s, https://en.wikipedia.org/wiki/Biological_Computer_Laboratory https://en.wikipedia.org/wiki/Biological_Computer_Laboratory > The focus of research at BCL was systems theory and specifically the area of self-organizing systems, bionics, and bio-inspired computing; that is, analyzing, formalizing, and implementing biological processes using computers. BCL was inspired by the ideas of Warren McCulloch and the Macy Conferences, as well as many other thinkers in the field of cybernetics. On cybernetics, https://www.pangaro.com/definition-cybernetics.html https://www.pangaro.com/definition-cybernetics.html > Artificial Intelligence (AI) grew from a desire to make computers smart, whether smart like humans or just smart in some other way. Cybernetics grew from a desire to understand and build systems that can achieve goals.. it connects control (actions taken in hope of achieving goals) with communication (connection and information flow between the actor and the environment).. Later, Gordon Pask offered conversation as the core interaction of systems that have goals.
- teja_nemana 2y ago> The "during" is hard work, and very lonely work. There are no promises of success, and indeed, the path is one where you can't see more than three feet ahead of you and you exist on the cliff's edge of extinction by any silly mishap. The work of "during" is exhausting, and it constantly holds you taut and alert, afraid of the shadows that lurk beyond the campfire's edge. Well said. All anyone can do is to do the lonely work till you can't anymore or you find friends to not be lonely at that work anymore.
- quest88 2y agoI've had related ideas lurking at the back of my mind for a while now. Essentially, I want to save more things locally and and interact with it. For example, I have a bunch of book notes stored in Bear. I'd like to be able to ask questions about those notes, and also show the pages of the book itself.
- squirrel 2y agoTry Zenfetch. It's designed for this use case.
- gabev 2y agoThanks for mentioning Zenfetch :) Happy to answer any questions
- bibliotekka 2y agoWhat is Zenfetch?
- gabev 2y agoPersonal RAG. Connect your existing bookmarks/web browsing/notes into a knowledge library with AI search and chat over top it
- openrisk 2y agoPersonal computing has stagnated for such a long time, it creates substantial uncertainty about what state it might evolve to if and when the next step actually happens. In this respect local LLM's are simply the tip of the iceberg, pointing out the vast amount of personal information processing that is available in principle but does not actually happen.
- walterbell 2y agoOne could argue that personal computing (desktop) software piracy lead to web-based SaaS subscription licensing. In theory, mobile app stores solved device software piracy, at the cost of high distribution fees, policy restrictions and telemetry. Thanks to Linux being used at scale in Android and WSL, it's now maintained and capable on the desktop, as a hypothetical foundation for personal computing innovation. But even there, native GUI toolkits took a backseat to web and CLI. Remember Chandler? http://www.osafoundation.org/ http://www.osafoundation.org/ Investors poured small fortunes into cauldrons of smart devices, wearables and AR/VR, with little to show as nascent ecosystems failed to achieve escape velocity, due to closed hardware and software that forestalled the experimentation which birthed personal computing. Apple Silicon has reinvigorated walled laptops. Hopefully next month's derivative Qualcomm SoC from PC OEMs can offer good price/performance/watt for Apple-competitive-yet-open Arm laptops and tablets that can run any Linux distro, with retail SSDs and RAM, plus AI silicon roadmap. A modular Framework Arm laptop would be a good start to rebooting PC innovation.
- idle_zealot 2y agoHow does slightly improved laptop hardware relate to re-invigorating desktop software? Surely desktop computing has stagnated because most users are primarily or exclusively mobile users. In Mac land Apple has been progressively dumbing down their interfaces, in Windows land Microsoft is more focused on extracting maximum value from their users than trying to meaningfully improve their platform. In Linux land there are some interesting things happening with Nix/Guix around declarative system configurations, and around Fedora with its layered images+Flatpak distros for making systems more reliable, and System76 may be doing something novel interface-wise with Cosmic marrying powerful tiling/tabbing window layouts with intuitive controls and the niceties of an all-in-one desktop environment. From my perspective desktop computing is definitely advancing, but only for hobbyists, not for mainstream desktop operating systems.
- A_D_E_P_T 2y agoI work in a very interdisciplinary, and somewhat niche, tech/engineering field. For the past 15 years, I've been saving every relevant PDF that I can find -- mostly studies of the sort published by Elsevier and Springer, but also books and presentations. I now have around 10k, which probably makes it the largest private library focused on this particular domain of expertise. It has been extremely useful, especially because it's text-searchable and the really important papers are properly categorized. A local LLM will make it 100x more useful. Also, it might not even need be "local." If I make it available via the web, I can probably sell access to other scientists and engineers in my field. Recent advances really benefit data hoarders out there. I'd add that these days it totally makes sense to download libgen's entire archive, because (1) storage has never been cheaper, and (2) you can use it to train local LLMs.
- hervature 2y ago> If I make it available via the web, I can probably sell access to other scientists and engineers in my field. Out of curiosity. Does this statement come from complete ignorance of or complete disregard to copyright of the author?
- throwaway11460 2y agoIf it's research it very probably is at least partially publicly funded. Regardless of whatever the law says, I don't think it's immoral to take it and offer better services around it that will be useful enough that someone decides to pay.
- hervature 2y agoDo you not see the hypocrisy in stating that someone should be able to take something partially publicly funded and profit from it while the creator of said work should not retain some rights over said profit? By extension of the transitive property of the nebulous "partially", the LLM wrapper should be provided for free with complete disregard to the wrapper's creator since it is a derivative of partially publicly funded work.
- detourdog 2y agoIf one remembers NeXT included all sorts of non-computer documents and literature. The idea of storing vasts amounts of data for personal use was at the dawn of the PC era.
- WillAdams 2y agoI miss Librarian.app --- it was quite useful for a project of mine: https://tug.org/TUGboat/Articles/tb24-2/tb77adams.pdf https://tug.org/TUGboat/Articles/tb24-2/tb77adams.pdf (basically used copies of _The Bible_ and The Works of Shakespeare to determine if a given set of letters appeared in the English language or no)
- WillAdams 2y agoNot too long ago, I managed to pretty much ruin the wiki for a small (and at that time opensource) CNC machine by using it as my personal notebook --- probably my usage of it thus was a big part of why it was left off-line when the person hosting it moved. You can see it on the Wayback Machine: https://web.archive.org/web/20211127090321/https://wiki.shapeoko.com/ https://web.archive.org/web/20211127090321/https://wiki.shap... In retrospect, I should have put some of that effort into: https://en.wikibooks.org/wiki/Hobbyist_CNC_Machining https://en.wikibooks.org/wiki/Hobbyist_CNC_Machining although since then, a machine owner worked up: https://shapeokoenthusiasts.gitbook.io/shapeoko-cnc-a-to-z https://shapeokoenthusiasts.gitbook.io/shapeoko-cnc-a-to-z I still regret a bunch of stuff I didn't keep copies of, esp. the scans of Barry Hughart's notes for his novels. The irony is that one can see a bit of the result of discussion of this sort of thing at the top of one's browser window --- the URL bar, where URL == "Uniform Resource Locator" --- the originally proposed term was "Universal Resource Locator", but the argument against that was that people were not librarians, and that unlike Ted Nelson's Xanadu, there wouldn't an over-arching data structure and organization, so a given document wouldn't have a single canonical location. Anyone interested in this sort of thing who hasn't read it, should read Tim Berner-Lee's book: https://www.w3.org/People/Berners-Lee/Weaving/Overview.html https://www.w3.org/People/Berners-Lee/Weaving/Overview.html
- dtagames 2y agoBest quote from the article: "...personal library science is focused on your relationship with your information. How do we store information so that it useful at a later date? How do we transform our information into new valuable assets in different creative domains? How do we do all of this while being flexible enough for the idiosyncrasies, proclivities, likes and dislikes of eight billion distinct individuals? How do we chronicle the information diet of a single person as they learn new things, interact with the world at different phases in their life? How do we make sure we can pass down our best knowledge to generations below?"
- EricE 2y agoMy personal favorite https://www.devontechnologies.com/apps/devonthink https://www.devontechnologies.com/apps/devonthink
- RecycledEle 2y agoYears ago I spent thousands of hours trying to figure out how to organize a digital library. My final answer was to use the Library of Congress catalog system. They need to add some sub-categories for how-to explanations. Then have a field for media type (video vs. PDF vs. image) Then note the style of presentation (academic vs. folksy vs. a manual vs. a dad showing you how to do this) Then note the language
- gumboshoes 2y agoFor me, on macOS, FoxTrot Professional has been my personal file data indexer. Its essential go-to feature for me is sophisticated searching, including a form of regex. Also, true wildcards, no stopwords, and proximity searches put it far out in front of anything else I have tried, including many recent LLM local-docs tools. I have millions of files in dozens of formats in hundreds of GB gathered over decades (many digitized by me) and it handles it all like a champ, though an SSD drive and a late-model Mac is a must at that size. And backups, cause ain't nobody want to lose that.