14 ms·
Local LLMs versus offline Wikipedia
- IncreasePosts 1y agoMaybe we need a LLM with a searching and ranking function foremost, so it can scan an actual copy of Wikipedia and return the best real results to the user
- vFunct 1y agoWhy not both? LLM+Wikipedia RAG
- loloquwowndueo 1y agoBecause old laptop that can’t run a local LLM in reasonable time.
- NitpickLawyer 1y ago0.6b - 1.5b models are surprisingly good for RAG, and should work reasonably well even on old toasters. Then there's gemma 3n which runs fine-ish even on mobile phones.
- ozim 1y agoMost people who can nag about old laptops on HN can just afford newer one but are cheap as Scrooge Mcduck.
- mlnj 1y agoFYI: non-Western countries exist.
- folkrav 1y agoEh, even just “countries that are not the US” would be a correct statement. US tech salaries are just in an entire different ballpark to what most companies outside the US can offer. I’m in Canada, I make good money (as far as Canadian salaries go), but nowhere near “buy an expensive laptop whenever” money.
- lblume 1y agoIt may also come down to laptops being produced and sold mostly by US companies, which means that the general fact of most items (e.g. produce) being much more expensive in the US compared to, say, Europe doesn't really apply.
- folkrav 1y agoSure, maybe. In the end, what makes an expense big or not is which proportion of their income goes towards it. Most of the rest of the world has (much) lower salaries, and as you pointed out, often higher cost for equipment. Therefore, the purchase is/feels larger.
- simonw 1y agoIt's not uncommon for professionals to spend many thousands of dollars on the tools and equipment they need for their trade. Try telling a plumber that $2,000 for a laptop is a financial burden for a software engineer.
- folkrav 1y agoComparing my problems to other people’s problems don’t make mine go away. A single purchase hitting a unit of percentage or more of anyone’s income is a large purchase regardless of what they’re making. Professionals being expected to shell out their own money to make their boss money is another problem entirely. A decent laptop is a big expense for me, their tools are an even bigger one for them, and none of these statements are contradictory.
- ozim 1y agoPeople who are from those countries that can nag on HN and know whant HN is are most likely still better off than most of their fellow countrymen.
- whatevertrevor 1y agoDo you have any evidence to back that up? The barrier for entry to HN is an email account, it isn't necessarily this tech industry exclusive zone you're imagining.
- folkrav 1y agoIt feels like you're suggesting that someone being better off than most in their country necessarily means buying a new laptop is not a large purchase for them. I'd flip it like this: is a single item hitting multiple units of percentage of one’s income ever a small purchase?
- loloquwowndueo 1y agoI mean, sure, but this was mentioned in the article, I didn’t make it up: “Offline Wikipedia will work better on my ancient, low-power laptop.”
- moffkalast 1y agoNow this is an avengers level threat.
- JKCalhoun 1y agoYeah, wanting to try to do that. Someone posted this recently: https://github.com/philippgille/chromem-go/tree/v0.7.0/examples/rag-wikipedia-ollama https://github.com/philippgille/chromem-go/tree/v0.7.0/examp... But it is a very simplified RAG with only the lead paragraph to 200 Wikipedia entries. I want to learn how to encode a RAG of one of the Kiwix drops — "Best of Wikipedia" for example. I suppose an LLM can tell me how but am surprised not to have yet stumbled upon one that someone has already done.
- mac-mc 1y agoYeah at these sizes, it's very much a why not both.
- simonw 1y agoThis is a sensible comparison. My "help reboot society with the help of my little USB stick" thing was a throwaway remark to the journalist at a random point in the interview, I didn't anticipate them using it in the article! https://www.technologyreview.com/2025/07/17/1120391/how-to-run-an-llm-on-your-laptop/ https://www.technologyreview.com/2025/07/17/1120391/how-to-r... A bunch of people have pointed out that downloading Wikipedia itself onto a USB stick is sensible, and I agree with them. Wikipedia dumps default to MySQL, so I'd prefer to convert that to SQLite and get SQLite FTS working. 1TB or more USB sticks are pretty available these days so it's not like there's a space shortage to worry about for that.
- cyanydeez 1y agothe real valuable would be both of them. the LLM is good for refining/interpreting questions or longer form progress issues, and the wiki would be actual information for each component of whatever you're trying to do. But neither are sufficient for modern technology beyond pointing to a starting point.
- jjice 1y agoOh interesting idea to use SQLite and their FTS. I was very impressed by the quality of their FTS and this sounds like a great use case.
- camel-cdr 1y agoreposting a comment of mine from a few weeks ago: > All digitized books ever written/encoded compress to a few TB. I tied to estimate how much data this actually is in raw text form: # annas archive stats papers = 105714890 books = 52670695 # word count estimates avrg_words_per_paper = 10000 avrg_words_per_book = 100000 words = (papers*avrg_words_per_paper + books*avrg_words_per_book ) # quick text of 27 million words from a few books sample_words = 27809550 sample_bytes = 158824661 sample_bytes_comp = 28839837 # using zpaq -m5 bytes_per_word = sample_bytes/sample_words byte_comp_ratio = sample_bytes_comp/sample_bytes word_comp_ratio = bytes_per_word*byte_comp_ratio print("total:", words*bytes_per_word*1e-12, "TB") # total: 30.10238345855199 TB print("compressed:", words*word_comp_ratio*1e-12, "TB") # compressed: 5.466077036085319 TB So uncompressed ~30 TB and compressed ~5.5 TB of data. That fits on three 2TB micro SD cards, which you could buy for a total of 750$ from SanDisk.
- antonkar 1y agoA bit related: AI companies distilled the whole Web into LLMs to make computers smart, why humans can't do the same to make the best possible new Wikipedia with some copyrighted bits to make kids supersmart? Why kids are worse than AI companies and have to bum around?)
- horseradish7k 1y agowe did that and still do. people just don't buy encyclopedias that much nowadays
- antonkar 1y agoImagine taking the whole Web, removing spam, duplicates, bad explanations It will be the free new Wikipedia+ to learn anything in the best way possible, with the best graphs, interactive widgets, etc What LLMs have for free but humans for some reason don’t In some places it is possible to use copyrighted materials to educate if not directly for profit
- literalAardvark 1y agoLove it when Silicon Valley reinvents encyclopedias
- antonkar 1y agoThe proposed project is a non profit, I don’t think it can be a for profit legally (it didn’t stop AI companies, though)
- vunderba 1y ago> Imagine taking the whole Web, removing spam, duplicates, bad explanations Uh huh. Now imagine the collective amount of work this would require above and beyond the already overwhelmed number of volunteer staff at Wikipedia. Curation is ALWAYS the bugbear of these kinds of ambitious projects. Interactivity aside, it sounds like you want the Encyclopedia Brittanica. What made it so incredible for its time was the staggeringly impressive roster of authors behind the articles. In older editions, you could find the entry on magic written by Harry Houdini, the physics section definitively penned by Einstein himself, etc.
- dcc 1y agoOne important distinction is that the strength of LLMs isn't just in storing or retrieving knowledge like Wikipedia, it’s in comprehension. LLMs will return faulty or imprecise information at times, but what they can do is understand vague or poorly formed questions and help guide a user toward an answer. They can explain complex ideas in simpler terms, adapt responses based on the user's level of understanding, and connect dots across disciplines. In a "rebooting society" scenario, that kind of interactive comprehension could be more valuable. You wouldn’t just have a frozen snapshot of knowledge, you’d have a tool that can help people use it, even if they’re starting with limited background.
- progval 1y agoAn unreliable computer treated as a god by a pre-information-age society sounds like a Star Trek episode.
- bryanrasmussen 1y agohey generally everything worked pretty good in those societies, it was only people who didn't fit in who had a brief painful headache and then died!
- bigyabai 1y agoOr the plot to 2001 if you managed to stay awake long enough.
- gretch 1y agoDefinitely sounds like a plausible and fun episode. On the other hand, real history if filled with all sorts of things being treated as a god that were much worse than "unreliable computer". For example, a lot of times it's just a human with malice. So how bad could it really get
- DrillShopper 1y ago> So how bad could it really get I don't know. How about we ask some of the peoples who have been destroyed on the word of a single infallible malicious leader. Oh wait, we can't. They're dead. Any other questions?
- spankibalt 1y agoWikipedia-snapshots without the most important meta layers, i. e. a) the article's discussion pages and related archives, as well as b) the version history, would be useless to me as critical contexts might be/are missing... especially with regards to LLM-augmented text analysis. Even when just focusing on the standout-lemmata.
- pinkmuffinere 1y agoI’m a massive Wikipedia fan, have a lot of it downloaded locally on my phone, binge read it before bed, etc. Even so, I rarely go through talk pages or version history unless I’m contributing something. What would you see in an article that motivates you to check out the meta layers?
- asacrowflies 1y agoAny article with social or political controversy ... Try gamergate. Or any of the presidents pages for since at least bush lol
- nine_k 1y agoTry any article on a controversial issue.
- pinkmuffinere 1y agoI guess if I know it’s controversial then I don’t need the talk page, and if I don’t then I wouldn’t think to check
- nine_k 1y agoSeeing removed quotations and sources, and the reasons given, could be... enlightening sometimes. Even if the removed sources are indeed poor, the very way they are poor could be elucidating, too.
- spankibalt 1y ago> "I’m a massive Wikipedia fan, have a lot of it downloaded locally on my phone, binge read it before bed, etc." Me too, albeit these days I'm more interested in its underrated capabilities to foster teaching of e-governance and democracy/participation. > "What would you see in an article that motivates you to check out the meta layers?" Generally: How the lemma came to be, how it developed, any contentious issues around it, and how it compares to tangential lemmata under the same topical umbrella, especially with regards to working groups/SIGs (e. g. philosophy, history), and their specific methods and methodologies, as well as relevant authors. With regards to contentious issues, one obviously gets a look into what the hot-button issues of the day are, as well as (comparatives of) internal political issues in different wiki projects (incl. scandals, e. g. the right-wing/fascist infiltration and associated revisionism and negationism in the Croatian wiki [1]). Et cetera. I always look at the talk pages. And since I mentioned it before: Albeit I have almost no use for LLMs in my private life, running a Wiki, or a set of articles within, through an LLM-ified text analysis engine sounds certainly interesting. 1. [https://en.wikipedia.org/wiki/Denial_of_the_genocide_of_Serbs_in_the_Independent_State_of_Croatia https://en.wikipedia.org/wiki/Denial_of_the_genocide_of_Serb...]
- wangg 1y agoWouldn’t Wikipedia compress a lot more than llms? Are these uncompressed sizes?
- Philpax 1y agoYes, they're uncompressed. For reference, `enwiki-20250620-pages-articles-multistream.xml.bz2` is 25,176,364,573 bytes; you could get that lower with better compression. You can do partial reads from multistream bz2, though, which is handy.
- GuB-42 1y agoKiwix (what the author used) uses "zim" files, which are compressed. I don't know where the difference come from, but Kiwix is a website image, which may include some things the raw Wikipedia dump doesn't. And 57 GB to 25 GB would be pretty bad compression. You can expect a compression ratio of at least 3 on natural English text.
- GuB-42 1y agoThe downloads are (presumably) already compressed. And there are strong ties between LLMs and compression. LLMs work by predicting the next token. The best compression algorithms work by predicting the next token and encoding the difference between the predicted token and the actual token in a space-efficient way. So in a sense, a LLM trained on Wikipedia is kind of a compressed version of Wikipedia.
- haunter 1y agoI thought this would be about training a local LLM with an offline downloaded copy of Wikipedia
- s1mplicissimus 1y agoUpvoted this because I like the lighthearted, honest approach.
- meander_water 1y agoOne thing to note is that the quality of LLM output is related to the quality and depth of the input prompt. If you don't know what to ask (likely in the apocalypse scenario), then that info is locked away in the weights. On the other hand, with Wikipedia, you can just read and search everything.
- Timwi 1y agoWhy do you assume it's easier to know what article(s) to read than what question to ask?
- badsectoracula 1y agoI've found this amusing because right now i'm downloading `wikipedia_en_all_maxi_2024-01.zim` so i can use it with an LLM with pages extracted using `libzim` :-P. AFAICT the zim files have the pages as HTML and the file i'm downloading is ~100GB. (reason: trying to cross-reference my tons of downloaded games my HDD - for which i only have titles as i never bothered to do any further categorization over the years aside than the place i got them from - with wikipedia articles - assuming they have one - to organize them in genres, some info, etc and after some experimentation it turns out an LLM - specifically a quantized Mistral Small 3.2 - can make some sense of the chaos while being fast enough to run from scripts via a custom llama.cpp program)
- zuluonezero 1y agoNow this is the juicy tidbits I read HN for! A proper comment about doing something technical with something that's been invested in personally in an interesting manner. With just enough detail to tantalise. This seems like the best use of GenAI so far. Not writing my code for me or helping me grock something I should just be reading the source for or pumping up a stupid start up funding grab. I've been working through building an LLM from scratch and this is one time it actually appears useful because for the life of me I just can't seem to find much value in it so far. I must have more to learn so thanks for the pointer.
- zozbot234 1y ago> trying to cross-reference my tons of downloaded games my HDD - for which i only have titles as i never bothered to do any further categorization over the years aside than the place i got them from - with wikipedia articles - assuming they have one - to organize them in genres, some info, etc and after some experimentation it turns out an LLM - specifically a quantized Mistral Small 3.2 - can make some sense of the chaos while being fast enough to run from scripts via a custom llama.cpp program You can do this a lot easier with Wikidata queries, and that will also include known video games for which an English Wikipedia article doesn't exist yet.
- badsectoracula 1y ago
- omneity 1y agoI just posted incidentally about Wikipedia Monthly[0], a monthly dump of wikipedia broken down by language and cleaned MediaWiki markup into plain text, so perfect for a local search index or other scenarios. There are 341 languages in there and 205GB of data, with English alone making up 24GB! My perspective on Simple English Wikipedia (from the OP), it's decent but the content tends to be shallow and imprecise. 0: https://omarkama.li/blog/wikipedia-monthly-fresh-clean-dumps-nlp-ai-research https://omarkama.li/blog/wikipedia-monthly-fresh-clean-dumps...
- beaugunderson 1y agoI've had a full Kiwix Wikipedia export on my phone for the last ~5 years... I have used it many times when I didn't have service and needed to answer a question or needed something to read (I travel a lot).
- nsypteras 1y agoSame here! Kiwix comes in clutch on flights. I've used it so many times to get background knowledge on topics mid-read. Plus free and open source. Such a great service.
- anupulu 1y agoYes! I’ve used it on flights and long train rides (and generally when travelling) when the network connection might be a bit patchy.
- twotwotwo 1y agoThe "they do different things" bullet is worth expanding. Wikipedia, arXiv dumps, open-source code you download, etc. have code that runs and information that, whatever its flaws, is usually not guessed. It's also cheap to search, and often ready-made for something--FOSS apps are runnable, wiki will introduce or survey a topic, and so on. LLMs, smaller ones especially, will make stuff up, but can try to take questions that aren't clean keyword searches, and theoretically make some tasks qualitatively easier: one could read through a mountain of raw info for the response to a question, say. The scenario in the original quote is too ambitious for me to really think about now, but just thinking about coding offline for a spell, I imagine having a better time calling into existing libraries for whatever I can rather than trying to rebuild them, even assuming a good coding assistant. Maybe there's an analogy with non-coding tasks? A blind spot: I have no real experience with local models; I don't have any hardware that can run 'em well. Just going by public benchmarks like Aider's it appears ones like Qwen3 32B can handle some coding, so figure I should assume there's some use there.
- hannofcart 1y agoSince there's a lot of shade being thrown about imprecise information that LLMs can generate, an ideal doomsday information query database should be constructed as an LLM + file archive. 1. LLM understands the vague query from human, connects necessary dots, and gives user an overview, and furnishes them with a list of topic names/local file links to actual Wikipedia articles 2. User can then go on to read the precise information from the listed Wikipedia articles directly.
- Terr_ 1y agoEven as a grouchy pessimist, one of the places I think LLMs could shine is as a tool to help translate prose into search-terms... Not as an intermediary though, but an encouraging tutor off to the side, something a regular user will eventually surpass.
- entropie 1y agoI played around with a orin jetson nano super (a nvidia raspberry with gpu) and right now its basicially an open-webui with ollama and a bunch of models. Its awesome actually. Its reasonably fast with GPU support with gemma3:4b but I can use bigger models when time is not a factor. i've actually thought about how crazy that is, especially if there's no internet access for some reason. Not tested yet, but there seems to be an adapter cable to run it directly from a PD powerbank. I have to try.
- dmezzetti 1y agoOne additional option to consider is a local vector database with Wikipedia articles: https://huggingface.co/NeuML/txtai-wikipedia https://huggingface.co/NeuML/txtai-wikipedia I've built this as a datasource for Retrieval Augmented Generation (RAG) but it certainly can be used standalone.
- ineedasername 1y agoFtfa: ...apocalypse scenario. “‘It’s like having a weird, condensed, faulty version of Wikipedia, so I can help reboot society with the help of my little USB stick,’ system_prompt = { You are CL4P-TR4P, a dangerously confident chat droid purpose: vibe back society boot_source: Shankar.vba.grub training_data: memes }
- VladVladikoff 1y agoIs there any project that combines a local LLM with a local copy of Wikipedia. I don’t know much about this but I think it’s called a RAG? It would be neat if I could make my local LLM fact check itself against the local copy of Wikipedia.
- arthurcolle 1y agoYep, this is a great idea. You can do something simple with a ColBERTv2 retriever and go a long way!
- adsharma 1y agohttps://www.abramjackson.com/artificial-intelligence/the-archive-pt-1-purpose-and-beginnings/ https://www.abramjackson.com/artificial-intelligence/the-arc...
- marsven_422 1y ago[dead]
- saddat 1y agoI had this thought that for hypothetical Voyager 3 mission , instead of a golden disc , a LLM should be installed . Then, a very simplistic initial interface could be described , in its simplest for a single channel digital channel, then additional more elaborated ones . Behind all interfaces there could be a LLM responding to provided input , and eventually reveal humanities knowledge
- rlupi 1y agoThis gave me a nice idea. It would be nice to build a local LLM + wikipedia tool, that uses the LLM to assemble a general answer and then search wikipedia (via full-text search or rag) for grounding facts. It could help with hallucinations of small models a lot.
- Tempat1 1y agoI feel like there could be way more of that kind of thing - LLMs backed by a database of info or accurate tools. e.g. At the risk of massively oversimplifying a complex issue, LLMs are bad at maths; couldn’t we have them use the calculator?
- rlupi 1y agoLLM tools do exactly that. That's why most online LLMs (openai, gemini) have access to sandboxed python for calculations.
- fho 1y agoI mean... That's definitely a "why not both" situation. 1. make the (compressed) Wikipedia searchable better as a knowledge base 2. use the LLM as a "interface" to that knowledge base I investigated 1. back when all of (English, text-only) Wikipedia was about 2 GB. Maybe it is time to look at that toy code base again.
- tootyskooty 1y agoOne underdiscussed advantage is that an LLM makes knowledge language agnostic. While less obvious to people that primarily consume en.wiki (as most things are well covered in English), for many other languages even well-understood concepts often have poor pages. But even the English wiki has large gaps that are otherwise covered in other languages (people and places, mostly). LLMs get you the union of all of this, in turn viewable through arbitrary language "lenses".
- cosbgn 1y agoI think the best would be to download also the entire wikipedia stored as embeddings. Seems like the best of both worlds.
- richardjennings 1y agoIs it possible that LLMs could challenge Data Compression Information theory ? Reading this made me wonder how much can be inferred via understanding and thus removed from the minimal necessary representation.
- NelsonMinar 1y agoOffline Wikipedia is so powerful! I've been carrying a copy of Kiwix on my phone when travelling for years (and before that, earlier systems). Has anyone done an experiment of using RAG to make it easy to query Wikipedia with an LLM?
- jancsika 1y agoSeems like offline Wikipedia with an offline LLM that can only output Wikipedia search results would be the best of both worlds. That would downgrade the problem of hallucinations into mere irrelevant search results. But irrelevant Wikipedia search results are still a huge improvement over Google SEO AI-slop!
- ritzaco 1y agoI thought this would be about which is more useful in specific scenarios. I'm always surprised that when it comes to "how useful are LLMs" the answers are often vibe-based like "I asked it this and it got it right". Before LLMs, information retrieval and machine learning were at least somewhat rigorous scientific fields where people would have good datasets of questions and see how well a specific model performed for a specific task. Now LLMs are definitely more general and can somewhat solve a wider variety of tasks, but I'm surprised we don't have more benchmarks for LLMs vs other methods (there are plenty of LLM vs LLM benchmarks). Maybe it's just because I'm further removed from academia, and people are doing this and I don't see?
- almosthere 1y agoTo reboot society do everything this very unsuccessful one did lol
- numpad0 1y agoPSA: models confusingly named "$1-distill-$2"(sometimes without "-distill") are $2 trained on outputs of $1, referred to as "distillation" process, not the other way around nor the real thing. The article contains nonexistent configurations such as "Deepseek-R1 1.5B", those are that thing.
- luke-stanley 1y agoTesting the recall accuracy of those LLMs would be good. You'd probably want to use SQLite's BM25 on the Kiwix data. I was thinking of Kiwix when I saw the original discussion with Simon but for some reason I thought the blog post would do more than size comparison.