11 ms·
Using generative AI as part of historical research: three case studies
- BeefWellington 2y agoGood read on what someone in a specific field considers to have been achieved (rightly or wrongly). It does lead me to wonder how many of these old manuscripts and their translations are in the training set. That may limit its abilities against any random sample that isn't included. Then again, maybe not; OCR is one of the most worked on problems, so the quality of parsing characters into text maybe shouldn't be as surprising. Off topic: it's wild to me that in 2025 sites like substack don't apply `prefers-color-scheme` logic to all their blogs.
- isitnow 2y ago[dead]
- satisfice 2y agoThe intractable problem, here, is that “LLMs are good historians” is a nearly useless heuristic. I’m not a historian. I don’t speak old spanish. I am not a domain expert at all. I can’t do what the author of this post can do: expertly review the work of an LLM in his field. My expertise is in software testing, and I can report that LLMs sometimes have reasonable testing ideas— but that doesn’t mean they are safe and effective when used for that purpose by an amateur. Despite what the author writes, I cannot use an LLM to get good information about history.
- amelius 2y agoYou __can__ get good information from an LLM, however you just have to backtrack every once in a while because the information turned out to be false.
- userbinator 2y agohowever you just have to backtrack every once in a while because the information turned out to be false. The problem is, how do you know? I've seen developers go completely off-course just from bad search engine results and one did admit he felt something wasn't right but kept going because he didn't know better; now imagine he's being told by a very confident but incorrect LLM, and you can see how hazardous that'll be. "You don't know what you don't know."
- sebmellen 2y agoUnless you have a very good understanding of the system you’re working on or the tools you’re using, it’s very possible to get knee deep in crap without knowing it. That’s one of the biggest risks of using LLMs as assistants.
- simonw 2y agoYou need to develop skills like critical thinking, metacognition, analytical reasoning - being able to get to a robust mental model from a bunch of different inputs, some of which may even contradict each other.
- userbinator 2y agoPeople were generally already horrible at that before AI.
- deleted 2y ago[deleted]
- mvdtnz 2y agoAnd therein lies the problem - if you're not already an expert there's no way to tell when is the right moment to backtrack.
- nithril 2y agoThe exact definition of a useful heuristic, "good enough"
- simonw 2y agoThis is similar to the problem with some of the things people have been doing with o1 and o3. I've seen people share "PhD level" results from them... but if I don't have a PhD myself in that subject it's almost impossible for me to evaluate their output and spot if it makes sense or not. I get a ton of value out of LLMs as a programmer partly because I have 20+ years of programming experience, so it's trivial for me to spot when they are doing "good" work as opposed to making dumb mistakes. I can't credibly evaluate their higher level output in other disciplines at all.
- xigency 2y agoThis begs the question, is this wave of LLM AI anything more than a fancy mirror? They're certainly very good at agreeing with people and following along, but, as many have noted, not really useful for anything acting on their own.
- britch 2y agoInteresting perspective. I appreciate that it tests the models at different "layers" of understanding. I have always felt that LLMs would fall apart beyond the summarization. Maybe they would be able to regurgitate someone else's analysis. The author seems to think there's some level of intelligent creativity at play I'm hopeful that the author is right. That truly creative thinking may be beyond the abilities of LLMs and be decades away. I think the author doesn't consider the implications of broad use of LLM societally. Will people be willing to fund human historian grad students when they can get a LLM for a fraction of the price? Will prospective historians have gained the training necessary if they've used an LLM through all of school? I believe the education system could figure it out over time. I'm more worried that LLMs like this will be used as further justification to defund or halt humanities research. Who needs a history department when I can get 80% for the cost of a few chatGPT queries?
- simonw 2y agoI'd love to read way more stuff like this. There are plenty of people writing about LLMs from a computer science point of view, but I'm much more interested in hearing from people in fields like this one (academic history) who are vigorously exploring applications of these tools.
- benbreen 2y agoThank you! Have been a big fan of your writing on LLMs over the past couple years. One thing I have been encouraged by over this period is that there are some interesting interdisciplinary conversations starting to happen. Ethan Mollick has been doing a good job as a bridge between people working in different academic fields, IMO.
- dr_dshiv 2y agoI’m working with Neo-Latin texts at the Ritman Library of Hermetic Philosophy in Amsterdam (aka Embassy of the Free Mind). Most of the library is untranslated Latin. I have a book that was recently professionally translated but it has not yet been published. I’d like to benchmark LLMs against this work by having experts rate preference for human translation vs LLM, at a paragraph level. I’m also interested in a workflow that can enable much more rapid LLM transcriptions and translations — whereby experts might only need to evaluate randomized pages to create a known error rate that can be improved over time. This can be contrasted to a perfect critical edition. And, on this topic, just yesterday I tried and failed to find English translations of key works by Gustav Fechner, an early German psychologist. This isn’t obscure—he invented the median and created the field of “empirical aesthetics.” A quick translation of some of his work with Claude immediately revealed concept I was looking for. Luckily, I had a German around to validate the translation… LLMs will have a huge impact on humanities scholarship; we need methods and evals.
- quantadev 2y ago[flagged]
- rationalfaith 2y ago[dead]
- cyrillite 2y agoNow the question is how can I, someone without a PhD in history but currently a PhD candidate in another discipline, use these tools to reliably interrogate topics of interest and produce at least a graduate level understanding of them? I know this is possible, but the further away I get from my core domains, the harder it is for me to use these tools in a way that doesn’t feel like too much blind faith (even if it works!)
- kozikow 2y ago> the harder it is for me to use these tools in a way that doesn’t feel like too much blind faith (even if it works!) I tend to ask multiple models and if they all give me roughly the same answer, then it's probably right.
- aquafox 2y ago> if they all give me roughly the same answer, then it's probably right. ... or they had a lot of overlapping training data in that area.
- energy123 2y agoAlso keeping context short. Virtually all my cases of bad hallucinations with o1 have been when I've provided too much context or the conversation has been going on for too long. Starting a new chat fixes it. You can see this effect in the ARC-AGI evals, too much context impacts even o3(high).
- otabdeveloper4 2y agoOr maybe they were just trained on the same (incorrect) dataset.
- aquafox 2y agoYou ask them for references and check yourself. They are good exploratory and hypothesis generating tools, but not more. Getting a sensible sounding answer should not be an excuse for you to confirm. Often, the devil is in the details.
- 2y ago
- option 2y agoYeah, especially the ones from China /s
- deleted 2y ago[deleted]
- abathur 2y agoSomeone who knows a lot of history is a history buff. A historian works with (and may even seek out in musty rooms) primary and secondary sources to produce novel research and interpretation. An AI is at best limited to ~reading sources that human historians/archivists/librarians have already identified and digitized. Certainly value to be had here wrt to finding needles in and making sense of already-digitized historical records, but that's more like a research assistant.
- AlotOfReading 2y agoA significant amount of historical work is re-analyzing existing, known material rather than seeking out novel sources. I do know some people working in classical literature that have been testing LLMs against untranslated sources and finding them perform reasonably well. It's completely within the scope of possibility to imagine them becoming more useful for academic work over time.
- otabdeveloper4 2y ago> Certainly value to be had here I don't agree. You won't cite an LLM in an academic paper as a source (since it's unverifiable and not reproducible), and claiming than an LLM's result is your own original would be fraud. So unless you never plan on publishing anything ever, what's the point?
- yannis 2y agoYes it is more like a research assistant. The "novel research and interpretation " part is your own synthesis, deserving to be published or awarded a degree and a research assistant can save you a lot of time. As AI companies throw more money into their training data or tools become available for researchers to easily enhance this, by uploading their own data the answers will become more "accurate" and more detailed.
- throwup238 2y ago> After all (he said, pleadingly) consciousness really is an irreducible interior fortress that refuses to be pinned down by the numeric lens (really, it is!) I love this line and the “flattening of human complexity into numbers” quote above it. It sums up perfectly how I feel about the whole LLM to AGI hype/debate (even though he’s talking about consciousness). Everyone who develops a model has to jump through the benchmark hoop which we all use to measure progress but we don’t even have anything approaching a rigorous definition of intelligence. Researchers are chasing benchmarks but it doesn’t feel like we’re getting any closer to true intelligence, just flattening its expression into next token prediction (aka everything is a vector).
- voidhorse 2y agoYeah precisely. Ever since the "brain as computer" metaphor was birthed in the 50s-60s the chief line of attack in the effort to make "intelligent" machines has been to continually narrow what we mean by intelligence further and further until we can divest it of any dependence on humanist notions. We have "intelligent" machines today more as a byproduct of our lowering the bar for what constitutes intelligence than by actually producing anything we'd consider remotely capable of the same ingenuity as the average human being.
- afthonos 2y agoI find this take strange. My observation has been the opposite. We used to say it would take human intelligence to play chess. Then Deep Blue came up and we said, no, not like that. Then it was go. Then AlphaGo came up and we said no, not like that. Along the way, it was recognizing images. And then AlexNet came along, and we said no, not like that. Then it was creating art, and then LLMs came along, and we said no, not like that. I agree a narrowing has happened. But the narrowing is to move us closer to saying "if it's not implemented in a brain, located inside a skull, in a body that was developed by DNA-coded cells replicating in a controlled manner over a period of years, it's not really AI." There's an emotional attachment to intelligence being what makes us human that causes people to lose their minds when machines approach our intelligence. Machines aren't humans. If we value humanity, we should recognize that distinction—even as machines become intelligent and even sentient. And we should definitely think twice, or, you know, many many many many more times, before building intelligent machines. But I don't think pretending we're not doing that right now is helpful.
- petermcneeley 2y agoBrings a whole new meaning to "history is written by the winners"
- tolerance 2y agoThis seems to sap the intrigue out of research. But I get it. My impression of academia is antiquated. People have Jobs to do. Capital J. And this is more convenient to them. Even though I think it makes them look sort of dumb. But that’s just me and I’m not an academic anyhow. While I welcome the rise of parallel shadow institutions as civilization grows spiritlessly utilitarian, the future for common sense looks bleak.
- voidhorse 2y agoYep, we are witnessing the climactic zenith of instrumental reason operationalized and distributed on a worldwide scale. At least there's a contingent of semi-humanist thinkers left, but the number is growing worrying slim.
- pelagicAustral 2y agoI'm not sure what good will a system that only focuses on targeted truths will ever do to humanity, we already live in a world were stats are only valid if they do not offend a single person. The reason AI's are so doctored are that sometimes we just do not want to hear the truth, and we dont.
- jolmg 2y ago> explicación poética > There are, again, a couple errors here: it should be “explicación phisica” [physical explanation] not “poetic explanation” in the first line, for instance. The image seems to say "phicica" (with a "c"), but that's not Spanish. "ph" is not even a thing in Spanish. "Physical" is "física", at least today, IDK about the 1700's. So, if you try to make sense of it in such a way that you assume a nonsense word is you misreading rather than the writer "miswriting", I can see why it assumes it might say "poética", even though that makes less sense semantically.
- benbreen 2y agoAuthor here, I agree that my read may not be correct either. It’s tough to make out. Although keep in mind that “ph” is used in Latin and Greek (or at least transliterations of Greek into the Roman alphabet) so in an early modern medical context (I.e. one in which it is assumed the reader knows Latin, regardless of the language being used) “ph” is still a plausible start to a word. Early modern spelling in general is famously variable - common to see an author spell the same word two different ways in the same text.
- jolmg 2y ago> So, if you try to make sense of it in such a way that you assume a nonsense word is you misreading > I agree that my read may not be correct either Just in case, by "you", I meant from the POV of the AI, not you the author. That's interesting to know about "ph". I didn't know it was present in Latin, and I wonder if that's also the case with Spanish.
- schoen 2y agoI just looked in the Corpus Diacrónico del Español https://corpus.rae.es/cordenet.html https://corpus.rae.es/cordenet.html and it found 33 hits for "phisica" and 99 for "phisico", mostly from the 1490s. Now some of these can be deceptive, like a few are from a bilingual Spanish-Latin book and occur in the Latin portions rather than the Spanish portions, but it seems like some authors in the 1400s wrote "ph" in some Spanish words, at least when they knew the Latin or Greek etymologies. I don't know when the Iberian languages first got their more phonetic orthographies, especially suppressing that h (that was originally in Latin digraphs used to transliterate Greek letters θ, φ, χ). Edit: There are also about two dozen hits for physico/physica, interestingly more from the 1700s rather than 1400s.
- ris 2y agoStill waiting for someone to train an LLM entirely from sources written before a chosen date and be able to discuss concepts with someone apparently lacking any knowledge of the world after that date.
- dataviz1000 2y agoIn the 1950s, most people believed that the Soviets made the biggest contribution to stopping the Nazis. However, today, most people think it was actually the Americans who played the biggest role in defeating the Nazis. > "In 1945, the French public said the Soviets did the most to defeat Nazi Germany - but in 2024 they're most likely to say it was the Americans"[0] [0] https://yougov.co.uk/politics/articles/49613-d-day-anniversary-britons-disagree-with-other-countries-on-who-did-the-most-to-defeat-the-nazis https://yougov.co.uk/politics/articles/49613-d-day-anniversa...
- kranke155 2y agoThe Soviets put the men. The Americans put the materiel. Stalin thought he would’ve lost if it wasn’t for Lend Lease.
- dataviz1000 2y ago[flagged]
- AdieuToLogic 2y ago> According to ChatGPT ... ChatGPT is not an authoritative source of knowledge in any way, shape, or form. To cite it as such is folly.
- deleted 2y ago[deleted]
- kranke155 2y agoI am not any of that. Good luck for asking chatGPT for world history.
- 3willows 2y agoOn the last point, why struggle with history: Robert Nozick (in Examined Life) asked how we feel if we found out, say, Beethoven seriously composed music based on a secret formula, which is entire mechanical and required no effort for him at all. Would we still appreciate the music in the same way? If not, does our appreciation really stem from the fact that we feel he has also struggled like we do, and nevertheless produced something incredible. I remember as a very small child watching figure skaters on TV and thinking "that's no big deal". And before I started programming: "it's just logic, all very straightforward". But that was before I first entered an ice rink or centre-d a div Maybe we don't really appreciate something unless we appreciate it is hard in a visceral way.
- conception 2y agoThe real value of almost everything is based on effort I think. The best gifts aren’t the ones that are the most expensive but the ones the giver put the most effort and time into. One of the reasons I like pre-cgi is the amount of skill and the effort put into FX is astonishing. Claymation and stop motion don’t look amazing - it’s the effort. And to your point, knowing how much effort really goes into something often requires a bit of experience to really appreciate it.
- pinoy420 2y agoThe only claymation that looks good is that of aardman animation. Everything else looks absolute garbage.
- spencerflem 2y agoAardman is incredible and their polish is wonderful, but this is not true fantasic mr fox, coraline, jack stauber's opal, etc. are also very beautiful
- jfengel 2y agoThe first two are stop motion, but not claymation. The third is claymation, but I'm hard pressed to call it "beautiful". Striking, to be sure, but also conspicuously ugly. At this point Aardman is also doing a lot of non clay stop motion, but it's still the core of their work.
- grobbyy 2y agoA basic problem is they're trained on the Internet, and take on all the biases. Ask any of them so purposed edX to MIT or wrote the platform. You'll get back official PR. Look at a primary source (e.g. public git history or private email records) and you'll get a factual story. The tendency to reaffirm popular beliefs would make current LLMs almost useless for actual historical work, which often involves sifting fact from fiction.
- dmix 2y agoCouldn’t LLMs cite primary sources much the same way as a textbook or Wikipedia? Which is how you circumvent the biases in textbooks and wikipedia summaries?
- simonw 2y agoA raw LLM is a bad tool for citations, because you can't guarantee that their model weights will contain accurate enough information to be citable. Instead, you should find the primary sources through other means and then paste them into the LLMs to help translate/evaluate/etc, which is what this author is doing.
- Almondsetat 2y agoCircumventing the bias would mean providing a uniform sampling of the primary sources, which is not guaranteed to happen
- bandrami 2y agoThey can, but they also hallucinate non-existent references: https://journals.sagepub.com/doi/10.1177/05694345231218454 https://journals.sagepub.com/doi/10.1177/05694345231218454
- afinlayson 2y agoThis also means they'll be excellent at changing history for those who wish history was more aligned with their views.
- sdesol 2y agoI actually think changing history will be harder in the future as it requires alignment across models.
- esafak 2y agoWhy wouldn't people use models of their preference, just as they do news sources?
- sdesol 2y agoThey would, but I think challenging facts will be easier. Instead of saying "I heard if from blah" and not having an easy way to fact check it, you just get LLMs to challenge one another. LLMs (today's, not future versions which could be drastically different) don't have a built in goal post mover.
- afinlayson 2y agoThat assumption presumes the existence of more than one model. Currently, numerous models are being developed frequently, and I hope the price decreases. However, if a single model captures 90% of the market, the others will cease to be updated, making it easier to control. What would transpire in an authoritarian country? Would they permit a disputed border to be included in that model? If that company were acquired by the first trillionaire, could they alter history to favor any wrongdoings they committed to achieve that status? Power is accumulating, not dispersing.
- thomashop 2y agoIt feels like we've been moving in the opposite direction, where more and more models from various countries are state-of-the-art. The idea that there will be one model to rule them seems very unlikely.
- sloproth 2y agoIn my experience, these AI models haven't been great with knowledge about one specific figure (like a President). I wonder if there's a movement to start introducing these AI models to books or e-books that aren't accessible online? I wish I could be able to discuss the less publicly known details of historical figures' lives or upbringings with AI, but it's clear that more niche information that you can only read about isn't available to it.
- urbandw311er 2y agoHow would one know that the translation of the Italian text (that he gives as an example) was not just already baked into the model’s training data?
- fencepost 2y agoGood tools for translations, etc? Sure! Good historians? Ehhhhhhh. The problem is one of trust, and it's very difficult to trust the output of LLMs to be correct/true vs "truthy" without extensive verification that may be either as laborious as doing the original research or that may be difficult or impossible without knowledge and understanding of the internals and sources that may not be available.
- rgmerk 2y agoI'm no professional historian, but every time I try this kind of thing I'm very disappointed in the results. A hobby of mine is editing Wikipedia articles about Australian motorsport (yes, I have an odd hobby, sue me). The vehicles in the premier domestic auto racing category in Australia, the Supercars Championship, are unique to the category. Like NASCAR, they're built on a dedicated space frame chassis with body panels that look like either a Mustang or a Camaro draped over the top. I'd seen occasional claims on forums that when the organising body was deciding on the design of the current generation of cars, they considered using the "Group GT3" rules that are used for a bunch of racing series around the world (including the German DTM championship, the GT World Challenge events raced across Europe, Asia, and Australia, and the IMSA GTD and GTD Pro categories). If true, it might be an interesting side note to the article about the Supercars Championship. So I asked Copilot (the paid model) to find articles in motor sport media about this (there are a number of professional online publications that cover the series extensively). It confidently claimed that yes, indeed, there was some interest in using GT3 cars in the Supercars championship, and pointed me to three articles making this case. The first was an article featuring quotes from the promoter of the DTM series saying what a good idea it was to have a common car across different national series. So the first article was relevant, but didn't actually show that anyone involved in the administration of the Supercars Championship was interested in the idea. The second and third references were articles about drivers and teams whose core business is the Supercars championship also running cars in the local GT3 championship (while not explicitly mentioned in the article, they do this for a large wad of cash from the rich hobbyists who co-drive and fund most GT3 racing). Copilot's interpretation of the articles was just flat-out wrong. Yes, this was a sample size of one historical query, but its response was very poor.
- daveguy 2y ago[flagged]
- simonw 2y agoDid you read the article or are you just reacting to the headline?
- daveguy 2y agoYes, I read the article. Using these black box jumbles of weights to interpret historical documents is anathema to historical study. "I’m told that OpenAI’s newish o1 model is genuinely helpful and creative when it comes to thinking through open problems in the sciences..." "Likewise, although my knowledge of Italian is not great, I can read it well enough to confirm that the translation it offers is good enough to use for research:" Using a translation for research that you couldn't perform yourself seems extraordinarily substandard for a historian. The only part I agree with is a simple search to identify sources that may be relevant that you had not considered. i.e. Primary sources to be examined directly -- the way historians have done it for millennia. I don't think history should be filtered through model hallucinations. It seems an invitation to mistakes. The reference to a "medallion, or seat of humors, or badge of office" is obviously something being held. The historian specifically says it is not any of these, but a "urine flask." A historian obviously should not take an LLM model as factual. Later, the author writes, "After all: when you get down to it, o1 talking about a panopticon and Foucault in the above snippet is very, very similar to what a first year history PhD student might produce." This is exactly the point. A mediocre average of writing that a first year student would produce. Sure it could be used for "I hadn't considered that," but it surely should not be used for any factual interpretation.
- simonw 2y ago"Using a translation for research that you couldn't perform yourself seems extraordinarily substandard for a historian." Are you saying historians should only ever consider sources in languages they are personally fluent in?
- dartos 2y agoThis is a showcase of exactly what LLMs are good at. Handwriting recognition, a classic neural network application, and surfacing information and ideas, however flawed, that one may not have had themselves. This is really cool. This is AI augmenting human capabilities.
- DennisP 2y agoIt was pretty neat seeing this because a recent paper found that AI models are bad historians: https://techcrunch.com/2025/01/19/ai-isnt-very-good-at-history-new-paper-finds/ https://techcrunch.com/2025/01/19/ai-isnt-very-good-at-histo... But the gist of its argument just seems to be that they don't know fine details of history, and make the same generalized assumptions that humans would make with only a cursory knowledge of a particular topic. This seems unavoidable for a model that compresses a broad swath of human knowledge down to a couple hundred gigabytes. Using AI as a research tool instead of a fact database is of course a whole different thing.
- trgn 2y agoOne thing I'd love if models would get to help me confirm a thing or find the source od soemthing I have a vague memory of and which may be right or wrong, I just don't know. E.g. I have this recollection of a quote, slightly pithy, from around the 19 hundreds about hobby clubs controlling social life, maybe from Mark twain, maybe not. I just cannot come up with the prompt that gets me the answer, instead I just get hallucination after hallucination, just confirming whatever I put in, like a student who didn't study for the test and is just going along with what the professor is asking at the oral exam.
- grammarnazzzi 2y ago[dead]
- gcanyon 2y agoI wonder (hope) that for any given issue, the majority of the internet/the training data, and therefore the model's output, will be fairly near to the truth. Maybe not for every topic, but most. E.g., the models won't report that unicorns are real because the majority of the internet doesn't report that unicorns are real. Of course, there may be issues (like ghosts?) where the majority of the internet isn't accurate?
- deleted 2y ago[deleted]
- otabdeveloper4 2y ago[flagged]
- simonw 2y agoHe's a historian. He's using LLMs to assist, a little, in the evaluation of those sources.
- otabdeveloper4 2y ago[flagged]
- dang 2y agoIf you familiarize yourself with Ben's work you'll soon discover that he is no fraud. Also, on HN please don't make it look like you're quoting someone when you're not. That's an internet snark trope and we're trying for something quite different here. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.
- dang 2y ago"Please don't post shallow dismissals, especially of other people's work. A good critical comment teaches us something." https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html Edit: I think part of the problem here was the title. I've replaced it with more representative language from the article.
- Animats 2y ago"LLMs, which are exquisitely well-tuned machines for finding the median viewpoint on a given issue..." That's an excellent way to put it. It's the default mode of an LLM. You can ask an LLM for biases, and get them, of course.
- astrange 2y agoI don't think there is any reason to believe this except that everyone seems to want it to be true. An easy way to make it not be true would be to emphasize some sources in pretraining by putting them in the corpus multiple times.
- dleeftink 2y agoMaybe not 'median' but rather 'sufficiently representative', as with all distributional semantics, given a large enough corpus we can approach the 'true' distribution of word/phrases in a given language.
- krainboltgreene 2y agoExcept the corpus itself is fractional of all media. This is like saying Twitter is sufficiently representative of all human history.
- dleeftink 2y agoWe are truly dealing with immense sizes here, see [0]. I am not saying it is representative of all things, but that large corpora are sufficiently representative of words and phrases in a language (and by extension, the words and phrases assoiciated with topics and ideas). In written form you may miss some slang, pidgins and dialects, but those would be underrepresented in other media as well. [0]: https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-scaling-based-on-human-generated-data https://epoch.ai/blog/will-we-run-out-of-data-limits-of-llm-...
- miki123211 2y agoA much better way is to RLHF the LLM until you get the behavior you want. As far as I know, modern LLMs try to strike a balance between being somewhat neutral, while not being too neutral on topics outside of the overton window. They'll give you a "both sides have their good points" argument on abortion, religion, guns or immigration, but won't do that for obvious racism or nazi viewpoints. Early LLMs had a problem with getting this balance right, I feel like many of them were a lot more left-leaning. I don't know how much of the change is caused by us understanding the technology better and how much is just the political winds shifting, though. I felt like we had a moment there when some models were a bit too "well it depends", even on very uncontroversial subjects.
- fumeux_fume 2y agoAre LLMs good historians? Of course not. These types of articles always have some click/rage-bait title declaring AI supremacy at whatever task. I have used ChatGPT 4o to help translate old high German from broadsides printed in the 15-16th centuries into English and it seems to work pretty well. I don't think I'm doing serious ground-breaking research, but I feel like LLMs open doors and expand access to many things that were once completely locked without specialized knowledge.
- dang 2y agoOk, I've replaced the title above with more representative language from the article.
- p3rls 2y agoLLMs are trained on far too many reddit posts and top 12 kangs of ancient history type blogs to even contend with wikipedia for anything beyond surface level. Thucydidean-level insights? Forget it.
- astrange 2y ago"Training on" a website doesn't mean the output will agree with the website. (eg: imagine pretraining where some of the documents are prepended with a "this is a bad example" token.)
- p3rls 2y agoI just mean the format and topics. When I use o1, it's great for things like calculating square mile comparisons based on your uploaded map of tribes or translating script like the OP's example, but as far as history itself goes, like drawing inferences and making sense of primary sources-- not so much. Ask a question about anything in-depth and it feels like you're interacting with Buzzfeed and not Gibbon.
- zwischenzug 2y agoI wrote this piece in 2023, which argues similarly that LLMs are a boon, not a threat to historians https://zwischenzugs.com/2023/12/27/what-i-learned-using-private-llms-to-write-an-undergraduate-history-essay/ https://zwischenzugs.com/2023/12/27/what-i-learned-using-pri...
- dang 2y agoDiscussed here! What I learned using private LLMs to write an undergraduate history essay - https://news.ycombinator.com/item?id=38813297 https://news.ycombinator.com/item?id=38813297 - Dec 2023 (81 comments)
- adamredwoods 2y ago>> One of the well-known limitations with ChatGPT is that it doesn’t tell you what the relevant sources are that it looked at to generate the text it gives you. This isn't a limitation, this is critically dangerous. Commercial AI is a centralized, controlled, biased LLM. At what point will someone train it to say something they want people to believe? How can it be trusted? Consensus based information is still best, and I don't feel LLMs will give us that.
- delichon 2y agoOn the contrary. The heart of an LLM is a next word predictor, based on statistics. They do much the same with concepts, making them essentially consensus distillation devices. They are zeitgeisters. They get weird mainly when their training data is too sparse to find actual consensus, so instead tell you to stick cheese to your pizza with glue.
- astrange 2y ago> They get weird mainly when their training data is too sparse to find actual consensus, so instead tell you to stick cheese to your pizza with glue. That's exactly not how that happened. That happened because Google's summaries are based on their search results and one of the search results contained that.
- eviks 2y agoFor a case study would be nice if the case were actually studied… > had unusually legible handwriting, but even “easy” early modern paleography like this is still the sort of thing that requires days or weeks of training to get the hang of. Why would you need weeks of training to use some OCR tool? No comparison to any used alternatives in the article. And only using "unusually legible" isn't that relevant for the… usual cases > This is basically perfect, I’ve counted at least 5 errors on the first line, how is this anywhere close to perfection??? Same with translation: first, is this an obscure text that has no existing translation to compare the accuracy to instead of relying on your own poor knowledge? Second, what about existing tools? > which I hadn’t considered as being relevant to understanding a specific early modern map, but which, on reflection, actually are (the Peter Burke book on the Renaissance sense of the past). How? > Does this replace the actual reading required? Not at all. With seemingly irrelevant books like the previous one, yes, it does, the poor student has a rather limited time budget
- carschno 2y agoI wanted to say this, but could not express it as well. I think what your points also reveal is the biggest success factor of ChatGPT: it can do many things that specialised tools have been doing (better), but many ChatGPT users had not known about those tools. I do understand that a mere user of e.g. OCR tooling does not perform a systematic evaluation with the available tools, although it would be the scientific way to decide for one. For a researcher, however, the lack of knowledge about the tooling ecosystem seems concerning.
- deleted 2y ago[deleted]
- pjc50 2y agoDo you know any OCR tools that work on early modern English handwriting?
- conjectures 2y agoI used to work for a historical records org. As of 10 years back, OCR was getting humans to transcribe such work. So whatever the limitations of genai, my prior is against there being a perfectly good old fashioned OCR solution to the 'obscure hisotrical handwriting' problem.
- socki 2y agoWow what an incredibly interesting article. Thank you for sharing.
- tptacek 2y agoThis was so good. I'm super curious to learn more about the strategies used to set up system prompts for the custom GPT that was set up here.