7 ms·
I am on the vesuvius challenge team that did the segmentation, unwrapping, and ink detection, so feel free to ask any questions.
by verditelabs 4mo ago
I am on the vesuvius challenge team that did the segmentation, unwrapping, and ink detection, so feel free to ask any questions.
- helterskelter 4mo agoGiven the current rate of progress, how long do you think it will take to decipher the entire collection?
- verditelabs 4mo agoThat's a tough one to give a strong estimate of. Some scrolls are easier or harder to unwrap and read for a multitude of different reasons, mostly due to how damaged the scroll was in the eruption, and how easy or not the ink is to read. IIRC from what we've scanned of the herculaneum collection, none of the ink is easily visible via spectrum alone, so we have to use a lot of ML and physically based rendering techniques to be able to find ink. That also requires unwrapping and segmentation _before_ any ink detection. For iron gall ink with high enough iron concentration, the ink stands out in the xray volume through simply masking off low values, such as was shown in our campfire scroll experiment a few years ago. No herculaneum scrolls show similar ink.
- helterskelter 4mo agoThanks!
- pimlottc 4mo agoDo you think this particular scroll is easier or harder to read that the others will be? Or about average?
- verditelabs 4mo agoPherc1667 was quite small and just so happened to have readable ink, so it was easier than I expect most others to be.
- superjan 4mo agoDo we known what ink is used?
- verditelabs 4mo agoMost of the evidence so far points towards carbon based ink. I am not sure if any of the scrolls we have scanned show strong evidence of iron gall based ink. I know that there are different types and preparation methods for different carbon based inks, but I do not know if it is possible to determine which kind(s) were used solely from inspecting the xrays. I am, though, not a papyrologist, so historical ink making, preparation, and usage are not my field.
- junon 4mo agoThanks for answering all the questions in here. Fascinating work.
- jimbob45 4mo agoAre the fragments destroyed in ‘69 and ‘80 available to be read similarly? Or were they disposed of?
- verditelabs 4mo agoI am unaware of those fragments in particular. Though we have scanned a dozen or so fragments, mostly to help guide ink detection, since the ink in them is often more visible in visible and/or near IR light, but can be hard to impossible to detect in the xray spectrum.
- adriand 4mo agoWhat are the wildest, most exciting but plausible things that might be discovered in these documents?
- verditelabs 4mo agoI am not a papyrologist or a classicist, rather I'm a computer scientist, so my expertise is unfortunately not in _what_ the scrolls say, rather how we get there. That being said I think and hope that there will be a trove of things that has no known provenance at all, completely lost works that elude the public memory.
- arikrahman 4mo agoWell what were your first thoughts when you decoded the script, besides the obvious Eureka, after making some sense of the texts?
- tremon 4mo agoProbably something along the lines of "finally, now it looks like a coherent piece of text. I wonder what it says".
- verditelabs 4mo agoOther members that were on the team before me had already proved it out before I came along so I knew it was possible. The cool thing for me though was specifically doing some physicically based rendering techniques. How well these work varies greatly, but on a few segments in one scroll they work extremely well. I whipped up some simple code to composite layers, did up a render, and without any ML at all was looking at multiple rows of text that no one had read for 2000 years. That was neat.
- readthenotes1 4mo agoYour response reminds me of Nigel Richards :) https://en.wikipedia.org/wiki/Nigel_Richards https://en.wikipedia.org/wiki/Nigel_Richards Congratulations, and thank-you!
- echelon 4mo agoDid anyone on the team come from a non-science, non-math, non-academia background? Did anyone working on this just teach themselves and start contributing?
- verditelabs 4mo agoYes. Sean, who was a co-winner of the 2024 prize, IIRC has no formal background in ML, computer science, AI, etc. He is one of our core researchers and the most productive team member.
- fintechjock 4mo agoI've been on the Discord for a couple of years now, and poking around with submissions as well. Sean and the entire team deserve so much praise for all of this work. It's easy to just read about the breakthrough and see it as one neat, linear line to get there, and hard to comprehend the hours, months and years that so many spent to get there. Big congrats to you, Sean, Nat and the entire team!
- echelon 4mo agoThat's incredibly impressive. Major kudos to all of you on your achievements! This is amazing work for anthropology and for society, and it's greatly appreciated.
- tsol 4mo agoHow do get to do that? As in what did you study to get the prerequisite knowledge, and how did you find this particular job? When I see interesting jobs I'm anyways curious what path lead there
- verditelabs 4mo agoI am a computer scientist. I studied CS in university, worked in the semiconductor industry for a while, got started as a participant in the challenge aspect of the Vesuivus Challenge. They were hiring, I sent in an application, interviewed, and was offered the job.
- matneyx 4mo agoThat last sentence is so perfect, like my dad answering the question of how he lost weight. "I ate less and exercised more."
- tsol 4mo agoVery cool that you got in just through your interests. It's anyways cool to see stories where that works out! Good to know it's possible to get that kind of job
- inglor_cz 4mo agoI don't have any questions, just a comment. You have a potential to rewrite the history of European Antiquity quite substantially. The Herculaneum set of scrolls is enormous and must contain a lot of hitherto unknown. That comes with a set of peculiar risks. Once your work starts producing something that contradicts previous work of Very Important People, they will lobby to stop you. Be prepared for that. Science should be neutral and always value new evidence. Scientists as humans are unfortunately subject to all sorts of passions.
- Rebelgecko 4mo agoWhat contradictions do you think the scrolls contain?
- inglor_cz 4mo agoI don't have any concrete tips. We have very little written material surviving from Rome, at least from the period before a codex (book) was invented, which was more durable that a scroll. Often, we only know of one source describing important events, and when it comes to political struggles and civil wars, the perspective of the defeated party often did not survive. The punishment of damnatio memoriae was practised and even among the early emperors, Caligula and Nero were subject to a form thereof. (This library in Herculaneum was buried 11 years after Nero's death.) I would be surprised if everything in the scrolls perfectly aligned with the record that survived for 2000 years and that was filtered by both random chance and political/religious censorship. Even Christians later destroyed some pagan texts. BTW personally, I would love for some textbook of Etruscan to emerge from there. This was once again a language whose teaching was banned in Rome.
- TheOtherHobbes 4mo agoNo questions, but I just want to say this is really exciting work!
- Dzugaru 4mo agoOutstanding work! I've participated in the challenge, but didn't get far. One of the questions I had at the time was - if I'm going to use ML to detect ink, could it invent hallucinated letters, or even parts of text, and how to prevent that?
- verditelabs 4mo agoYes, it's quite possible for ML to hallucinate ink, though it is on a much more local scale, like predicting a slightly longer stroke, filling in more of a character than is actually in the data, etc. Perhaps enough to change a reading of a character or show where ink isnt. It is difficult for ink detection to hallucinate grammatical and idiomatic greek and latin.
- im3w1l 4mo agoWhat is the input to the ML algorithm? Does it know the surrounding context so that it has a chance to deduce "if this stroke is slightly longer then the end result will be idiomatic greek and latin"?
- verditelabs 4mo agoThe input is 3d chunks of reconstructed CT data from our scans. I can't remember the specifics but maybe enough voxels for .5mm^3 at a time or so? They're all available for free from https://registry.opendata.aws/vesuvius-challenge-herculaneum-scrolls/ https://registry.opendata.aws/vesuvius-challenge-herculaneum... . Our trained models are all available at https://huggingface.co/scrollprize https://huggingface.co/scrollprize
- cwnyth 4mo agoNot all machine learning is generative AI.
- mc32 4mo agoTrue but like regular document scanning software there can be errors in detection.
- BiraIgnacio 4mo agoAmazing work, fantastic!
- deleted 4mo ago[deleted]
- tomcam 4mo agoAbsolutely incredible work. This is one of the most amazing news articles I’ve encountered in decades. Congratulations team!
- temp987 4mo agothis is überragend. by many means!
- 2ap 4mo agoI'm interested to know about the approaches that you tried with the ML, and then decided to not use. In practice, the options are so many. How did you come up with the final approach - and was there a systematic way to decide which options to go for?
- verditelabs 4mo agoI am not on the research team, rather on the production side of things, so my knowledge on that is pretty limited. I think one of the main takeaways from a lot of the research, though, on both the segmentation side and the ink detection side, is that it's a lot less about what models and techniques and such you use, but how good your training data is. Gathering ground truth is hard, and if you don't have a lot of good ground truth, it doesn't matter if your code is perfect, you'll never get results.
- gekoxyz 4mo ago> it's a lot less about what models and techniques and such you use, but how good your training data is. Ah, the good old bitter lesson strikes again
- rossdavidh 4mo agoThat is a general truth of most ML; many models _can_ find the information in the data, if the data is good enough. If it is not, then likely no model can.
- EvanAnderson 4mo agoYou brought up what I'm most curious about: Where does the ground truth come from for this work since you can't just to unwrap a scroll to tell if the model got it right or, presumably, make a facsimile scroll and wrap it up.
- verditelabs 4mo agoThe ground truth comes from manual work. The scrolls can be unwrapped virtually, manually, through extensive pointing and clicking by a human on the boundaries of the scroll. This, in and of itself, is not particularly hard in sections of the scroll that are preserved well, but is extremely tedious and slow and error prone. We have a team of annotators who do manual annotation and refinement through custom software we've written, mostly improving on automatically generated segmentations and unwrappings. Once you have some unwrapped papyrus, you can render it to an image and look for ink. Ink leaves a certain texture that can be identified by the naked eye and labeled. Between these two processes you get the segmentation and ink detection ground truth. Segments can be flattened virtually through existing software and algorithms.
- NooneAtAll3 4mo agohow many scrolls have been scanned so far? what's the main limitation on scan amount? have any attempts (or just ideas) been made to recreate such charring on known texts?
- verditelabs 4mo ago30 scrolls, maybe? Something like that. I scanned Pherc Paris 4 and Pherc Paris 3 at Beam line 18 at ESRF back in March. The team did "the campfire scroll" experiment a few years ago to replicate carbonization, unrolling, and ink detection. That is the only case I am aware of. It proved the method could work but it's not a source of say training data; it varies too much from the real scrolls. The main limitation is time and cost. We have to scan on what is AFAIK the most powerful x-ray beam line in the world. It is not cheap
- CGMthrowaway 4mo agoYou had to pay? I understand the machine cost many hundreds of millions of dollars, but I would have thought for academic researchers doing open science, the beamtime is free (funded by the govt / science trusts).
- verditelabs 4mo agoThe beam time is unfortunately not free. I scanned Pherc Paris 4 and Pherc Paris 3 in March and had the final shift on the beam. As I was removing the scroll from the scanning pedestal the next team of scientists were already in the lab getting their samples ready. It's a well oiled machine and they've got customers.
- prox 4mo agoWhat other type of stuff gets scanned? I can’t imagine a whole industry waiting to x-ray something?
- CGMthrowaway 4mo ago
- deleted 4mo ago[deleted]
- negergreger 4mo agoHow fast is the process? Could it be automated to the point where it's faster to scan a book closed than opened?
- verditelabs 4mo agoWe've been trying to automate since the beginning. A lot of it is automated but it's mostly the easier and less damaged parts of the scrolls. Scanning takes a few days for the biggest scrolls but the amount of human refinement is still a multi month process.
- itsthecourier 4mo agomay you please tell us how much effort goes into each type of task in those months? where else do you think these techniques be applied?
- verditelabs 4mo agoWe are a core team of about 10 researchers and developers working full time on work that applies to all of the scrolls. We also ahve 4 full time annotators that tend to work on one scroll at a time. The amount of time spent on any given scroll varies with how difficult and large it is. There is an extremely large overlap between a lot of the work we do with medical imaging, CT scanning, XRay technology, and such. A lot of the ML models and frameworks we have used and adapted for our purposes originated in the medical field for things like cancer detection or segmenting different body parts.
- fph 4mo agoRandom shower thought: I wonder if it would be better in the long term to stop digging out archeological findings. The more we excavate, the more damage we do for future archaeologists who will have the superpower of reading these texts without even needing to dig the scrolls free and open them.
- flir 4mo ago
- nkoren 4mo agoMassive kudos to the whole team. I've been waiting 30 years for this announcement, ever since I first heard about the scrolls. Fantastic work!
- dogscatstrees 4mo agoWhat is your origin story? How did you end up doing this and how can I do the same?
- verditelabs 4mo agoBS in CS from a big state school in the USA. I have a hobby interest in history. I learned about the challenge on YouTube. Got involved contributing because I needed money. Then they put out a job posting. I applied, interviewed, and was hired.
- Refreeze5224 4mo agoWhat a cool job, and congrats on great work!
- ghghgfdfgh 4mo agoI understand that the complexity of the project has increased over the years. How difficult is it for a newcomer to get into it?
- verditelabs 4mo agoIt has gotten harder, unfortunately. One of the barriers to entry is simply the massive amounts of data; not everyone can set aside $100s worth of HDD or SSD space to play around. That said I have done a lot of work to dramatically reduce the amount of storage and bandwidth needed. We unfortunately get a lot of slop submissions, which is unfortunate. I think a _really_ good place to start is simply joining the discord and looking at the data we've published and trying to replicate something or anything really. We understand that not everyone is a researcher that can jump in making awesome immediately applicate submissions. Granted, that's pretty specifically for people that want to submit for prizes and prize money. Everyone on the team absolutely loves to talk shop and interact with real people with real interest, so if you show it in the discord we are all more than happy to help, engage, fix bugs, gvmive advice, etc. I would personally love to see more open source and contributed papyrology and translation, musing on difficult readings etc. For the more technically inclined, testing software, pointing out bugs, and actually running and trying to fix things is a huge positive that we like. We get a lot of slop submissions that are just someone pasting an issue on our GitHub into codex or Claude. We don't want to encourage that. We can do that ourselves.
- amluto 4mo agoDo you know what kinds of features the model is picking up on to distinguish ink from papyrus? And did you have any labeled data (images where a human expert has identified ink or perhaps a scan of a burnt scroll with known content) to help train it? Certainly my Mark 1 eyeballs would not obviously perform better than random guessing at this task. Although my eyeballs are, if nothing else, nerfed by only being able to see a 2D slice of the data.
- verditelabs 4mo agoYes. Most of the ink we have come across is carbon based. This leaves a certain texture on the scrolls that is recoverable and viewable with fairly basic physically based rendering, though how much ink is recoverable varies greatly from one character to the next. I don't have links handy but we just published updates to our data viewer page on our website. Pherc.Paris.4 I believe has the best overlay of ink. A lot of labeled data is available on our ftp server which has public access
- londons_explore 4mo agoI assume that's because the writer probably sometimes shortly after re-inking the writing instrument was putting down a 10x thicker layer...
- amluto 4mo agoWhen you say "physically based rendering" do you mean that one could build a PBR model based on the (unrolled?) xray data, render that model, and be able to see the ink? edit: I found this: https://scrollprize.org/data_browser#/samples/PHercParis4/segments https://scrollprize.org/data_browser#/samples/PHercParis4/se... The JSON seems to suggest that I'm mostly looking at ink detection output, but I could easily be using the tool wrong. But I also found this awesome explanation: https://scrollprize.org/data_fragments https://scrollprize.org/data_fragments I guess I bunch of the training was done by using fragments of scrolls where ground truth data is available using IR photography. Also... that xray resolution is absolutely amazing!
- 4mo ago
- eboy 4mo ago[dead]
- Izmaki 4mo agoHow awesome do you feel right now? This is HUUUGE! To think that a scroll was unreadable for so, so long, until we invented machines that let us read it slice by slice. It's such an unfathomable achievement - we made machines that let us read 2000+ year olds fragile scrolls without ever opening them - and you helped do just that. Hats off!
- verditelabs 4mo agoIn March I went to Beam Line 18 at the European Synchrotron Radiation Facility. I had to swap out the scrolls on the xray pedestal. Scrolls that were presented as a diplomatic gift to Napoleon and Josephine by King Ferdinand. France has 2 of the 6 that they were given still in tact. I had to handle both of them. I have never felt more stressed in my life and have never and will probably never again handle such a priceless artifact. I feel the opposite of that feeling and am immensely proud of everything that the core challenge team has accomplished
- _boffin_ 4mo agoI am floored at these achievements. Such amazing work. If I may ask, when you started thinking about achieving this, what were the first attempts, ideas on how to go about it? What were some of the obstacles that had to be overcome to achieve this ?
- verditelabs 4mo agoThe process of trying to read the scrolls has been going on for about 275 years or so, now. Doing it nondestructively via CT scanning and virtual unrolling and reading has been in the works for 25 years or so, so it's a lot of building on previous work. Virtual unrolling and reading are not terribly hard to do manually, they are just not feasable on a large scale. Like years and years of human time spent tediously clicking on papyrus and labelling ink in renders, so a large amount of automation is required. A lot of difficulty has come from the first step: xraying the scrolls. It's hard and expensive and difficult to get right. The efforts since this all began with CT scanning 25 years ago has been kneecapped by the data simply not being good enough. We xray on what is AFAIK literally the most powerful xray beamline in the world and we would still like for it to be more powerful and faster. Not to mention the massive amounts of data. For Pherc Paris 3, our largest scroll, the raw reconstructed data is 260 terabytes. That's a lot of data to have to deal with.
- thom 4mo agoDo we have a sense for what proportion of text is actually retrievable from these scrolls?
- verditelabs 4mo agoThat varies greatly on the state of preservation of the scroll. For some of the scrolls we can recover entire columns of text. But this is a best case. Plenty of scrolls, or portions of scrolls, are extremely damaged and warped to where our current methods cannot unroll them through any combination of automated and human driven unrolling. Both of these still have massive headroom for improvement, but achieving that headroom is hard as the preservation gets worse. To give numbers, for ideal portions of scrolls, we can read 100% of the characters. In nonideal portions of scrolls, we can read 0% of the characters. It's not really possible to quantify how much we could theoretically recover of that 0% through better methods, and how much is truly destroyed.
- ex-aws-dude 4mo agoI'm curious prior to this has there been any research/attempts at chemical methods to strengthen the structure and allow it to be unrolled?
- verditelabs 4mo agoVarious physical methods of unrolling, including the use of chemicals, have been attempted over the past 275 years, but none have proven to not destroy the scrolls. As far as I know no physical unrolling has been attempted since the 80s and I believe that now only non destructive methods are being employed. For fragments and already shattered and opened, unrolled scrolls a variety of imaging techniques exist and are still being improved upon by teams and research groups we are not associated with. For unrolled scrolls, I believe at this point no one will ever attempt physical unwrapping ever again.
- equalbeforegod 4mo agoWhere are the shattered scrolls? Is there a list somewhere?
- thatoneengineer 4mo agoImagine a worst case scenario: the Herculaneum scrolls turn out to be just the works of this one mediocre pet philosopher. What would we still expect to learn from them, and what would the next step be?
- verditelabs 4mo agoBeats me; I am a programmer, not a classicist. Though I have an interest in Old Norse and I spend a lot of time reading Scandinavian runestones. > 90% of them are grave markers for a dead father, mother, brother, sister, cousin, etc. If I've learned anything from that, it's that people across time and space all lead lives as real and complex as anyone else's. Their joys were as high as mine have been and their sorrows as low as mine have been.
- manbash 4mo agoIt's so humbling to realize that the human mind hasn't changed a lot. Only our environment.
- msuniverse2026 4mo agoHow many more scrolls exist?
- verditelabs 4mo agoThat have been dug up? I think 600 or so still exist. Perhaps about 2000 or so have ever been excavated. We have scanned about 30 of them. Still underground? I've seen various counts. Maybe more than 10000?
- quotemstr 4mo agoShame there's a modern city over most of Herculaneum. I'd love to excavate the remainder. Now that we can read what we find, there's a good scientific reason to do so now instead of waiting.
- mygooch 4mo ago[dead]
- Jeaye 4mo agoI am researching for a talk on the philosophy of code, the similarities of engineering and art, and why we enjoy reading old code. This amazing work you folks have done may be an interesting tangent. The biggest question I have for you is why you imagine we are so interested in reading these old scrolls. Surely some of it is to see whether or not, technically, we can. Surely some of it is to get a glimpse into the human expression inscribed on them. Are we looking to learn anything, or just to connect with our ancestors? I'd like to hear your take on it, both for why you think it's important and, if you know, why your colleagues feel similarly.
- verditelabs 4mo agoI wrote this as an answer to a different question but I think it applies to what you're asking as well > Though I have an interest in Old Norse and I spend a lot of time reading Scandinavian runestones. > 90% of them are grave markers for a dead father, mother, brother, sister, cousin, etc. If I've learned anything from that, it's that people across time and space all lead lives as real and complex as anyone else's. Their joys were as high as mine have been and their sorrows as low as mine have been. A VSauce video I watched a long time ago described that realization as "chronosonder". I think trying to understand those that came before us and why they made the decisions that they did given the circumstances they were in can help better inform us of the things we choose to do given our own circumstances. Otherwise, I think that a lot of things are worth doing just to see if it's possible. I like to lift weights and I'm training to lift the Dinnie Stones one day; a pair of stones that are a combined ~730 pounds. The physical and mental benefits of exercise and training are well documented and great but at the end of the day I just _really_ wanna pick up 2 stones. There's nothing more to it than that, and that's ok with me. One of the things we said a lot in 2023 was "We just wanna read the scrolls" but that slogan has unfortunately fallen a bit by the wayside as the goal and path got longer and initial hype started to fade, but I think it perfectly encapsulates why: The scrolls are there. They can be read. Why not read them?
- card_zero 4mo ago1. Why is that a realization, are there really people who say "Scandinavians are just mechanical" or "9th century people were made out of wood"? Why would their lives be assumed not to be "real", what even is that mindset? 2. "Real and complex lives" doesn't mean "just the same as ours", mind you.
- gadders 4mo agoThe science to get the text is cool, but where is the best place to read discussion of the text in the scroll, it's context, meaning etc?
- verditelabs 4mo agoThe core challenge team is focused on the technology side to provide the images of ink to our team of papyrologists and they do the transcription, translation, reading, and scholarship. This announcement was part of a larger conference being put on by Frederica Nicolardi, our lead papyrologist. The livestream of each day are available at: https://www.youtube.com/@cispemgigante/streams https://www.youtube.com/@cispemgigante/streams .
- Barbing 4mo agoMind-bending achievement from you all - thank you!
- SidewaysView 4mo ago[flagged]
- verditelabs 4mo ago5/7 trolling; not bad
- deleted 4mo ago[deleted]
- SidewaysView 4mo agoYou don't like what I have to say? Fine by me. Guess we'll see.
- verditelabs 4mo agoI can't hear you over all this cool stuff I'm discovering
- SidewaysView 4mo agoI hear Haiti has some pretty cool history too. But I'm sure you've heard it before.
- equalbeforegod 4mo agoCan you tell me how many scrolls there were to begin with, how many have been confirmed to have been lost, and how many could he salvageable if they are in broken / dessicated form? And how many there are left? Thank you so much. :)