4 ms·
Parsing huge XML files with Go
- fleitz 14y agoStreaming parsers are key when dealing with XML files this big. Used to have a C# parser that would parse about 1 TB of XML per day the biggest files were > 200GB. It was impossible with out rewriting everything to use a SAX style parser.
- duaneb 14y agoAs much as I like hearing about Go, SAX parsers are not exactly new.
- willvarfar 14y agoPersonally, I have a penchant for writing my own pull parsers. Its a mind-expanding exercise. The neat thing about Go is that parsers can return functions that consume the next token. Rob Pike has an excellent video about this: http://www.youtube.com/watch?v=HxaD_trXwRE http://www.youtube.com/watch?v=HxaD_trXwRE
- pjmlp 14y agoAs in any language that supports functions as first class objects.
- signa11 14y ago> Rob Pike has an excellent video about this: http://www.youtube.com/watch?v=HxaD_trXwRE http://www.youtube.com/watch?v=HxaD_trXwRE thank you ! this is an excellent talk. having concurrent implementation of lexer & parser as co-routines communicating over message channels is very, very cool.
- exim 14y agoIn the first place, why should you have huge XML files? (Except those wikipedia dump files :))
- willvarfar 14y agoXML is often used when migrating datasets, large and small. Interchange between disparate systems is the very thing its good for.
- archangel_one 14y agoOpenStreetMap is another example that uses huge XML files. I'm not sure I really like the idea, but it does happen, and if you need the data then you have to be able to deal with it somehow even if you don't like the format.
- human_error 14y agoIt happens sometimes. I had written a multithreaded parser in C++ to parse around 800MB per day so another team could build up the rest of the project based on the data. Someone had thought it'd be a better idea to store all fetched data in XML.
- mercurial 14y agoI had to deal with relatively large XML dumps containing dictionary data.
- pradeepprabakar 14y agoI had to do a similar task of parsing the huge wikipedia dump and rewriting the Wikipedia XML (I had to add a couple of other tags to the main "page" tag) I used a SAX parser in Python and rewrote the dump. I found SAX parsers very simple to deal with huge XML Streams.