3 ms·
Most definitely no offence meant, but if you are talking about gigabytes in the context of JSON, XML, or any other text based format you are doing something wro
by datalist 7y ago
Most definitely no offence meant, but if you are talking about gigabytes in the context of JSON, XML, or any other text based format you are doing something wrong. And yes, in this case I will stand by my "lack of experience" concerning your entire team, I am sorry.
However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes.
It doesnt really matter what you are using, Java was just one example. If the XML parser you employ has similar issues you must not be surprised if the outcome is similar. And you seem to be coming back over and over to software support (dependencies). Yes, particularly Java was poor when it came to that but as I said quite some time ago, that is an issue with that software not the document format.
I will disregard the 10M+ tests, but could you publish somewhere the results of the 5MB files?
Again, JSON and XML are way too similar to be anywhere close to what you described and aforementioned benchmark highlighted that. Yes, its dataset is average but I am sure you'll be able to extrapolate that for larger sets.
Apart from the apparent improper use for data of that magnitude, I could only imagine you used an XML parser that simply was not fit for the task and if you do that you shouldnt be surprised that it does not work.
- rs23296008n1 7y agoWow. I predict any kind of collaboration would involve a painful set of further interactions with little benefit. For the benefit of the probably only two others in the studio audience (who are probably currently both facepalming), we tested with multiple libraries, multiple languages, multiple OS and multiple data subsets. We found in our particular experience that JSON worked the best across our criteria using a representative sample of our datasets. Nowhere did I say XML is always the wrong choice for others. I vaguely recall I wrote I'm now none too keen on XML but have used it in the past. For some things I'd actually choose TSV over XML but thats on fairly, hopefully, obvious cases. I think XML's verbosity is actually its strength but that it has tradeoffs which are quite real. This should not come as a revelation to anyone. I shared a necessarily limited snapshot of an experience I had and an opinion I formed based on it. I think others can do their own testing as I expect they will anyway. They will confirm or deny based on what they are doing. Especially the opposite case of large imports in XML being faster than everything else. That's completely fine by me. You've definitely made too many assumptions based on too little data. You didn't even ask what industry this was for. Or what kind of data it was. Or even what disparate systems were involved such that we'd end up with something you state are inappropriately large compressed text files. You disregarded the use of "keywords" such as gigabytes or even compression in general as if those should be unimportant to us. Or why we would use JSON at all. Then you make judgements. Fairly condescending ones at that. This shows a general lack of awareness across several aspects of life in general. For the sake of both of those other people still following this chain, I'll finish here. Life is too short.
- datalist 7y agoEhm, I did not ask? I very much did so > However you havent really addressed your use case anyhow but you just threw keywords around - gigabyte, IO, compressed, etc. You might want to elaborate on where you have to use XML files of the size of 50 gigabytes. I even asked if you could provide that one 5 megabyte file. I take your response as you cant. I really have the feeling we are going in circles here and you seem to want to resort to ridicule at this point, which will make the discussion pointless. I believe I have made my point very clear from the start, elaborated more than once what my stance on this subject is, and even dug out some benchmarks. If none of that pleases you or makes you understand what I was actually trying to say, then I am terribly sorry but it is pointless. And I'd appreciate if you could point out where I was "condescending", as I would object to that, except for the "lack of experience" and I still stand by that given the information you have revealed so far.