3 ms·
First off, great to see interest in binary data in JS :) We looked into this problem years ago as part of our spreadsheet parser/writer library (open source ve
by sheetjs 8y ago
First off, great to see interest in binary data in JS :)
We looked into this problem years ago as part of our spreadsheet parser/writer library (open source version https://github.com/sheetjs/js-xlsx https://github.com/sheetjs/js-xlsx) and ultimately opted for lower-level functions that act on Buffers/ArrayBuffers/Arrays. Performance-wise, creating a new ArrayBuffer / Buffer for each field (which you are currently doing in https://github.com/francisrstokes/construct-js/blob/master/src/index.js#L20 https://github.com/francisrstokes/construct-js/blob/master/s...) is incredibly expensive. When we did performance tests years ago, allocating in 2KB blocks and manually orchestrating writes was nearly 10x faster than individual field-level allocations and concatenations both in the browser and in NodeJS.
Ironically, the ZIP file example referenced in the README alludes to another pitfall in the approach: the actual DEFLATE algorithm used in compression actually requires unaligned bit writes. See section 5.5 of the current APPNOTE.TXT for more details: https://pkware.cachefly.net/webdocs/casestudies/APPNOTE.TXT https://pkware.cachefly.net/webdocs/casestudies/APPNOTE.TXT
- FrancisStokes 8y agoThanks! When it comes to performance I can understand choosing a more low level approach. construct-js trades performance for declarativeness (though I suspsect that I can do a lot of optimising under the hood). As for DEFLATE, I'm working on an automatic bit level structure at the moment which would allow for unaligned structures. Should be in the lib in a couple of days.
- vanderZwan 8y agoIIRC, when I last checked this (admittedly this was years ago) a single TypedArray object had 200 bytes of overhead - not counting the backing buffer for the array itself - due to the complexity of the one backing buffer being potentially used by multiple arrays and so on. By comparison regular object had about a dozen bytes of overhead. Because of this, tiny typed arrays are almost always worse in terms of performance than objects or plain arrays. Especially if we are talking about "type stable" objects, that is: prototypical objects with a hidden class that does not change. Things have probably improved in Buffer/DataView/TypedArray land thanks to the push for WASM, but the allocation overhead still will be a lot higher and involve a lot more work to allocate them.