4 ms·
Ouch. They're probably encoding in XML and several nested layers of base64!! (yes I have actually seen people doing this in a finance company's integration pl
by csmuk 13y ago
Ouch.
They're probably encoding in XML and several nested layers of base64!!
(yes I have actually seen people doing this in a finance company's integration platform).
- waterlesscloud 13y agoTweets come with a ton of metadata these days.
- csmuk 13y agoYeah but 32x the original data size...
- jzwinck 13y agoThey're storing all the metadata, and for better or worse they are storing each string field as an "object" whereas they could have used a fixed-width field. The former of course is something like a reference count and a pointer to a string elsewhere; the latter would have some overhead when strings are of varying lengths. If memory usage were really a concern, decent savings could be realized by changing (most of?) the 13 float64 columns to float32, and some of the object (string) columns to fixed-width. For example, "lang" is two characters always, and usernames (which are stored in several fields) would be more efficiently stored as char[15]. Easy to try, just not the default behavior; I bet you could cut the total size by a third (too bad they didn't).
- csmuk 13y agoIck. I wonder if anyone has actually heard of ASN.1 BER before? Seems like ridiculous encoding and structures are the in thing at the moment.
- jzwinck 13y agoHave you used Numpy/Scipy/Pandas? It works a bit differently. It basically creates a matrix as in C--the trick is choosing the right column types. I don't see what relevance ASN.1 has here.
- csmuk 13y agoYes SciPy (although I usually end up with Mathematica).That's a generic encoding. ASN.1 is a specific and compact encoding of the data.