Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
chaokunyang
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
61.
▲
Apache Fury serialization Framework 0.7.1 released: better compatibility
(github.com)
1 points
by
chaokunyang
2y ago
|
0 comments
62.
▲
by
chaokunyang
2y ago
The code are located at: 1. https://github.com/apache/fury/blob/main/java/fury-core/src/... 2. https://github.com/apache/fury/blob/main/java/fury-c
63.
▲
Apache Fury 0.6.0 Released: 6x faster and 1/2 payload smaller than protobuf
(fury.apache.org)
1 points
by
chaokunyang
2y ago
|
2 comments
64.
▲
by
chaokunyang
2y ago
JSON/Protobuf used a KV layout when serialization, it will write field names/types multiple times for multiple objects of same type. And the sparse layout is not friendly for CPU cache and compression. We proposed a scoped meta pa
65.
▲
by
chaokunyang
2y ago
Pretty cool, multiple level IS is great. Does KQIR support predicate pushdown?
66.
▲
Apache Fury Serialization 0.5.1 released
(github.com)
1 points
by
chaokunyang
2y ago
|
0 comments
67.
▲
by
chaokunyang
2y ago
We've already support such dict encoding, we let users register class with an id. And write class by id, id will be encoded as a varint, which uses only 1~5 bytes. But not every users like the registration. Meta string encoding here is
68.
▲
by
chaokunyang
2y ago
Meta string is used only for ascii chars, so every char is in range of a byte. We can just random access them. But your reminder is good, I just find out that we didn't check the passed string are ascii string, although our case always
69.
▲
by
chaokunyang
2y ago
Glad to see H2 database engine benifits from this. I use H2 too.
70.
▲
by
chaokunyang
2y ago
Yes, if we can. The users of Fury may be able to do this. Fury may only provide such an interface to users to let them pass such an dict/zstd/huffman or something like
71.
▲
by
chaokunyang
2y ago
Yes, we only use this encoding when it brings benifits. `wire format` is a good term to describe such cases, which the context are constrained to current message only. So the data for compression are small, and many statistical methods won&
72.
▲
by
chaokunyang
2y ago
The thing is that we can't store a precomputed dictionary. Fury is just a serialization framework, we don't know which data will be serialized
73.
▲
by
chaokunyang
2y ago
Hi, it's not “sent only once” . It's millions of “sent only once” . The thing here are that all those RPC are stateless, so the context for compression are the one message itself. i.e. compress object like `Point(1,2)` only withou
74.
▲
by
chaokunyang
2y ago
Sadly, we don't have such a training corpus of representative data. Fury is just a serialization framework, we can't assume any string distribution. I thought about scan the code of apache ofbiz, and use the domain objects in thi
75.
▲
by
chaokunyang
2y ago
Dictionary encoding is already used in fury. We will encode same string as an varint when it's seen later. The thing here is that many cases the string don't have repeat. Imagine you send `record Point(into x, int y)` in an rpc. Y
76.
▲
by
chaokunyang
2y ago
This is interesting, thanks for sharing this. I will take a deep look at it.
77.
▲
by
chaokunyang
2y ago
We have such options in Fury too. We support register classes. In such cases, such string will be encoded with a varint id. It's kind of dict encoding. But not all users like to register classes. Meta string will try to reduce space co
78.
▲
by
chaokunyang
2y ago
We can use `|` instead, `|` is not used in path. Currently we have only 2 chars can be extended. I haven't decided what to inlucde
79.
▲
by
chaokunyang
2y ago
Good suggestion, I was planing it as an optional option. We can crawl some github repos, and extract most types, extract their class names, paths, fields names, and use such data as the corpus
80.
▲
by
chaokunyang
2y ago
The thing is that we don't the frequency about every char happens
81.
▲
by
chaokunyang
2y ago
And Fury meta string are not used to compress data plane , it's used to compress meta in data plane.
82.
▲
by
chaokunyang
2y ago
Protobuf only support varint encoding, fury supports that too. But IMO, protobuf use utf-8 for string encoding. Could you share more details how protobuf compress string.
83.
▲
by
chaokunyang
2y ago
Yes, in many rpc/serialization systems, we need to serialize class name, field names, package names. So we can supprrt dynamic deserialization, type forward/backward compatibility. In such cases, those meta may took more bits tha
84.
▲
by
chaokunyang
2y ago
Meta string is not designed for encoding arbitrary strings, they are used to encode limited enumerated string only, such as classname/fieldname/packagename/namespace/path/modulename. This is why we name it as meta
85.
▲
by
chaokunyang
2y ago
We use meta string mainly for object serialization. Users passed an object to Fury, we return a binary to users. In such cases, the serialized binary are mostly in 200~1000 bytes. Not big enough for zstd to work. For dictionary encoding, F
86.
▲
by
chaokunyang
2y ago
It's not designed to beat zstd, those are used for different scenarios. zstd is used for data compression, the meta string encoding here is used for meta meta encoding. Meta data here are classname/fieldname/packagename/
87.
▲
by
chaokunyang
2y ago
Yes, totally agree. I tried gzip first, but it introduce about 10~20 bytes cost first. So I proposed such meta encoding here. If we have a much bigger string, gzip will be better, but all string we want to compress here are just classname&#
88.
▲
by
chaokunyang
2y ago
Using gzip to the whole stream will introduce extra cpu cost, we try to make out serialization implementation as fast as possible. If we apply gzip only on such small meta string, it emits bigger size. I belive some stats are written in the
89.
▲
by
chaokunyang
2y ago
We first consider using some lossless compression algorithms, but since we are just encoding meta strings such as classname/fieldname/packagename/path. There doesn't have much duplciation pattern to detect, the compressi
90.
▲
by
chaokunyang
2y ago
We tried huffman first, But we can't know which string will be serialized. If using huffman, we must collect most strings and compute a static huffman tree. If not, we must send huffman stats to peer, which the cost will be bigger
More ›