Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
oertl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
Techniques to beat Arrays.hashCode(byte[]) using Java's own means
(dynatrace.com)
13 points
by
oertl
1y ago
|
0 comments
2.
▲
by
oertl
2y ago
The algorithm is simple, which is why it is also suitable for textbooks. However, it is far from representing the state of the art. In case you are interested in the latest development, I have recently published a new distinct counting algo
3.
▲
ExaLogLog: Approximate distinct counting with 43% less space than HyperLogLog
(arxiv.org)
3 points
by
oertl
3y ago
|
1 comments
4.
▲
UltraLogLog: A more space-efficient alternative to HyperLogLog
1 points
by
oertl
4y ago
|
0 comments
5.
▲
Dynatrace Hash Library for Java
(github.com)
2 points
by
oertl
4y ago
|
0 comments
6.
▲
SetSketch: Filling the Gap Between MinHash and HyperLogLog
(arxiv.org)
2 points
by
oertl
6y ago
|
0 comments
7.
▲
DynaHist: A Dynamic Histogram Library for Java
(github.com)
1 points
by
oertl
6y ago
|
0 comments
8.
▲
Fast Locality-Sensitive Hash Algorithms for the (Probability) Jaccard Similarity
(arxiv.org)
2 points
by
oertl
7y ago
|
0 comments
9.
▲
by
oertl
7y ago
FYI, already in 2015, I have proposed exactly the same idea as improvement to HdrHistogram. See https://github.com/HdrHistogram/HdrHistogram/issues/54 and corresponding code https://github.com/
10.
▲
by
oertl
9y ago
Sure, send me an email.
11.
▲
by
oertl
9y ago
I have not yet checked their ML approach. Regarding register updates you are right. I have also seen that they did only 10000 simulation runs to determine the standard deviation to be 1.00/sqrt(m). I think more simulations would be nec
12.
▲
by
oertl
9y ago
The tailcut approach breaks one important property of cardinality estimation algorithms: Adding the same value multiple times should change the state of the sketch only once when adding it for the first time. However, due to the chance of o
13.
▲
by
oertl
10y ago
If you are interested in new cardinality estimation algorithms for HyperLogLog sketches you could also have a look on the paper I am currently working on: http://oertl.github.io/hyperloglog-sketch-estimation-paper/ The