4 ms·
This is interesting, but why don’t they link to the paper? Or at least give its title? The paper is «“Low-Resource” Text Classification: A Parameter-Free Class
by AlbertoGP 3y ago
This is interesting, but why don’t they link to the paper?
Or at least give its title?
The paper is «“Low-Resource” Text Classification: A Parameter-Free Classification
Method with Compressors»:
ACL Anthology, has PDF link: https://aclanthology.org/2023.findings-acl.426/ https://aclanthology.org/2023.findings-acl.426/
Arxiv, has PDF link: https://arxiv.org/abs/2212.09410 https://arxiv.org/abs/2212.09410
arXiv Vanity, HTML render of it: https://www.arxiv-vanity.com/papers/2212.09410/ https://www.arxiv-vanity.com/papers/2212.09410/
ResearchGate, has PDF link: https://www.researchgate.net/publication/366423949_Less_is_More_Parameter-Free_Text_Classification_with_Gzip https://www.researchgate.net/publication/366423949_Less_is_M...
This is Listing 1: ”Python Code for Text Classification with gzip.”
import gzip
import numpy as np
for (x1, _) in test_set:
Cx1 = len(gzip.compress(x1.encode()))
distance_from_x1 = []
for (x2, _) in training_set:
Cx2 = len(gzip.compress(x2.encode())
x1x2 = " ".join([x1, x2])
Cx1x2 = len(gzip.compress(x1x2.encode())
ncd = (Cx1x2 - min(Cx1,Cx2)) / max(Cx1, Cx2)
distance_from_x1.append(ncd)
sorted_idx = np.argsort(np.array(distance_from_x1))
top_k_class = training_set[sorted_idx[:k], 1]
predict_class = max(set(top_k_class), key=top_k_class.count)
Abstract: In this paper, we propose a non-parametric
alternative to DNNs that’s easy, lightweight,
and universal in text classification: a combi-
nation of a simple compressor like gzip with
a k-nearest-neighbor classifier. Without any
training parameters, our method achieves re-
sults that are competitive with non-pretrained
deep learning methods on six in-distribution
datasets. It even outperforms BERT on all five
OOD datasets, including four low-resource lan-
guages. Our method also excels in the few-shot
setting, where labeled data are too scarce to
train DNNs effectively. Code is available at
https://github.com/bazingagin/npc_gzip https://github.com/bazingagin/npc_gzip
Ah, this was already mentioned a few days ago in HN:
73 comments: “Gzip and KNN Outperforms Transformers on Text Classification” https://news.ycombinator.com/item?id=36707193 https://news.ycombinator.com/item?id=36707193
7 comments: “Plain old gzip+kNN outperforms BERT and other DNNs” https://news.ycombinator.com/item?id=36705472 https://news.ycombinator.com/item?id=36705472
7 comments: “A Parameter-Free Classification Method with Compressors” https://news.ycombinator.com/item?id=36707509 https://news.ycombinator.com/item?id=36707509