Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
cubie
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
cubie
7mo ago
I'm a big fan of their work as well, good shout.
2.
▲
by
cubie
1y ago
That's awesome to hear! It's been growing a lot in the background, still useful as ever, especially for retrieval/semantic search.
3.
▲
by
cubie
2y ago
Looks very solid; I'm excited for finetuned variants for retrieval and reranking.
4.
▲
by
cubie
2y ago
Spot on
5.
▲
by
cubie
2y ago
Not yet - these are base models, or "foundational models". They're great for molding into different use cases via finetuning, better than common models like BERT, RoBERTa, etc. in fact, but like those models, these ModernBERT
6.
▲
by
cubie
2y ago
Beyond what the others have said about 1) ModernBERT-base being 149M parameters vs BERT-base's 110M and 2) most LLMs being decoder-only models, also consider that alternating attention (local vs global) only starts helping once you
7.
▲
by
cubie
2y ago
On a very high level, for NLP: 1. an encoder takes an input (e.g. text), and turns it into a numerical representation (e.g. an embedding). 2. a decoder takes an input (e.g. text), and then extends the text. (There's also encoder-decode
8.
▲
A Replacement for BERT
(huggingface.co)
348 points
by
cubie
2y ago
|
75 comments
9.
▲
Training and Finetuning Embedding Models with Sentence Transformers v3
(huggingface.co)
2 points
by
cubie
2y ago
|
0 comments
10.
▲
Embedding Quantization: 25-45x retrieval speedup, 32x or 4x less memory usage
(huggingface.co)
4 points
by
cubie
3y ago
|
0 comments
11.
▲
by
cubie
3y ago
That is exactly correct
12.
▲
by
cubie
3y ago
By "irrespective of their relevance to the language modeling task", the authors mean that the semantic meaning of the tokens is not important. These 4 tokens can be completely replaced by newlines (i.e. tokens with no semantic mea
13.
▲
Attention Sinks in LLMs for endless fluency
(huggingface.co)
12 points
by
cubie
3y ago
|
6 comments
14.
▲
by
cubie
3y ago
Various experiments on the recent Window Attention with Attention Sinks/StreamingLLM approach indicate that the approach certainly improves inference fluency of pretrained LLMs, while also improving the VRAM usage from linear to consta