Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
Jasssss
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Show HN: I found why my coding agent got expensive and dumb, so I built a fix
(github.com)
2 points
by
Jasssss
2mo ago
|
0 comments
2.
▲
Show HN: Catcher – An AI web testing tool where most tests never hit the API
(github.com)
2 points
by
Jasssss
5mo ago
|
0 comments
3.
▲
by
Jasssss
5mo ago
Nice! Mistral 7B v0.1 is sliding_window: 4096 in the HuggingFace config.json (though v0.2 sets it to null). Gemma 2 alternates sliding window (4096) and full attention every other layer. Both have the field in the model config so maybe you
4.
▲
by
Jasssss
5mo ago
The plan command is clever. How do you handle the VRAM estimation for models with sliding window attention vs full context? Something like Mistral at 32k context uses way less KV cache than Llama at the same context length, but from the REA