Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
DARSHANFOFADIYA
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
DARSHANFOFADIYA
5mo ago
Most "sparse training" in PyTorch today isn't actually sparse. A binary mask gets multiplied into a dense weight matrix, which means the zeros still consume memory, still move through the cache, and still get multiplied. That
2.
▲
Show HN: SparseLab–real sparse training(CSR+custom kernel) in PyTorch, CPU-first
(news.ycombinator.com)
1 points
by
DARSHANFOFADIYA
5mo ago
|
1 comments
3.
▲
by
DARSHANFOFADIYA
6mo ago
Full paper: arXiv:2603.28846. Two circuits for ECDLP-256 — one at <1,200 logical qubits / 90M Toffoli gates, one at <1,450 / 70M. ~20x reduction in physical qubit requirements over prior estimates. Notably, Google withheld t
4.
▲
TurboQuant: KV Cache Quantization to 3.5 Bits with Zero Accuracy Loss- ICLR 2026
(darshanfofadiya.com)
1 points
by
DARSHANFOFADIYA
6mo ago
|
0 comments
5.
▲
Is this the right approach to go from zero to designing quantum algorithms?
(darshanfofadiya.com)
1 points
by
DARSHANFOFADIYA
7mo ago
|
0 comments
6.
▲
by
DARSHANFOFADIYA
7mo ago
As we scale to 1MN context length (inference) the biggest bottleneck is memory and to tackle that at scale we pay the price of communication overhead. Now fortunately the gpus are smartly fetching data for the next step while the previous s
7.
▲
by
DARSHANFOFADIYA
7mo ago
I've been working on optimizing training for long-context models (70B+) and found that while Tensor Parallelism is well-documented, the newer "Unified" Sequence Parallelism techniques (like DeepSpeed Ulysses) are often treate
8.
▲
Visualizing DeepSpeed Ulysses: Sequence Parallelism for 1M Context Windows
(darshanfofadiya.com)
1 points
by
DARSHANFOFADIYA
7mo ago
|
3 comments
9.
▲
Idea-Gated Transformers: open-source semantic gating trick (2025)
(arxiv.org)
1 points
by
DARSHANFOFADIYA
10mo ago
|
0 comments