Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
t55
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
10 ms
·
121.
▲
by
t55
2y ago
Right, that's a good point. I'll adjust the intro a bit. We wanted to provide a more holistic overview on what MLA is, what came before it, and why it matters :) hope it was useful!
122.
▲
by
t55
2y ago
Thank you so much, glad you liked it
123.
▲
by
t55
2y ago
DeepSeep proposed the multi-head latent attention technique! :) As far as I know, they are the only ones using it so far
124.
▲
by
t55
2y ago
The KV cache is typically stored in a data structure external to the trained weights—often a buffer or set of tensors kept alongside the model’s forward pass (e.g., in PyTorch, one might store it in a dictionary-like container). It’s not ba
125.
▲
by
t55
2y ago
Fair point! Do you prefer a different format or tone? We really like the concise bullet point format :)
126.
▲
by
t55
2y ago
True! We love bullet points! :)
127.
▲
by
t55
2y ago
Thank you! Great question. "Infinite-length sequence processing" in StreamingLLM refers to handling much longer sequences than the model's training window (e.g., millions of tokens), by combining a sliding window for recent t
128.
▲
by
t55
2y ago
Thanks for reading! In most contexts (including this one), seq length encompasses both the initial input (prompt) tokens and the output tokens the model generates. It’s the total length of all tokens processed by the model so far.
129.
▲
DeepSeek's multi-head latent attention and other KV cache tricks
(pyspur.dev)
292 points
by
t55
2y ago
|
72 comments
130.
▲
VideoRAG: Retrieval-Augmented Generation over Video Corpus
(arxiv.org)
4 points
by
t55
2y ago
|
0 comments
131.
▲
by
t55
2y ago
I really like day-time raves
132.
▲
by
t55
2y ago
this sounds pretty cool for quick prototyping! can i see some logging/ statistics on how often my function was called and with what parameters?
133.
▲
by
t55
2y ago
Not if nobody knows!
134.
▲
Cryptoscammers Impersonated and Hacked Us – Now What?
(pyspur.dev)
7 points
by
t55
2y ago
|
2 comments
135.
▲
by
t55
2y ago
Thank you! We haven't added observability yet, is this something you would like to see?
136.
▲
by
t55
2y ago
This is excellent feedback, thank you so so much! We will simplify the Dockerfiles soon. Please let us know if we can help you get started!
137.
▲
Comparing Llama 3.2 vs. Gemma 2 vs. Mistral on philosophical questions
(old.reddit.com)
1 points
by
t55
2y ago
|
0 comments
138.
▲
by
t55
2y ago
We built several LLM-powered applications that collectively served thousands of users. The biggest challenge we faced was ensuring reliability: making sure the workflows were robust enough to handle edge cases and deliver consistent results
139.
▲
Show HN: Graph-Based Editor for LLM Workflows
(github.com)
8 points
by
t55
2y ago
|
5 comments
140.
▲
by
t55
2y ago
Sorry for that. Here's the video: https://www.youtube.com/watch?v=NIQDnWlwYyQ&ab_channel=OpenA... Concretely, they release three features inside ChatGPT's Advanced Voice Mode: 1) Stream live video (eg. with you
141.
▲
ChatGPT's Advanced Voice Mode adds Santa Mode, Live Video, Screensharing
(openai.com)
43 points
by
t55
2y ago
|
7 comments
142.
▲
Visual Autoregressive Modeling: Image Generation via Next-Scale Prediction
(openreview.net)
2 points
by
t55
2y ago
|
0 comments
143.
▲
Canvas
(openai.com)
168 points
by
t55
2y ago
|
95 comments
144.
▲
by
t55
2y ago
1. Available for Everyone 2. Python Code 3. Custom GPTs
145.
▲
Luigi Mangione's Storyline
(defenderofbasic.github.io)
2 points
by
t55
2y ago
|
0 comments
146.
▲
by
t55
2y ago
https://arxiv.org/abs/2411.02844 this paper is for you
147.
▲
How to profile CUDA kernels in PyTorch [video]
(youtube.com)
1 points
by
t55
2y ago
|
0 comments
148.
▲
Three senior Ex-DeepMind researchers about to open OpenAI's Zurich Office
(twitter.com)
7 points
by
t55
2y ago
|
0 comments
149.
▲
How to Tell Great Stories
(julian.com)
13 points
by
t55
2y ago
|
0 comments
150.
▲
by
t55
2y ago
these links don't work
More ›