Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
snyhlxde
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
JetSpec Enables Up to 9.64x Lossless LLM Inference Speedup with Up to 1000TPS
(haoailab.com)
4 points
by
snyhlxde
3mo ago
|
1 comments
2.
▲
by
snyhlxde
3mo ago
We find speculative decoding can push LLM generation latency to extreme by co-optimizing drafting cost and drafting quality with causal parallel tree drafting. JetSpec reaches up to 9.64× end-to-end speedup on MATH-500 and 4.58× on open-end
3.
▲
by
snyhlxde
10mo ago
Today’s best LLMs mostly decode autoregressively from left-to-right, which gives great quality but is terribly slow. Diffusion LLM can decode many tokens in parallel thanks to their non-casual, any-order generation, but they must be trained
4.
▲
by
snyhlxde
10mo ago
Diffusion large language models (dLLMs) promise things that autoregressive LLMs cannot: parallel decoding, error correction, and random-order generation. Over the past year, a wave of papers has pushed this vision, and closed-source systems
5.
▲
by
snyhlxde
1y ago
Pokémon Red is also becoming a go-to benchmark for testing the agentic abilities of advanced AI models. But is Pokémon Red actually a good eval for LLMs? We study this problem in a standardized setting and identify three big issues: 1⃣ With
6.
▲
Claude-3.7 outperforms other models in realtime Super Mario Bros
(x.com)
3 points
by
snyhlxde
2y ago
|
1 comments
7.
▲
by
snyhlxde
2y ago
AI gaming agents perform surprising well in real time, and very easy to deploy We see a very scalable way of AI evaluations ahead. Check out how we did it!
8.
▲
by
snyhlxde
2y ago
Interesting work. How sensitive would the probing prompt affect certainty? i.e. replacing "Oh, I suddenly got the answer to the whole problem, Final Answer: boxed{" with some other probing text
9.
▲
by
snyhlxde
2y ago
I think it's totally possible. Multimodal reasoning eval would be fun to consider too
10.
▲
by
snyhlxde
2y ago
It's funny that reasoning models sometime speaking nonsense and perform worse than well-aligned models like claude-3.5-sonnet in multi-turn games like Akinator. I think it's one current weak point of applying longCoT RL vs. instr
11.
▲
by
snyhlxde
2y ago
When super intelligence comes, it would be very interesting to see multi-party game play among AI too. What role humans play in this story is unclear. Maybe humans can't directly engage in the games neither as they are too naive and wi
12.
▲
Evaluating LLM Reasoning Through Live Computer Games
(lmgame.org)
24 points
by
snyhlxde
2y ago
|
14 comments
13.
▲
by
snyhlxde
2y ago
Challenge yourself with latest reasoning LLMs and checkout our latest leaderboard!
14.
▲
by
snyhlxde
2y ago
Yes, we consider both domain-specific applications (spider for text2SQL, gsm8k for math, codesearchnet for python) as well as open-domain conversational applications (ShareGPT). We use test set from each application to evaluate CLLMs’ perfo
15.
▲
by
snyhlxde
2y ago
The only similarity between Medusa and CLLM is both train and adapt LLMs for fast inference. But they use completely different training technique, decoding technique and as you pointed out CLLMs don't need extra parameters or configuri
16.
▲
by
snyhlxde
2y ago
In some conversations, maybe it's easier to form complete sentences. In some others, the best we can do is: have a rough draft about what to say in mind and then refine it word by word while speaking.
17.
▲
by
snyhlxde
2y ago
A bit more intuition on how training works: in natural language processing, some phrases/collocations, for example "remind ... of ...", "make a decision", "learn a skill" etc. are used together. We can ask
18.
▲
by
snyhlxde
2y ago
lol I take that as a compliment. Good try but sadly no LLM in this writing :)
19.
▲
by
snyhlxde
2y ago
from CLLM authors: Thank you guys for the great questions and insights! We have made a Twitter posts with some more details and we invite you to engage with us on Twitter as well. https://twitter.com/haoailab/status
20.
▲
by
snyhlxde
2y ago
Thanks for interesting in our work! Yes we found training with consistency loss + AR loss on even a subset of a dataset results in significant speedup (0.01% pre-training cost). Training on more data permits even further speedup: the model
21.
▲
by
snyhlxde
2y ago
Yes this is a great question! We are actively working on supporting other sampling strategies other than greedy sampling. In the context of CLLM training, instead of mapping to a static fixed point obtained from Jacobi decoding as the train
22.
▲
by
snyhlxde
2y ago
Hi we are CLLM authors and thanks for sharing your experience and insights! I can see this drawing skill refining process echoes with the training process in CLLM, the only thing is at this point stressor in CLLM training is not getting pro
23.
▲
CLLMs: LLMs can be taught to parallel decode with up to 3.5x speedup
(twitter.com)
3 points
by
snyhlxde
2y ago
|
0 comments
24.
▲
Transforming LLMs into parallel decoders boosts inference speed by up to 3.5x
(hao-ai-lab.github.io)
7 points
by
snyhlxde
2y ago
|
0 comments