Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
lhl
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
lhl
1y ago
Apple actually makes a lot more acquisitions than you think, but they are rarely very high profile/talked about: https://en.wikipedia.org/wiki/List_of_mergers_and_acquisitio...
32.
▲
by
lhl
1y ago
For fungible things, it's easy to cost out. But not all things can be broken down just in token cost, especially as people start building their lives around specific models. Even beyond privacy just the availability is out of your cont
33.
▲
by
lhl
1y ago
You're right on ratios, but actually the ratio is much worse than 6:1 since they are MoEs. The 20B has 3.6B active, and the 120B has only 5.1B active, only about 40% more!
34.
▲
by
lhl
1y ago
A few people have mentioned looking a the vLLM docs and blog (recommended!). I'd also recommend SGLang's docs and blog as well. If you're interested in a bit of a deeper dive, I can highly recommend reading some of what DeepS
35.
▲
by
lhl
1y ago
In Linux, you can allocate as much as you want with `ttm`: In 4K pages for example: options ttm pages_limit=31457280 options ttm page_pool_size=15728640 This will allow up to 120GB to be allocated and pre-allocate 60GB (you cou
36.
▲
by
lhl
1y ago
Slimbook (Spanish OEM) basically sells the same ODM designs as Tuxedo and is an option. They have a few cobranded (KDE etc) versions that contribute to the development teams. Otherwise at this point I believe the Framework laptops have pret
37.
▲
by
lhl
1y ago
Strix Halo does not run a 70B Q6 dense model at real-time conversational speed - it has a real-world MBW of about 210 GB/s. A 40GB Q4 will clock just over 5 tok/s. A Q6 would be slower. It will run some big MoEs at a decent spee
38.
▲
by
lhl
1y ago
You can go much lower: https://gpulist.ai/
39.
▲
by
lhl
1y ago
For anyone interested in tracking max achievable matmul FLOPS for hardware and unaware, I highly recommend tracking Stas Bekman's mamf-finder results: https://github.com/stas00/ml-engineering/tree/master&
40.
▲
by
lhl
1y ago
On the corp side you have FB w/ PyTorch, xformers (still pretty iffy on AMD support tbt) and MS w/ DeepSpeed. But let's see about some others: Flash Attention: academia, 2y behind for AMD support bitsandbytes: academia, 2y be
41.
▲
by
lhl
1y ago
Last year I had issues using MI300X for training, and when it did work, was about 20-30% slower than H100, but I'm doing some OpenRLHF (transformers/DeepSpeed-based) DPO training atm w/ latest ROCm and PyTorch and it seems to
42.
▲
by
lhl
1y ago
Yeah, I think personalized evals will definitely be a thing. Besides reviewing way too much Arena, WildChat and having now seen lots of live traces firsthand, there's a wide range of LLM usage (and preferences), which really don't
43.
▲
by
lhl
1y ago
For the Claude issues, I'm referring to the claude.ai frontend. While I use some Codex, Aider, and other agentic tools, I found Claude Code to be not to my taste - for my uses it tended burn a lot of tokens and gave relatively medioc
44.
▲
by
lhl
1y ago
I've been using o3 extensively since release (and a lot of Deep Research). I also use a lot of Claude and Gemini 2.5 Pro (most of the times, for code I'll let all of them go at it and iterate on my fav results). So far I've o
45.
▲
by
lhl
1y ago
Yeah, actually for my batch usage, I usually push to 256+ concurrency, but on H100s at least, currently 64-128 is about the bend of the curve for where latency starts going out of control (this depends a lot on your context length and kvcac
46.
▲
by
lhl
1y ago
No, let's add it. The cost for an inference provider to deploy a trained and weights available existing model is $0 (or whatever you want to add for the HF download of the weights). Open weight models simply exist now. Deal with it? If
47.
▲
by
lhl
1y ago
Rather than speculating another option is to just measure things. I churned through billions of tokens for evals and synthetic data earlier this year, so I did some of that. On an H100 node, a Llama3 70B FP8 at concurrency=128 generated at
48.
▲
by
lhl
1y ago
On my shotengai there are many cash only shops and some cashless shops right next to it. I've also seen a lot of shops that are cash except for PayPay (presumably incentivized, or maybe they can support it without additional hardware).
49.
▲
by
lhl
1y ago
Well I've seen some other writeups like https://atadistance.net/2020/06/13/transit-gate-evolution-wh... that have also been referenced on HN. Discussions like https://news.ycombinator.com/
50.
▲
by
lhl
1y ago
It depends on the vendor and whether they are willing to pay for global licensing. For Garmin devices for example, only the APAC version have NFC-F support.
51.
▲
by
lhl
1y ago
This is a great writeup and reminded me of some others I've seen in the past. For those interested on the topic, I used Deep Research to generated a report on turnstile/ticketing systems compared to others like Hong Kong, Taiwan,
52.
▲
by
lhl
1y ago
I think it's about equal for utility - Japanese Suica/Pasmo cards are also usable in every single konbini, at all the stations stores, across most regional transportation and taxis, and at maybe half of Tokyo shops/restaurant
53.
▲
by
lhl
1y ago
Here's some discussion here: https://www.reddit.com/r/LocalLLaMA/comments/1jzocoo/finally... Ollama appears to not properly credit llama.cpp: https://github.com/ollama/ollama&#x
54.
▲
by
lhl
1y ago
Basically anything llama.cpp (Vulkan backend) should work out of the box w/o much fuss (LM Studio, Ollama, etc). The HIP backend can have a big prefill speed boost on some architectures (high-end RDNA3 for example). For everything els
55.
▲
by
lhl
1y ago
Maybe of interest, I built and open-sourced a similar (web-based) end-to-end voice project last year for an AMD Hackathon: https://github.com/lhl/voicechat2 As a submission for an AMD Hackathon, one big thing is that I
56.
▲
by
lhl
1y ago
While Llama 4 had a pretty bad launch (the LM Arena gaming in particular is terrible), having run my own evals on it (using the April 5 v0.8.3 vLLM release - https://blog.vllm.ai/2025/04/05/llama4.html , so b
57.
▲
by
lhl
1y ago
This is just someone's personal blog/opinion. I wouldn't read too much into it... "The site is run by Zygmunt Zajc (pronounced “Ziontz”). ... An economist by education"
58.
▲
by
lhl
2y ago
Since no one specifically answered your question yet, yes, you should be able to get usable performance. A Q4_K_M GGUF of DeepSeek-R1 is 404GB. This is a 671B MoE that "only" has 37B activations per pass. You'd probably expec
59.
▲
by
lhl
2y ago
I think this framing isn't quite right either. DeepSeek's R1 isn't very different from what OpenAI has already been doing with o1 (and that other groups have been doing as well). As for distilling - the R1 "distilled&quo
60.
▲
by
lhl
2y ago
It's great to see vLLM getting faster/better for DeepSeek. I tested vLLM vs SGLang a couple weeks ago and SGLang's DeepSeek support was much better/faster (on 2 x p5 H100 nodes). It's great that no one's standi
More ›