Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
alexellisuk
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
How and Why We Bought 4x DGX Sparks
(blog.alexellis.io)
2 points
by
alexellisuk
18d ago
|
0 comments
2.
▲
by
alexellisuk
19d ago
Author here: No flashy clones of Diablo, or CoD here - just hardware deployed for local AI, in a business, with use cases explained, and how a GPU became 2x Sparks, then 4x. The switchless NCCL finding in this post is worth 1200-1500 GBP al
3.
▲
How and Why We Bought 4x DGX Sparks
(blog.alexellis.io)
3 points
by
alexellisuk
19d ago
|
1 comments
4.
▲
by
alexellisuk
26d ago
I was hoping for a whole catalogue of providers - doesn't amp have their own offering here "Orbs"? So, the interesting part for me was: "Who runs the fleet Anthropic, its own E2B, rented, third party" Like that
5.
▲
by
alexellisuk
1mo ago
This is the culmination of several weeks' worth of R&D and testing. It builds on OSS components and we credit others. Could a 200-400G switch be faster? Potentially. Sparks are compute/memory bound and we've had at least
6.
▲
Show HN: Our GLM-5.3 Flash Switchless recipe is now out for 4x DGX Sparks
(github.com)
4 points
by
alexellisuk
1mo ago
|
1 comments
7.
▲
by
alexellisuk
1mo ago
The "mia" persona on X has a specific vaguepost: https://x.com/MiaAI_lab/status/2090736338328748220?s=20 > "I've got a confirmation on what model is Ox Alpha, but I can't share it yet.
8.
▲
by
alexellisuk
2mo ago
I also very fond memories of the finger command - though mainly used on MUDs and the odd Linux host: https://blog.alexellis.io/the-90s-unix-command-fell-out-of-f... https://news.ycombinator.com/item?id=44943
9.
▲
by
alexellisuk
3mo ago
Additionally: https://www.openwall.com/lists/oss-security/2026/07/06/7
10.
▲
Januscape vulnerability CVE-2026-53359 mitigations available (KVM breakout)
(ubuntu.com)
3 points
by
alexellisuk
3mo ago
|
1 comments
11.
▲
by
alexellisuk
3mo ago
Funnily enough - I built this (delegation) over the weekend with Fable for a local voice chat running 100% on local LLMs, Parakeet and Kokoro. I say "...ask the thinking model..." and that redirects it to Qwen 3.6 27B on vLLM. Can
12.
▲
by
alexellisuk
3mo ago
Yeah, I'm surprised Justin posted this like it was new(s). Wasn't it doing the rounds on the 22nd when it launched?
13.
▲
by
alexellisuk
3mo ago
For self-hosting, have a look at what we're building with SlicerVM.com (disclosure: I'm the founder). Also runs just as well on Apple Silicon. We run quite a few Slicer instances on mini PCs and Ryzen builds - also on Hetzner (and
14.
▲
by
alexellisuk
4mo ago
This is clever work, especially given that Proxmox is already a very viable VMware replacement and wasn’t originally designed around microVMs as the primary abstraction. I’m glad this is working well for you. We’ve been on a similar journey
15.
▲
by
alexellisuk
4mo ago
Thanks for the comment ZDR is mentioned in the post - in particular many the coding plans that are not from the two major leaders have questionable IP/ownership claims on inputs/outputs :) And ZDR is still data sharing with a thir
16.
▲
by
alexellisuk
4mo ago
1. On the technical: The cache only makes generation fast, it doesn't influence what gets chosen next. The loops that hurt the most (point 2 below) are when the model re-decides to do the same thing in different words, which is much ha
17.
▲
by
alexellisuk
4mo ago
vLLM is great at continuous batching and model serving in production, but it's a very different beast and much less versatile for the prosumer category (where we sit for our usage) Dismissed is a strong term, but let me give you some m
18.
▲
by
alexellisuk
4mo ago
We did run vLLM on the 3090s — measured ~3 tok/s slower on generation for our single-to-few-user pattern, plus less flexibility on quant and slower startup (actual minutes vs single digit seconds). We may do more with it again in the f
19.
▲
by
alexellisuk
4mo ago
Fair enough, that sentence was fairly compressed. I’ve reworded it - the meaning remains the same. The post is not AI generated, I use AI for code generation and write my own articles. Which part of the post are you struggling with? This is
20.
▲
by
alexellisuk
4mo ago
I think that's quite telling Gorgi replied that he uses Qwen with 131k context. https://x.com/ggerganov/status/2067539416436867230?s=20 We also use it with 200-256k (native) context length. The issue could be
21.
▲
by
alexellisuk
4mo ago
Ha, you underestimate how dogged you need to be to get this stuff working well. The RTX 3090 in question was used from eBay, no way to return it. The RTX 6000 Pro is the "new card" in question here. The 3090s remain an interesting
22.
▲
by
alexellisuk
4mo ago
The important thing about MoEs which I mention in the conclusion is that they carry fewer (way fewer) active tokens during inference/generation. 35B-A3B is what we started out with in the days of only having the 3090, but the quality i
23.
▲
by
alexellisuk
4mo ago
One of the things I mentioned in the post: > Local models can quickly read and explain codebases, even if they can't write them - this is a superpower Might have been buried lower down. And yes latency of local on a fast card with M
24.
▲
by
alexellisuk
4mo ago
Author here. Thanks for the question. I'll answer assuming this is a question you have for me. As explained in the post - the 3090s were what were the test bed that proved the investment was worth it. Customer support, architecture rev
25.
▲
by
alexellisuk
4mo ago
Hi - the author of the post here. I wanted to write up something that was a bit more than "Qwen is the goat" or "Cancelled Claude, run everything local now" or even "The model organised my CD collection, so it'
26.
▲
Local Qwen isn't a worse Opus, it's a different tool
(blog.alexellis.io)
4 points
by
alexellisuk
4mo ago
|
2 comments
27.
▲
by
alexellisuk
4mo ago
What quant?
28.
▲
Stop driving Slicer by hand – give your agent the wheel
(slicervm.com)
3 points
by
alexellisuk
4mo ago
|
0 comments
29.
▲
by
alexellisuk
4mo ago
I was thinking about the RPi 6 yesterday whilst realising I couldn't set up my RPi Zero 2W anymore - the OS has become burdensome - tied strictly to an imager, that gives me an allergic reaction. Yes - they did all this for the uniniti
30.
▲
Look Ma No HTTP_proxy
(slicervm.com)
2 points
by
alexellisuk
4mo ago
|
0 comments
More ›