3 ms·
Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
- RIshabh235 2mo agogreat work
- vforno 2mo agoThanks for the support!
- ptsneves 2mo agoP2P inference sounds really nice and it would bring us back to the 2000s culture if not for the fact inference is already so cheap. Even so it is an insurance policy if the cloud providers or governments have ideas of further censoring and monitoring chats.
- Scaled 2mo agoWasn't a p2p ai model how skynet got started? J/k, this looks cool :)
- snovv_crash 2mo agoCool idea. How do you handle temperature in the verification?
- vforno 2mo agoHo thanks for the comment. Verification does not depend on temperature. Expert execution is deterministic (pure matmul). LUMABRI_VERIFY=N re-runs N% of the calls on a second replica and requires byte-identical output. Temperature (and sampling) happens only on the chatter, after the experts return their activations. So it can be any value (0, 0.7, 1.2…) without affecting the verification contract.
- stymaar 2mo ago> Expert execution is deterministic (pure matmul). Isn't that only true in theory but wrong in practice due to floating points?
- vforno 2mo ago[dead]
- nikunjbjj 2mo agoThe harder case isn't same-hardware determinism; it's that lumabri wants CPU and GPU peers in the same swarm. Accumulation order differs across hardware, not just across runs.
- snovv_crash 2mo agoThere are fixed-point models too which can be deterministic. But for floating point you are very instruction-set dependent, never mind floating point operations not being commutative.
- robertJk 2mo ago[dead]
- brainless 2mo agoI am sorry I did not understand all of it. But, would this allow running large MoE LLMs on a local network with experts spread out over multiple cheaper GPUs (or even CPUs)? This would perhaps be more useful than over the Internet, within offices for example.
- vforno 2mo agoThat’s one of the strongest use-cases. On a local network (office, lab, home cluster) the RTT is a few milliseconds instead of 20-50 ms, so the expert-offloading becomes much more practical. You can spread the experts across several cheaper GPUs or even CPUs, keep only the dense parts + router on the machine you’re chatting from, and the whole thing stays private inside your LAN. No internet required, no cloud, just the machines you already have.
- s2l 2mo agoWhat if one wishes to use various busybox nodes within the house? All the iot devices contributing to matmul but within a LAN?
- vforno 2mo agoBecause everything stays local and latency is tiny, even modest always-on devices can contribute. A few Raspberry Pi 5s, old mini-PCs, or stronger IoT-style boards can each hold and run a handful of experts. The protocol doesn’t care if the peer is a big GPU or a small ARM box, as long as it can load the expert weights and do the matmul. Pure busybox-class sensors are usually too limited in RAM and compute for current MoE experts, but the broader “every half-decent always-on box in the house joins the swarm” vision works well and keeps everything private inside your LAN.
- SCHiM 2mo agoThis looks really interesting, and if I understand what this does properly: it was high time someone built this! Without diving into an experiment myself, it would be amazing if you could add some stats or experiment logs, if it's not too much problem and you have them. For example: Given model XYZ, every assuming 5 donors with a uniform 32GB each, each forward pass shunts xGB over the link. Each pass takes nMS, etc. etc. Resulting in n T/s, assuming latency of n ms. Do you have such stats? Or perhaps I missed them in the repo?
- vforno 2mo ago[flagged]
- bicepjai 2mo agoGreat work. Love the idea. I have been thinking along the same way, but for training. For inference, the waiting time might be turn off for users
- vforno 2mo agoThanks really thanks for support!
- krautsauer 2mo agoOn the surface it seems similar to https://meshllm.cloud/ https://meshllm.cloud/, but now I notice that I don't understand either.
- vforno 2mo ago[dead]
- gebdev 2mo agoI’ve been thinking about this same idea recently, so I’m glad it exists now! The biggest benefit I see is to enable RAM constrained GPUs to perform inference of large parameter models with surprisingly high throughput. Because only a single expert is resident, the memory to compute ratio over the network is limited only by the activations, not the weights. For an Moe like kimi k3 where active parameters are 103B, we might expect to achieve performance limited only by ~5 effective tok/s per Tflop and ~1 tok/s per 100GB/s. The more members of the network, the smaller your resident parameters are required to be. I’m not sure whether we can split layer inference into arbitrary chunks, but if so you’d be able to increase memory throughput by storing everything in GPU caches. Of course, we expect latency to be relatively high, but that’s a tradeoff that's fine for certain circumstances. I’m not sure whether there are any issues more with this idea, but it’s a fun one nonetheless :)
- vforno 2mo ago[flagged]
- cyanydeez 2mo agoso looking at the start docs, you should integrating the same llamacpp setup where they just bake in huggingface support. second, don't use vague names of models; point to the actual repos you're testing on huggingface. There's enough diversity and specialization that even if you're smart enough to know that colibri is doing something that's particularly applicable to a type of model, it's easy to get lost in all the acronyms. third, this looks like a fun tool to unite the diversity of random hardware people have, which is always going to win for local inference.