Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
fatihturker
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
fatihturker
3mo ago
Let me share a secret about the future: If LLMs are monoliths, SLMs are microservices. Then serverless models will come next — but SLMs will still be the backbone. I’m starting this project today. Open-source, agent-driven, built in public,
2.
▲
by
fatihturker
7mo ago
Author here. Thank you for all the good and curious comments. For 72B models, around *36GB memory works fine* by the way. I ran the benchmark and shared the results on the website: https://opengraviton.github.io/index.html
3.
▲
by
fatihturker
7mo ago
One question I'm interested in exploring: If models become heavily compressed and streamed from SSD, where do people think the real bottleneck moves to — storage bandwidth, memory bandwidth, or kernel efficiency?
4.
▲
Show HN: Efficient LLM Architectures for 32GB RAM (Ternary and Sparse Inference)
(github.com)
2 points
by
fatihturker
7mo ago
|
1 comments
5.
▲
Show HN: Run 500B+ Parameter LLMs Locally on a Mac Mini
(github.com)
17 points
by
fatihturker
7mo ago
|
10 comments
6.
▲
by
fatihturker
7mo ago
It’s inspired by ideas similar to BitNet, but I wouldn’t call it “next-gen BitNet.” BitNet focuses mainly on model representation, while OpenGraviton is about inference — pushing the limits of running large models efficiently on consumer ha
7.
▲
by
fatihturker
7mo ago
One question I’m particularly curious about: At what point does SSD bandwidth become the main bottleneck for inference when weights are heavily compressed? If anyone has experience with streaming layers or low-bit runtimes, would love to he
8.
▲
Could ternary weights make 500B models runnable on consumer hardware?
(opengraviton.github.io)
7 points
by
fatihturker
7mo ago
|
6 comments
9.
▲
by
fatihturker
7mo ago
I've been thinking about whether extreme weight compression could fundamentally change the hardware requirements for large language models. Most LLM deployments assume large GPU clusters mainly because of memory constraints (VRAM /
10.
▲
by
fatihturker
7mo ago
Author here. I'm currently working on further speed improvements — it's already around 8× faster in some cases, but there’s still potential for more optimization. Since this is an open-source project, community support is very imp
11.
▲
by
fatihturker
7mo ago
Happy to help if needed. The project is already tested and benchmarked with several models and everything is working as expected. If you run into any specific issues, feel free to open an issue or PR.
12.
▲
by
fatihturker
7mo ago
Author here. The architecture page explains how ternary quantization, dynamic sparsity, and mmap layer streaming work together to push models far beyond normal RAM limits. Happy to answer questions about the implementation or benchmarks.
13.
▲
Show HN: OpenGraviton – Run 500B+ parameter models on a consumer Mac Mini
(opengraviton.github.io)
13 points
by
fatihturker
7mo ago
|
5 comments