Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
OsamaJaber
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Agent loops stop re-reading your codebase on every turn
(runinfra.ai)
2 points
by
OsamaJaber
1mo ago
|
0 comments
2.
▲
GLM 5.3 Flash faster and cheaper
(runinfra.ai)
1 points
by
OsamaJaber
1mo ago
|
0 comments
3.
▲
The fastest and cheapest GLM 5.3 Flash endpoint
(runinfra.ai)
3 points
by
OsamaJaber
1mo ago
|
0 comments
4.
▲
DeepSeek V4 Pro at 207 tok/s with the full 1M context, no quantization
(runinfra.ai)
3 points
by
OsamaJaber
2mo ago
|
0 comments
5.
▲
Fastest Inference in MENA
(runinfra.ai)
2 points
by
OsamaJaber
2mo ago
|
0 comments
6.
▲
by
OsamaJaber
2mo ago
soon :)
7.
▲
by
OsamaJaber
2mo ago
you can give it a try
8.
▲
DeepSeek V4 Flash at 278 tok/s, full precision, no quantization
(runinfra.ai)
10 points
by
OsamaJaber
2mo ago
|
4 comments
9.
▲
The fastest full-precision Nemotron 3.5 Lightning endpoint (540 tok/s)
(runinfra.ai)
5 points
by
OsamaJaber
2mo ago
|
0 comments
10.
▲
by
OsamaJaber
2mo ago
The comparison set is Gemma4-31B and Qwen3.6-27B, not the current Qwen Fair on size, but the headline numbers are against a model a generation back
11.
▲
With software alone, one B200 beats the LPU and gets close to Cerebras
(runinfra.ai)
11 points
by
OsamaJaber
2mo ago
|
1 comments
12.
▲
by
OsamaJaber
2mo ago
Expanding free access is mostly an inference cost Serving cheaply at that scale means routing, batching, and cache hits, not a better model :D
13.
▲
Kimi K3 2.78T on One CPU with 8GB RAM
(github.com)
4 points
by
OsamaJaber
2mo ago
|
0 comments
14.
▲
Mixture-of-Kittens: An MoE training megakernel for NVL72
(twitter.com)
3 points
by
OsamaJaber
2mo ago
|
0 comments
15.
▲
Lossless Inference
(runinfra.ai)
4 points
by
OsamaJaber
2mo ago
|
0 comments
16.
▲
DeepSeek V4 Flash 2.98x faster, lossless
(runinfra.ai)
1 points
by
OsamaJaber
2mo ago
|
1 comments
17.
▲
DeepSeek-V4-Flash 2.98x faster on 4x B200, lossless
(twitter.com)
3 points
by
OsamaJaber
2mo ago
|
0 comments
18.
▲
What LLM Inference Costs
(twitter.com)
2 points
by
OsamaJaber
2mo ago
|
0 comments
19.
▲
Kimi k3 run on RTX 5090
(github.com)
4 points
by
OsamaJaber
2mo ago
|
0 comments
20.
▲
Kimi k3 now runs on one consumer GPU
(twitter.com)
8 points
by
OsamaJaber
2mo ago
|
0 comments
21.
▲
by
OsamaJaber
3mo ago
The gap I've hit generating GPU kernels with agents code that compiles and runs fine but is slower than the baseline. Validator says pass, result is useless. Speed targets have to be part of the check, not just correctness
22.
▲
The AGI Compiler "Auto"
(github.com)
2 points
by
OsamaJaber
3mo ago
|
0 comments
23.
▲
Compiler for LLMs, world models, and AGI
(arxiv.org)
2 points
by
OsamaJaber
3mo ago
|
0 comments
24.
▲
Auto: The AGI Compiler
(github.com)
4 points
by
OsamaJaber
3mo ago
|
2 comments
25.
▲
by
OsamaJaber
3mo ago
thread: https://x.com/Akashi203/status/2074495867449434389
26.
▲
Build You Own Model
(runinfra.ai)
2 points
by
OsamaJaber
3mo ago
|
0 comments
27.
▲
RunInfra: Optimize any open model down to the kernel, deploy in 5 min
(runinfra.ai)
2 points
by
OsamaJaber
3mo ago
|
0 comments
28.
▲
Compiles any HuggingFace model into a single persistent megakernel
(twitter.com)
2 points
by
OsamaJaber
4mo ago
|
0 comments
29.
▲
Mega Kernels, Written by Agents
(arxiv.org)
2 points
by
OsamaJaber
4mo ago
|
0 comments
30.
▲
AutoMegaKernel: Compiling a LLM into a single CUDA kernel
(arxiv.org)
3 points
by
OsamaJaber
4mo ago
|
0 comments
More ›