Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
medicis123
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
medicis123
2mo ago
We did something similar - Streaming experts. Maintaining an expert cache, optimizing it to simulate running a multi-model agentic workflow on a 2-DGC Spark Cluster. The models we ran were: DeepSeek V4 Flash, Gemma 4 26B A4B, and Nemotron 3
2.
▲
New Inference Server for DGX Spark: large model C4:55-90 tok/s no spec decode
4 points
by
medicis123
2mo ago
|
0 comments
3.
▲
Show HN: Stop GPU pods placement getting bottlenecked by reserved VRAM
2 points
by
medicis123
7mo ago
|
0 comments
4.
▲
A New Approach to GPU Sharing: Deterministic, SLA-Based GPU Kernel Scheduling
1 points
by
medicis123
10mo ago
|
0 comments
5.
▲
Show HN: Disaggregating GPU compute from CPU in ML job execution to scale GPUs
(woolyai.com)
1 points
by
medicis123
10mo ago
|
0 comments
6.
▲
Show HN: Run PyTorch on CPU boxes, offload kernels to remote GPUs
1 points
by
medicis123
11mo ago
|
0 comments
7.
▲
Running Nvidia CUDA PyTorch container project/pipelines on AMD with no changes
1 points
by
medicis123
1y ago
|
0 comments
8.
▲
GPU-accelerated code on CPU-only environments -Remote GPU Kernel Execution
(youtube.com)
1 points
by
medicis123
1y ago
|
1 comments
9.
▲
by
medicis123
1y ago
We built this feature in our GPU Hypervisor, which cleanly separates your user-space ML environment from the GPU runtime, so you can code locally, run remotely, and execute kernels remotely on GPU hosts. Would love to hear how this impacts
10.
▲
Sharing base model in GPU VRAM across multiple inference stack process [video]
(youtube.com)
7 points
by
medicis123
1y ago
|
1 comments
11.
▲
by
medicis123
1y ago
We have just published a short demo of the WoolyAI GPU Hypervisor, showcasing VRAM memory sharing/deduplication. Load a single base model once, then run multiple isolated LoRA stacks or VLLM stacks on the same GPU. Why this matters Hig
12.
▲
by
medicis123
2y ago
Looks interesting. Will give it a try. I have been following AI startups developing various agents(Agentic phenomenon). Are you using OPenAi and any other LLM service APIs or you fine-tuned an open source model?
13.
▲
Sharing actual GPU core and VRAM utilization metrics for query on 10 LLM models
(woolyai.com)
1 points
by
medicis123
2y ago
|
1 comments
14.
▲
by
medicis123
2y ago
We ran it on WoolyAI Acceleration Service https://docs.woolyai.com/getting-started/running-your-first-... There are some interesting insights just looking at these numbers. Environment Details Wooly Client: Linux non-G
15.
▲
Show HN: WoolyAI-CUDA Abstraction Layer to Decouple Kernel Shader Exec on GPU
(woolyai.com)
4 points
by
medicis123
2y ago
|
0 comments
16.
▲
Locally delivered and centrally managed macOS envs for privileged access setup
(veertu.com)
1 points
by
medicis123
7y ago
|
0 comments
17.
▲
by
medicis123
8y ago
Wanted to share that we developed a GitLab CI Runner/executor for Anka Build and have made it public (It was done for one of our user). It basically enables you to run your iOS/macOS GitLab CI pipelines/jobs on Anka build mac
18.
▲
Shopify scaling iOS CI with Anka
(engineering.shopify.com)
1 points
by
medicis123
8y ago
|
0 comments