Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sleepyeldrazi
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
sleepyeldrazi
6mo ago
I haven't honestly dug around to figure out if there's a hardware reason for it, but prompt processing has always been a lot slower for me on macs in general. I mostly use MLX on my 24GB M4 Pro though, so I will pull llama.cpp on
32.
▲
by
sleepyeldrazi
6mo ago
I have been testing and using Qwen3.6 27B (running from my 3090) since it dropped and I genuinely think this is the first consumer hardware-grade model that can actually replace frontiers for a lot of workloads. I ran 8 tests on a variety o
33.
▲
by
sleepyeldrazi
6mo ago
I have been getting good results with IQ4_NL and TurboQuant at 4bits on 24gb (3090). It easily fits 256k with that setup, but it starts slowing down quite a bit after 80-100k. Quality in my testing is also still good: - Coding task test: h
34.
▲
by
sleepyeldrazi
6mo ago
I ran 3 prompts (short versions, full version in the repo): - Implement a numerically stable backward pass for layer normalization from scratch in NumPy. - Design and implement a high-performance fused softmax + top-k kernel in CUDA (or CUD
35.
▲
The only non-LLM-generated file in my repo
(github.com)
2 points
by
sleepyeldrazi
6mo ago
|
1 comments
36.
▲
by
sleepyeldrazi
5y ago
If he can comfortably move his head, might I suggest something VR related? Most of the first VR games that came out for mobile (i.e. using your phone as a VR screen) couldn't utilise a controller so they used head movement for controll