Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
intothemild
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
intothemild
3mo ago
No mistakes
32.
▲
by
intothemild
3mo ago
That's the thing. You can't change it. They do not accept PRs from what I've seen. Whilst it's open source it's only in name. There may be the odd PR accepted from outside the team. But very very rarely. It's h
33.
▲
by
intothemild
3mo ago
This is all lovely and I wish them the best. But please don't use ollama, or their quants. Not only is the app itself slower than pure llamacpp. But their quants are often no where near the best. I really hope people start with somethi
34.
▲
by
intothemild
3mo ago
> That is a very astute and concise way to explain everything about how the frontier labs are behaving and how they're trying to push more people to pay token rates for the best models. Are they really the best models? Like take ant
35.
▲
by
intothemild
3mo ago
Exactly.. he knows exactly what people want .. to not hear the dialogue
36.
▲
by
intothemild
3mo ago
A rare event you see a wild "HTTP I'm a little Teapot"
37.
▲
by
intothemild
3mo ago
About time. My biggest gripe with Java has always been the tooling ecosystem, almost every java engineer I know solves this with the attitude of "just use intellij". No. I won't, I refuse to use that, and I refuse to give a c
38.
▲
by
intothemild
3mo ago
Q6 can run with 256k at Q4 on 32gb easy. 200k @ K : Q5_0 V: 4_1 (which is a bit of a sweet spot)
39.
▲
by
intothemild
4mo ago
Since I started running my own inference server, I've had zero degradation that I didn't do myself. Basically the only time I see it get worse is if I drop one of the quants. Which is what I suspect the providers are doing to fit
40.
▲
by
intothemild
4mo ago
I think the majority here have stated the same... That CLAUDE.md or AGENTS.md effectively do this. Either that or the readme. The only tip I can give is that your skill that builds or wraps up work. You should have it update those files if
41.
▲
by
intothemild
4mo ago
It's that anyone who is keeping count can see that the amount of Palestinian civilians that have been killed in this war is a very very large number.
42.
▲
by
intothemild
4mo ago
> to reduce civilian casualties. I am in shock you wrote this.
43.
▲
by
intothemild
4mo ago
Well. Right now buying hardware to run your own models tops off at about 32gb VRAM at any price point that's not insane. Sure you can get a Mac mini, or a PC equivalent. But the problem is RAM. More RAM means bigger models, which means
44.
▲
by
intothemild
4mo ago
I get 50-60t/s tg on my r9700 with the dense, unsloth MTP quant UD-Q5_K_XL, K@8/V@4 256k context. Using Vulkan backend. ``` llama-server -fa on -t 7 -ngl 999 --mlock --fit off --kv-offload --no-webui --metrics --chat-template-kwar
45.
▲
by
intothemild
4mo ago
You should enable MTP now that its available. LLamaCPP has had some massive updates in the last week or so.
46.
▲
by
intothemild
5mo ago
Sure. It's just an old I7 8700 (non-k), 64gb ram. Running proxmox. But recently I put an AMD R9700 AI Pro, in there which is a 32gb inference focused card, think of it as a 32gb version of a 9070xt. All the inference happens on that ca
47.
▲
by
intothemild
5mo ago
Yes. I run local models, Qwen3.6-27B and IMHO the massive level up was the agents and skills files that I've worked on. Basically I run a flow Brainstorming > Create Spec > Review Spec* > Create Plans > Review Plan* > Ex
48.
▲
by
intothemild
5mo ago
I've spent the last month bringing in a small demo of what the future could be like, running Qwen, Gemma, and Deepseek, behind LiteLLM so we can monitor token usage, and instead of some dumb ass "tokenmaxxing" we're acti
49.
▲
by
intothemild
5mo ago
Same. Having experienced the growth of computing in those eras, the show itself had a very well researched yet very nostalgic sense of "oh yes. I'd forgotten about that".
50.
▲
by
intothemild
5mo ago
The best part of Silicon Valley was that it had a very south park quality to it.. in that things that were actually happening at the time were parodied on the show.
51.
▲
by
intothemild
5mo ago
Exactly. People love trackpoint because it's right there in the middle of the keyboard, and you don't have to move your hands. Any variation of trackpoint where you have to move your hand away from the keyboard, is a failure IMHO
52.
▲
by
intothemild
5mo ago
Well considering right now MTP support is being developed, there was a conversation in that that seemed to throw around the idea of separating the MTP model out of the main GGUF, like with Mmproj. This was rejected. Which I'm happy for
53.
▲
by
intothemild
5mo ago
There's a percentage of people who love to question how the open models were trained.. they are almost always going to try and make some argument about using the closed frontier models for distillation as some form of theft. Just total
54.
▲
by
intothemild
5mo ago
That's already happening. Qwen3.6 and Gemma4. Basically small and medium models that are crazy well trained for their sizes. Then we have a lot of specular decoding stuff like MTP and others coming to speed up responses, and finally be
55.
▲
by
intothemild
5mo ago
Don't forget to update the gguf you have too. The templates in them were updated recently too
56.
▲
by
intothemild
5mo ago
I like it, only one problem.. the fix it now types also are the same ones that didn't read anything.
57.
▲
by
intothemild
5mo ago
If only they flapped. Maybe they'd still be in the air.
58.
▲
by
intothemild
5mo ago
> We're talking about code that users can modify themselves to solve their own problems. That's it. I don't need to hear about the struggle. That's exactly the kind of attitude that this discusses. You create somethin
59.
▲
by
intothemild
5mo ago
How many pieces of flair is the minimum?
60.
▲
by
intothemild
6mo ago
I only have raw RAM, pastured RAM is wrong. I get my DRAM needs at the RAM ranch.
More ›