Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
npodbielski
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
npodbielski
16d ago
When I changed the number of draft tokens to 3 in both, it helped and they Draft is actually performing a bit better: - draft: 67.17 - MTP: 64.18 Why they used those examples? Seems strange.
2.
▲
by
npodbielski
16d ago
Which was not he point because I was testing their solution for MPT and it was just funny addition. But of course in internet you always will find some 'well akchually' person straight from the meme.
3.
▲
by
npodbielski
16d ago
Well I tested it on 7900XTX with the same prompts and their draft model gave me about 30t/s. Their own snippet of code with regular MTP model gave me 60t/s. Also model with their draft answered incorrectly. With MTP it answered c
4.
▲
by
npodbielski
16d ago
I am using Flash Next for few weeks and it is very capable model. I just wish there would a way to have faster prefill because reloading longer sections of session sometimes can take even 2h. I stopped using Qwen 3.8 27B completely on my St
5.
▲
by
npodbielski
16d ago
Yes, for fun I tried OVH AI Endpoint and they do not have cache read at all. They bill you every time you send a prompt regardless if you are hit cache or not. One agent session was like 80M input and 300K output and I paid 30$ for that. Or
6.
▲
by
npodbielski
20d ago
what would be the usecase? sharing account with kids? it is easier to just install something else and register on throwaway email?
7.
▲
by
npodbielski
24d ago
Or two gorgon halos?
8.
▲
by
npodbielski
26d ago
Looks like really great tool to generate some graphs and diagrams for static file blogs.
9.
▲
by
npodbielski
1mo ago
Wow. I just read couple, but... That seems terrible! People do that? On the other hand my own father believes now he is an alien from outer space because someone generated stupid youtube video...
10.
▲
by
npodbielski
1mo ago
I am running it on 32GB and I did not saw model loosing it context even after 4-5 compactions in pi. I am running sessions for few days sometimes. I think it looped once, but loop police extension stopped it. The only problem I have know is
11.
▲
by
npodbielski
1mo ago
I bought several cameras and they unvr early this year. It works great but I have no AI features because you have to buy they hardware just to accept some license (!). I mean... OK but no. If they would allow me to send them email saying: &
12.
▲
by
npodbielski
1mo ago
What about Raspberry Pi is not open?
13.
▲
by
npodbielski
1mo ago
As you said: everything works on llama.cpp Why it does not work on vllm? Of course you can say that it is AMD fault but there was an issue of abysmal performance of models on Strix Halo, that is open for half a year ( https://gith
14.
▲
by
npodbielski
2mo ago
And it fails on rocm of course. This engine is such a hassle on AMD.
15.
▲
by
npodbielski
2mo ago
Anybody was able to run this model in a server?
16.
▲
by
npodbielski
2mo ago
If it is not hard, with an experience of this guy, he could came up with a better example? It is your own words so I am sure you will agree? He did alright job so he could exercise gray matter a bit and came up with some API for saving cust
17.
▲
by
npodbielski
2mo ago
It spits out password in logs according to gif. No thank you.
18.
▲
by
npodbielski
2mo ago
What I do not like about that kind of example is its abstractiveness. Yes sure you can argue about testing `doSomething()` and it all falls apart when there is an actual business scenario to test.
19.
▲
by
npodbielski
2mo ago
It is. I am running it on R9700
20.
▲
by
npodbielski
2mo ago
On the other hand I am running this model to write some tests for my hobby project for two days now and it is able to deduce and fix errors and bugs that Qwen 3.6 was not able to. Yes, it thinks a lot but this makes reasoning about problem
21.
▲
by
npodbielski
2mo ago
This is not true. As author points out his domain is invalid unless you install something. So the whole thing is misleading.
22.
▲
by
npodbielski
2mo ago
What a terrible website to open on your phone. Text and image is clipped, it jumps up and down, loads for several seconds, refreshes entire page several times... I opened it and have no idea what this hardware supposed tonbe.
23.
▲
by
npodbielski
2mo ago
so it is just llamafile with few additional application in bundle?
24.
▲
by
npodbielski
2mo ago
Yes, the best way to land yourself a good gig is to have a friend in the company already and will put a good word.
25.
▲
by
npodbielski
2mo ago
In my case I would say they are comparable but moe models are looping and getting lost a lot more than dense models. On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lo
26.
▲
by
npodbielski
2mo ago
How people do such nice animation? Is there is some nice program you can use? I looked and I could not find anything that seemed really easy to use. Llms can sometimes generate something usefull and sometimes something attrocious. I remembe
27.
▲
by
npodbielski
2mo ago
And it is fun!
28.
▲
by
npodbielski
2mo ago
Anybody have good recommendation of a watch that you can write your own apps for?
29.
▲
by
npodbielski
2mo ago
Yeah makes everyone agree on banning Chinese models and then force them to use mine models. How I could misunderstand *that*.
30.
▲
by
npodbielski
2mo ago
very addictive :) thanks!
More ›