Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
dust42
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
by
dust42
7mo ago
You can spend every euro or dollar only once. If you consider CO2 emissions a critical problem, then you should spend every single dollar as efficiently as possible. Obviously independence of fossil fuels has a value too, as the current sit
32.
▲
by
dust42
7mo ago
I once read an article that in Berlin the sewage system is flushed with fresh water because too many people have installed water saving toilet flushers. So plenty of people bought these water savers and now the price of water has gone up be
33.
▲
by
dust42
7mo ago
I really think this is a security disaster waiting to happen, landing right in time for all the agentic terminal apps: printf '\e]8;;http://evil.com\e\\https://good.com\e]8;;\e\\\n' The next step would b
34.
▲
by
dust42
7mo ago
I have to say correctly so. It is a case of "ROMANES EUNT DOMUS" [1]. What is "Lick eggs Merz" supposed to mean? In order to be a proper political message on a proper demonstration of proper school kids, it should at lea
35.
▲
by
dust42
7mo ago
Sounds very plausible to me too. Because even if you refocus the business unit it makes no sense to lay off a highly capable team. Finding new people, integrating them into the team - all that costs a lot of time and money and there is no g
36.
▲
by
dust42
7mo ago
Cynical or not, I think it was an absolutely brilliant move: "Mass domestic surveillance of Americans constitutes a violation of fundamental rights". I think they placed their bets on Sama signing a contract with the DoD and here
37.
▲
by
dust42
8mo ago
I'd counter that at this level of capital, if the CEO doesn't well align with the capital, then super-control shares will be overpowered by super-lawyers and if there is need some super-donations. OpenAI was a public interest comp
38.
▲
by
dust42
8mo ago
Exactly. At this level you don't just put out a statement of your personal opinion. This is run through PR and coordinated with the investors. Otherwise the CEO finds himself on the street by tomorrow. Whatever their motives are, it is
39.
▲
by
dust42
8mo ago
Whenever you see a Youtube video from a restaurant kitchen you can almost be sure to see some pans where the teflon has been scrubbed down to the pure metal. Probably not that healthy...
40.
▲
by
dust42
8mo ago
I use MLX server directly from the MLX community project (by Apple). 42 tps is with 0-5000 token context. Starts to drop from there, I have never seen 60. Yesterday I tested the latest llama.cpp and the result is that PP has made a huge jum
41.
▲
by
dust42
8mo ago
320 tok/s PP and 42 tok/s TG with 4bit quant and MLX. Llama.cpp was half for this model but afaik has improved a few days ago, I haven't yet tested though. I have tried many tools locally and was never really happy with any.
42.
▲
by
dust42
8mo ago
From the article: "Ljubisa Bajic desiged video encoders for Teralogic and Oak Technology before moving over to AMD and rising through the engineering ranks to be the architect and senior manager of the company’s hybrid CPU-GPU chip des
43.
▲
by
dust42
8mo ago
I am actually doing now a good part of dev with Qwen3-Coder-Next on an M1 64GB with Qwen Code CLI (a fork of Gemini CLI). I very much like a) to have an idea how much tokens I use and b) be independent of VC financed token machines a
44.
▲
by
dust42
8mo ago
I haven't found any end-to-end voice chat models useful. I had much better results with separate STT-LLM-TTS. One big problem is the turn detection and having inference with 150-200ms latency would allow for a whole new level of qualit
45.
▲
by
dust42
8mo ago
Low latency inference is very useful in voice-to-voice applications. You say it is a waste of power but at least their claim is that it is 10x more efficient. We'll see but if it works out it will definitely find its applications.
46.
▲
by
dust42
8mo ago
> Don’t forget that the 8B model requires 10 of said chips to run. Are you sure about that? If true it would definitely make it look a lot less interesting.
47.
▲
by
dust42
8mo ago
This is not a general purpose chip but specialized for high speed, low latency inference with small context. But it is potentially a lot cheaper than Nvidia for those purposes. Tech summary: - 15k tok/sec on 8B dense 3bit quant (ll
48.
▲
by
dust42
8mo ago
Interesting hardware but I wonder if it is capable of KV caching. Thus (only) useful for applications that have short context but would benefit from very low latency. Voice-to-voice applications may be a good example.
49.
▲
by
dust42
8mo ago
Actually Pavlov did research about the digestive system for which he got the Nobel prize of medicine a few years earlier. > Did they interview the dogs and ask them if they actively and consciously decide to produce saliva? > Is "
50.
▲
by
dust42
8mo ago
"It showed that dogs process information from their environment and use it to make predictions" Exactly that is not what the experiment is about because we all know that dogs will quickly learn the connection between bell and food
51.
▲
by
dust42
8mo ago
Interesting stuff but it hurts so much that the writer has the common misconception of pavlov's dog doing a circus trick. Sure the dog also consciously understands the connection between bell and food. But the physiological reaction of
52.
▲
by
dust42
8mo ago
> I can only conclude that you're not using these things in any kind of practical way. I burn about 100M tokens per month. LLMs are like knives, the outcome of cooking depends on the cook and for 99% of purposes not on the knife. Th
53.
▲
by
dust42
8mo ago
They are all just token generators without any intelligence. There is so little difference nowadays that I think in a blind test nobody will be able to differentiate the models - whether open source or closed source. Today's meme was t
54.
▲
by
dust42
8mo ago
Not even pocket change compared to the billions of VC money burnt every month to keep the show running.
55.
▲
by
dust42
8mo ago
Same experience here with Whisper, medium is often not good enough. The large-turbo model however is pretty decent and on Apple silicon fast enough for real time conversations. The addition of the prompt parameter can also help with transcr
56.
▲
by
dust42
8mo ago
The brand new Qwen3-Coder-Next runs at 300Tok/s PP and 40Tok/s on M1 64GB with 4-bit MLX quant. Together with Qwen Code (fork of Gemini) it is actually pretty capable. Before that I used Qwen3-30B which is good enough for some qui
57.
▲
by
dust42
8mo ago
Well, he destroyed Gawker. Not that I think they were good people. But it was definitely a personal vendetta.
58.
▲
by
dust42
8mo ago
There are now quite a few cases in Europe where the EU or local govs been de-banking individuals. No court, no judge needed. Much more efficient way to shut down critics. We ain't need no people who delegitimize those in power.
59.
▲
by
dust42
8mo ago
He's Chinese and if you had looked into his comment history you'd know this is not someone who uses LLMs for karma farming and looking at his blog he has a long history of posting about database topics going back before there was
60.
▲
by
dust42
8mo ago
It is the buffer implementation. [u1 10kTok]->[a1]->[u2]->[a2]. If you branch between the assistant1 and user2 answers then MLX does reprocess the u1 prompt of let's say 10k tokens while llama.cpp does not. I just tested with
More ›