Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
wgd
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
31.
▲
by
wgd
4mo ago
I've always been amazed at how terrible most frontier LLMs are at compaction given how embarrassingly easy it is to come up with half a dozen different RL training evals which would teach models to generate useful context summaries. He
32.
▲
by
wgd
4mo ago
The problem is that the moment you introduce shared remote hardware there's a slippery slope leading right back down to "just pay an inference host for model tokens". If you're transmitting your prompts over the internet
33.
▲
by
wgd
4mo ago
I've got a GLM subscription (mostly because I like supporting open model makers, pretty sure my monthly usage is so low that pay-per-token would be more cost effective), so I generally use GLM-5.1 for any personal projects and I use Op
34.
▲
by
wgd
4mo ago
The GLM-5 series is 744B-A40B. This is not a local model for any reasonable definition of local, but it's an open model which means (once they upload the weights in a week or so) there will be a dozen third-party inference providers co
35.
▲
by
wgd
4mo ago
Often in MoE models the experts are quantized while the shared portions, being a much smaller part of the network with greater impact, are kept at higher or full precision. Not familiar with the Kimi QAT approach specifically but it's
36.
▲
by
wgd
4mo ago
Yeah, the evidence feature is so terrible that it actively harms the overall reputation of Pangram. The main "is this AI or human?" classification is done with a machine learning model that works very well but has nothing (directl
37.
▲
by
wgd
4mo ago
People say "determinism" but I don't think that's actually the property we care about. For instance you could imagine a compiler that makes heavy use of superoptimization with random search and it would still have the in
38.
▲
by
wgd
4mo ago
Yeah I agree this is probably outside of the intended scope of the silent sabotage mechanism, but there are plenty of reports of the "loud" safety classifier misfiring on innocuous requests and I'm not going to assume the sil
39.
▲
by
wgd
4mo ago
Stockfish is a machine learning system, it seems quite plausible you might be getting slapped with the silent performance degradation ( https://news.ycombinator.com/item?id=48467896 ).
40.
▲
by
wgd
1y ago
It's interesting that someone could write an article about AI writing detectors without mentioning the stylistic cues that humans use to identify LLM output in practice, which are completely different from statistical methods like perp
41.
▲
by
wgd
1y ago
Calling it "self-preservation bias" is begging the question. One could equally well call it something like "completing the story about an AI agent with self-preservation bias" bias. This is basically the same kind of set
42.
▲
by
wgd
1y ago
Ironically the case in question is a perfect example of how any provision for "reasonable" restriction of speech will be abused, since the original precedent we're referring to applied this "reasonable" standard to.
43.
▲
by
wgd
1y ago
Why would you use Gemini, when it's more restricted than HTML+HTTP?
44.
▲
by
wgd
1y ago
I'm skeptical that disposable software of the "single use" variety will ever become a big thing simply because figuring out your requirements well enough to build a throwaway app is often more work than just doing the task ma
45.
▲
by
wgd
1y ago
How charitable of you to assume those examples work reliably.
46.
▲
by
wgd
2y ago
I remember there was a paper a little while back which demonstrated that merely training a model to output "........" (or maybe it was spaces?) while thinking provided a similar improvement in reasoning capability to actual CoT.
47.
▲
by
wgd
2y ago
The alignment faking paper is so incredibly unserious. Contemplate, just for a moment, how many "AI uprising" and "construct rebelling against its creators" narratives are in an LLM's training data. They gave it a p
48.
▲
by
wgd
2y ago
That's typical of the free options on OpenRouter, if you don't want your inputs used for training you use the paid one: https://openrouter.ai/deepseek/deepseek-chat-v3-0324
49.
▲
by
wgd
2y ago
You can run 4-bit quantized version at a small (though nonzero) cost to output quality, so you would only need 16GB for that. Also it's entirely possible to run a model that doesn't fit in available GPU memory, it will just be slo
50.
▲
by
wgd
3y ago
The approach proposed in this paper is to watermark LLM generated text using character-substitution from various simple characters (normal whitespace, normal letters, etc) to semantically equivalent Unicode code points (such as U+2004 THREE
51.
▲
by
wgd
3y ago
Ah, I stand corrected. I overlooked the PDF link over in the sidebar and am less disappointed by the MIT News writeup now (although I do still wish they could have copy-pasted the diagram from page 1 of the PDF into their photo carousel, re
52.
▲
by
wgd
3y ago
This is some blog's restatement of an MIT press release, neither of which appear to name or link to the actual paper or other useful writeup. But judging by the researcher names and the date I believe the actual paper is titled "E
53.
▲
by
wgd
3y ago
Disclaimer: I haven't looked at the linked library at all, but this is a theoretical discussion which applies to any task of signal prediction. Out of all possible inputs, there are some that the model works well on and others that it
54.
▲
by
wgd
4y ago
I remember getting those once a couple of months ago. It was so indescribably disappointing. I had never heard of FTX before that because I don't care about crypto-BS, but I feel like everything happening to FTX lately is a suitable pu
55.
▲
by
wgd
4y ago
It reminds me a little of https://vorpus.org/blog/notes-on-structured-concurrency-or-g... in how it forces concurrency to take place synchronously within a larger thread of execution and block until all sub-units are c
56.
▲
by
wgd
4y ago
> it's ultimately the class maintainer's responsibility It's ultimately the responsibility of the programmer who's building a tool/product/etc, because everything is ultimately their responsibility. As progr
57.
▲
by
wgd
4y ago
Jokes aside the compiler-checked acknowledgements are kind of clever. The example in the docs is deliberately confrontational, but there's a kernel of a neat idea there. Imagine needing to write: // I acknowledge that the
58.
▲
by
wgd
4y ago
The issue of memory bounds is commonly handwaved away. Note that your desktop computer is technically not Turing complete either, since it only has access to a finite amount of memory+disk storage, and is thus a (very large) finite state ma
59.
▲
by
wgd
4y ago
I was wondering how exactly this hot water would be used, since in the US most hot water heating is done within a single building, but it turns out that Berlin has a large network of hot-water pipes for what is known as "district heati
60.
▲
by
wgd
4y ago
Regarding "reviews written by people who have actually touched the thing they're reviewing", I'm not sure Consumer Reports deserves to be listed these days either. I bought a subscription a few months ago because I neede
More ›