Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aesthesia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
7 ms
·
121.
▲
by
aesthesia
4mo ago
Hallucination rate scores are a little tricky to interpret because they're conditional on the model not knowing the answer. That means they don't measure the probability of your encountering a hallucination in everyday use, since
122.
▲
by
aesthesia
4mo ago
One reason to use XML-like formatting is that it makes the beginning and end of sections explicit. This is less of an issue when the model is generating text but can still be helpful when using templated prompts.
123.
▲
by
aesthesia
4mo ago
You can see their general approach to guardrail classifiers in these posts: https://www.anthropic.com/research/constitutional-classifier... https://www.anthropic.com/research/next-generation-consti
124.
▲
by
aesthesia
4mo ago
Disappointing they don't actually say how their sparse attention mechanism works.
125.
▲
by
aesthesia
4mo ago
From the linked docs page: > Requests go directly from your app to the Claude API; Apple is not in the request path and does not see prompts or responses. Usage is billed to your Anthropic account at standard API pricing. Your app decide
126.
▲
by
aesthesia
4mo ago
LLM-isms are much less prevalent in base models, which is what GPT-2 was. It had significant problems with maintaining coherence, but GPT-2 generated text did not have the obvious tells of today's LLMs.
127.
▲
by
aesthesia
4mo ago
They can certainly enforce that you answer the survey. But it's very difficult to enforce a requirement that people answer questions accurately, particularly when they perceive that doing so will expose them to danger.
128.
▲
by
aesthesia
4mo ago
I'm skeptical that you're going to be able to reliably exfiltrate ~10TB of model weights using TEMPEST. Which is not to say weights are secure, just that this isn't the threat model I would be concerned about.
129.
▲
by
aesthesia
4mo ago
This is not legislation.
130.
▲
by
aesthesia
4mo ago
Come on, no one was worried that GPT-2 would help people engineer viruses. The concern was generating misinformation and spam.
131.
▲
by
aesthesia
4mo ago
Moolenaar's quote: "The AI models these companies use are trained by China’s censorship regime and introduce hidden vulnerabilities that put Americans’ data and businesses at risk." That is, Americans using Chinese-trained AI
132.
▲
by
aesthesia
4mo ago
Word was originally released for the Mac in 1985, so the deal was not that Office would be ported, just that MS would keep developing Office for the Mac.
133.
▲
by
aesthesia
4mo ago
This is a neat little trick, but I wonder if you could do substantially the same thing by just prompting/LoRA finetuning the model to produce a single-token output ("yes" or "no"). This only requires a single model
134.
▲
by
aesthesia
4mo ago
I appreciate the extremely low fuss interface, but I'm always a little disappointed by chord progression ear training that just plays triads one after another with no thought for voice leading. Generating a nice voice leading for an ar
135.
▲
by
aesthesia
4mo ago
Using only 3/2 ratios can sound pretty bad in just intonation as well. Major thirds tuned to 81/64 are off (by a ratio of 81/80) compared with the standard 5/4 tuning, and they don't sound great. This difference is
136.
▲
by
aesthesia
4mo ago
If you really want to see fully open training pipelines for modern LLMs, Olmo and to a lesser extent Nemotron are what you should look at. https://github.com/allenai/OLMo https://github.com/NVIDIA-NeMo&
137.
▲
by
aesthesia
4mo ago
> They are asking for FAA style preclearance and third party audits. That literally means no new AI startup can emerge. Do they not know that audits cost money? Training frontier AI models costs money, orders of magnitude more than third
138.
▲
by
aesthesia
4mo ago
Yep, in their analysis depreciation meant "get no useful work out of the GPU after this point," though.
139.
▲
by
aesthesia
4mo ago
One of the main purposes of model cards, from the beginning, has been to outline the ways that a model could be harmful or dangerous, and mitigations that can be or have been taken to reduce those risks. How do you expect labs to publish mo
140.
▲
by
aesthesia
4mo ago
Oh, just noticed one other very significant error: they evaluate revenue using input token pricing while counting capacity using generated tokens per second. There's a big gap between input and output token pricing, and between prefill
141.
▲
by
aesthesia
4mo ago
There are some glaring local errors that make this analysis less than trustworthy. For instance, an assumption that corporate income tax applies directly to revenue, or a supposedly generous assumption that GPUs will fully depreciate after
142.
▲
by
aesthesia
4mo ago
I mean, they do actually describe what that extra work was, and people elsewhere in this thread are complaining about the effects of those safeguards. So it's not like this is purely empty rhetoric.
143.
▲
by
aesthesia
4mo ago
Because it's not a user manual? The idea of a model card originated in 2018 (see https://arxiv.org/abs/1810.03993 ) as a summary of important facts about a model. At the time, this was typically an image classifier
144.
▲
by
aesthesia
4mo ago
Not quite an answer to your question, but you might find this interesting. The Olmo Hybrid paper has some results on relative complexity of problems that can be solved by transformers and RNNs. They don't look at size, just solvability
145.
▲
by
aesthesia
4mo ago
Sorry, I should have said this explicitly in the original comment: I think you're likely _correct_ that there isn't a clear increase in the rate of bugs attributable to LLM-authored code in rsync. Your analysis provides evidence i
146.
▲
by
aesthesia
4mo ago
I don't have a dog in this fight, but a few points that look a little suspicious: - The release with the highest number of attributed bugs is the release _right before_ the first release with Claude-coauthored commits, released in Janu
147.
▲
by
aesthesia
4mo ago
LLM inference can be implemented in a way where nondeterminism depends only on the random seed, but that's not common. It ends up being more efficient/easier to implement kernels whose exact results depend on how many other prompt
148.
▲
by
aesthesia
4mo ago
Audio is 1 dimensional so the usual RoPE position encoding should handle it like it does for text tokens. You only need extra position encoding for higher-dimensional stuff like images.
149.
▲
by
aesthesia
4mo ago
What's interesting is that although they don't seem to be releasing the model weights, they have published a technical report ( https://microsoft.ai/wp-content/uploads/2026/06/main_2026060... ) t
150.
▲
by
aesthesia
4mo ago
Worth looking at the followup post that evaluates the current version of Mythos, which solves one of the main tasks that GPT-5.5-Cyber does not. https://www.aisi.gov.uk/blog/how-fast-is-autonomous-ai-cyber...
More ›