Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sigmar
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
sigmar
3mo ago
Qualified it with "100%" because claude4 models show the first few lines of the chain of thought: >On Claude 4 models, the first few lines of thinking output are more verbose, providing detailed reasoning that's particular
32.
▲
by
sigmar
3mo ago
Fable/mythos are the first models from anthropic that hide 100% of reasoning tokens. So it seems to me like we're about to get a lot more data about to what extent Chinese model progress has been a consequence of distillation tech
33.
▲
by
sigmar
3mo ago
>Some administration officials have said that a resolution should include an acknowledgment on Anthropic’s part that its rollout of Fable and communication with the White House could have been improved, people familiar with the talks sai
34.
▲
by
sigmar
3mo ago
I should have contextualized the quote- "chat is dead" is from an openai employee which was describing how they're shifting focus to more agentic consumer products, and putting less focus on the back-and-forth chatbot interfa
35.
▲
by
sigmar
3mo ago
I like that "chat is dead" framing I heard recently because too many people are having interpersonal relations with these LLMs and want to tune their "emotions"/tone. Humanity would be in a better place if we though
36.
▲
by
sigmar
3mo ago
All science-based conclusions come with uncertainty. Only ideologues (and siths) write in absolute terms.
37.
▲
by
sigmar
3mo ago
Wacky decision. I know they want ads but it's an endurance sport! Imagine if they chopped a sport like 10km run into quarters.
38.
▲
by
sigmar
3mo ago
Did you read the blog post where they explained why there was a temporary block on all biology-related questions?
39.
▲
by
sigmar
4mo ago
Who foots the bill?
40.
▲
by
sigmar
4mo ago
I'd love to see Anthropic (or someone with mythos access) create a cybersecurity version of this. So that I could create a pool that says "find security concerns in this github repo." Then the report from mythos gets sent to
41.
▲
by
sigmar
4mo ago
Agree with this. Strange to me to frame the "training recall" as cheating (33 of the 38 cheating instances). Most people think of "cheating" as breaking rules. How is the LLM model supposed to not use what was put into t
42.
▲
by
sigmar
4mo ago
>As long as Codex remains so affordable and useful they do not have to slash prices, just keep Codex usable. I imagine they track usage and can see whether their habitual users are switching to something else and aren't going to sla
43.
▲
by
sigmar
4mo ago
It's temporary. From the fable blogpost: >To release the model both safely and quickly, we’ve tuned these safeguards conservatively—they’ll sometimes catch harmless requests, though they trigger, on average, in less than 5% of sessi
44.
▲
by
sigmar
4mo ago
The system card is 319 pages, at what point do we call it a "book" instead of a "card"? There's a quote from a METR report on page 52: >We ran [Mythos 5] on 38 of our hardest software tasks, including tasks cente
45.
▲
by
sigmar
4mo ago
I think it is intended to sound like Sam Altman. Would look exactly like a tweet of his if it didn't have capitalized characters.
46.
▲
by
sigmar
4mo ago
Not unexpected. Is anyone tracking which episode we're on in the Pantheon timeline?
47.
▲
by
sigmar
4mo ago
>So what context would cause me to seriously consider the possibility that engineers had created a computer program that is conscious and an intentional user of language? Let me outline one potential sequence of steps. The first requirem
48.
▲
by
sigmar
4mo ago
>I can only hope the doomer narrative dominates until I can get a few shares at a reasonable valuation. I conjecture that some amount of the "doomer posting" is a consequence of other people realizing what you realized here and
49.
▲
by
sigmar
4mo ago
"capacity ramping" denotes that compute is increasing, which doesn't read like a discount, it reads like prorating.
50.
▲
by
sigmar
4mo ago
imho Anthropic publicly posting accurate information about their revenue and operations would be a step in a healthy direction for the economy/markets if there's an "AI bust blast" coming. This filing is movement towards
51.
▲
by
sigmar
4mo ago
Seems like AI is the Ozempic of tech. IE token generation keeps soaring, yet if you ask any individual- many swear they aren't touching it.
52.
▲
by
sigmar
4mo ago
>My recommendation is also to not choose an extreme approach (e.g. by completely banning LLM-related discourse) unless you feel very strongly about it. Organizers are allowed to ban the mention of certain programming topics? I could unde
53.
▲
by
sigmar
4mo ago
lol. Solid idea. Going to add an email signature with "Emailing me is billed at the following rates: $20k/M token input, $100k/M token output"
54.
▲
by
sigmar
4mo ago
That video is insane. Is the guy so arrogant that he couldn't conceive of being wrong? How could he possibly continue to write a ticket after this interaction?
55.
▲
by
sigmar
4mo ago
>But lately I’ve been thinking if it is just a class issue? This cohort of people likely have a cushion that softens the concussive blows they are doling out right now. They perhaps have the luxury of a somewhat functioning government an
56.
▲
by
sigmar
4mo ago
It was Mythos >Our engineers, working together with Mythos Preview, built a working exploit in five days. https://news.ycombinator.com/item?id=48139219
57.
▲
by
sigmar
4mo ago
Lots of interesting data in this. I'm surprised China is at 16.4%, I thought that would be higher
58.
▲
by
sigmar
4mo ago
UK's 'indefinite leave to remain' (closest thing to a green card) requires that you apply from within the UK. https://www.gov.uk/long-residence/apply-to-settle
59.
▲
by
sigmar
4mo ago
>it has to do with being battle tested over time. If a team of humans had rewritten it in a week, I wouldn't trust or use it either. "it was made in a week" gets repeated a lot on HN, but the PR wasn't a release. They
60.
▲
by
sigmar
4mo ago
Plateaued? Lol. Based on what? Pg 18 and 45 on that link are not showing a plateau.
More ›