Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
irthomasthomas
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
irthomasthomas
5d ago
And Claude often identifies as Qwen or Deepseek when prompted in Chinese.
2.
▲
by
irthomasthomas
5d ago
Except for all the exceptions, which are many, like the Jack the Ripper police files which where denied to the public.
3.
▲
by
irthomasthomas
7d ago
I doub't it. The main issue is not cost, though they do get expensive as context grows, but intelligence. A frontier model like fable becomes as dumb as haiku after 200k tokens. They have been stuck at ~1M context/200k useful cont
4.
▲
by
irthomasthomas
9d ago
I see, thanks. It reminds me a little of those spam emails which are sprinkled, intentionally, with obvious spelling errors. They aren't trying to trick the average person, they are trying to filter for a much smaller, more valuable a
5.
▲
by
irthomasthomas
9d ago
What are your favourite quotes from it?
6.
▲
by
irthomasthomas
9d ago
It is getting a lot harder for those people to justify using OpenAI to assist such endeavours. Afterall, OpenAI might just front-run you if they hear a rumor you solved some marquee problem that they can brag about in PR campaigns.
7.
▲
by
irthomasthomas
10d ago
artificialanalysis just updated their benchmark after the release of GPT-6. They removed old, saturated benchmarks and replaced them with new, until GPT-6 floated to the top with the cream. One of those new benchmarks is AutomationBench-AA,
8.
▲
by
irthomasthomas
12d ago
It took them two years to finally get him out with the help of the government in Shenzhen https://www.nme.com/news/arm-china-finally-ousts-rogue-ceo-t...
9.
▲
by
irthomasthomas
14d ago
> For instance the task naming in the task file starts with an optimistic 1, 2, 3, 5, 5a but then eventually gets to 8a, 8a1, and then ends up with 8b2c2b3 and “8b2c2b2b checkpoint1”. The code that it produced got ever more wild. I don’
10.
▲
by
irthomasthomas
14d ago
Astra scores the same on DeepSWE 1.1 (~75%) as Gemini Flash 3.8 and Deeepseek Flash 4.1 So general coding ability has plateaued, for now. Also consider the context windows. 1M token models where a breakthrough two years ago. Today they are
11.
▲
by
irthomasthomas
16d ago
Is there a reason they scoped that so narrowly to Buckmaster/codex/2 months two people worked on this for a year before the breakthrough. Perhaps that earlier work reduced the search space sufficiently to brute force the problem w
12.
▲
by
irthomasthomas
16d ago
Not on par, but in the same league. Astra is way ahead on visual tasks, but scores the same as gemini and deepseek on DeepSWE.
13.
▲
by
irthomasthomas
16d ago
Openai said that a new model became available to them during this. But that could mean anything from a big new base model to a LoRA, fine-tuned on a few dozen prompts...
14.
▲
by
irthomasthomas
16d ago
hmm I'm hoping there is a bug on their API because my first impression is not good. I asked it to return bash code between <bash></bash> tags. It is failing frequently and writing it's own tool calling format inst
15.
▲
by
irthomasthomas
16d ago
Quite a flex calling their GPT-6 competitor "Flash"! But it is faster than their last flash model due to a combination of architectural innovations including engrams and a new encoder/decoder design that uses 8B parameters f
16.
▲
by
irthomasthomas
17d ago
Has the method for extracting the COT been blocked, now? Otherwise why could we not generate some fresh samples?
17.
▲
by
irthomasthomas
17d ago
Chutes.ai models are served from a Trusted Execution Environment, so the GPU owners can't see your prompts.
18.
▲
by
irthomasthomas
17d ago
And deliberate or not it is still plagiarism by the sound of it.
19.
▲
by
irthomasthomas
18d ago
The researcher told them it was an independent effort, and they still pushed ahead with it.
20.
▲
by
irthomasthomas
18d ago
This should make an excellent choice for arbiter in llm-consortium, mercury-2 was pretty good. One of the main drawbacks of the multi-model system is the added latency of the llm judge, but having a model run at 1100tps goes a long a way to
21.
▲
by
irthomasthomas
18d ago
If the goal was not to scoop them, why did openai put a massive team on this, working weekends, only after they heard rumors of the solution?
22.
▲
by
irthomasthomas
18d ago
Or they trained a LoRA on the victims chats in order to launder their plagiarism.
23.
▲
by
irthomasthomas
18d ago
It can still be academic plagiarism even if they ticked the box to allow training on their prompts.
24.
▲
by
irthomasthomas
18d ago
Doesn't that count as plagiarism?
25.
▲
by
irthomasthomas
18d ago
"When a further trained version of our internal model became available over the course of the effort, we updated our agents to that model." woah, this gives some credit to the rumor that openai finetuned a model over the course of
26.
▲
by
irthomasthomas
18d ago
They where working on the problem for a year using codex.
27.
▲
by
irthomasthomas
18d ago
Attack is the best form of defence
28.
▲
by
irthomasthomas
18d ago
Is it opensource?
29.
▲
by
irthomasthomas
21d ago
You need to literally review the review with another llm pass to push back on the first. Ask it to do something like reassess the severity claims and only surface real P0 to P2 issues.
30.
▲
by
irthomasthomas
22d ago
I don't think so: https://news.ycombinator.com/item?id=49538217
More ›