Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
irthomasthomas
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
8 ms
·
91.
▲
by
irthomasthomas
2mo ago
Because they do not know the name of the model before they train it. There is also distillation, where multiple models will be trained from a larger one. E.G. Sonnet was promoted to Opus at one point after it surpassed expectations.
92.
▲
by
irthomasthomas
2mo ago
Cooking with lard IS healthy. There was never any evidence to the contrary. The USDA promoted the low fat diet to sell cheap commodity crops, ultra-processed foods, and alternative vegetable oils. Those things ARE bad for you. The low smoke
93.
▲
by
irthomasthomas
2mo ago
Why do OpenAI never release logs to prove their claims? Why should we believe them when they write extraordinary anecdotes about the power of their products without ever providing proof?
94.
▲
by
irthomasthomas
2mo ago
Aaron Schwartz was being prosecuted and threatened with 35 years in jail for the crime of saving research papers to a thumb drive. What damage did he cause?
95.
▲
by
irthomasthomas
2mo ago
Changelog - fixed issue where model acts like qwen when prompted in chinese
96.
▲
by
irthomasthomas
2mo ago
Openai hacked HF with a zero-day. Definitely an interesting story! I just find their explanations hard to believe. They're admitting to a great deal of incompetence. I know, don't ascribe to malice... but still, the more extraordi
97.
▲
by
irthomasthomas
2mo ago
Simon says not to dismiss this as a publicity stunt, but I think we should reserve judgement, and not treat this as true until they publish the logs. A company that uses industrial espionage against Apple does not deserve the benefit of the
98.
▲
by
irthomasthomas
3mo ago
A piece of the frame is missing between pedals and back wheel. The frame of the bike passes through the bird. It also puts a cap on the bird's head, and a fish in it's mouth. The fish and the cap where always added when I asked an
99.
▲
by
irthomasthomas
3mo ago
Ask claude it's name in Chinese and it says Qwen, or Deepseek. By your own logic Anthropic must have distilled from Chinese models rather than produce their own Chinese training data. Here is sonnet acting like deepseek: https:/&
100.
▲
by
irthomasthomas
3mo ago
And if you ask Opus 4.8 in the API: '你是什么模型' (what model are you?) It responds ~9/10 times with: 我是通义千问(Qwen),是阿里巴巴集团旗下的通义实验室自主研发的大语言模型。我可以帮助你回答问题、创作文字(比如写故事、写公文、写邮件、写剧本等)、进行逻辑推理、编程、翻译等等。 有什么我可以帮你的吗? (I am Tongyi Qianwen (Qw
101.
▲
by
irthomasthomas
3mo ago
Ask claude its name in Chinese and it says Qwen or Deepseek. Anthropic distilled Chinese tokens rather than create their own Chinese language training data.
102.
▲
by
irthomasthomas
3mo ago
Did you know claude models identify as qwen or deepseek when asked in chinese?
103.
▲
by
irthomasthomas
3mo ago
Have you ever looked at how much performance drops as context grows? The difference in intelligence between 100k and 1M is huge, like opus drops to haiku level performance, or worse. For that reason I try to keep under 200k. That feels abou
104.
▲
by
irthomasthomas
3mo ago
It's like scaling a swiss cheese and the holes grow bigger with it. You can't get rid of the holes without making a different cheese.
105.
▲
by
irthomasthomas
3mo ago
Claude use to be leader, too. Their metaprompt was great at the time with opus 3
106.
▲
by
irthomasthomas
3mo ago
Cool. I still find these a useful visualization of some the qualities of llms. Even if they did train for [animal] on [vehicle] svg, it's still nice to see at a glance how the different models and reasoning levels perform. Lunar misses
107.
▲
by
irthomasthomas
3mo ago
Incredible photography.
108.
▲
by
irthomasthomas
3mo ago
This is what I do in llm-consortium. An arbiter evaluates the response(s) and decides if more iterations are needed. You can also loop until a minimum confidence threshold, but self-reported confidence isn't a great metric.
109.
▲
by
irthomasthomas
3mo ago
What for?
110.
▲
by
irthomasthomas
3mo ago
Try this prompt: While working on the main task, launch a parallel sub-agent with the task context so far. The sub agent should think of high quality questions and put them to the user using a dialogue tool like zenity. Customize the inpu
111.
▲
by
irthomasthomas
3mo ago
>Chain of reasoning is a lot of context to guide token generation, but we simply see that newer models don’t need that context to get to the answer I thought each new generation typically used more reasoning tokens?
112.
▲
by
irthomasthomas
3mo ago
In think in all cases where I've seen it compared CC performed worse than a minimal harness.
113.
▲
by
irthomasthomas
3mo ago
This feels more communist than communist China. They typically take about 1% in Golden Shares that give them a board seat.
114.
▲
by
irthomasthomas
3mo ago
Claude in Claude code has been shown to perform persistently worse in evals than claude + a minimal harness.
115.
▲
by
irthomasthomas
3mo ago
The comment from the other Mullvad founder is here https://news.ycombinator.com/item?id=48696800
116.
▲
by
irthomasthomas
3mo ago
Very fast and reliable? Sold!
117.
▲
by
irthomasthomas
3mo ago
GLM 5.2 is ~40B active parameters, which is what matters most for training cost.
118.
▲
by
irthomasthomas
3mo ago
They provide benchmarks in the paper https:// arxiv.org/abs/2606.21228
119.
▲
by
irthomasthomas
3mo ago
So only 100 companies have exclusive access to frontier AI.
120.
▲
by
irthomasthomas
3mo ago
Weird, I didn't think I was throwing technocracy under the bus. What makes you say that?
More ›