Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ahmadawais
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
1.
▲
by
ahmadawais
2mo ago
Introducing GLM-5.3: Built to Code. Ready for Cyber Defense. - Top-tier coding and agentic capabilities, achieved through post-training on the 743B base model - A major leap in cybersecurity, setting a new standard among open models Tech Bl
2.
▲
by
ahmadawais
2mo ago
We’re launching DeepSeek-V4-Pro today! Major Agent upgrades with strong production gains! Flexible reasoning effort for V4-Pro & V4-Flash: low for simple tasks, high for daily Agent workflows, max for complex tasks. Native OpenAI Resp
3.
▲
by
ahmadawais
2mo ago
i thought that's what we're supposed to do. it has the link too btw.
4.
▲
Show HN: Command Code GOAT, $10/month for $70 of credits across 30 models
(twitter.com)
6 points
by
ahmadawais
2mo ago
|
2 comments
5.
▲
by
ahmadawais
2mo ago
can't wait for it. i think they'll do 35B variant as well.
6.
▲
Qwen 3.8-Max and Qwen3.8-27B are going open weight
(twitter.com)
5 points
by
ahmadawais
2mo ago
|
4 comments
7.
▲
by
ahmadawais
2mo ago
> Next week, the open weights of Qwen3.8-Max will be released, and Qwen3.8-27B is also going open-weights to meet you all!
8.
▲
by
ahmadawais
3mo ago
see that's what i was saying. in v1 we'll have a super strong extension api — week out probably.
9.
▲
by
ahmadawais
3mo ago
Thanks! I've seen many developers use the entire tweet as a prompt to improve their harness. If you trace errors, you can find tool call failures per billion tokens. I've seen like many versions of this on pi, oc, or whatnot. This
10.
▲
by
ahmadawais
3mo ago
probably yep, i doubt they'll be able to maintain that well. we did a lot of work improving MiMo's cache thrashing, it was pretty bad early on. Now we're seeing 97% cache rate on their models after lots of prefix work.
11.
▲
How did we make DeepSeek outperform Opus
(twitter.com)
34 points
by
ahmadawais
3mo ago
|
7 comments
12.
▲
by
ahmadawais
3mo ago
hey HN, sharing harness engineering deep dive on tool calling repairs for open models. i've been thinking about why "open model bad at tool calling" is almost always a harness problem, not a model problem. spent time looking
13.
▲
by
ahmadawais
2y ago
Have you seen any CS bots those are agents. We built an email agent example among many other at https://Langbase.com/docs
14.
▲
by
ahmadawais
2y ago
This was fun to read. Thanks for sharing. Our memory agents are built somewhat like this but way too much going on in there to make them prod ready.
15.
▲
by
ahmadawais
2y ago
Couldn’t agree more. Btw what’s your use case?
16.
▲
by
ahmadawais
2y ago
Spending 100+ hours on this, getting to know people who read this was the point after all. Pretty standard practice everywhere? What did I miss?
17.
▲
by
ahmadawais
2y ago
I was surprised by the Mistral usage — i think Llama has taken over the open-source LLM scene lately.
18.
▲
Show HN: State of AI Agents 2024 – 184B tokens · 786M runs analyzed
(langbase.com)
6 points
by
ahmadawais
2y ago
|
4 comments