Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ubermon
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
ubermon
2mo ago
it is indeed surprising how much a small model can do! As you said, for specific tasks like coding, the training data quality and new architecture of newer models probably beats model size
32.
▲
by
ubermon
2mo ago
yes. I think there is a threshold of some sort where if a model pass it can tolerate less system prompts. For example, same short prompt could work with deepseek v4 flash but not so much with Qwen 3.6 27B.
33.
▲
Claude Code Cut Their System Prompt by 80%. Does That Work for Small Models Too?
(antigma.ai)
7 points
by
ubermon
2mo ago
|
4 comments
34.
▲
by
ubermon
2mo ago
yeah, I think per user turn makes sense. But I am more inclining to only include it at session start.
35.
▲
by
ubermon
2mo ago
If you are interested in truly battle tested one, check this https://ante.run/start/philosophy
36.
▲
by
ubermon
3mo ago
For some frontier models like Fable 5 it doesn't matter, but for models less trained on long horizon tasks it very useful.
37.
▲
by
ubermon
3mo ago
I think it is more on distrusting python/typescript in general. Those dynamic scripting language are too easy to poison.
38.
▲
by
ubermon
3mo ago
we need a real agent that represent our own interest and deal with a more powerful general agents. And we can delegate answering this question to those more personal agents.
39.
▲
by
ubermon
3mo ago
lol same here, this is what I use https://github.com/AntigmaLabs/ante-preview
40.
▲
by
ubermon
3mo ago
I think this is mostly being overengineered. At this stage, a curated MEMORY.md is enough: short, readable at session start, and easy to correct when it becomes stale. The session history is another database like any other database, can be
41.
▲
by
ubermon
3mo ago
> Both scores come from a fixed fallback rule: Opus 4.8 and Sol xhigh run first; games scoring below 80 are rerun with Fable 5 and Sol max, respectively, and the higher per-game score is retained. hmm, this is like pass@n until you get t
42.
▲
by
ubermon
5mo ago
Thanks!
43.
▲
by
ubermon
5mo ago
Totally agree! I think we are very early on discovering the full potential of local models
44.
▲
by
ubermon
5mo ago
Thank you! I think there is a lot to dive deep later with different hardware, inference engine, prompt/harness setup etc.
45.
▲
by
ubermon
5mo ago
Not yet, will conduct a more comprehensive one later
46.
▲
Open-weight 27B hits 38% on Terminal-Bench 2.0 (Opus 4.1 hit 38% in Aug 2025)
(antigma.ai)
6 points
by
ubermon
5mo ago
|
9 comments
47.
▲
How to Achieve #1 on Terminal Bench
(twitter.com)
1 points
by
ubermon
7mo ago
|
1 comments
48.
▲
by
ubermon
7mo ago
and Why We Can't Have Nice Things): A Story
49.
▲
by
ubermon
1y ago
very good articulation on the problem. My bet is that it is probably fine in the sense that the mastery itself is not going to be relevant. I want to view AI coding as invention of new coding tools, at least in the way you described. I (hop
50.
▲
by
ubermon
3y ago
Give Copilot a try, it completely shift coding experience in Rust. Especially this: > I always have to search libraries and how to do things. Once you pass the initial curve with crutch like Copilot, then you can be almost as productive
51.
▲
by
ubermon
3y ago
haha was thinking the same. The other day openAI's api hiccup gave me a small panic attack
52.
▲
by
ubermon
3y ago
Moat does not come from compile time but runtime. The company, the operation, and the accumulated data, the brand, the trust and the reputation.
53.
▲
by
ubermon
6y ago
you seem really uninformed about what is happening in this space. Just to name a few, incredible fast settlement with lending and yielding; Automatic Market Maker; Decentralized Exchange guaranteed with no wash trade, insolvency and max tra
54.
▲
by
ubermon
6y ago
Curious on what is the general view of Indian population, twitter is full of trolls and paid campaign. From what I gather in chinese forum, I don't think any one believes China was the aggressive side, their most aggression is on getti
55.
▲
by
ubermon
6y ago
I think two important factors in play here, one is like what others said about better trust on the legal system in US to be relatively more independent; second I guess is just US market is too lucrative and significant to just let go.
56.
▲
by
ubermon
6y ago
I remember using this technique to get US content on Netflix while sitting in Canada
57.
▲
by
ubermon
6y ago
I think not manager but peers. They are mostly over achievers and somewhat competitive.
58.
▲
by
ubermon
6y ago
haha very possible, but it is still quite shocking to me that my mom, who is almost 60, keep telling me how fun it is...
59.
▲
by
ubermon
6y ago
Fueled by a recent Indian-China border conflicts, a political decision.
60.
▲
by
ubermon
6y ago
Could be some over zealous employees inspired by recent Indian-China conflicts and TikTok ban and decided within their own org to do it without realizing how big of a news it would be.
More ›