Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ramoz
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
ramoz
5d ago
Even for coding agents. MCP is the only unified interface INTO harnesses. This pattern has not struck yet for most, but will soon in the coming months given the MCP spec. MCP has been about models calling tools., it is about to become some
2.
▲
by
ramoz
5d ago
Bad take. And what's coming down the pipe with MCP will cement it. Has less with the model calling tools and will become about tools calling models. There is no unified integration into harnesses to support this other than MCP.
3.
▲
Jev Jailbreak Benchmark
(backnotprop.com)
3 points
by
ramoz
6d ago
|
1 comments
4.
▲
by
ramoz
7d ago
There's no benchmarking because it's not very intelligent at all. Right now everybody's being hype-shotted into believing you can use it for intelligent decisions. https://backnotprop.com/blog/jev-poker&#
5.
▲
by
ramoz
7d ago
> Think shared Claude Artifacts that don’t live @ Anthropic. Shameless plug, I've built a self-hosted capability here with things like live collaboration for humans and agents. There is a native cloudflare deplyoment and integration
6.
▲
Pace. The Final Fronteir
(twitter.com)
2 points
by
ramoz
8d ago
|
0 comments
7.
▲
Typesafe's Jev is the fish at the poker table
(backnotprop.com)
2 points
by
ramoz
9d ago
|
1 comments
8.
▲
The Gap Will Grow
(backnotprop.substack.com)
1 points
by
ramoz
10d ago
|
0 comments
9.
▲
by
ramoz
10d ago
Rereading some things and because there's no official benchmarks, I misspoke about the ballpark comparison., but the open model's still a useful foundation to work with
10.
▲
by
ramoz
10d ago
GLiClass is performant, and its zero-shot classification scores are in the same ballpark as the Terra-level results Jev points to. https://github.com/knowledgator/gliclass
11.
▲
by
ramoz
10d ago
Guess I'm a bit less impressed seeing that for some of the more intelligent driven+action work -- splitting requests in the video -- they had to kick out to an anthropic model.
12.
▲
The Operational Limits of Agent Security (Dec 2025)
(cupcake.eqtylab.io)
3 points
by
ramoz
18d ago
|
0 comments
13.
▲
Artifact Server: open-source Claude Artifacts alternative
(github.com)
2 points
by
ramoz
23d ago
|
1 comments
14.
▲
by
ramoz
23d ago
Artifact Server is an open source alternative to Claude Code Artifacts; enabling teams to own, share, and collaborate on artifacts throughout the development lifecycle. https://github.com/plannotator/artifact-server h
15.
▲
Claude Code session limits are jacked up with Fable 5.1
(twitter.com)
3 points
by
ramoz
24d ago
|
1 comments
16.
▲
by
ramoz
24d ago
I don't have 28 accounts but am experiencing the same with my handful others: https://github.com/anthropics/claude-code/issues/91289 https://github.com/anthropics/claude-code/is
17.
▲
by
ramoz
25d ago
I bought the game so I could study the code. Unfortunately it is wasm & obfuscated. I'm gonna see what my handy dandy big giant brain buddies can do with it.
18.
▲
Show HN: Annotate Anything in the Terminal
(plannotator.ai)
3 points
by
ramoz
26d ago
|
1 comments
19.
▲
by
ramoz
28d ago
What does this mean for Fable limits exclusively?
20.
▲
Show HN: Herdr plugin to annotate anything in the terminal
(github.com)
2 points
by
ramoz
1mo ago
|
0 comments
21.
▲
by
ramoz
1mo ago
bro is good https://github.com/backnotprop/bro/blob/main/skills/bro/SKIL...
22.
▲
by
ramoz
1mo ago
I see. What would the proper solution then look like for when we should consider moving our teams? Sounds like you you said they're all current ones are missing ergonomics. Do you think it's slack code?
23.
▲
by
ramoz
1mo ago
Why did you move from on-dev-machine coding to @claude in GitHub if it's a weaker capability? Im not convinced of any of these cloud or tagging solutions where I get to the point of moving coding away from my dev's machines.
24.
▲
Slack Code Is Live
(twitter.com)
3 points
by
ramoz
1mo ago
|
0 comments
25.
▲
by
ramoz
1mo ago
I created /bro a couple model iterations prior. I use it everyday many times a day and it's getting worse. I was actually using it more with GPT models but they have been getting better. https://github.com/backnotp
26.
▲
by
ramoz
1mo ago
ah interesting
27.
▲
by
ramoz
1mo ago
"It reads the screen, not the DOM", but is built completely on Playwright? At first it made me think you have some visual model at play, but doesn't actually seem that way.
28.
▲
by
ramoz
1mo ago
You're underselling how good/reliable 5.6 sol is. It's in the fable discussion league not Opus4.8
29.
▲
by
ramoz
1mo ago
Physical custody of our work is going to be an important discussion for industry. ie where does your code live and how do you access it? Every agent platform is moving toward you living with them, logging into them, working on their remote
30.
▲
by
ramoz
1mo ago
Source required here
More ›