Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jbellis
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
jbellis
12d ago
For the mj adversarial review we use a configurable coordinator (I recommend Fable) with six subagents (I recommend Luna) for specific specialties like code complexity. Narrowing down what a smaller model like Luna needs to look for /
2.
▲
by
jbellis
12d ago
Oh, cool. I used tailscale to provide a private web app for my cross-harness orchestrator, too. ( https://github.com/BrokkAi/mjolnir/ )
3.
▲
by
jbellis
12d ago
While I'm generally sympathetic to the idea that the public coding benchmarks are inadequate, to the point that I've written my own in the past and will likely do so again, the complaint here is that "these tasks don't m
4.
▲
by
jbellis
12d ago
> I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands. And I think that the takes on AI killing us all by creati
5.
▲
by
jbellis
13d ago
One thing we've noticed is that if you have model X implement AND review the same change, the reviewer often shares the same blind spots as the implementer. https://blog.brokk.ai/mjolnir-automated-cross-vendor-adversa..
6.
▲
by
jbellis
13d ago
If you're more of a SWE than a data scientist, check out Mjolnir: https://github.com/BrokkAi/mjolnir/ Built in containerization and EC2 support, allows moving sessions across harnesses/profiles/mach
7.
▲
by
jbellis
14d ago
This is only correct if you ignore subscriptions. I ran the numbers on what you get for your subscriptions last week: https://blog.brokk.ai/a-coding-subscription-tier-list/
8.
▲
Don't use musl if you care about performance
(blog.brokk.ai)
109 points
by
jbellis
29d ago
|
77 comments
9.
▲
by
jbellis
1mo ago
This is why I stop at xhigh.
10.
▲
by
jbellis
1mo ago
No, you need more compute for those use cases. That's why everyone trains on GPUs.
11.
▲
by
jbellis
1mo ago
Love to see this. I started something similar a couple years ago but didn't quite have the patience to make the browser extensions really bulletproof. Planning to contribute significant improvements to the vector search side here.
12.
▲
by
jbellis
1mo ago
Unfortunately we seem to live in an era where your mainstream choices for a cross platform app are "Electron taking 1GB when idle" and TUI.
13.
▲
Mjolnir: Automated Cross-Vendor Adversarial Review
(blog.brokk.ai)
3 points
by
jbellis
1mo ago
|
0 comments
14.
▲
by
jbellis
1mo ago
If it's 2x faster than rustc, I'll switch.
15.
▲
Muninn: Code localization model in the world fits on your phone
(blog.brokk.ai)
1 points
by
jbellis
1mo ago
|
0 comments
16.
▲
It's time to stop doing code reviews
(blog.brokk.ai)
6 points
by
jbellis
1mo ago
|
9 comments
17.
▲
by
jbellis
1mo ago
Because there are better ways now to solve the problems that code review exists to solve.
18.
▲
by
jbellis
1mo ago
And that's it, that's the last lab releasing models worth coding with that didn't have a first party harness that its models are trained to use.
19.
▲
by
jbellis
2mo ago
I'm so old I can remember when ars technica published actual technical content.
20.
▲
by
jbellis
2mo ago
Why less workflow friction?
21.
▲
by
jbellis
2mo ago
The only scenario is if you have enough work to do batch inference. Using a tiny fraction of GPU capacity to decode a single request at a time just doesn't make sense, as you say.
22.
▲
by
jbellis
2mo ago
Yes, look up REAP.
23.
▲
by
jbellis
2mo ago
Apparently this is an unpopular opinion on HN, but a job isn't an output that companies produce, it's an agreement to compensate an employee in exchange for services, and just as the CEO shouldn't take it personally when one
24.
▲
by
jbellis
2mo ago
It's roughly equal on price and intelligence as GLM 5.2 while being ~8x faster.
25.
▲
by
jbellis
2mo ago
I guess you didn't read ellie's reply directly underneath that?
26.
▲
by
jbellis
2mo ago
It tops benchmarks because it uses them in its training data. https://x.com/eliebakouch/status/2077425801633427919
27.
▲
by
jbellis
2mo ago
This just isn't true. On balance, data centers are turning out to be more like the "anchor tenant" of the power grid, financing improvements for everyone. Overview article with links to actual studies: https://city
28.
▲
by
jbellis
3mo ago
voyage 4 nano is sota at the next size up and if you really want the best teacher models it's probably the voyage commercial APIs
29.
▲
by
jbellis
3mo ago
FWIW -- Granite r2 small is a 30M model, still small enough to run on CPU, and a good baseline for fine tunes.
30.
▲
by
jbellis
3mo ago
Yes, disappointing given that Chronicle actually does have legit expertise here. (Chronicle Map is not very well-known even in the Java space but it's by far the best larger-than-memory Map available.)
More ›