Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zuzuen_1
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
zuzuen_1
3mo ago
"We make a tentative calibration of the self-sustaining ac celeration condition using the existing data that is available, measuring AI capabilities using the Epoch Capabilities Index (Ho et al., 2025). We find that the condition is me
2.
▲
by
zuzuen_1
4mo ago
I think we need better classification and taxonomy on erroneous LLM behaviors than the catch-all term "hallucinate"..
3.
▲
Twelve Factor App Method (2011)
(12factor.net)
2 points
by
zuzuen_1
4mo ago
|
0 comments
4.
▲
by
zuzuen_1
4mo ago
Tbh, I think troubleshooting / investigation is really a lot different from feature development. Wrong fixes upstream are really bad because software downstream starts relying on these. You have to be unafraid of deep dives, but at the
5.
▲
by
zuzuen_1
5mo ago
Wait, this is insane! And the method they use to classify an outage seems solid as well. Imagine being a Github dev working in this environment.
6.
▲
Startup Mode Engineering
(moduloware.ai)
2 points
by
zuzuen_1
11mo ago
|
0 comments
7.
▲
Tips for building performant LLM applications
(moduloware.ai)
4 points
by
zuzuen_1
11mo ago
|
1 comments
8.
▲
by
zuzuen_1
11mo ago
I've been building Modulo AI for the past year - an AI system that fixes GitHub issues. Early versions took 5+ minutes to analyze a single issue. After months of optimization, we're now sub-60 seconds with better accuracy. This pr
9.
▲
by
zuzuen_1
1y ago
One pain point such a PL could address is encoding tribal knowledge about optimal prompting strategies for various LLMs, which changes with each new model release.
10.
▲
by
zuzuen_1
1y ago
Perhaps when LLMs introduce a lot more primitives for modifying behvavior such a programming language would be necessary. As such for anyone working with LLMs, they know most of the work happens before and after the LLM call, like doing RES
11.
▲
by
zuzuen_1
1y ago
I would be more interested in Qodo's performance on the swe-bench-multilingual benchmark. Swe-bench-verified only includes bugs related to python breakages. The best submission is swe-bench-multilingual is Claude 3.7 Sonnet which solve
12.
▲
by
zuzuen_1
1y ago
Does anyone have a benchmark on the effectiveness of using embeddings for mapping bug reports to code files as opposed to extensive grepping as Qodo, Cursor and a number of tools I use do to localize faults?