Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
anotherpaulg
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
anotherpaulg
1y ago
Fun project with a nice compact code base. Agreed that ad-hoc scripting agents can be very powerful. Aider has had support for scripting [0] in python or via the command line for a long time. I made a screencast [1] recently that included a
32.
▲
by
anotherpaulg
1y ago
I've long been fascinated by AI's ability to do the reverse: generate photos with lots of highly relevant content when the prompt includes a location. Terrain, plants, buildings, landmarks, coastlines and lots of details are inclu
33.
▲
by
anotherpaulg
1y ago
In aider, instead of “ultrathink” you would say: /thinking-tokens 32k Or, shorthand: /think 32k
34.
▲
by
anotherpaulg
1y ago
Mostly Gemini 2.5 Pro lately. I get asked this often enough that I have a FAQ entry with automatically updating statistics [0]. Model Tokens Pct Gemini 2.5 Pro 4,027,983 88.1% Sonnet 3.7 518,708 11.3
35.
▲
by
anotherpaulg
1y ago
I just finished updating the aider polyglot leaderboard [0] with GPT-4.1, mini and nano. My results basically agree with OpenAI's published numbers. Results, with other models for comparison: Model Score C
36.
▲
by
anotherpaulg
1y ago
Aider author here. Based on some DMs with the Gemini team, they weren't aware that aider supports a "diff-fenced" edit format. And that it is specifically tuned to work well with Gemini models. So they didn't think to tr
37.
▲
by
anotherpaulg
2y ago
The graph in the article about Quasar Alpha’s coding skill is taken from my tweet [0]. It shows QA’s results on the aider polyglot coding benchmark [1]. QA seems to be a skilled coder, and is very fast. Aider supports Quasar Alpha as of v0.
38.
▲
by
anotherpaulg
2y ago
Llama 4 Maverick scored 16% on the aider polyglot coding benchmark [0]. 73% Gemini 2.5 Pro (SOTA) 60% Sonnet 3.7 (no thinking) 55% DeepSeek V3 0324 22% Qwen Max 16% Qwen2.5-Coder-32B-Instruct 16% Llama 4 Maverick [0] https
39.
▲
by
anotherpaulg
2y ago
I actually get asked for screencasts a lot, so I made recently made some [0]. The recording of adding support for 100+ new coding languages with tree-sitter [1] shows some pretty advanced usage. It includes using aider to script downloading
40.
▲
by
anotherpaulg
2y ago
Gemini 2.5 Pro set a wide SOTA on the aider polyglot coding leaderboard [0]. It scored 73%, well ahead of the previous 65% SOTA from Sonnet 3.7. I use LLMs to improve aider, which is >30k lines of python. So not a toy codebase, not green
41.
▲
by
anotherpaulg
2y ago
It does have fairly low adherence to the edit format, compared to the other frontier models. But it is much better than any previous Gemini model in this regard. Aider automatically asks models to retry malformed edits, so it recovers. And
42.
▲
by
anotherpaulg
2y ago
Gemini 2.5 Pro set the SOTA on the aider polyglot coding leaderboard [0] with a score of 73%. This is well ahead of thinking/reasoning models. A huge jump from prior Gemini models. The first Gemini model to effectively use efficient di
43.
▲
by
anotherpaulg
2y ago
Yup. It’s a complex enough set of in/excludes that I think that would get unwieldy for my use case. Details here: https://github.com/Aider-AI/aider/blob/main/scripts/blame.py Again, nice work o
44.
▲
by
anotherpaulg
2y ago
This is great. I do this sort of git-blame accounting to track how much code is written by AI versus humans in each release of my app. My "blame script" has been slowing down as the repo size increases. I was just about to add cac
45.
▲
by
anotherpaulg
2y ago
Here’s my write up about the unique ways that aider uses uv as an installer. https://aider.chat/2025/01/15/uv.html Offering these install methods dramatically reduced the number of GitHub issues from users wi
46.
▲
by
anotherpaulg
2y ago
You can use —no-auto-commits or see these docs for all the git settings: https://aider.chat/docs/git.html
47.
▲
by
anotherpaulg
2y ago
Thanks for trying aider! Sorry to hear you had trouble with the install. Installing Python cli tools is tricky, which is why I recommend using the uv-based installer. The docs do list a variety of other install methods: https://a
48.
▲
Aider: Using Uv as an Installer
(simonwillison.net)
39 points
by
anotherpaulg
2y ago
|
14 comments
49.
▲
by
anotherpaulg
2y ago
GPT-4.5 Preview scored 45% on aider's polyglot coding benchmark [0]. OpenAI describes it as "good at creative tasks" [1], so perhaps it is not primarily intended for coding. 65% Sonnet 3.7, 32k think tokens (SOTA) 60% S
50.
▲
by
anotherpaulg
2y ago
Good point! This is easy to picture if you imagine widely spreading out the equipment used for the eraser experiment. If the signal hitting the screen and idler hitting one of the detectors are space-like separated events... the OP's e
51.
▲
by
anotherpaulg
2y ago
This is a very interesting, novel take on explaining the delayed choice quantum eraser. I think I can summarize it as follows: The signal photon hits the screen, which is a measurement. The entangled idler's wave function is thereby co
52.
▲
by
anotherpaulg
2y ago
Here are the current docs for changing the thinking token limits. https://aider.chat/docs/llms/anthropic.html#thinking-tokens I'll make this less clunky soon.
53.
▲
by
anotherpaulg
2y ago
I try not to let perfect be the enemy of good. All benchmarks have limitations. The Exercism problems have proven to be very effective at measuring an LLM's ability to modify existing code. I receive a lot of feedback that the aider be
54.
▲
by
anotherpaulg
2y ago
Using up to 32k thinking tokens, Sonnet 3.7 set SOTA with a 64.9% score. 65% Sonnet 3.7, 32k thinking 64% R1+Sonnet 3.5 62% o1 high 60% Sonnet 3.7, no thinking 60% o3-mini high 57% R1 52% Sonnet 3.5
55.
▲
by
anotherpaulg
2y ago
Claude 3.7 Sonnet scored 60.4% on the aider polyglot leaderboard [0], WITHOUT USING THINKING. Tied for 3rd place with o3-mini-high. Sonnet 3.7 has the highest non-thinking score, taking that title from Sonnet 3.5. Aider 0.75.0 is out with s
56.
▲
by
anotherpaulg
2y ago
Thanks for letting folks know about aider's /copy-context command. To add some more detail, aider has a mode/UX that is optimized for "copy and paste" coding with LLM web chats. The "big brain" LLM in the
57.
▲
by
anotherpaulg
2y ago
"Improving forecasting ability" is a central plot point of the recent fictional account of How AI Takeover Might Happen in 2 Years [0]. It's an interesting read, and is also being discussed on HN [1]. ... [T]hese researchers
58.
▲
by
anotherpaulg
2y ago
For AI coding, o3-mini scored similarly to o1 at 10X less cost on the aider polyglot benchmark [0]. This comparison was with both models using high reasoning effort. o3-mini with medium effort scored in between R1 and Sonnet. 62% $186 o
59.
▲
by
anotherpaulg
2y ago
Every metric has limitations, but git blame line counts seem pretty uncontroversial. Typical aider changes are not like autocompleting braces or reformatting code. You tell aider what to do in natural language, like a pair programmer. It th
60.
▲
by
anotherpaulg
2y ago
It's "total" tokens, input plus output. I'd guess more than two-thirds of them are input tokens.
More ›