Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pcwelder
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
31.
▲
by
pcwelder
1y ago
Agree. To reduce costs: 1. Precompute frequently used knowledge and surface early. For example repository structure, os information, system time. 2. Anticipate next tool calls. If a match is not found while editing, instead of simply failin
32.
▲
by
pcwelder
1y ago
>I found, a required sample size for just one thousand people would be 278 It's interesting to note that for a billion people this number changes to a whopping ... 385. Doesn't change much. I was curious, with 22 sample size (a
33.
▲
by
pcwelder
1y ago
I get similar accuracy to claude code using claude desktop app with a file+bash mcp (different tools same performance). My guess for why GPT5 scores more on benchmarks is that they evaluate on well defined tasks with all instructions given
34.
▲
by
pcwelder
1y ago
To be absolutely honest, this wasn't a very conscious choice :-) I don't think a direct similarity with domain specific languages is evident to me. I rather find the messaging similar to some "agents" from other domains.
35.
▲
Why build a domain-specific agent for front end tasks?
(kombai.com)
8 points
by
pcwelder
1y ago
|
9 comments
36.
▲
LLMs are bad at returning code in JSON
(aider.chat)
5 points
by
pcwelder
1y ago
|
1 comments
37.
▲
by
pcwelder
1y ago
PSA: don't generate code using tools (and MCPs) if you're using Gemini or Openai; both ask LLMs to generate JSON directly for function calling. Claude uses XML, so it escapes the issue.
38.
▲
by
pcwelder
1y ago
Losing the sense of cwd is the reason why I append it in the output of each command run in wcgw mcp [1] It rarely does it incorrectly after that. I won't be surprised if claude code does the same soon. However, they do have an env flag
39.
▲
by
pcwelder
1y ago
Yes, thanks for clarifying. I specified in the system prompt that they're built by xAI and other system instructions from Grok 4.
40.
▲
by
pcwelder
1y ago
> My best guess is that Grok “knows” that it is “Grok 4 buit by xAI”, and it knows that Elon Musk owns xAI, so in circumstances where it’s asked for an opinion the reasoning process often decides to see what Elon thinks. I tried this hyp
41.
▲
by
pcwelder
1y ago
Impressive demo! Is there any rate limit (TPM)?
42.
▲
by
pcwelder
1y ago
You've essentially just trained your own LM instead of using a pretrained large LM. Speaking generically -- any place in your workflow you feel the task is not hard, you can use smaller and cheaper LM. Smaller LMs come with accuracy re
43.
▲
by
pcwelder
1y ago
This script has been sufficient for me to configure gpu drivers on fresh ubuntu machines. It's just uv add torch after this. https://cloud.google.com/compute/docs/gpus/install-drivers-g... (NOTE: not gc
44.
▲
by
pcwelder
1y ago
``` try: answer = chain.invoke(question) # print(answer) # raw JSON output display_answer(answer) except Exception as e: print(f"An error occurred: {e}") chain_no_parser = prompt | llm raw_output
45.
▲
by
pcwelder
1y ago
Back when GANs were popular, I'd train generator-discriminator models for image generation. I thought a lot about it and realised discriminating is much easier than generating. I can discriminate good vs bad UI for example, but I can&#
46.
▲
by
pcwelder
1y ago
Tool annotations are missing. (Are they of any use though?)
47.
▲
by
pcwelder
1y ago
I'm sorry but AGI is one of those loaded words which would lose substance with just a few rounds of the rationalist's taboo. If it just means human level intelligence, then world modeling isn't needed as argued. Simply becaus
48.
▲
by
pcwelder
1y ago
I believe it's not possible to restrict an LLM from executing certain commands while also allowing it to run python/bash. Even if you allow just `find` command it can execute arbitrary script. Or even 'npm' command (whic
49.
▲
by
pcwelder
1y ago
> other prompts yours get batched with Why would batching lead to variance?
50.
▲
by
pcwelder
1y ago
OpenAPI definitions are verbose and exhaustive. In MCPs you can remove a lot of extra material, saving tokens. For example in [1], whole `responses` schema can be eliminated. The error texts can instead be surfaced when they appear. You als
51.
▲
by
pcwelder
1y ago
I believe this applies to all AI use cases to varying degrees. If AI can't use X, then there is something wrong with X. X in { website, codebase, function, language, library, mcp, ...}
52.
▲
by
pcwelder
1y ago
Good MCP clients avoid having LLMs generate JSON. Claude for example uses XML to generate mcp tool usage. At least top level strings don't need to be json encoded.
53.
▲
by
pcwelder
1y ago
In the latest update they've replaced "Allow for this chat" with "Always Allow".
54.
▲
by
pcwelder
1y ago
Cool. If I'm not wrong you don't detect prompt injection done in the tool results? Any plans for that?
55.
▲
by
pcwelder
1y ago
Did some quick tests. I believe its the same model as Quasar. It struggles with agentic loop [1]. You'd have to force it to do tool calls. Tool use ability feels ability better than gemini-2.5-pro-exp [2] which struggles with JSON sche
56.
▲
by
pcwelder
1y ago
Can someone explain to me why we should take Aider's polyglot benchmark seriously? All the solutions are already available on the internet on which various models are trained, albeit in various ratios. Any variance could likely be due
57.
▲
Show HN: Chat.md
(github.com)
1 points
by
pcwelder
1y ago
|
0 comments
58.
▲
Claude avoids JSON accuracy issue in tool calling
(arcfu.com)
1 points
by
pcwelder
1y ago
|
0 comments
59.
▲
Show HN: Chat.md – file as chat interface with editable history [MCP-client]
(github.com)
2 points
by
pcwelder
1y ago
|
0 comments
60.
▲
by
pcwelder
2y ago
Why not have a single mcp server that takes in the repo path or url in the tool call args? Changing config in claude desktop is painful everytime.
More ›