Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jameswhitford
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
Agents feel free to take desktop screenshots?
(wecreatethis.com)
5 points
by
jameswhitford
2mo ago
|
1 comments
2.
▲
by
jameswhitford
2mo ago
I set my agent (Sonnet 5) to work on my KOReader extension on auto mode. It started autonomously taking screenshots of my screen and reading them, without my consent or instruction. Screenshots were being fired on whatever my screen was foc
3.
▲
by
jameswhitford
2mo ago
Awesome game! A potential UX update: if you fail a guess, the letters go back down below in the same configuration as your guess. But others might disagree!
4.
▲
by
jameswhitford
2mo ago
Eish, probably needs some bug hunting. During testing I had some issues with less capable models, if you have your own key for a more capable model it should perform better.
5.
▲
by
jameswhitford
3mo ago
Agreed! The world is a little unbalanced in this way right now unfortunately.
6.
▲
by
jameswhitford
3mo ago
This is sick where have you been all my life!
7.
▲
by
jameswhitford
3mo ago
This is awesome! I love mermaid but have struggled to get my agents to be any good at it. Maybe I just need to context engineer a little better
8.
▲
by
jameswhitford
3mo ago
It would be great if it could have examples of your own diagrams in its context, so it knows what points you like to map out and how you like to visualize them
9.
▲
by
jameswhitford
3mo ago
That’s awesome! My inspiration was the IBM YouTube channel where they write on the board in front of them, maybe you could use the white canvas as like a green screen or something to get a similar effect?
10.
▲
by
jameswhitford
3mo ago
The potential use cases was my favorite part of designing it, it’s not perfect now but it’s so fun to think what people might do with these kinds of designs in the future
11.
▲
by
jameswhitford
3mo ago
Yes I would love to see some work on the design, because I think it could be useful, just needs some speed and accuracy improvements, and maybe some design towards specific use cases
12.
▲
by
jameswhitford
3mo ago
Yes it could get a lot better with some design improvements, maybe live streaming the transcription, and maybe after some trial and error testing the best + fastest model for the job
13.
▲
by
jameswhitford
3mo ago
Facts!
14.
▲
Show HN: Agent Draw: An agent draws while you talk, built on TLDraw
(techstackups.com)
53 points
by
jameswhitford
3mo ago
|
16 comments
15.
▲
by
jameswhitford
3mo ago
That is a great suggestion that I am definitely going to look into, thanks!
16.
▲
by
jameswhitford
3mo ago
I hear you
17.
▲
by
jameswhitford
3mo ago
Cool to hear, what kind of tasks have you been using GLM for? And what other models have you found useful through Ollama?
18.
▲
by
jameswhitford
3mo ago
I see your point. Just the fact that one model does have vision and one does not might be an interesting point of comparison, however.
19.
▲
by
jameswhitford
3mo ago
This is excellent feedback thank you! These LLMisms in writing are a challenge I am living with currently and trying to improve on. The technical writing industry is taking a huge knock right now with companies demanding more work in less t
20.
▲
by
jameswhitford
3mo ago
Hi, author here, can you link? I would love to read about this.
21.
▲
by
jameswhitford
3mo ago
Yes I agree 100%. My next guide would do better to use identical harnesses.
22.
▲
by
jameswhitford
3mo ago
GLM 5.2 is text only, not multi modal. And Opus is multi modal.
23.
▲
by
jameswhitford
3mo ago
Hi, author here, I cannot give an exact number for how many token the verification step took, but the verification GLM 5.2 ran was very stupid and definitely a waste of time. It read the pixel color data to try and verify the scene rendered
24.
▲
by
jameswhitford
3mo ago
Yes I 100% agree. Time-taken can be improved (with harnesses, subagent workflows etc.) and varies based on task.
25.
▲
by
jameswhitford
3mo ago
Yes, part of the reason I chose the one-shot test was really to test long-running tasks. A lot of people seem to be experimenting with this format, for example in the now trending loop-writing workflows. And really I am interested in diving
26.
▲
by
jameswhitford
3mo ago
I appreciate the feedback!
27.
▲
by
jameswhitford
3mo ago
Yes this is true. This test was run on a $20 pro Claude subscription. I would definitely love to try use both models on the highest plans for a whole month and compare the two, great format for a future head-to-head comparison.
28.
▲
by
jameswhitford
3mo ago
Hi, I am the author, I completely agree! I set out to run a vibe test on this one, not a benchmark, the real benchmarks are listed. My test shows what the models can do when both tasked with a long-running, technically difficult, one-shot t
29.
▲
Claude is skeptical about OpenClaw
(wecreatethis.com)
2 points
by
jameswhitford
5mo ago
|
2 comments
30.
▲
by
jameswhitford
5mo ago
I asked Claude Code to research Openclaw. It spawned a subagent, got back detailed results, and then flagged them as unreliable and/or hallucinated before I could read them. TL;DR: Claude isn't trained on openclaw data due to its
More ›