Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
thegeomaster
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
9 ms
·
61.
▲
by
thegeomaster
1y ago
This is kinda frustrating to read. The style is very busy, and it lacks a clear structure. It's basically an information dump without any acknowledgment of what's important or not. Big-O notation is provided for a lot of operation
62.
▲
by
thegeomaster
1y ago
Gemini 2.5 Pro is getting thinking budgets when it GAs in June (at least that's the promise).
63.
▲
by
thegeomaster
1y ago
> Since then OpenAI has refreshed the weights behind the “gpt-4.1” alias a couple of times, and one of those updates fixed the em-dash miss. I don't know where you are getting this information from... The only snapshot of gpt-4.1 is
64.
▲
by
thegeomaster
1y ago
My point still stands --- the reasoning tokens being consumed to interpret the abstracted llms.txt could have been used for solving the problem at hand. Again, I'm not saying the solution doesn't work well (my intuition on LLMs ha
65.
▲
by
thegeomaster
1y ago
What is absolutely essential to present here, but is missing, is a rigorous evaluation of task completion effectiveness between an agent using this format vs the original format. It has to be done on a new library which is guaranteed not to
66.
▲
by
thegeomaster
1y ago
I slightly tweaked your baseline em dash example and got 100% success rate with GPT-4.1 without any additional calls, token spend, or technobabble. System prompt: "Remove every em-dash (—) from the following text while leaving other ch
67.
▲
by
thegeomaster
1y ago
Absolutely stellar for 0-to-1-oriented frontend-related tasks, less so but still quite useful for isolated features in backends. For larger changes or smaller changes in large/more interconnected codebases, refactors, test-run-fix-loop
68.
▲
by
thegeomaster
1y ago
This is exactly what frameworks like Claude Code, OpenAI Codex, Cursor agent mode, OpenHands, SWE-Agent, Devin, and others do. It definitely does allow models to do more. However, the high-level planning, reflection and executive function s
69.
▲
by
thegeomaster
1y ago
Tested: https://chatgpt.com/share/68188ed9-e25c-8002-b4f8-460a5d1efe...
70.
▲
by
thegeomaster
1y ago
The LLM provides a correct answer to the questions above: https://chatgpt.com/share/68188ed9-e25c-8002-b4f8-460a5d1efe...
71.
▲
by
thegeomaster
1y ago
Thank you for your contribution!!! Does min_p help with non-creative-writing tasks as well, such as maths or coding? Is there any way to improve/tune the performance for these with sampling? Do you have a recommendation/guide on t
72.
▲
by
thegeomaster
1y ago
Is it possible to share the picture? I've been looking for exactly that kind of jump the other day when playing around.
73.
▲
by
thegeomaster
1y ago
Try this prompt to give it a CoT nudge: Where exactly was this photo taken? Think step-by-step at length, analyzing all details. Then provide 3 precise most likely guesses. Though I've found that it doesn't even need that f
74.
▲
by
thegeomaster
1y ago
For all of the images I've tried, the base model (e.g. 4o) already has a ~95% accurate idea of where the photo is, and then o3 does so much tool use only to confirm its intuition from the base model and slightly narrow down. For OP
75.
▲
by
thegeomaster
1y ago
See my other comment: https://news.ycombinator.com/item?id=43804041 4o can do it almost as well in a few seconds and probably 10-50x fewer tokens: https://chatgpt.com/share/680ceeff-011c-8002-ab31-d6b4c
76.
▲
Deepfake porn is destroying real lives in South Korea
(cnn.com)
12 points
by
thegeomaster
1y ago
|
1 comments
77.
▲
by
thegeomaster
1y ago
It's a mix between the Transformer architecture and diffusion, shown to provide better output results than simple autoregressive image token generation alone: https://arxiv.org/html/2408.11039v1 Of course, nobody
78.
▲
by
thegeomaster
1y ago
Well, there's also gemini-2.0-flash-exp-image-generation. Also autoregressive/transfusion based.
79.
▲
by
thegeomaster
1y ago
Is it possible that a community of people who are constantly pushing LLMs to their limits would be most aware of their limitations, and so more inclined to think they are junk? In terms of business utility, Google has had great releases eve
80.
▲
by
thegeomaster
2y ago
Same here. There was a fair amount illicit copying of textbooks too. But this teacher doesn't even want to share his lecture slides.
81.
▲
by
thegeomaster
2y ago
https://archive.is/5Hekg
82.
▲
The Government Knows AGI Is Coming
(nytimes.com)
3 points
by
thegeomaster
2y ago
|
1 comments
83.
▲
by
thegeomaster
2y ago
> I don't know if it's ignorance or an attempt at a publicity stunt on the author's part, but it isn't at all what they claim. I let the author know on Twitter too: https://x.com/thegeomaster/stat
84.
▲
by
thegeomaster
2y ago
This is total bullshit. It's clear by spending 2 minutes with the output, located on https://github.com/ghuntley/claude-code-source-code-deobfusc... . The AI has just made educated guesses about the functionality,
85.
▲
by
thegeomaster
2y ago
Thank you to the team. Looks like a great release. Already switching existing prompts to Claude 3.7 to see the eval results :)
86.
▲
by
thegeomaster
2y ago
This amounts to a cost-saving measure - you can generate arbitrarily many tokens by appending the output and re-invoking the model.
87.
▲
by
thegeomaster
2y ago
...by piping it through the world's most inefficient echo function.
88.
▲
Carthagine: Pixel Perfect Figma to React AI
(carthagine.ai)
2 points
by
thegeomaster
2y ago
|
0 comments
89.
▲
by
thegeomaster
2y ago
Every bubble has a narrative.
90.
▲
Vision language models are blind (2024) [pdf]
(arxiv.org)
3 points
by
thegeomaster
2y ago
|
0 comments
More ›