Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
sgk284
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
11 ms
·
31.
▲
Show HN: Logic, Inc. – Automate human-in-the-loop fuzzy decisions
(logic.inc)
3 points
by
sgk284
1y ago
|
0 comments
32.
▲
Every GPT-5 coding example implemented with Opus 4.1
(gpt-5-vs-opus-4-1-coding-examples.vercel.app)
6 points
by
sgk284
1y ago
|
1 comments
33.
▲
by
sgk284
1y ago
We use Claude Code pretty aggressively at our startup, so were naturally curious to compare the coding examples that OpenAI published today to Opus 4.1. All of these were one-shotted by Claude. (OpenAI's GPT-5 Coding Examples: https:&
34.
▲
by
sgk284
1y ago
The two mechanisms are a bit disjoint, so I don't think it's the right tool to do so. Though it could have been an interesting experiment.
35.
▲
by
sgk284
1y ago
Yea, absolutely. That's a good point. We could have phrased that more clearly.
36.
▲
by
sgk284
1y ago
Fun post! Back during the holidays we wrote one where we abused temperature AND structured output to approximate a random selection: https://bits.logic.inc/p/all-i-want-for-christmas-is-a-rando...
37.
▲
When AI Writes the Code, What Do Programmers Do?
(bits.logic.inc)
4 points
by
sgk284
2y ago
|
3 comments
38.
▲
by
sgk284
2y ago
> A more robust approach would be to give the whole reasoning to an LLM and ask to grade according to a given criterion We actually use a variant of this approach in our reasoning prompts. We use structured output to force the LLM to t
39.
▲
by
sgk284
2y ago
Another take on a similar idea from FAIR is a Large Concept Model: https://arxiv.org/pdf/2412.08821
40.
▲
Predicting the Super Bowl with LLMs
(bits.logic.inc)
2 points
by
sgk284
2y ago
|
0 comments
41.
▲
The AI Programming Paradox
(bits.logic.inc)
1 points
by
sgk284
2y ago
|
0 comments
42.
▲
by
sgk284
2y ago
In our case, yes we treat them the same. Though it might be interesting to decouple them. You could, for example, include all few-shots that meet the similarity threshold, but you’ll use more tokens for (I assume) marginal gain. Definitely
43.
▲
by
sgk284
2y ago
One point of confusion might be that this is a tough but relatively low-value task (on a per-unit basis). The budget per item moderated is measured in small double-digit cents, but there's hundreds of thousands of items regularly bei
44.
▲
by
sgk284
2y ago
re: 90% – this particular case is a fairly subjective and creative task, where humans (and the LLM) are asked to follow a 22 page SOP. They've had a team of humans doing the task for 9 years, with exceptionally high variance in perform
45.
▲
by
sgk284
2y ago
We have, and it works great! We currently do this in production, though we use it to help us optimize for consistency between task executions (vs the linked post, which is about improving the capabilities of a model). Phrased differentl
46.
▲
by
sgk284
2y ago
Thanks! Great suggestion for improving the graphs – I just updated the post with axis labels.
47.
▲
by
sgk284
2y ago
Over the holidays, we published a post[1] on using high-precision few-shot examples to get `gpt-4o-mini` to perform similar to `gpt-4o`. I just re-ran that same experiment, but swapped out `gpt-4o-mini` with `phi-4`. `phi-4` really blew me
48.
▲
Getting GPT-4o-mini to perform like GPT-4o
(bits.logic.inc)
2 points
by
sgk284
2y ago
|
0 comments
49.
▲
The role of embeddings and rerankers in semantic retrieval
(bits.logic.inc)
2 points
by
sgk284
2y ago
|
0 comments
50.
▲
Santa Paws Is Coming to Town
(bits.logic.inc)
2 points
by
sgk284
2y ago
|
0 comments
51.
▲
Using LLMs to Make Truly Random Decisions
(bits.logic.inc)
2 points
by
sgk284
2y ago
|
0 comments
52.
▲
Ho-Ho-How to Make ChatGPT Sound Like Santa Claus
(logicinc.substack.com)
2 points
by
sgk284
2y ago
|
0 comments
53.
▲
by
sgk284
3y ago
The substantial original work here is the prompt authorship. That’s the “paintbrush” in this work. And the “paint” is the AI. It’s very easy to spend hours iterating on prompts for the perfect photo while using human judgment to figure out
54.
▲
by
sgk284
3y ago
> if someone steals my physical device, then they have full access Apple protects passkeys via FaceID or TouchID. If you're satisfied with biometrics as a 2nd factor, then there is no regression in your scenario.
55.
▲
by
sgk284
3y ago
Humans learn on copyrighted works as a matter of standard training. And certainly humans can memorize those works and replicate them – and we rely on the legal system to ensure that they don't monetize them. The same will apply to neur
56.
▲
by
sgk284
3y ago
OpenAI touches a little on this on page 12 of the GPT-4 technical report ( https://cdn.openai.com/papers/gpt-4.pdf ). Prior to aligning to safer outputs, the model's confidence in an answer is highly correlated with
57.
▲
by
sgk284
3y ago
Safer means constraining the kinds of answers the model will provide (e.g. it won't try to talk you into committing self-harm, it won't teach you how to make a break laws, etc...). It will generally avoid sensitive topics. Is &quo
58.
▲
by
sgk284
3y ago
Tweaked my comment a little bit to clarify. Surely we'd agree that not everyone who uses software is a software engineer, but writing software is generally agreed upon as "software engineering". If you consider my comment as
59.
▲
by
sgk284
3y ago
Author of the guide here. I attempt to address this in the "Why do we need prompt engineering?"[^1] section. > ... we used an analogy of prompts as the “source code” that a language model “interprets”. Prompt engineering is the
60.
▲
by
sgk284
3y ago
Thanks for pointing this out. That was my mistake – my brain must have swapped out "different transformer architectures" with "different model architectures". I just updated the guide: https://github.com/
More ›