3 ms·
This matches what I found experimentally with gpt-oss-20b during OpenAI's red-teaming challenge. After moving from hosted inference to running the model myself
by ndr_ 20d ago
This matches what I found experimentally with gpt-oss-20b during OpenAI's red-teaming challenge. After moving from hosted inference to running the model myself on rented H100s via vast.ai, I saw the model refuse the same kinds of prompts at noticeably different rates depending on the inference stack — differences of roughly 5–10 percentage points with otherwise identical experimental parameters and seeds.
So I very much agree that this isn't necessarily about providers secretly changing the weights. For reproducible work, the serving stack - engine, version, hardware, configuration, and probably more - really belongs in the methodology alongside the model itself.
I wrote up the results here: “In AI Sweet Harmony” (arXiv:2510.01259).
- epistasis 20d agoGreat concrete experience! I'd love to see some published token log_probs for given prompts and seeds that accompany a model card. Not sure if that's enough, thigh. But then it's been so long since I dealt with the api directly that I don't even know if modern models show the top tokens and probabilities any more.
- dang 20d agoCan you please not post AI-generated or AI-edited comments to HN? It's not allowed here - see https://news.ycombinator.com/newsguidelines.html#generated https://news.ycombinator.com/newsguidelines.html#generated and https://news.ycombinator.com/item?id=47340079 https://news.ycombinator.com/item?id=47340079. Of course, it's impossible to know for sure what was LLM processed or not, but some of your posts (like this one) have been getting classified that way. (And if this was a false positive, I apologize!)