3 ms·
Very good point, it's about exploiting inference engine itself, and not the agentic stuff. I found it lacking details. All these things do is split the input i
by drillsteps5 1mo ago
Very good point, it's about exploiting inference engine itself, and not the agentic stuff.
I found it lacking details. All these things do is split the input into tokens, encode them, feed to the input, run the inference, collect the output, decode, convert to tokens, assemble the final text, and output (leaving aside multi-modal capabilities for now). What exactly can be exploited here? Aside from usual things like buffer overruns etc.
It does mention that vLLM at some point ran eval() on final output (?) which was confusing, why would it do that? It's an inference engine, not an agent.
So interesting topic, but lacks details.
Edit: many (all?) of them have http endpoints, so that obviously can be exploited, but I don't think it would qualify as exploiting inference, it's just hacking the http service.
- angry_octet 1mo agoYou get to the engine via http. Also once you've exploited the engine instance / host you can C2 via http. We're not talking about vulnerabilities in http. There is lots of surface inside the engine, see links in https://news.ycombinator.com/item?id=49441417 https://news.ycombinator.com/item?id=49441417
- yencabulator 1mo ago> It does mention that vLLM at some point ran eval() on final output (?) which was confusing, why would it do that? It's an inference engine, not an agent. There's a couple of different conventions for what the LLM generates for tool calls, I think the code in vLLM is converting it from whatever Qwen3 was trained for to whatever convention the HTTP API wants to expose. That part about using eval in 2025 got me to add "#naive" to my notes about vLLM. Total WTF. This should never have been done. https://github.com/vllm-project/vllm/pull/21396#discussion_r2223397938 https://github.com/vllm-project/vllm/pull/21396#discussion_r...