3 ms·
I was hoping it would work with vLLM (openai compatible) to test it, does anyone know a similar proxy for local coding models?
by somethingsome 5mo ago
I was hoping it would work with vLLM (openai compatible) to test it, does anyone know a similar proxy for local coding models?
- DeathArrow 5mo agoCheck this: https://github.com/antoinezambelli/forge/tree/az/vllm https://github.com/antoinezambelli/forge/tree/az/vllm
- zambelli 5mo agoYeah I got it working as a quick test run to confirm a model issue vs backend issue on a consumer app. It worked on my dual-5070 Ti rig, but I didn't have time to formalize all the way and merge it in. Thanks for linking it!
- somethingsome 5mo agoThanks, I just tried, for me it worked on 2x L40S with vLLM. I had some issues due to the model name, forge was forwarding 'default' instead of the real model name 'Qwen2.5-Coder-14B-Instruct'. If someone else struggle on this step, I added in vLLM args: --served-model-name "Qwen2.5-Coder-14B-Instruct" --served-model-name "default" So default becomes an alias. I didn't yet test Forge, I was just happy that it worked at the moment ;)
- zambelli 5mo agoOh that's a good find, I'll book ark this for a GitHub issue. Glad to hear it's working!