4 ms·
I spin up a gpu instance in a cloud, run my model via vllm, connect to it via an ssh tunnel. done.
by blinded 3mo ago
I spin up a gpu instance in a cloud, run my model via vllm, connect to it via an ssh tunnel. done.
- stuxnet79 3mo agoCan you elaborate on the first step? Which cloud and which service? What's the cost outlay if you are just having a convo and not doing anything 'agentic'?
- blinded 3mo agoHey sure. It depends, but usually spin up an h100 on lambda.ai or coreweave. They have capacity and their UIs/APIs are nice. I spin it up for an hour or two, believe it was 6~ dollars an hour. Once the gpu instance is up, you need to run vllm and a model, ie https://docs.lambda.ai/education/large-language-models/deploying-nemotron-3-nano/ https://docs.lambda.ai/education/large-language-models/deplo.... Then you can connect your pi.dev, openwebui, etc etc to vllm and interact with it like normal.