2 ms·
LLM inference takes up some scarce GPU time and many people are trying to use free entrypoints to build services instead of the intended paid APIs, so I underst
by debugnik 3y ago
LLM inference takes up some scarce GPU time and many people are trying to use free entrypoints to build services instead of the intended paid APIs, so I understand why those services want to put limits on usage.
Programming playgrounds however are freely available for pretty much every mildly popular language, and these days many toolchains can even be compiled and run with JS or WASM so one could just serve some static files to host it. This is definitely more suspicious than what OpenAI and other ML companies are doing.