3 ms·
Since they don’t seem to be able to give a simple answer: the inference does not run in the worker. It connects to external GPUs.
by pseg134 3y ago
Since they don’t seem to be able to give a simple answer: the inference does not run in the worker. It connects to external GPUs.
- eastdakota 3y agoI think the confusion is what is meant by "in the Worker." From a hardware perspective, the GPU may be in the same machine as the CPU that's powering the Worker. Or they may be across different machines in our network. We are not routing requests to some third party. And we will try to run the inference task as close as possible to who/whatever requested it. The whole idea of "serverless" is you shouldn't have to worry about what machine where runs whatever unless you're on the team building the scheduling and routing logic at Cloudflare.
- thegagne 3y agoI think his question is more about does the worker directly access the GPU and thus require js tooling to handle the GPU somehow (no), or does it make subrequests to a separate GPU service not running the worker runtime (yes).