3 ms·
They didn't do this before because the memory overhead of having a separate worker for each document would have made the infrastructure costs exorbitant, presum
by apendleton 8y ago
They didn't do this before because the memory overhead of having a separate worker for each document would have made the infrastructure costs exorbitant, presumably (since you'd have to pay for whatever fixed overhead costs Node has). But the lower overhead of using a Rust process per doc instead of a JS process per doc has allowed them to move to that model.
- hsaliak 8y agoThere can be an M:N mapping - An update queue can be consumed by a fixed pool of workers that can lock a file and commit the update. This is crude but wont explode the number of workers.. if you have a hashing scheme for the worker pool and implement linear probing, you can achieve some degree of preference for the same worker. If their system was such that a worker maintains state for a bunch of docs, and persists them at checkpoints, while also being doing compute intensive tasks per document on a single threaded node instance, I would ask why it was designed like that in the first place.