3 ms·
What if instead of trying to evaluate these models privately, which tends to have a lot of overhead, we instead try to mix lots of user queries and send them in
by blintz 4y ago
What if instead of trying to evaluate these models privately, which tends to have a lot of overhead, we instead try to mix lots of user queries and send them in batches?
We could use enclaves to do the mixing, and while there’s be added latency, we could achieve model outputs that are (by definition) at the current state of the art for LLMs. We would not hide the contents of queries, but we’d at least hide who is making which queries.