2 ms·
> Databricks is often managing a million-ish Spark-sub tasks for various users. They couldn't do that using traditional operating system scheduling techniques:
by arhyth 3y ago
> Databricks is often managing a million-ish Spark-sub tasks for various users. They couldn't do that using traditional operating system scheduling techniques: they needed something that could scale. The obvious answer was to put all scheduling information into a database. That's exactly what the Databricks guys did: they put it all in a PostgreSQL database, and then started whining about Postgres performance," says Stonebrake.
anyone know of a talk/paper discussing this scheduling "hack" in more detail
- sdave 3y agohttps://arxiv.org/abs/2007.11112 https://arxiv.org/abs/2007.11112 The link is referenced in the article.