3 ms·
Spark sits on top of YARN/Mesos, and is used for data processing scalability that pandas can't handle. Personally, I think two areas often lacking are software
by jwilbs 9y ago
Spark sits on top of YARN/Mesos, and is used for data processing scalability that pandas can't handle.
Personally, I think two areas often lacking are software development skills and general statistics knowledge. The former is necessary for writing production-quality code, assisting with an sort of data engineering pipeline, writing reliable, reusable code, and creating custom solutions. Unfortunately, the latter is often skimped on (if not skipped entirely) in favor of more 'hot' fields like ml/dl, with the result being a fuzzy understanding across the board. (You'd be amazed at the quantity of candidates lacking fundamental knowledge about glm's, basic nonparametric stats, popular distributions, etc).