3 ms·
Didn’t know about it. Is it possible to use only that subset of apache spark lib?
by mmalek06 7y ago
Didn’t know about it. Is it possible to use only that subset of apache spark lib?
- anon176 7y agoYes, you can run apache spark as a single node really easily. Then once its running you can fire up the Scala or Python shell that it comes with. After that, it is just a matter of issuing the statements to setup the data set then issues queries against it.
- khc 7y agoNot many people know this but Databricks offers https://community.cloud.databricks.com/ https://community.cloud.databricks.com/ for free which allows you to run simple spark notebooks. Disclosure: works for Databricks but not on spark
- mmalek06 7y agoStill, it seems like a lot of work. My case was that I just wanted to feed JSON document into some mechanism that will give me back what I want. I wouldn't like to embed whole Spark framework into my app. I will read about it though :)
- EdwardDiego 7y agoI imagine you could reuse Catalyst to generate queries against JSON derived Datasets with some work. Starting point for some reading maybe? https://github.com/apache/spark/blob/e2d3983de78f5c80fac066b7ee8bedd0987110dd/sql/core/src/main/scala/org/apache/spark/sql/SparkSession.scala#L603 https://github.com/apache/spark/blob/e2d3983de78f5c80fac066b...
- yami-sama 7y agoIf you need to parse semi-structured documents using Spark, Quenya-DSL is much better than SQL, specially if you want to flatten the data. https://github.com/music-of-the-ainur/quenya-dsl https://github.com/music-of-the-ainur/quenya-dsl