4 ms·
Pretty cool. You can do this in apache spark already but this looks a lot quicker to get going with.
by anon176 7y ago
Pretty cool. You can do this in apache spark already but this looks a lot quicker to get going with.
- mmalek06 7y agoDidn’t know about it. Is it possible to use only that subset of apache spark lib?
- anon176 7y agoYes, you can run apache spark as a single node really easily. Then once its running you can fire up the Scala or Python shell that it comes with. After that, it is just a matter of issuing the statements to setup the data set then issues queries against it.
- khc 7y agoNot many people know this but Databricks offers https://community.cloud.databricks.com/ https://community.cloud.databricks.com/ for free which allows you to run simple spark notebooks. Disclosure: works for Databricks but not on spark
- mmalek06 7y agoStill, it seems like a lot of work. My case was that I just wanted to feed JSON document into some mechanism that will give me back what I want. I wouldn't like to embed whole Spark framework into my app. I will read about it though :)
- EdwardDiego 7y agoI imagine you could reuse Catalyst to generate queries against JSON derived Datasets with some work. Starting point for some reading maybe? https://github.com/apache/spark/blob/e2d3983de78f5c80fac066b7ee8bedd0987110dd/sql/core/src/main/scala/org/apache/spark/sql/SparkSession.scala#L603 https://github.com/apache/spark/blob/e2d3983de78f5c80fac066b...
- yami-sama 7y agoIf you need to parse semi-structured documents using Spark, Quenya-DSL is much better than SQL, specially if you want to flatten the data. https://github.com/music-of-the-ainur/quenya-dsl https://github.com/music-of-the-ainur/quenya-dsl