3 ms·
yes, the size of the data you can process is limited. I had reasonable results with something like 50MiB, after tweaking jvm parameters. One limitation I find
by febeling 7y ago
yes, the size of the data you can process is limited. I had reasonable results with something like 50MiB, after tweaking jvm parameters.
One limitation I find annoying is that you can't easily strip your transformation from the data and apply it to similar but different data. It's all together "a project."
So incrementally transforming data works well, applying that transformation elsewhere not really. That bothered me more than the size limitation, which I think is not limiting most use cases with real “messy” data anyway. Maybe processing large volume event / log data would still need something else.
I used it to look at a json export from a SaaS tool, and to convert it to table structure. Cleaning field contents, which where following certain conventions, but which evolved over time, things like that. For such use-cases it's powerful.
- _frkl 7y ago> So incrementally transforming data works well, applying that transformation elsewhere not really. That bothered me more than the size limitation, which I think is not limiting most use cases with real “messy” data anyway. Maybe processing large volume event / log data would still need something else. Agreed, this would be the one feature I would really need. Would be nice to be able to setup (and refine over time) pipelines to automatically clean up new data.