4 ms·
Incredible piece of software. I've used it in production at my last two jobs. You can build almost anything in NiFi once you get into the mindset of how it work
by monstrado 6y ago
Incredible piece of software. I've used it in production at my last two jobs. You can build almost anything in NiFi once you get into the mindset of how it works.
A good way to get started with NiFi is to use it as a highly available quartz-cron scheduler. For example, running "some process" every 5 seconds.
Disclaimer: I'm an Apache NiFi committer.
An article you might find interesting about it's ability to scale.
https://blog.cloudera.com/benchmarking-nifi-performance-and-scalability/ https://blog.cloudera.com/benchmarking-nifi-performance-and-...
Disclaimer v2: I used to work at Cloudera
- nightowl_games 6y agoThis page is classic Apache project in that I read it and have no idea what it does. Can you high level explain what this thing is really for?
- taftster 6y agoAgreed. So here's an attempt to describe NiFi at a high level. Fundamentally NiFi is a "dataflow engine", a system that can be used to automate data transfer from different and varying types of sources and sinks. It has a fairly usable UI that enables a "dataflow manager" (end user) to perform transformation, routing and delivery of data using a "drag-n-drop" configuration approach. Getting data into or out of your application/system, or performing simple schema transformations, is a common (maybe tedious) task that most developers face. NiFi helps connect the dots, so to speak, and decouples the receipt/delivery of data away from your application. NiFi comes with a set of "batteries included" connectors for almost every transport protocol you would generally need. And it's modular so you can create your own processing components as well. NiFi is fundamentally modeled after what's called "Flow-Based Programming"[1], which is a style of programming that facilitates composition of black-box processing units. It can run at an enterprise or IoT level, depending on where that decomposition best fits into your architecture. [1] https://en.wikipedia.org/wiki/Flow-based_programming https://en.wikipedia.org/wiki/Flow-based_programming (disclaimer: I'm affiliated with the NiFi project)
- _JamesA_ 6y agoIs this a layer on Apache Camel [1] or something completely different? [1]: https://camel.apache.org/ https://camel.apache.org/
- monstrado 6y agoIt was built from scratch at the NSA and open sourced a few years ago.
- zmmmmm 6y agoI'm curious how you would compare it to Apache Camel - when would you use Camel and when would you use this? We use camel with DSLs to make programmatic workflows that connect data flows together. However Camel itself doesn't typically carry the data. Sometimes it SFTPs files around etc., but mostly, it is just a messaging layer. Is that the main difference here?
- PeterStuer 6y agoSo is this comparable to something like the old Microsoft Biztalk?
- mindentropy 6y agoIs this similar to NodeRed for IoT? I was trying to bring up something similar for IoT with CEP rules and source and sink connections on the Edge.
- monstrado 6y agoI fine that explaining NiFi tends to be difficult, especially if you're an advanced user. This is because NiFi is extremely general purpose. I've used it for many different types of projects. That being said, here's a few use cases that are relatively easy to get started. 1. Move data from A to B 2. Move data from A to B, but perform some intermediate processing in-between (ETL) 3. Grab data from A, write it to [INSERT SYSTEM HERE] (Mongo, Kafka, many others) 4. Periodically run an arbitrary script or program (like cron but you can schedule in seconds) Those are just some use cases that people tend to start off doing. It's important to note...the above doesn't sound very impressive until you think about the following features: 1. Build these pipeline completely from a UI. You can leverage custom code, but the number of processors that come out of the box is astounding. If you can think of it, it most likely already has something built-in. 2. By default safety nets. Data is backed by a high performance WAL that allows for recovering in the case of a failure. 3. Ability to pause and resume specific areas of a pipeline at any time. For example, you may have a processor that's receiving data and perhaps the database you write the data to is offline (or has moved). Data is automatically buffered on disk so when the DB comes back online, you can backfill the data that you've been consuming. Further, none of the data was dropped in this process. 4. Tunable and verifiable back pressuring out of the box (_very_ hard to get right when writing custom code) 5. Easy to prioritize certain events over others (e.g. higher priority, latency sensitive data) Lastly, when you start getting comfortable using NiFi, you will realize how general purpose the engine is. Data is represented as a "FlowFile". A "FlowFile" is essentially just a Array<Byte>. This means you can operate on virtually anything. Further, you can operate on large data since its backed by a WAL and NiFi facilitates the ability to stream data from disk when you want to process it. As long as you don't read the entire thing into memory (of course). Easiest way to get started is to just download the binaries and run it. Then go to the http://localhost:8080/nifi http://localhost:8080/nifi and just mess around. There's plenty of tutorials online (Blogs, YouTube, etc). Once you get comfortable using NiFi, people will be blown away with how fast you can get something up and running. Things that would normally take days or weeks to get into production, you can routinely do it in hours or minutes. Hope this helps! edit: You can download from https://nifi.apache.org/download.html https://nifi.apache.org/download.html. Just unzip/untar the package and run "./bin/nifi.sh run"
- 6y ago
- Random_ernest 6y agoThanks for your honesty, I thought I was the only one.
- chirau 6y agoQuick question, what does the role of Data Engineer at Epic Games entail and what technologies are you working with?