4 ms·
Is anyone running Marathon in production? Real production. The kind where any downtime means lost money. I see a lot of intro-level tutorials, but almost nothi
by Wilya 11y ago
Is anyone running Marathon in production? Real production. The kind where any downtime means lost money.
I see a lot of intro-level tutorials, but almost nothing on the more advanced side.
My (completely casual) experience with Marathon is pretty bad, with the main process crashing quite regularly even under no load, so I'm wondering if people who write about these systems have actually used them for non-trivial tasks. And for something as critical as Marathon, which is supposed to handle... well... all my services, I'd rather be sure that the system is rock solid.
(This is specifically about Marathon. Mesos itself has proven more reliable)
- brndnmtthws 11y agoYes, many people have in fact run it in "real production". Go ahead and Google my name for credibility. There were indeed issues with the 0.7.x series of Marathon, but we've made a big effort to focus on stability and performance in 0.8.x, and onward. As with any new software project, there tend to be issues in early releases.
- steve0ps 11y agoI've been running Marathon in production (real production) to power more than 100 applications for the past six months. I chose it because it seemed like the most stable thing at the time; however, quickly found it was not production ready. While many of the original issues I encountered in 0.7.x were resolved with the 0.8.x release, 0.8.x brought new issues such as stuck deployments, etc. Additionally, I have found the upgrade path to be obtrusive and frankly scary. I am actively moving away from Marathon because of these issues. Marathon does not make using Docker or building microservices simple. There are many important pieces that Marathon does not provide. Sure your operations team can tie in Mesos-DNS / Bamboo / Consul / whatever else, but it's going to take time, requires a specialized team, and leaves you feeling nervous about what happens if everything crashes in the middle of the night. Even when tying in these third party tools, it is likely you will have to make significant code updates to utilize features such as service-discovery / SRV records. You will inevitably end up with a hobbled-together system that needs serious support from your operations team. I am fairly frustrated as a whole with Mesosphere, and expected more from a company who raised so much capital.
- nuschk 11y agoMight I ask: If you're moving away from Marathon, where are you moving to? Kubernetes?
- nemothekid 11y agoI wouldn't expect Marathon to do service discovery for you, as I believe that is better left to something like Mesos-dns/Consul which marathon can supervise for you, and docker integration has been fairly simple. In any case I found my marathon was not without issues, like failover causing every application to restart (I think this was fixed in 0.8.2), or the fact that marathon tends to use 2x as much RAM as Zookeeper or Mesos-Master (I run the 3 on the same node). Have you seen the aurora apache project? It solves the same problems as marathon, and its creators claim it was built to handle stability. I originally chose marathon as JSON configuration over REST was easier to wrap my head around, but was this something you tried and how did it work for you?
- tjkells 11y agoRunning a complete hosted telephony service using nearly the exact stack defined in this article - https://developers.corvisa.com/ https://developers.corvisa.com/ It has worked remarkably well and allowed us to scale up/down during peak hours or unexpected high traffic peaks.