5 ms·
MarkLogic is an XML document store like Oracle is a CSV file store. We store compressed trees. We require memory - like all databases - because we actually h
by arosenbaum 13y ago
MarkLogic is an XML document store like Oracle is a CSV file store.
We store compressed trees. We require memory - like all databases - because we actually have rich indexes.
Why? Scaling by requiring massive servers is Oracle's story, not MarkLogic. We scale out on commodity servers in a cluster just fine. However, we do not magically create more capacity when adding additional VM's on top of over-subscribed network and storage systems.
We don't require proprietary shared storage systems (local, NFS, HDFS - just give us bandwidth) or proprietary networks (TCPIP please, 10G preferred, don't need FC).
Cheapest way to scale is a cluster of decent sized machines with a decent amount of IO capacity each.
{Q: What do you call a database without indexes, transactions or security?
A: A file system!}
- hga 13y agoOuch. As in, per this NYT article: http://www.nytimes.com/2013/11/23/us/politics/tension-and-woes-before-health-website-crash.html?hpw&rref= http://www.nytimes.com/2013/11/23/us/politics/tension-and-wo... the contractors weren't familiar with its paradigm. They got started so late (not until February in earnest from what we've heard), it sounds like they just didn't have the time to learn it ... in a project (micro-)managed by political and bureaucratic types in the government. Not learning it well enough would likely go along with not provisioning the necessary resources. Usual lesson of "don't try to do too many new things if you're on a tight time and/or money budget". Unfortunately, given that this is almost entirely a political exercise, I don't see how you're going to avoid getting some scapegoating.
- arosenbaum 13y agoFrom report issued Sunday morning: http://www.hhs.gov/digitalstrategy/sites/digitalstrategy/files/pdf/healthcare.gov-progress-report.pdf http://www.hhs.gov/digitalstrategy/sites/digitalstrategy/fil... "...the root causes for these site flaws to be hundreds of software bugs, insufficient hardware and infrastructure." The infrastructure flaws cannot be chalked up to "not learning it well enough"....There is no database - not Oracle, not PostgreSQL, not MongoDB....and not MarkLogic - that would have handled the traffic on this hardware and infrastructure.
- hga 13y ago"There is no database ... that would have handled the traffic on this hardware and infrastructure" That's what I meant by "not provisioning the necessary resources". I'm a programmer who dabbles in small scale systems building (from scratch, as in mount CPU on motherboard, etc.), so I'm perhaps not using the word "provisioning" as domain experts do. But to the extent MarkLogic was the major database used (I'm getting that impression the more I look into this), unfamiliarity plus all the bad management could have contributed to not procuring beefy enough infrastructure. Or helped contribute to the widespread magical thinking, i.e. an experienced Oracle DBA could say with authority "this won't work" but not have as much weight when talking about MarkLogic. And I'm sure you had at least one field engineer helping them ... or I should say trying to help them. Ditto, I'm sure, the people or contractors in/for CMS who were already using MarkLogic, assuming the ones who really knew their stuff were even consulted. Engineers not being listened to/respected in favor of magical thinking is obviously one of the biggest problems with this project. What can you say when the integration testing is delayed for the last 2 weeks before launch, proves it can't work the week before, and launches anyway?
- arosenbaum 13y agoNot that much controversy over that. HHS said most of the same this morning...I (and many folks) disagree with the premise that MarkLogic had anything to do with it. If the exact same processes and analysis were applied to a LAMP stack or an Oracle Exa-stack, the results would have likely been the same. I think the public news about the firewall sizing (4G instead of 50G) supports my claim. So far, there have been two contractors that have been changed out - QSSI took over from CGI to lead rebuild and, recently, Terramark to be replaced by HP. Maybe this has absolutely nothing to do with not being familiar with MarkLogic. http://kellblog.com/2013/12/01/the-pillorying-of-marklogic-why-selling-disruptive-technology-to-the-government-is-hard-and-risky/ http://kellblog.com/2013/12/01/the-pillorying-of-marklogic-w... There was a healthcare exchange built 100% on Oracle in Oregon (Oracle team, Oracle packaged software including Siebel, Peoplesoft, IDM, Oracle integration SW + people,Oracle infrastructure, Oracle hosting). It's not going particularly well. I don't think familiarity of technology has one thing to do with it. I do think that MarkLogic's ability to be agile - programming (EasyApp), infrastructure (speed of transition) and performance - have a great deal to do with the speed of the team being able to deal with the software bugs above MarkLogic and the weak infrastructure around us. (For those joining this conversation in progress, I'm a product manager at MarkLogic - I'm in charge of infrastructure like storage, performance monitoring and cloud platforms.)
- acdha 13y ago> MarkLogic is an XML document store like Oracle is a CSV file store. Sorry, I wasn't trying to say that was a negative – just different in ways which many developers don't expect. With similar systems (e.g. zodb, MongoDV) I've seen developers use ORM-like patterns where they store properties in separate records or do reports by walking millions of records and then complain about performance rather than using it wrong. As for memory, that was awhile back so it might have been an old version or poor usage. I just heard about trying to max out server RAM as a key requirement.
- fennecfoxen 13y ago> {Q: What do you call a database without indexes, transactions or security? A: A file system!} Mmmm... I'll take ZFS over MongoDB any day of the week. :)