6 ms·
Black Hat SEO Case Study: How Mahalo Makes Black Look White
- whyenot 17y agoWhy isn't it copyright infringement for Mahalo to scrape content like claimed in the article? I don't see how either fair use or DMCA safe harbor would apply (but I'm no lawyer). This seems like a lawsuit just waiting to happen.
- wmf 17y agoThey're just scraping page titles and occasionally short excerpts; most people believe that's fair use.
- jfornear 17y agoI honestly don't understand how this is anything new? Mahalo always has been a sketchy SEO scam that only a shameless self-promoter could pull off...
- rman666 17y agoSorry for the spam, but anyone on this thread interested in the domain, "blackhatsystems.com"? Send me an email (clint{dot}laskowski{at}gmail{dot}com).
- vaksel 17y agogood article, never noticed the the scraped content part
- tomh- 17y agoThis pretty much proves once again that gaming search engines is here to stay. There is still a lot of research to be done to make it harder to get away with this type of websites, but luckily there are more and more ways, other than Google, to find the content you are looking for.
- ojbyrne 17y agoGreat article. "the willingness to lie just to get a bit of media ink" very succinctly captures what I most disliked about my experiences amongst the movers and shakers of California.
- aristus 17y agoJason, any comment?
- johng 17y agoIn for the comment as well. I doubt you'll get a true response because it's quite clear Mahalo is just auto-generated SPAM 99% of the time.
- vaksel 17y agomy guess it's "oh shit!", followed by a prayer that Google doesn't do anything. I don't think he has anything to worry about for that last part, Google is notorious for letting big sites off the hook(remember Target?) + here Google is also making their share of money off adsense
- vladmato 17y agoAnd Scribd.com, exact same model, scrape content regardless of copyrights, straps adsense on it, makes money.
- slig 17y agoYou forgot related search queries, which is basically bogus search queries that produces even more visits and bogus search queries.
- patio11 17y agoScribd ceased doing related search queries some time ago. Their CEO described it, to Techcrunch, as "reducing the aggressiveness of our SEO, which reduces total traffic in the near term but increases the relevancy of Scribd links in search engine results." Scuttlebutt among SEOs whose opinion I respect suggests that it is highly probable they got a backchannel from Google telling them that either they could drastically reduce the footprint of their pages in the SERPs or that Google's search quality team would assist them in doing so. Anyhow, their traffic went down by about 50% in a month, if you trust Compete et al. Search for [Scribd "aggressive SEO"] if you want the whole tale.
- qeorge 17y agoYou actually don't need all that much authority to get away with ranking scraped content in Google. Despite their FUD, Google's duplicate content detection algorithm seems to be largely non-existent. For example, check out the Google results for http://hackerne.ws http://hackerne.ws, which is a page-for-page duplicate of news.ycombinator: http://www.google.com/q=site%3Ahackerne.ws http://www.google.com/q=site%3Ahackerne.ws 10,000 pages indexed, not a single word of original content. Note: I know hackerne.ws is not trying to be spammy, and merely parked the domain improperly. If the owner is reading, all it would take is a simple 301 redirect to fix.
- blasdel 17y agoAll it would take to fix is for pg to make news.arc less shitty: actually check the HTTP/1.1 Host header, and respond with your own 301.
- dminor 17y agoOr include a rel=canonical link in the head.
- patio11 17y agorel=canonical won't work if the domain Google sees you on is different from the one specified as canonical. This is to prevent people capable of content injection from hijacking entire websites in a subtle manner.
- dminor 17y agoGoogle says otherwise: http://googlewebmastercentral.blogspot.com/2009/12/handling-legitimate-cross-domain.html http://googlewebmastercentral.blogspot.com/2009/12/handling-... No doubt they must use other indicators to ensure the authoritative source.
- merraksh 17y agoYour link didn't work for me, but this http://www.google.com/search?hl=en&site=q%3Dsite%3Ahackerne.ws&q=hackerne.ws&btnG=Search http://www.google.com/search?hl=en&site=q%3Dsite%3Ahacke... shows 255,000 results, of which hackerne.ws is the first, news.ycombinator.com is second :-| [edit: luckily, Google's duplicate content detection algorithm didn't work here...]
- kyro 17y agoJesus, that's about as s(c/p)ammy as you can get. Your business is rooted in theft and trickery and deception, Jason.
- callmeed 17y agoHere's a weird experience I had: I tweeted about a startup idea of a "woot.com site for travel" http://twitter.com/callmeed/status/1422601143 http://twitter.com/callmeed/status/1422601143 Later that day, it somehow got turned into a Mahalo question (I didn't submit it). I thought it was interesting that Jason himself commented on it, but it still seemed strange. Now, when you google "Woot.com for travel" or "Woot for travel", that Mahalo page comes up on position 1 or 3.
- vaksel 17y agothat's the beauty of authority sites, you can rank for stuff just by mentioning them once.
- javery 17y agoThis is the flaw in how Google does things that competitors need to exploit. Authority should be more topic based then site-wide.
- vaksel 17y agoyeah I remember reading a blog a few days ago, it was a PR8, and the guy did an experiment. Just added a link for something viagra related. Just a single link. And within a week he managed to get on a front page in Google results.
- deleted 17y ago[deleted]
- vaksel 17y agoPage Rank 8
- deleted 17y ago[deleted]
- NZ_Matt 17y agoGreat article! Just a few days ago I landed on a Maholo page from a google search. My exact thoughts were "where is the content".
- fuzzmeister 17y agoThis article makes it clear that Mahalo is in many ways quite similar to another questionably ethical startup, Demand Media. Here's the fascinating Wired article: http://www.wired.com/magazine/2009/10/ff_demandmedia/all/1 http://www.wired.com/magazine/2009/10/ff_demandmedia/all/1
- atamyrat 17y agoOne important difference is that Demand Media actually produces their content themselves.
- patio11 17y agoRight. Demand Media has a distributed, virtualized workforce of freelancers. (Read the Wired article on it. That is some of their best reporting. Ever.) Mahalo used to have in-house editors before they moved to mostly outsourced "editors" before they realized editors cost a lot of money and firing them didn't decrease revenues in the slightest. At the moment their editorial staff is a thin pretense maintained to keep the site from getting bounced out of the index. Disclaimer: As with most other massive content plays which have large audiences of unsophisticated Internet users, I indirectly subsidize Mahalo through AdSense expenditures. To the tune of probably over a hundred bucks last year, but I don't have my numbers in front of me. Like I mentioned in my blog post earlier today, they send great traffic (i.e. it is cheap and converts well) because my ads are the content on their pages. That is disquieting to me in some ways. I could ban them and start chopping off heads from the Demand Media hydra in my AdSense account, but that would consume vast amounts of my time and just cost me money.
- krtl 17y agoRight. Mahalo had a call out for 17 or so "interns/volunteers" a few months back... thats who replaced the editorial staff.
- kareemm 17y agoi think it's interesting that jason - who manages his online reputation exceptionally well (and speedily) - has yet to comment.
- jcromartie 17y agoWho would have thought that the guy who said (paraphrased) "want to have a life? then work somewhere else!" would be slimy? I really don't understand how or why Jason Calacanis has any credibility or notoriety today. Point me to Mahalo and I see an utterly worthless spammy waste of a website that I and all of my peers avoid at all costs which was built with exploitative labor practices.
- monkeygrinder 17y agoInteresting article. Mahalo won't be the last site to exploit these methods. There are a number of issues here: 1. My boss said something interesting the other day: As Google has already conquered search and is diversifying its business into different areas, it is not paying as much diligence into its search algorithm to weed out those sites that exploit it. Google makes money from AdSense, so why would it be in a big hurry to take down sites that exploit dodgy SEO practices. 2. As for scraping content without any backlinks, the media industry seems to have very little protection when it comes to copyright. Existing copyright law is woefully unable to get to grips with digital copying and display. If the content had been music, or films, the RIAA would have clamped down so fast, Jason's head would be spinning. But we are talking about digital publishing industry, where content has very little protection at all. 3. Even if we decide that taking the first paragraph is fair use, not back linking or citing your source is still a copyright issue (not to mention bad Internet etiquette).