Wednesday, May 2, 2018

12 Free Tools For Creating an Effective Editorial Calendar

12 Free Tools For Creating an Effective Editorial Calendar

If it were not for my editorial calendar I would be completely blind. I mean it, my work would not only be unproductive but there is no way I could keep up with the demands of running a blog and a business at the same time. I would lose my mind.

No matter how dedicated or naturally good at organizing we are, maintaining an editorial calendar is a game changer when it comes to productivity and effective content management. It is also remarkably simple and makes the rest of the posting month so much easier in comparison to how it is without one.

They say you cannot improve something that cannot be measured. I say you cannot improve something that cannot be organized. Getting organized makes you more efficient, helps you see the big picture and keeps you on point.

Plenty of experts will give you tips on making a great editorial calendar and I am sure I will cover that myself in the future.

https://ift.tt/2FA6dM5

How to Retarget People Who Click on Curated Content

Do you share curated content on social media? Wondering how to retarget people who click on the third-party content you share? In this article, you’ll discover two ways to retarget people who click links to content you share, whether that content is yours or someone else’s. What Is Link Retargeting? Retargeting gives you the power

This post How to Retarget People Who Click on Curated Content first appeared on Social Media Examiner.

https://ift.tt/2I3C78Z

Efficient Link Reclamation: How to Speed Up & Scale Your Efforts

Posted by DarrenKingman

Link reclamation: Tools, tools everywhere

Every link builder, over time, starts to narrow down their favorite tactics and techniques. Link reclamation is pretty much my numero-uno. In my experience, it’s one of the best ROI activities we can use for gaining links particularly to the homepage, simply because the hard work — the "mention" (in whatever form that is) — is already there. That mention could be of your brand, an influencer who works there, or a tagline from a piece of content you’ve produced, whether it’s an image asset, video, etc. That’s the hard part. But with it done, and after a little hunting and vetting the right mentions, you’re just left with the outreach.

Aside from the effort-to-return ratio, there are various other benefits to link reclamation:

  1. It’s something you can start right away without assets
  2. It’s a low risk/low investment form of link building
  3. Nearly all brands have unlinked mentions, but big brands tend to have the most and therefore see the biggest routine returns
  4. If you’re doing this for clients, they get to see an instant return on their investment

Link reclamation isn’t a new tactic, but it is becoming more complex and tool providers are out there helping us to optimize our efforts. In this post, I’m going to talk a little about those tools and how to apply them to speed up and scale your link reclamation.

Finding mentions

Firstly, we want to find mentions. No point getting too fancy at this stage, so we just head over to trusty Google and search for the range of mentions we’re working on.

As I described earlier, these mentions can come in a variety of shapes and sizes, so I would generally treat each type of mention that I’m looking for as a separate project. For example, if Moz were the site I was working on, I would look for mentions of the brand and create that as one "project," then look for mentions of Followerwonk and treat that as another, and so on. The reasons why will become clear later on!

So, we head to the almighty Google and start our searches.

To help speed things up it’s best to expand your search result to gather as many URLs as you can in as few clicks as possible. Using Google’s Search Settings, you can quickly max out your SERPs to one hundred results, or you can install a plugin like GInfinity, which allows you to infinitely scroll through the results and grab as many as you can before your hand cramps up.

Now we want to start copying as many of these results as possible into an Excel sheet, or wherever it is you’ll be working from. Clicking each one and copying/pasting is hell, so another tool to quickly install for Chrome is Linkclump. With this one, you’ll be able to right click, drag, and copy as many URLs as you want.

Linkclump Pro Tip: To ensure you don’t copy the page titles and cache data from a SERP, head over to your Linkclump settings by right-clicking the extension icon and selecting "options." Then, edit your actions to include "URLs only" and "copied to clipboard." This will make the next part of the process much easier!

Filtering your URL list

Now we’ve got a bunch of URLs, we want to do a little filtering, so we know a) the DA of these domains as a proxy metric to qualify mentions, and b) whether or not they already link to us.

How you do this bit will depend on which platforms you have access to. I would recommend using BuzzStream as it combines a few of the future processes in one place, but URL Profiler can also be used before transferring your list over to some alternative tools.

Using BuzzStream

If you’re going down this road, BuzzStream can pretty much handle the filtering for you once you’ve uploaded your list of URLs. The system will crawl through the URLs and use their API to display Domain Authority, as well as tell you if the page already links to you or not.

The first thing you’ll want to do is create a "project" for each type of mention you’re sourcing. As I mentioned earlier this could be "brand mentions," "creative content," "founder mentions," etc.

When adding your "New Project," be sure to include the domain URL for the site you’re building links to, as shown below. BuzzStream will then go through and crawl your list of URLs and flag any that are already linking to you, so you can filter them out.

Next, we need to get your list of URLs imported. In the Websites view, use Add Websites and select "Add from List of URLs":

The next steps are really easy: Upload your list of URLs, then ensure you select "Websites and Links" because we want BuzzStream to retrieve the link data for us.

Once you’ve added them, BuzzStream will work through the list and start displaying all the relevant data for you to filter through in the Link Monitoring tab. You can then sort by: link status (after hitting "Check Backlinks" and having added your URL), DA, and relationship stage to see if you/a colleague have ever been in touch with the writer (especially useful if you/your team uses BuzzStream for outreach like we do at Builtvisible).

Using URL Profiler

If you’re using URL Profiler, firstly, make sure you’ve set up URL Profiler to work with your Moz API. You don’t need a paid Moz account to do this, but having one will give you more than 500 checks per day on the URLs you and the team are pushing through.

Then, take the list of URLs you’ve copied using Linkclump from the SERPs (I’ve just copied the top 10 from the news vertical for "moz.com" as my search), then paste the URLs in the list. You’ll need to select "Moz" in the Domain Level Data section (see screenshot) and also fill out the "Domain to Check" with your preferred URL string (I’ve put "Moz.com" to capture any links to secure, non-secure, alternative subdomains and deeper level URLs).

Once you’ve set URL Profiler running, you’ll get a pretty intimidating spreadsheet, which can simply be cut right down to the columns: URL, Target URL and Domain Mozscape Domain Authority. Filter out any rows that have returned a value in the Target URL column (essentially filtering out any that found an HREF link to your domain), and any remaining rows with a DA lower than your benchmark for links (if you work with one).

And there’s my list of URLs that we now know:

1) don’t have any links to our target domain,

2) have a reference to the domain we’re working on, and

3) boast a DA above 40.

Qualify your list

Now that you’ve got a list of URLs that fit your criteria, we need to do a little manual qualification. But, we’re going to use some trusty tools to make it easy for us!

The key insight we’re looking for during our qualification is if the mention is in a natural linking element of the page. It’s important to avoid contacting sites where the mention is only in the title, as they’ll never place the link. We particularly want placements in the body copy as these are natural link locations and so increase the likelihood of your efforts leading somewhere.

So from my list of URLs, I’ll copy the list and head over to URLopener.com (now bought by 10bestseo.com presumably because it’s such an awesome tool) and paste in my list before asking it to open all the URLs for me:

Now, one by one, I can quickly scan the URLs and look for mentions in the right places (i.e. is the mention in the copy, is it in the headline, or is it used anywhere else where a link might not look natural?).

When we see something like this (below), we’re making sure to add this URL to our final outreach list:

However, when we see this (again, below), we’re probably stripping the URL out of our list as there’s very little chance the author/webmaster will add a link in such a prominent and unusual part of the page:

The idea is to finish up with a list of unlinked mentions in spots where a link would fit naturally for the publisher. We don’t want to get in touch with everyone, with mentions all over the place, as it can harm your future relationships. Link building needs to make sense, and not just for Google. If you’re working in a niche that mentions your client, you likely want not only to get a link but also build a relationship with this writer — it could lead to 5 links further down the line.

Getting email addresses

Now that you’ve got a list of URLs that all feature your brand/client, and you’ve qualified this list to ensure they are all unlinked and have mentions in places that make sense for a link, we need to do the most time-consuming part: finding email addresses.

To continue expanding our spreadsheet, we’re going to need to know the contact details of the writer or webmaster to request our link from. To continue our theme of efficiency, we just want to get the two most important details: email address and first name.

Getting the first name is usually pretty straightforward and there’s not really a need to automate this. However, finding email addresses could be an entirely separate article in itself, so I’ll be brief and get to the point. Read this, and here’s a summary of places to look and the tools I use:

  • Author page
  • Author’s personal website
  • Author’s Twitter profile
  • Rapportive & Email Permutator
  • Allmytweets
  • Journalisted.com
  • Mail Tester

More recently, we’ve been also using Skrapp.io. It’s a LinkedIn extension (like Hunter.io) that installs a "Find Email" button on LinkedIn with a percentage of accuracy. This can often be used with Mail Tester to discover if the suggested email address provided is working or not.

It’s likely to be a combination of these tools that helps you navigate finding a contact’s email address. Once we have it, we need to get in touch — at scale!

Pro Tip: When using Allmytweets, if you’re finding that searches for "email" or "contact" aren’t working, try "dot." Usually journalists don’t put their full email address on public profiles in a scrapeable format, so they use "me@gmail [dot] com" to get around it.

Making contact

So, because this is all about making the process efficient, I’m not going to repeat or try to build on the other already useful articles that provide templates for outreach (there is one below, but that’s just as an example!). However, I am going to show you how to scale your outreach and follow-ups.

Mail merges

If you and your team aren’t set in your ways with a particular paid tool, your best bet for optimizing scale is going to be a mail merge. There are a number of them out there, and honestly, they are all fairly similar with either varying levels of free emails per day before you have to pay, or they charge from the get-go. However, for the costs we’re talking about and the time it saves, building a business case to either convince yourself (freelancers) or your finance department (everyone else!) will be a walk in the park.

I’ve been a fan of Contact Monkey for some time, mainly for tracking open rates, but their mail merge product is also part of the $10-a-month package. It’s a great deal. However, if you’re after something a bit more specific, YAMM is free to a point (for personal Gmail accounts) and can send up to 50 emails a day.

You’ll likely need to work through the process with the whatever tool you pick but, using your spreadsheet, you’ll be able to specify which fields you want the mail merge to select from, and it’ll insert each element into the email.

For link reclamation, this is really as personable as you need to get — no lengthy paragraphs on how much you loved [insert article related to my infographic] or how long you’ve been following them on Twitter, just a good old to the point email:

Hi [first name],

I recently found a mention of a company I work with in one of your articles.

Here’s the article:[insert URL]

Where you’ve mentioned our company, Moz, would you be able to provide a link back to the domain Moz.com, in case users would like to know more about us?

Many thanks,
Darren.
If using BuzzStream

Although BuzzStream’s mail merge options are pretty similar to the process above, the best "above and beyond" feature that BuzzStream has is that you can schedule in follow up emails as well. So, if you didn’t hear back the first time, after a week or so their software will automatically do a little follow-up, which in my experience, often leads to the best results.

When you’re ready to start sending emails, select the project you’ve set up. In the "Websites" section, select "Outreach." Here, you can set up a sequence, which will send your initial email as well as customized follow-ups.

Using the same extremely brief template as above, I’ve inserted my dynamic fields to pull in from my data set and set up two follow up emails due to send if I don’t hear back within the next 4 days (BuzzStream hooks up with my email through Outlook and can monitor if I receive an email from this person or not).

Each project can now use templates set up for the type of mention you’re following up. By using pre-set templates, you can create one for brand mention, influencers, or creative projects to further save you time. Good times.

I really hope this has been useful for beginners and seasoned link reclamation pros alike. If you have any other tools you use that people may find useful or have any questions, please do let us know below.

Thanks everyone!


Sign up for The Moz Top 10, a semimonthly mailer updating you on the top ten hottest pieces of SEO news, tips, and rad links uncovered by the Moz team. Think of it as your exclusive digest of stuff you don't have time to hunt down but want to read!

https://ift.tt/2rg8xm0

Tuesday, May 1, 2018

The 5 Best Influencer Marketing Studies of 2018 So Far

The 5 Best Influencer Marketing Studies of 2018 So Far

The field of marketing is in constant flux, with marketers constantly analyzing emerging trends and data to determine optimal strategies. As a result, marketing studies only a year old can be outdated. Fortunately, a fresh crop of influencer marketing studies published in 2018 provide marketers with fresh insight that can positively impact their campaigns.

1. Increasing Marketing Budgets

The vast majority of marketers are increasing their budgets in 2018. Only five percent of marketers report plans to decrease their investment in marketing, per Linqia’s State of Influencer Marketing 2018 study.

Budget increases in marketing have been a trend for the past several years, with more viable channels than ever available for marketing efforts. The cost of managing influencer marketing is one noteworthy reason for increasing budgets, as well.

Another big reason for the continuing increase in marketing budgets is the rise of big data, which confirms the immense value of marketing to anyone previously doubting its importance. Today, brands compete to provide superior customer experience across many digital platforms. Buyers are increasingly seeking out products on their own, whereas in the past, the majority of products were broadcast to them. Customer journeys are becoming increasingly self-driven, which raises the importance of effective marketing.

Influencer marketing has the potential to captivate audiences with influencers they know and trust. As a result, marketers are finding influencer investments to be worthwhile and necessary. Marketers are also choosing to embrace turnkey providers for managing influencer marketing programs, especially since market research costs can vary depending on a variety of factors.

As long as customers continue to browse the web and pursue a self-guided approach to discovering content, marketing budgets are likely to continue increasing.

Customer journeys are becoming increasingly self-driven. #CX
Click To Tweet
2. Challenges in Managing Influencer Marketing Programs

Marketers in 2018 report that one of their biggest challenges is incorporating influencer marketing into their busy schedules. Influencer marketing has immense value in 2018, so marketers do not have much choice. However, many marketers are finding viable solutions to managing influencer marketing programs. According to Linqia, 42 percent of marketers report using a managed service or “turnkey service provider” to help run the programs, with an additional 18 percent outsourcing influencer marketing to an agency.

Agencies that act as turnkey service providers and influencers of marketing middle-servicing are increasing in popularity among marketers, who recognize the necessity of scaling their campaigns in 2018. Even if marketers do not have resources on their own to manage influencer marketing, they are getting it done nonetheless with outsourcing and turnkey service providers.

3. Influencers Continue to Engage

One of the most prominent reasons for increasing investments in the influencer marketing sphere is the strength of influencer engagement. The average influencer engagement rate across industry verticals is 5.7 percent. Comparatively, brands on Instagram have an average engagement rate between two and three percent.

Why do influencers have about double the engagement of brands? The answer may lie in influencers’ greater feeling of authenticity. Consumers are well aware that brands are promoting with the incentive to sell. An influencer, however, may be promoting something because they enjoy the product or service. Although consumers are aware that influencers often receive payment for promotion, they also know that influencers are unlikely to damage their reputation by promoting a product or service that’s low-quality or undesirable.

Influencers attract double the engagement

4. Video Content Continues to Reign Supreme

The dominance of video content is nothing new, though it’s notable that 2018 may see video reach a whole new level of engagement. Experts estimate that by 2020, video consumption will reach 80 percent of global online traffic.

86 percent of marketers use video content, which makes even more sense when you consider that viewers absorb 95 percent more messaging via video. Video’s staying power is also clear, with 99 percent of marketers already using video saying they will continue to use it throughout 2018.

In a 2018 study by Wyzowl, 97 percent of marketers reported that video helped increase users’ understanding of their product or service. Combine that with the fact that 72 percent of consumers prefer video over text when learning about a product or service, and it’s easy to see why marketers embrace video content.

It’s worth noting that influencers are rampant throughout video-centric social media platforms, like YouTube and Instagram. Video content’s rise in popularity alongside the ascent of influencers is no coincidence.

Video will constitute 80 percent of traffic by 2020

5. Marketers Are Honing in on Instagram and Blogs

In 2018, Instagram reigns supreme as the most important social network for influencer marketing, according to 92 percent of marketers in Linqia’s The State of Influencer Marketing 2018. The majority of social networks support video content, though Instagram stands out due to its enormous user base and easily digestible, often concise video content.

The Kardashian family alone provides a taste of Instagram’s potential, with Kim Kardashian, Kendall Jenner, Khloe Kardashian, and Kourtney Kardashian attracting over 26,000,000 likes and comments on Instagram from January 1 to March 15, 2018. When an influencer achieves success on Instagram, it’s usually a massive success.

Instagram is the most important network for influencers

Don’t overlook the rising importance of blogs, either. Although blogs still lag behind Instagram and Facebook as the most important social networks for influencer marketing, they are a close third and have risen in popularity since 2017. Influencers on social media are embracing the blog form for more word-heavy content, especially the influencers launching their own e-commerce lines. These blogs showcase products and more informative content beyond these influencers’ social media presences.

2018 is an exciting year for marketers, who are seizing emerging forms of content like video and embracing influencers in their marketing campaigns. As consumers become more wary of conventional marketing techniques and gain more awareness of market research, influencer marketing offers a fresh and (hopefully) authentic approach that provides consumers a new perspective on why a particular product or service may be a good fit for their needs.

https://ift.tt/2HCpXnU

Social Media and Trusting Your Instinct

Social Media and Trusting Your Instinct

Follow your instincts. That’s where true wisdom manifests itself. – Oprah Winfrey

As a business leader, we constantly face new challenges.

We steadily encounter interesting contacts, read inspiring stories, and experience that doing the right thing is not always easy.

“We spend our workdays in our outer world. We’re interacting with our team members and clients. We don’t have enough time in our inner world where we can reflect on those experiences and listen to what our gut might have to say,” says the professional development coach Hana Ayoub.

We are bombarded from different platforms with social media noise about how to fix our biggest career problems and with a large variety of information about how to be successful – and therefore it’s not surprising that many of us doubt our natural instincts and decision-making strategies. Instead of listening to our inner voice we quickly tune it out.

https://ift.tt/2HM68a0

3 Overlooked Facebook Lookalike Audiences That Will Improve Your Ad Results

Want better results from your Facebook lookalike audiences? Wondering which custom audiences yield the best-performing lookalike audiences? In this article, you’ll discover how to create three highly tuned Facebook lookalike audiences from your most valuable custom audiences. #1: Reach a Lookalike Audience of People Who Spend Time on Your Website’s Product Pages Website traffic lookalike

This post 3 Overlooked Facebook Lookalike Audiences That Will Improve Your Ad Results first appeared on Social Media Examiner.

https://ift.tt/2rccatq

Big, Fast, and Strong: Setting the Standard for Backlink Index Comparisons

Posted by rjonesx.

It's all wrong

It always was. Most of us knew it. But with limited resources, we just couldn't really compare the quality, size, and speed of link indexes very well. Frankly, most backlink index comparisons would barely pass for a high school science fair project, much less a rigorous peer review.

My most earnest attempt at determining the quality of a link index was back in 2015, before I joined Moz as Principal Search Scientist. But I knew at the time that I was missing a huge key to any study of this sort that hopes to call itself scientific, authoritative or, frankly, true: a random, uniform sample of the web.

But let me start with a quick request. Please take the time to read this through. If you can't today, schedule some time later. Your businesses depend on the data you bring in, and this article will allow you to stop taking data quality on faith alone. If you have questions with some technical aspects, I will respond in the comments, or you can reach me on twitter at @rjonesx. I desperately want our industry to finally get this right and to hold ourselves as data providers to rigorous quality standards.

Quick links:

  1. Home
  2. Getting it right
  3. What's the big deal with random?
  4. Now what? Defining metrics
  5. Caveats
  6. The metrics dashboard
  7. Size matters
  8. Speed
  9. Quality
  10. The Link Index Olympics
  11. What's next?
  12. About PA and DA
  13. Quick takeaways

Getting it right

One of the greatest things Moz offers is a leadership team that has given me the freedom to do what it takes to "get things right." I first encountered this when Moz agreed to spend an enormous amount of money on clickstream data so we could make our keyword tool search volume better (a huge, multi-year financial risk with the hope of improving literally one metric in our industry). Two years later, Ahrefs and SEMRush now use the same methodology because it's just the right way to do it.

About 6 months into this multi-year project to replace our link index with the huge Link Explorer, I was tasked with the open-ended question of "how do we know if our link index is good?" I had been thinking about this question ever since that article published in 2015 and I knew I wasn't going to go forward with anything other than a system that begins with a truly "random sample of the web." Once again, Moz asked me to do what it takes to "get this right," and they let me run with it.

What's the big deal with random?

It's really hard to over-state how important a good random sample is. Let me diverge for a second. Let's say you look at a survey that says 90% of Americans believe that the Earth is flat. That would be a terrifying statistic. But later you find out the survey was taken at a Flat-Earther convention and the 10% who disagreed were employees of the convention center. This would make total sense. The problem is the sample of people surveyed wasn't of random Americans — instead, it was biased because it was taken at a Flat-Earther convention.

Now, imagine the same thing for the web. Let's say an agency wants to run a test to determine which link index is better, so they look at a few hundred sites for comparison. Where did they get the sites? Past clients? Then they are probably biased towards SEO-friendly sites and not reflective of the web as a whole. Clickstream data? Then they would be biased towards popular sites and pages — once again, not reflective of the web as a whole!

Starting with a bad sample guarantees bad results.

It gets even worse, though. Indexes like Moz report our total statistics (number of links or number of domains in our index). However, this can be terribly misleading. Imagine a restaurant which claimed to have the largest wine selection in the world with over 1,000,000 bottles. They could make that claim, but it wouldn't be useful if they actually had 1,000,000 of the same type, or only Cabernet, or half-bottles. It's easy to mislead when you just throw out big numbers. Instead, it would be much better to have a random selection of wines from the world and measure if that restaurant has it in stock, and how many. Only then would you have a good measure of their inventory. The same is true for measuring link indexes — this is the theory behind my methodology.

Unfortunately, it turns out getting a random sample of the web is really hard. The first intuition most of us at Moz had was to just take a random sample of the URLs in our own index. Of course we couldn't — that would bias the sample towards our own index, so we scrapped that idea. The next thought was: "We know all these URLs from the SERPs we collect — perhaps we could use those." But we knew they'd be biased towards higher-quality pages. Most URLs don't rank for anything — scratch that idea. It was time to take a deeper look.

I fired up Google Scholar to see if any other organizations had attempted this process and found literally one paper, which Google produced back in June of 2000, called "On Near-Uniform URL Sampling." I hastily whipped out my credit card to buy the paper after reading just the first sentence of the abstract: "We consider the problem of sampling URLs uniformly at random from the Web." This was exactly what I needed.

Why not Common Crawl?

Many of the more technical SEOs reading this might ask why we didn't simply select random URLs from a third-party index of the web like the fantastic Common Crawl data set. There were several reasons why we considered, but chose to pass, on this methodology (despite it being far easier to implement).

  1. We can't be certain of Common Crawl's long-term availability. Top million lists (which we used as part of the seeding process) are available from multiple sources, which means if Quantcast goes away we can use other providers.
  2. We have contributed crawl sets in the past to Common Crawl and want to be certain there is no implicit or explicit bias in favor of Moz's index, no matter how marginal.
  3. The Common Crawl data set is quite large and would be harder to work with for many who are attempting to create their own random lists of URLs. We wanted our process to be reproducible.

How to get a random sample of the web

The process of getting to a "random sample of the web" is fairly tedious, but the general gist of it is this. First, we start with a well-understood biased set of URLs. We then attempt to remove or balance this bias out, making the best pseudo-random URL list we can. Finally, we use a random crawl of the web starting with those pseudo-random URLs to produce a final list of URLs that approach truly random. Here are the complete details.

1. The starting point: Getting seed URLs

The first big problem with getting a random sample of the web is that there is no true random starting point. Think about it. Unlike a bag of marbles where you could just reach in and blindly grab one at random, if you don't already know about a URL, you can't pick it at random. You could try to just brute-force create random URLs by shoving letters and slashes after each other, but we know language doesn't work that way, so the URLs would be very different from what we tend to find on the web. Unfortunately, everyone is forced to start with some pseudo-random process.

We had to make a choice. It was a tough one. Do we start with a known strong bias that doesn't favor Moz, or do we start with a known weaker bias that does? We could use a random selection from our own index for the starting point of this process, which would be pseudo-random but could potentially favor Moz, or we could start with a smaller, public index like the Quantcast Top Million which would be strongly biased towards good sites.

We decided to go with the latter as the starting point because Quantcast data is:

  1. Reproducible. We weren't going to make "random URL selection" part of the Moz API, so we needed something others in the industry could start with as well. Quantcast Top Million is free to everyone.
  2. Not biased towards Moz: We would prefer to err on the side of caution, even if it meant more work removing bias.
  3. Well-known bias: The bias inherent in the Quantcast Top 1,000,000 was easily understood — these are important sites and we need to remove that bias.
  4. Quantcast bias is natural: Any link graph itself already shares some of the Quantcast bias (powerful sites are more likely to be well-linked)

With that in mind, we randomly selected 10,000 domains from the Quantcast Top Million and began the process of removing bias.

2. Selecting based on size of domain rather than importance

Since we knew the Quantcast Top Million was ranked by traffic and we wanted to mitigate against that bias, we introduced a new bias based on the size of the site. For each of the 10,000 sites, we identified the number of pages on the site according to Google using the "site:" command and also grabbed the top 100 pages from the domain. Now we could balance the "importance bias" against a "size bias," which is more reflective of the number of URLs on the web. This was the first step in mitigating the known bias of only high-quality sites in the Quantcast Top Million.

3. Selecting pseudo-random starting points on each domain

The next step was randomly selecting domains from that 10,000 with a bias towards larger sites. When the system selects a site, it then randomly selects from the top 100 pages we gathered from that site via Google. This helps mitigate the importance bias a little more. We aren't always starting with the homepage. While these pages do tend to be important pages on the site, we know they aren't always the MOST important page, which tends to be the homepage. This was the second step in mitigating the known bias. Lower-quality pages on larger sites were balancing out the bias intrinsic to the Quantcast data.

4. Crawl, crawl, crawl

And here is where we make our biggest change. We actually crawl the web starting with this set of pseudo-random URLs to produce the actual set of random URLs. The idea here is to take all the randomization we have built into the pseudo-random URL set and let the crawlers randomly click on links to produce the truly random URL set. The crawler will select a random link from our pseudo-random crawlset and then start a process of randomly clicking links, each time with a 10% chance of stopping and a 90% chance of continuing. Wherever the crawler ends, the final URL is dropped into our list of random URLs. It is this final set of URLs that we use to run our metrics. We generate around 140,000 unique URLs through this process monthly to produce our test data set.

Phew, now what? Defining metrics

Once we have the random set of URLs, we can start really comparing link indexes and measuring their quality, quantity, and speed. Luckily, in their quest to "get this right," Moz gave me generous paid access to competitor APIs. We began by testing Moz, Majestic, Ahrefs, and SEMRush, but eventually dropped SEMRush after their partnership with Majestic.

So, what questions can we answer now that we have a random sample of the web? This is the exact wishlist I sent out in an email to leaders on the link project at Moz:

  1. Size:
    • What is the likelihood a randomly selected URL is in our index vs. competitors?
    • What is the likelihood a randomly selected domain is in our index vs. competitors?
    • What is the likelihood an index reports the highest number of backlinks for a URL?
    • What is the likelihood an index reports the highest number of root linking domains for a URL?
    • What is the likelihood an index reports the highest number of backlinks for a domain?
    • What is the likelihood an index reports the highest number of root linking domains for a domain?
  2. Speed:
    • What is the likelihood that the latest article from a randomly selected feed is in our index vs. our competitors?
    • What is the average age of a randomly selected URL in our index vs. competitors?
    • What is the likelihood that the best backlink for a randomly selected URL is still present on the web?
    • What is the likelihood that the best backlink for a randomly selected domain is still present on the web?
  3. Quality:
    • What is the likelihood that a randomly selected page's index status (included or not included in index) in Google is the same as ours vs. competitors?
    • What is the likelihood that a randomly selected page's index status in Google SERPs is the same as ours vs. competitors?
    • What is the likelihood that a randomly selected domain's index status in Google is the same as ours vs. competitors?
    • What is the likelihood that a randomly selected domain's index status in Google SERPs is the same as ours vs. competitors?
    • How closely does our index compare with Google's expressed as "a proportional ratio of pages per domain vs our competitors"?
    • How well do our URL metrics correlate with US Google rankings vs. our competitors?

Reality vs. theory

Unfortunately, like all things in life, I had to make some cutbacks. It turns out that the APIs provided by Moz, Majestic, Ahrefs, and SEMRush differ in some important ways — in cost structure, feature sets, and optimizations. For the sake of politeness, I am only going to mention name of the provider when it is Moz that was lacking. Let's look at each of the proposed metrics and see which ones we could keep and which we had to put aside...

  1. Size: We were able monitor all 6 of the size metrics!

  2. Speed:
    • We were able to include this Fast Crawl metric.
    • What is the average age of a randomly selected URL in our index vs. competitors?
      Getting the age of a URL or domain is not possible in all APIs, so we had to drop this metric.
    • What is the likelihood that the best backlink for a randomly selected URL is still present on the web?
      Unfortunately, doing this at scale was not possible because one API is cost prohibitive for top link sorts and another was extremely slow for large sites. We hope to run a set of live-link metrics independently from our daily metrics collection in the next few months.
    • What is the likelihood that the best backlink for a randomly selected Domain is still present on the web?
      Once again, doing this at scale was not possible because one API is cost prohibitive for top link sorts and another was extremely slow for large sites. We hope to run a set of live-link metrics independently from our daily metrics collection in the next few months.
  3. Quality:
    • We were able to keep this metric.
    • What is the likelihood that a randomly selected page's index status in Google SERPs is the same as ours vs. competitors?
      Chose not to pursue due to internal API needs, looking to add soon.
    • We were able to keep this metric.
    • What is the likelihood that a randomly selected domain's index status in Google SERPs is the same as ours vs. competitors?
      Chose not to pursue due to internal API needs at the beginning of project, looking to add soon.
    • How closely does our index compare with Google's expressed as a proportional ratio of pages per domain vs our competitors?
      Chose not to pursue due to internal API needs. Looking to add soon.
    • How well do our URL metrics correlate with US Google rankings vs. our competitors?
      Chose not to pursue due to known fluctuations in DA/PA as we radically change the link graph. The metric would be meaningless until the index became stable.

Ultimately, I wasn't able to get everything I wanted, but I was left with 9 solid, well-defined metrics.

On the subject of live links:

In the interest of being TAGFEE, I will openly admit that I think our index has more deleted links than others like the Ahrefs Live Index. As of writing, we have about 30 trillion links in our index, 25 trillion we believe to be live, but we know that some proportion are likely not. While I believe we have the most live links, I don't believe we have the highest proportion of live links in an index. That honor probably does not go to Moz. I can't be certain because we can't test it fully and regularly, but in the interest of transparency and fairness, I felt obligated to mention this. I might, however, devote a later post to just testing this one metric for a month and describe the proper methodology to do this fairly, as it is a deceptively tricky metric to measure. For example, if a link is retrieved from a chain of redirects, it is hard to tell if that link is still live unless you know the original link target. We weren't going to track any metric if we couldn't "get it right," so we had to put live links as a metric on hold for now.

Caveats

Don't read any more before reading this section. If you ask a question in the comments that shows you didn't read the Caveats section, I'm just going to say "read the Caveats section." So here goes...

  • This is a comparison of data that comes back via APIs, not within the tools themselves. Many competitors offer live, fresh, historical, etc. types of indexes which can differ in important ways. This is just a comparison of API data using default settings.
  • Some metrics are hard to estimate, especially like "whether a link is in the index," because no API — not even Moz — has a call that just tells you whether they have seen the link before. We do our best, but any errors here are on the the API provider. I think we (Moz, Majestic, and Ahrefs) should all consider adding an endpoint like this.
  • Links are counted differently. Whether duplicate links on a page are counted, whether redirects are counted, whether canonicals are counted (which Ahrefs just changed recently), etc. all affect these metrics. Because of this, we can't be certain that everything is apples-to-apples. We just report the data at face value.
  • Subsequently, the most important takeaway in all of these graphs and metrics is direction. How are the indexes moving relative to one another? Is one catching up, is another falling behind? These are the questions best answered.
  • The metrics are adversarial. For each random URL or domain, a link index (Moz, Majestic, or Ahrefs) gets 1 point for being the biggest, for tying with the biggest, or for being "correct." They get 0 points if they aren't the winner. This means that the graphs won't add up to 100 and it also tends to exaggerate the differences between the indexes.
  • Finally, I'm going to show everything, warts and all, even when it was my fault. I'll point out why some things look weird on graphs and what we fixed. This was a huge learning experience and I am grateful for the help I received from the support teams at Majestic and Ahrefs who, as a customer, responded to my questions honestly and openly.

The metrics dashboardhttps://ift.tt/2jhS1yp