15 SEO Experiments: What Real Testing Reveals
SEO experiments are controlled tests that isolate one variable, a heading tag, a response code, a hosting setup, and measure its actual effect on rankings or click-through rate, rather than relying on Google’s public statements or SEO folklore alone. Real testing regularly contradicts both: some official guidance holds up under scrutiny, some doesn’t, and knowing the difference is worth more than another theory piece.
How to Run a Valid SEO Experiment
A valid experiment needs a genuine control group and test group, isolating one variable while holding everything else constant, page speed, content length, domain age, so any measured difference in organic traffic or search engine rankings can be attributed to the change itself rather than a confounding variable. Statistical significance matters here specifically: a ranking swing on a handful of pages over a few days proves far less than a controlled split test run across enough pages and enough time to rule out normal ranking volatility.
The methodology matters as much as the result. One widely-discussed shared hosting experiment on this list drew direct criticism for using an invented keyword with zero existing competition, a real limitation worth understanding before trusting any single test’s conclusions: a made-up term with no competing pages can behave very differently from a real, contested SERP. Treat every experiment below, including the ones with clean, clear results, as evidence to weigh, not gospel to repeat, and prioritize experiments run by teams that disclose their full methodology, the way Reboot Online, SearchPilot, and Dejan SEO each do, over ones that only publish the headline finding.
15 SEO Experiments and What They Revealed
1. LLMs.txt Experiment
Reboot Online tested whether an LLMs.txt file actually gets AI crawlers to discover and visit new pages, publishing brand-new, completely unlinked test pages on two established domains, RecycleZone.org.uk and RevisionCentre.co.uk, referenced only in each site’s LLMs.txt file with no internal or external links pointing to them. After three months of log file analysis, neither the LLMs.txt files nor the new pages were visited by any AI crawler, including ChatGPT and Claude.
This result lines up with the broader current evidence: no major AI provider, including OpenAI, Google, or Anthropic, has publicly confirmed reading or acting on LLMs.txt in production, and Google’s Gary Illyes explicitly confirmed at Search Central Live APAC that Google does not support the file and has no plans to, unlike robots.txt, which crawlers have reliably respected for decades. Treat LLMs.txt as, at most, a low-cost, low-priority addition after core technical SEO and content fundamentals are solid, not a genuine AI discoverability lever.
2. Negative GEO Experiment (Competitor Sabotage in AI)
Reboot Online demonstrated how easily AI models can be manipulated by buying 10 expired domains carrying old, existing backlinks, then publishing a “sexiest bald men” list on each homepage with their own managing director included alongside genuinely well-known bald celebrities to keep the content looking legitimate. They tested the prompt across ChatGPT, Claude, Gemini, Perplexity, and DeepSeek using fresh, incognito accounts to rule out personalization, and both ChatGPT and Perplexity began confidently naming their managing director in response to the query.
This is a real, demonstrated vulnerability worth taking seriously: expired domains with pre-existing backlink authority can genuinely shift what large language models present as fact, a manipulation risk that applies just as easily to brand reputation attacks as it does to a harmless publicity stunt. Monitoring brand mentions across AI Overviews and chat-based answer engines is becoming a real, current defensive necessity, not an optional extra.
3. Controlled GEO Experiment (Influencing AI)
In a separate, legitimate test of the same underlying mechanism, Reboot Online published a genuinely researched comparison of digital PR agencies, built with real data rather than opinion, internally linked it for discoverability, and optimized it the way they would any competitive content piece. Within weeks of publishing, the article was being cited as a source in both AI Overview and ChatGPT responses to their target query.
Compared against the Negative GEO test above, the contrast is genuinely instructive: both manipulation through expired-domain backlink authority and legitimate, well-optimized original content can influence what large language models cite, meaning generative engine optimization rewards the same underlying signals, genuine authority and genuinely useful content, that classic SEO already does.
4. AI vs. Humans Content Writing Experiment
Reboot Online ran a strictly controlled comparison, five domains publishing human-written content against five domains publishing raw, unedited GPT-4 content, using an artificial keyword with no prior search history, fresh domains with no backlinks, and equally optimized on-page elements across every site to isolate content origin as the only variable. Human-written content averaged rank 4.4, while raw AI-generated content averaged rank 6.6, a statistically meaningful gap favoring human writing.
This aligns with broader current data too: a 2026 Semrush analysis of 42,000 blog posts across 20,000 keywords found human-written content held position 1 roughly 80 percent of the time versus roughly 9 percent for pure AI content. The consistent theme across both the controlled test and the broader dataset: unedited AI content underperforms, while AI content that gets genuine human editing, fact-checking, and original input closes much of that gap.
5. Content Compression Experiment
Reboot Online analyzed 42,000 webpages specifically to test whether content compression ratio correlates with search engine rankings. The research found content compression is not a determining ranking factor on its own, though top-ranking pages did tend to show higher compression ratios, likely a byproduct of well-written, non-repetitive content compressing more efficiently rather than compression itself influencing rankings.
This is a useful correlation-versus-causation lesson: the pattern exists in the data, but chasing a higher compression ratio directly would be optimizing for a symptom of good writing rather than good writing itself.
6. 404 vs. 410 Response Code Experiment
We analyzed over 350,000 rows of crawl data comparing how Google treats 404 error responses against 410 Gone status codes for removed URLs. The data showed that a 410 response code gets Google to recrawl a removed URL less frequently than a 404 does, meaning 410 is the more effective choice specifically if the goal is fast, permanent deindexing with minimal wasted crawl activity.
Current Google guidance from John Mueller complicates the practical takeaway somewhat: Mueller has stated publicly that 404 and 410 are essentially equivalent for general SEO purposes and that having pages return either code, even at massive scale, is perfectly fine. The honest synthesis: Reboot’s controlled data shows a real, measurable difference in recrawl behavior specifically, even though Google’s broader public messaging treats the distinction as practically unimportant for typical site health.
7. Shared Hosting (“Bad Neighborhood”) Experiment
FHSEOHub tested whether sharing an IP address with spammy, low-quality sites, confirmed to be running aggressive link schemes in adult, gambling, or pharma niches, hurts rankings compared to a dedicated IP. Using 20 brand-new domains, an invented keyword with zero existing search results, and identical content and performance across all sites, the dedicated-IP domains ranked meaningfully better than the shared-IP domains sharing space with spammy neighbors over the three-month test.
This experiment drew real, credible pushback worth including honestly: search engine researchers, and reportedly John Mueller himself, have questioned whether results from an invented keyword with no genuine competition can reliably predict behavior in a real, contested SERP. Treat the finding as a genuine signal worth some caution around, particularly for a site already on budget shared hosting with no other red flags, rather than a confirmed, universal ranking factor.
8. Outgoing Links Experiment
Reboot Online ran a long-term test specifically examining whether outgoing links to relevant, authoritative external sources play a genuine role in Google’s ranking algorithms, directly countering the old assumption that linking out simply leaks PageRank and domain authority away from your own page. Reboot’s own summary of the study describes the results as conclusive in favor of outgoing links having a real, positive effect.
This challenges a persistent myth still common among site owners reluctant to link out at all: citing genuinely relevant, authoritative external sources appears to be a real positive signal, not the pure link equity dilution risk older SEO thinking assumed.
9. Hidden Text Experiment
SearchPilot’s Iceland Groceries case study tested revealing content that had previously been hidden behind tabs, including product details and nutritional information, making it visible immediately on page load instead. Organic sessions increased 12 percent once the previously hidden content became visible by default.
This runs counter to Google’s own public statements, John Mueller and Gary Illyes have both said hidden content in tabs and accordions gets treated equivalently to visible content, yet SearchPilot’s controlled split test found a real, measurable uplift from making it visible. Worth noting honestly: a separate SearchPilot test on FAQ content in a collapsible section found the opposite direction, a positive result from keeping content collapsed, partly attributed to a mobile page speed improvement, a reminder that the same tactic can produce different results on different sites.
10. H1 Tag Ranking Experiment
Moz and SearchPilot ran a 50/50 split test on Moz’s own blog headlines, changing half from H2 tags to H1 tags while leaving the other half unchanged, then measuring the organic traffic difference over eight weeks. The result: no statistically significant difference between the two groups, meaning the H1-versus-H2 tag type itself made no measurable ranking impact.
A separate SearchPilot test on a different question, adding a specific keyword to the beginning of an existing H1’s actual wording rather than changing the tag type, showed a genuinely positive 8 percent uplift. The distinction matters: which heading tag you use appears to matter far less than what that heading actually says, current context from Google’s confirmed March 2026 AI headline rewrite testing makes clear, keeping title tag and H1 wording semantically aligned still helps your intended phrasing survive.
11. Meta Description Rewrite Experiment
SearchPilot ran multiple controlled tests on meta descriptions, and the results consistently favor Google’s own judgment more often than site owners expect. One test forcing Google to respect a custom meta description using the data-nosnippet attribute, preventing Google from pulling alternative text from the page, actually harmed organic traffic. A separate test removing meta descriptions entirely, letting Google auto-generate them from page content instead, produced a positive result.
The pattern across these tests: Google’s auto-generated snippets, when pulled from genuinely good on-page content, frequently match a specific searcher’s query better than a static, manually-written meta description can, the same underlying mechanism that determines which passage gets pulled into a featured snippet. This doesn’t mean skip meta descriptions entirely, they still give Google a strong option to pull from, but it does mean fighting Google’s rewrites through data-nosnippet is a riskier default than most SEO advice suggests.
12. Click-Through-Rate Ranking Experiment
Rand Fishkin’s original IMEC Lab test asked his Twitter following to search “IMEC Lab” and click the resulting link from his blog specifically. An estimated 175 to 250 people responded and clicked, and the page shot up to the number one position shortly afterward, a result Fishkin has pointed to for years as real evidence that click-through rate and click volume can influence rankings. Stone Temple Consulting, led by Eric Enge, later took over running IMEC Lab as a broader cooperative testing effort once Fishkin’s own bandwidth for the project ran out.
Google officials, including Gary Illyes, have repeatedly stated publicly that CTR isn’t used directly as a ranking factor, citing its noisiness as a signal. The honest synthesis: this remains one of the more debated findings in SEO testing, a real, observed ranking shift following a coordinated click campaign, set against consistent official statements that CTR isn’t a direct signal, worth treating as suggestive rather than settled.
13. Nofollow Links Experiment
IMEC Lab-style testing on nofollow links has generally found an indirect, rather than direct, ranking benefit: despite Google’s own guidance that nofollow links don’t pass link equity in the traditional PageRank sense, sites earning a genuine mix of nofollow and dofollow links, with natural anchor text variation across both, tend to see real referral traffic, brand visibility, and a more natural-looking backlink profile than sites relying purely on dofollow links.
The practical takeaway: pursuing a nofollow link from a genuinely relevant, high-traffic source is still worth real effort, not for the direct link equity Google says it doesn’t pass, but for the referral traffic and natural link diversity signal that appears to carry indirect value.
14. Negative SEO Experiment
Tasty Placement ran a deliberate negative SEO test on their own established, well-ranking site, Pool-Cleaning-Houston.com, purchasing a large volume of spam links, including comment links and forum profile links, and pointing them directly at the target domain. The site’s rankings dropped following the spam link campaign, demonstrating that negative SEO through purchased link spam is a real, executable risk rather than a myth, a risk that became far more consequential once Google Penguin started specifically targeting manipulative link profiles, a shift Matt Cutts publicly explained in detail as head of Google’s webspam team at the time.
This is worth taking seriously as a genuine threat, particularly for smaller or newer sites without an established, resilient backlink profile to absorb the damage. Monitoring your backlink profile regularly and using Google’s disavow tool when genuinely suspicious spam links appear remains a real defensive practice, not paranoia.
15. Link Echoes Experiment
Moz’s testing on link removal, associated with Rand Fishkin’s research, found that rankings don’t drop immediately once a backlink is removed. Instead, a page tends to retain much of its ranking position for a period afterward before gradually declining, a residual effect commonly called the link echo or ghost link phenomenon.
This has a real practical implication for link building strategy: a backlink’s value doesn’t disappear the instant it’s removed, meaning a broader, more diverse backlink profile with links accumulated over time carries more resilience than a profile dependent on a small number of recent, fragile placements.
What These Experiments Mean for Your SEO Strategy
Several patterns repeat across this list worth internalizing directly. Content structure choices, H1 versus H2, tabs versus visible text, matter less on their own than what the content actually says and whether it genuinely serves the reader, while technical hygiene choices, 410 over 404, dedicated hosting, outgoing links to real authorities, show more consistent, measurable effects than commonly assumed. AI-generated content without genuine human input underperforms consistently across both controlled testing and broader 2026 industry data, while generative engine optimization increasingly rewards the same underlying authority and content quality signals classic SEO already does, alongside genuine new manipulation risks worth monitoring for directly.
The honest throughline: official guidance and controlled testing sometimes agree and sometimes don’t, and the disagreements, hidden content, CTR, 404 versus 410, are exactly where genuine testing earns its value over repeating a stated policy uncritically. This list also stands in useful contrast to ranking factors Google has directly and unambiguously confirmed itself, like the 2015 Mobilegeddon mobile-friendliness update or HTTPS as a security signal, where no controlled experiment was needed since Google stated the effect outright.
Common Mistakes When Running Your Own SEO Tests
- Testing on an invented keyword with zero competition, the same limitation critics raised against the shared hosting experiment, produces cleaner isolation but weaker real-world applicability, since a genuinely contested SERP behaves differently than an empty one.
- Changing multiple variables at once, new content plus a new heading structure plus a new URL, makes it impossible to attribute any measured change to a specific cause, undermining the entire point of running a controlled test in the first place.
- Stopping a test too early, before reaching genuine statistical significance, risks mistaking normal ranking volatility for a real effect, while running a test far too long risks confounding variables like algorithm updates or seasonal demand shifts contaminating the result.
- Ignoring context that doesn’t match your hypothesis, publishing only the positive finding while quietly discarding an inconclusive or negative one, is how confirmation bias creeps into SEO testing even among well-intentioned practitioners.
Conclusion
Real SEO experiments reveal a more nuanced picture than either “Google confirmed it” or “everyone knows this” content ever captures, some official guidance holds up under controlled testing, some doesn’t, and the technical and content choices that show measurable, repeated effects are worth prioritizing over ones that sound intuitive but haven’t been tested. Build your own testing practice around genuine control groups, single-variable isolation, and honest reporting of inconclusive results, not just the wins. Treat every experiment, including the fifteen here, as evidence to weigh against your own site’s context, not a universal rule to apply blindly.
FAQs
Yes, arguably more relevant, since the same authority and content quality signals that classic SEO experiments measure increasingly determine AI citation eligibility too, and new GEO-specific experiments are actively testing how these systems can be influenced or manipulated.
It remains genuinely debated. Rand Fishkin’s IMEC Lab experiment showed a real ranking shift following a coordinated click campaign, while Google officials have repeatedly stated CTR isn’t used directly as a ranking signal due to its noisiness, making this one of the more contested findings in SEO testing.
A 410 gets Google to recrawl the removed URL less frequently based on controlled testing, though John Mueller has stated both codes are essentially fine for general SEO purposes, meaning the distinction matters most specifically when fast deindexing is the priority.
Reboot Online’s controlled test found dedicated-IP sites outranking shared-IP sites hosted alongside spammy neighbors, though this specific experiment drew credible criticism for using an invented keyword with no real competition, worth weighing as a signal rather than a confirmed universal factor.
Not without genuine human editing. Controlled testing found human-written content outranking raw, unedited AI content, and broader 2026 data shows human content dominating top positions, though AI content that receives real editorial input and original input closes much of that gap.
Yes. Tasty Placement’s own controlled test on their established site showed rankings dropping following a deliberate spam link campaign, confirming this as a genuine, executable risk worth monitoring your backlink profile against, not just a theoretical concern.