Back to blog

Percent of What? How an Entire Industry Quoted the Wrong Column for a Year

Reddit is not 46.7% of Perplexity's citations. It is 6.6%. The bigger number is a different column of the same table, and the year it spent going unchallenged says more about how marketers handle numbers than it does about AI search.

Aziz Banihashemi
Aziz Banihashemi
· 10 min read
Percent of What? How an Entire Industry Quoted the Wrong Column for a Year

Three numbers about AI search have been repeated for a year. All three are misreadings of a single table. The AI search part is the hook. The habit underneath it is the story, and it runs straight through your discount reporting.

The three numbers you have already seen

If you have read anything about answer engine optimization in the past year, you have met these:

  • Reddit is 11.3% of ChatGPT's citations.

  • Reddit is 21% of Google AI Overviews' citations.

  • Reddit is 46.7% of Perplexity's citations.

None of them is true. Not rounded wrong, not directionally wrong. Wrong by a factor of six to ten.

The actual figures come from the same study, the same page, and the same day: 1.8%, 2.2%, and 6.6%.

The receipt

The source is Profound's citation patterns study, written by Nick Lafferty, published June 5, 2025 and updated that August. It covers 680 million citations from August 2024 to June 2025. It is the largest public dataset in the space.

The page contains two different kinds of table, and the industry has been quoting the wrong one.

Platform

Reddit's share of ALL citations

Reddit's share of the TOP 10 sources

ChatGPT

1.8%

11.3%

Google AI Overviews

2.2%

21.0%

Perplexity

6.6%

46.7%

Profound labels the second kind of table plainly:

"These percentages do not represent overall citation volume, but rather how citations are distributed among the leading sources for each platform."

The denominator of the second column is ten websites. Not the internet. When someone tells you Reddit is 46.7% of Perplexity's citations, they are telling you that Reddit is 46.7% of a ten item list that Reddit already sits at the top of by construction.

You can prove it with a calculator

This is not a matter of interpretation. The second column is just the first column divided by the sum of the top ten.

Perplexity's ten most cited domains, as a share of all citations, are Reddit 6.6%, YouTube 2.0%, Gartner 1.0%, Yelp 0.8%, LinkedIn 0.8%, Forbes 0.7%, NerdWallet 0.6%, TripAdvisor 0.6%, G2 0.6%, and PCMag 0.5%. They add up to 14.2% of all citations. Reddit's 6.6% divided by that 14.2% is 46.5%.

The reported figure is 46.7%. Run the same arithmetic on Google AI Overviews and you get 21.4% against a reported 21.0%. Run it on ChatGPT and you get 11.0% against a reported 11.3%. Every one reproduces to within a few tenths of a point, which is exactly what you expect from source figures rounded to one decimal place.

So 46.7% was never a second measurement. It is the same 6.6%, with the internet taken out of the denominator and nine other websites put in.

Profound did not make this mistake

Worth being fair about, because it sharpens the point rather than blunting it. Profound published both tables, labelled the difference, and added an explicit caveat. The company that had the most to gain from the flattering number is the one that printed the honest one next to it.

The error is entirely downstream, in the blog posts, agency decks, and LinkedIn threads that reprinted the bigger figure because it made the better headline. Nobody checked the denominator, because checking would have killed the pitch.

Five denominators, one phrase

Here is the part that should worry anyone who reports on marketing performance. At least five different quantities are circulating, and all five get spoken aloud as "Reddit's citation share."

What is measured

Percent of what

Reported value

Source

Share of all citations

Every citation in the dataset

1.8% (ChatGPT)

Profound, 680M citations

Share of the top 10 sources

Ten domains

11.3% (ChatGPT)

Profound, same page

Share of cited sources

The sources one tracker saw

14.29% (early Aug 2025)

Spotlight

Share of responses citing Reddit at all

Prompt responses

Close to 60% (early Aug 2025)

Semrush, 230K prompts

Share of social citations only

Social platforms only

44% (AI Overviews, Jan 2026)

Tinuiti

From 1.8% to 60%. All five describe Reddit. All five describe AI citations. No two of them measure the same thing.

The Semrush number deserves its own warning, because it is the most quotable and the most misquoted. Semrush found that ChatGPT cited Reddit in close to 60% of prompt responses in early August, falling to around 10% by mid September. That is the percentage of answers containing at least one Reddit link. It is not Reddit's share of citations. Anyone who writes "Reddit fell from 60% of citations to 2%" has bolted one study's numerator onto another study's denominator and produced a sentence that describes nothing that exists.

Screenshot 2026-07-13 at 10.52.16 PM.png

The Tinuiti figure needs the same care. 44% is Reddit's share of social citations in AI Overviews, and social media was only around 13% of AI Overviews citations. Reddit is therefore about 5.7% of all citations there, not 44%.

Even the correction commits the error

In March 2026, Animalz published "Why We Gave Up On Reddit For AEO", and it deserves real credit: it catches the top ten problem cleanly.

"Maybe you've seen a stat that says Reddit accounts for around 47% of Perplexity's citations, but that's a misrepresentation. That number is also from Profound's data, but it says Reddit is 46.7% of Perplexity's top 10 most-cited sources, not its share of all citations."

Correct, and well caught. The same article quotes Profound's 1.8% figure for ChatGPT. Then, a few paragraphs later, it says the recovery is "still below pre-crash levels of 9-14%."

Both sentences describe Reddit in ChatGPT before the crash. They differ by a factor of five to eight. The article never reconciles them, never flags the jump, and never states which denominator the second number uses. The 14 has probably wandered in from Spotlight's 14.29% share of cited sources, which is a different measurement entirely, but the piece does not say so.

The article that exists to catch this error commits it four paragraphs later. That is not a takedown of one writer. It is evidence that the habit is systemic, and that catching it in someone else's work does not inoculate you against it in your own.

One more thing about that piece, offered as a caution rather than an accusation. Its conclusion is to invest in your own domain instead of Reddit, and it presents no comparative citation data for owned domains at all. Animalz sells Answer Engine Optimization as a service line. The single recommendation with no numbers behind it is the one the publisher monetizes. That does not make it wrong. It does mean the claim that needed the most evidence arrived with the least.

The citation supply chain

Trace the sources back and the independence starts to evaporate.

  • One widely shared post attributes its decline figures to Seshes.

  • Seshes publishes no original measurement of its own. It aggregates PromptWatch, Similarweb, and RBC, and says its own index is still to come.

  • RBC's much repeated fall from 29.2% to 5.3% is a bank relaying an unnamed "third-party study," with no denominator stated anywhere.

  • Tinuiti's citation reports are built on the Profound platform, so the "independent corroboration" of Profound's data is, in part, Profound's data.

What looks like a dozen sources converging on a conclusion is a handful of measurements being passed hand to hand, losing a denominator at every step.

The numbers nobody can source

These four claims circulate constantly. You will meet them in LinkedIn posts and Reddit threads, used as settled inputs to technical arguments about where to spend your content budget:

  • ChatGPT cites positive Reddit sentiment at 5% and negative sentiment at 6.1%.

  • The average Reddit post cited by AI models was written about a year earlier, and 4% date from 2019 or older.

  • Reddit's citation share grew at least 73% between October and January.

  • That growth is concentrated in technology and electronics.

Each is attributed to a named research source. I went looking for the primary documents, and I could not stand any of them up. The Profound page these first two are credited to contains no sentiment analysis and no content age analysis of any kind. The Tinuiti report the others are credited to now returns a 403, and its link serves a different quarter's report instead. What Tinuiti has published openly puts the highest Reddit share in apparel and lists technology among the biggest decliners, which points the other way.

I cannot confirm these numbers and I cannot refute them. I am including them precisely because you will see them quoted with confidence. If you are about to build a strategy on one of them, find the primary source first. I could not.

What actually survives

Strip out everything unverifiable and a real story remains.

The crash happened. Reddit's presence in ChatGPT fell sharply in mid September 2025. Every tracker agrees on the direction, and Reddit's stock fell in the weeks that followed.

The mechanism is plausible. Google removed its num=100 search parameter around September 10, 2025. As Kevin Indig explains, OpenAI does not crawl Google directly but buys search data from third parties who relied on that parameter, and per Ahrefs, 57.8% of Reddit's keywords rank outside the top 20. Lose results 11 through 100 and you lose most of Reddit. It is coherent.

It is not settled. Semrush's own head of organic and AI visibility says he does not think the parameter removal is the root cause, "or at least, not the only one," and suggests OpenAI may simply be rebalancing away from over-cited domains. Semrush itself declines to endorse a cause. The tidiest explanation in the space is doubted inside the company holding one of the biggest datasets on it.

The magnitude is unknown. Not disputed. Unknown. The reported declines cannot be compared because the things being measured are not the same thing.

Now go and look at your own dashboard

Here is why this belongs on a Shopify blog rather than an SEO one. The exact same failure lives in discount reporting, and it costs merchants real margin.

"Our Black Friday discount drove a 40% lift." Lift over what? The previous week, which was artificially flat because customers were waiting for the sale? The same week last year, when the catalog was a third smaller? The baseline decides the answer, and the baseline is almost never stated.

"Discount codes drove 60% of revenue." No. Sixty percent of revenue had a code attached to it. Attached is not caused. This is the ecommerce version of reading a top ten share as a total share: a number that is technically accurate and completely misleading.

"Conversion rate rose 25% during the promotion." Conversion rate is a fraction. Discount hunting traffic changes the denominator. The ratio can move without a single extra sale being caused by anything you did.

And the question almost nobody asks: how many of those discounted orders would have happened anyway, at full price? A discount handed to a customer who was already going to buy is not a win. It is a margin transfer, out of your pocket and into theirs.

This is the entire reason we built profit analytics into Discount Prime the way we did, showing net margin per campaign against your real Shopify costs. A margin floor is not a safety feature. It is a device that forces the profit question to be asked before the campaign runs, instead of being reconstructed afterwards from whichever metric happens to flatter the result.

Three rules

1. Ask "percent of what." If a source does not state its denominator, the number is not conservative and it is not directional. It is unusable. Discard it.

2. Choose the denominator before the campaign, not after. A metric picked once the results are in will always be the one that makes the results look good. That is not dishonesty. It is gravity.

3. Keep the negative results. The reason a bad number ran unchallenged through an entire industry for a year is that almost nobody publishes what did not work.

The 46.7% spread and the 6.6% did not, and the only difference between them was which one justified the budget.

AEOGEOAI searchmeasurementanalyticsdiscount reporting
Aziz Banihashemi

About the author

Written by the team who design and build Discount Prime for Shopify merchants. We write about commerce infrastructure, profit-aware pricing, and the ideas behind what we ship.

Frequently asked questions

Is Reddit really 46.7% of Perplexity's citations?

No. Reddit is 6.6% of Perplexity's total citations. The 46.7% figure is Reddit's share of Perplexity's ten most-cited sources, a list Reddit already tops. Profound published both numbers on the same page and explicitly noted that the top-ten percentages do not represent overall citation volume.

What is Reddit's actual share of AI citations?

Across 680 million citations measured between August 2024 and June 2025, Reddit accounted for 1.8% of ChatGPT citations, 2.2% of Google AI Overviews citations, and 6.6% of Perplexity citations. On ChatGPT, Wikipedia led at 7.8%. These are shares of all citations, not shares of a top-ten list.

Why did Reddit citations in ChatGPT drop in September 2025?

The cause is not settled. Google removed its num=100 search parameter around September 10, 2025, and since OpenAI buys search data from third parties who relied on it, Reddit lost visibility because 57.8% of its keywords rank outside the top 20. However, Semrush's head of organic and AI visibility publicly doubts this is the root cause.

Why do different sources report wildly different Reddit citation numbers?

Because they measure different things and call them the same name. At least five denominators circulate: share of all citations, share of the top ten sources, share of one tracker's cited sources, share of prompt responses containing a Reddit link, and share of social citations only. Reported values run from 1.8% to 60%.

What does an AI citation statistic have to do with discount reporting?

Both fail the same way. Saying discount codes drove 60% of revenue means 60% of revenue had a code attached, not that the code caused the sale. Attached is not caused. Like reading a top-ten share as a total share, it is technically accurate and completely misleading, and it hides whether a discount was incremental at all.

Run profit-first promotions on Shopify

Discount Prime brings eight discount types, margin analytics, and conflict detection into one Shopify-native app.

See pricing

Or see it live in our demo store and watch every feature working on a real Shopify store.

Related articles