← Back to blog/Blog·August 17, 2026·50 min

Where Should You Publish Content to Get Cited by AI? Evidence from 11 Engines and 7,317 Citations

Five answers from testing 11 AI engines and 7,317 citations: the number of sources provided by different engines varies tenfold; engines with their own content platforms cite their own ecosystems significantly more often; only 30 sources are genuinely universal across engines, accounting for 1.9% of all domains; and choosing the right platform does not mean your content will be cited. Includes a list of platforms that accept contributions, a comparison of citation concentration, and an explanation of the data’s limitations.

Y
YinJen GEO Team
Generative Engine Optimization · YinJen

Five Questions This Article Answers

  1. How many sources does an AI provide when answering a question? How much do different AI engines vary?
  2. Do different AI engines have the same source preferences? Are there identifiable patterns?
  3. Are there any “universal sources”—platforms recognized by most AI engines?
  4. If you publish on these high-coverage platforms, are you certain to be cited?
  5. How can you determine whether a platform is worth investing in?

We will answer them one by one below. Every conclusion comes from the same set of empirical data. The method and definitions are set out in the first section so the results can be reproduced.


1. Where This Data Came From

The question itself is not complicated: when an AI answers a recommendation-oriented question, where do the sources it provides come from?

The difficult part is that this cannot be answered in general terms—change the category or the wording of the question, and the answer may be completely different. The only option is to count everything clearly within a specific sample.

Our method:

  • Question set: 45 fixed Chinese-language questions centered on the single category of “AI visibility monitoring / Generative Engine Optimization,”

divided into five tiers—core category, long-tail Q&A, region-specific, technical and integration, and competitor comparison

  • Execution: each question was asked once on each of 11 AI engines, and every source URL appearing in the answers was recorded
  • Window: 2026-07-28 to 08-11
  • Scale: 821 answers containing citations, yielding 7,317 citations across 1,610 unique domains

Three definitions need to be clear upfront so the findings are not mistaken for universal conclusions:

  1. This is an empirical test of a single category. All 45 questions concern the single topic of GEO. If the topic changed to “How should I choose a robot vacuum?”, the source structure would probably be completely different. Every conclusion in this article applies only within this category.
  2. Citation counts are based on the number of URL appearances. If the same URL appears repeatedly under different questions or on different engines, every appearance is counted. Therefore, “a platform was cited 378 times” does not mean “378 different articles.”
  3. There are 11 engines, not 12. One engine did not run during this window, while another completed only about 60% of the runs. Its data is weaker and will be identified separately in the article.

2. Q1: Different AI Engines Provide Tenfold Different Numbers of Sources

This is the first thing to examine because it determines how many opportunities there are to be cited.

EngineAverage Sources per AnswerMedianMaximum in One Answer
ERNIE Bot20.02020
Doubao16.817.520
DeepSeek11.41215
Tencent Yuanbao10.91520
Zhipu AI9.28.520
StepFun9.0817
Qwen7.3715
Kimi5.9616
ChatGPT4.6413
Perplexity3.739
Gemini1.924

The difference between ERNIE Bot and Gemini is tenfold.

ERNIE Bot’s mean is exactly 20, its median is 20, and its single-answer maximum is also 20. That pattern shows that it fills to the maximum limit: whenever source material is available, it takes as many items as permitted. Doubao follows nearly the same pattern.

Conversely, Perplexity, with a mean of 3.7, and Gemini, with a mean of 1.9, follow a model of selecting only a few sources.

The direct implication for content distribution is this: for the same article, the probability structure of appearing in an ERNIE Bot answer is inherently much more favorable than that of appearing in Gemini, because the former has 20 slots to fill each time while the latter fills only 2. If your goal is the act of “being cited” itself, put resources into engines that provide more sources first; the expected hit rate is higher.


3. Q2: Every AI Engine Has Very Different Tastes, and Chinese Engines Clearly Favor Their Own Ecosystems

3.1 Most-Cited Sources for Each Engine

EngineTop 5
ERNIE BotBaijiahao 217, Bilibili 118, CSDN 101, Zhihu Columns 79, China Development Network 61
DeepSeekGitHub 63, Baidu Baike 46, Alibaba Cloud Community 35, Tencent Cloud Community 32, Newrank 31
Tencent YuanbaoWeChat Official Accounts 106, Tencent News 39, Toutiao 27, China Economic Net 26, Volcano Engine Community 23
DoubaoDouyin 187, Toutiao 124, CNBlogs 48, Sohu 37, CSDN 33
StepFunShenma Search 117, CNBlogs 60, NetEase 31, CSDN 31, ITHome 30
Zhipu AICSDN 45, CNBlogs 43, Volcano Engine Community 33, Toutiao 32, Tencent Search 27
KimiCNBlogs 77, Volcano Engine Community 25, ITHome 20, Sohu 15
QwenJiemian News 90, ITHome 56, China Business Network 38, Toutiao 30, China Economic Net 23
PerplexityCNBlogs 48, Tencent Cloud Community 36, NetEase 29, Volcano Engine Community 17, Alibaba Cloud Community 15
ChatGPTarXiv 40, Reddit 18, Volcano Engine Community 18, TechRadar 12
GeminiPhoenix Finance 21, Sina Finance 12, SimilarWeb 6, YouTube 5

Several differences are immediately visible:

  • DeepSeek favors technical sources the most: GitHub ranks first, followed by the developer communities of several cloud providers.
  • Qwen favors news media the most: Jiemian News, ITHome, and China Business Network rank at the top, while technical communities account for very little.
  • ChatGPT uses a completely different set: arXiv and Reddit lead, combining English-language academic sources and community forums, with no overlap with any of the Chinese engines.
  • ERNIE Bot is the only one to cite a video platform heavily: Bilibili was cited 118 times.

3.2 Engines with Their Own Content Platforms Cite Their Own Ecosystems Significantly More Often

Classifying each engine’s citations by whether they belong to a content platform owned by the same company:

EngineTotal CitationsShare from Own EcosystemMain Sources
Doubao (ByteDance)93836%Douyin 187, Toutiao 124
ERNIE Bot (Baidu)127822%Baijiahao 217, other Baidu properties
Yuanbao (Tencent)107715%WeChat Official Accounts 106, Tencent News 39
Qwen (Alibaba)4382%Alibaba Cloud Community 9
DeepSeek / Kimi / Zhipu AI / StepFun0%These companies do not have their own content platforms

One-third of Doubao’s citations come from Douyin and Toutiao. That proportion is too high to dismiss as noise.

But note the counterexamples: DeepSeek, Kimi, Zhipu AI, and StepFun each have a 0% share from their own ecosystem because they simply do not have their own content platforms. This suggests that the preference above looks less like “models favoring their own side” and more like a difference in data availability: when a ready-made content repository is at hand, it is easier to retrieve material from there.

Qwen’s share of only 2% also supports this explanation: Alibaba has a content ecosystem, but there is little Alibaba-affiliated content in this particular category.


4. Q3: There Are Only 30 Truly “Universal Sources,” Accounting for 1.9% of All Domains

This is the most counterintuitive and useful finding in the entire dataset.

Among the 1,610 cited domains, the distribution by the number of engines that cited each domain is:

Number of Engines Citing the DomainNumber of DomainsShare
1030.2%
930.2%
870.4%
770.4%
6100.6%
5231.4%
4271.7%
3694.3%
219512.1%
Only 11,26578.6%

Nearly 80% of domains were cited by only one engine.

In other words, the great majority of content can win over only one engine. If DeepSeek cites you on a small website, that does not mean Doubao or ERNIE Bot will look in the same place—they probably will not.

There are only 30 “universal sources” cited by 6 or more engines. Of those, the ones where ordinary people can contribute include:

  • Technical communities: CNBlogs, CSDN, Volcano Engine Community, Tencent Cloud Community, Alibaba Cloud Community, SegmentFault
  • Media accounts: NetEase, Sohu
  • News media (not open to contributions; listed only for reference): ITHome, Jiemian News, China Economic Net, The Beijing News

If the goal is to “make yourself visible to as many AI engines as possible,” the available range is actually very narrow.

4.1 Platforms That Accept Contributions, Ranked by Engine Coverage

PlatformEngines CoveredTotal Citations
Volcano Engine Developer Community10/11149
CNBlogs9/11378
CSDN9/11287
Tencent Cloud Developer Community9/11161
NetEase Media9/11111
Sohu Media8/11123
Alibaba Cloud Developer Community8/1186
SegmentFault8/1137
Toutiao5/11214
Juejin5/1120
WeChat Official Accounts2/11111
Zhihu Columns2/1184

Zhihu Columns and WeChat Official Accounts both cover only 2/11 engines.

These are two of the largest content platforms on the Chinese internet, but in the citation pool for this question set, only a few engines cite them: WeChat Official Accounts are cited mainly by Tencent’s own Yuanbao, while Zhihu Columns are cited mainly by ERNIE Bot and DeepSeek; the other engines draw from them very little.

This does not mean that these two platforms have no value—it means that if your purpose in investing in them is “to be cited by AI,” that objective will probably fail. When they serve functions such as branding, owned-audience development, or conversion, they must be evaluated against those functions instead of being mixed into the same ledger.

4.2 Citation Concentration: Which Engines Can New Content “Break Into”?

Although all of these sources are cited, some engines repeatedly use the same few sources while others are much more dispersed:

EngineTop 10 Domains as a Share of All Its CitationsTotal Domains
Qwen69.9%70
ERNIE Bot62.4%195
Gemini60.2%47
StepFun59.3%120
Doubao56.3%256
Perplexity47.4%126
Zhipu AI44.0%225
ChatGPT39.2%144
Kimi36.1%228
Tencent Yuanbao28.1%483
DeepSeek25.1%413

Qwen recognizes only 70 domains, and the top 10 account for 70% of its citations—this is an almost closed pool that new content will find difficult to enter.

DeepSeek and Yuanbao are the opposite: they cite 413 and 483 domains, respectively, and their top 10 account for only 20% to 30%, leaving a very long tail. Structurally, the barrier is much lower for new content on these two engines.

Combining Q1 with this section produces an interesting ranking: Yuanbao and DeepSeek both provide many sources (10.9 / 11.4) and draw from dispersed pools (28% / 25%), making them the two easiest engines for new content to break into. Qwen, by contrast, provides fewer sources (7.3) and draws from a highly concentrated pool (70%), making it the hardest.


5. Q4: Choosing the Right Platform Does Not Mean You Will Be Cited

This is the most important section in the article and the easiest one to skip.

All the rankings above describe which platforms the engines cited. That does not mean that publishing on those platforms will make your content cited.

There are still two steps in between:

  1. Your article must first be indexed by a search engine;
  2. After it is indexed, it must still be selected for a specific question.

We ourselves got stuck at these two steps.

We published content on four platforms—two technical communities, one Q&A community, and one news platform—for a total of 6 articles. We matched the URLs of these 6 articles one by one against all 7,317 citations. The result was:

0 citations. Not even once.

During the same window, the domains of those four platforms appeared hundreds of times in the citation pool—every citation was to someone else’s content.

One point here is very easy to misread and deserves to be stated separately:

A domain entering the citation pool ≠ your article entering the citation pool.

We once made this exact analytical error. When we saw that the domain of a technical community had been cited 89 times, it was natural to ask, “How many of those were ours?” We used a path wildcard to tag “our articles”—but that classified the articles of everyone on that community as ours, producing the conclusion that “we were cited 22 times.” After switching to exact, article-by-article matching against real article IDs, the actual result was 0.

One result makes you think “this path works,” while the other shows that “not a single article got in.” Those conclusions lead to completely opposite allocations of resources.

Therefore, any determination that “my content was cited by AI” must use exact URL- or ID-level matching. A domain match or path wildcard is not sufficient.

5.1 So What Does Cited Content Look Like?

If the platforms were right but the content did not make it into the pool, the next step is to examine what did.

We analyzed the most-cited Toutiao article in the citation pool (cited 9 times). Its characteristics were:

  • The title is the question itself: “Which Full-Stack GEO Optimization Service Provider Is Reliable in 2026? Real-World Tests of Leading Vendors, Common Pitfalls, and a Complete Partnership FAQ”—almost synonymous with the question “Which brand AI visibility monitoring tool is best?” in the question set
  • Nearly 18,000 Chinese characters long
  • The first section opens by listing “the five key questions this article answers,” and then answers them one by one
  • Published on 08-04 and cited 9 times within one week—showing that old content does not exclusively occupy the slots; new content can also enter the pool quickly

The fourth point is especially important: it rules out the explanation that “we simply have not reached the front of the queue.” New content can enter the pool within a week. If our content was still at 0 after two weeks, time was not the cause.

By comparison, our 6 articles shared the following characteristics: their titles were declarative statements (“Issue N of an Experiment”), they were 2,000 to 3,000 Chinese characters long, they went directly into the main text, and each covered only one slice of the data.

The conclusion is not that “you must write that kind of ranking article”—that involves a separate set of tradeoffs. But four practices can be adopted without evaluating any competitor or promoting yourself:

Practice to AdoptExplanation
Use the exact question a user would ask as the titleAI matches questions, not topics
Make the article long enough to cover an entire question clusterFinish answering a topic, rather than presenting one issue of data
State at the beginning which questions the article answersThis structure is written for machines to read
Cover the complete question cluster in one articleInstead of slicing it by week and discussing only one part in each issue

This article itself follows these four practices. Whether they work must be judged from the citation-pool data in the next issue; it is too early to draw a conclusion now.


6. Q5: How to Determine Whether a Platform Is Worth Investing In

Combining the previous sections yields a decision sequence that does not depend on intuition:

Step 1: Check whether it is a universal source. Only 30 domains (1.9%) were cited by 6 or more engines. If the target platform is outside this set, content published there can reach no more than one or two engines. This step eliminates the great majority of platforms.

Step 2: For the engine you want to target, examine how many sources it provides and how dispersed its pool is. If you want to target only one engine—for example, the main traffic source for a particular category—look at its citation density and concentration: engines that provide many sources and use dispersed pools (Yuanbao: 10.9 sources / 28%; DeepSeek: 11.4 / 25%) are easier to enter; engines that provide few sources and use concentrated pools (Qwen: 7.3 / 70%) are difficult.

Step 3: Confirm that your content can be indexed. This is the step most easily skipped and most likely to become a dead end. A page that has not entered a search index cannot enter the citation pool. After publishing, check actual indexation; do not assume that “published” means “indexed.”

Step 4: After publishing, verify by exact URL; do not look only at the domain. This is the only method that can truly answer whether your content has been cited.

Of the four steps, the first two rely on data and the last two rely on discipline. Most people complete only an intuitive version of the first step (“this platform is large, so it should be useful”) and then skip the final three.


7. Appendix: Different Types of Questions Also Have Different Source Structures

For the same group of engines, the source distribution changes when questions at different tiers are asked:

Question TierTop 5 Sources
Core category (needs-based questions without a brand name)CNBlogs 206, CSDN 128, ITHome 97, Baijiahao 95, Douyin 89
Competitor comparison (directly asking about a product)ITHome 61, CNBlogs 44, CSDN 40, Volcano Engine Community 39, China Economic Net 36
Long-tail Q&A (conversational questions about specific scenarios)Baijiahao 46, WeChat Official Accounts 39, CNBlogs 35, CSDN 30
Technical and integration (implementation-oriented)CSDN 43, Baijiahao 42, Tencent Cloud Community 40, Douyin 34, Bilibili 31
Region-specific (with a geographic qualifier)CNBlogs 65, Baidu Baike 24, Toutiao 20, Baijiahao 16

Two points are worth noting:

  • The leading source for competitor-comparison questions is a news outlet (ITHome), not a technical community. When users ask directly what a product is like, engines tend to draw from media reviews rather than community discussions.
  • For technical and integration questions, Tencent Cloud Community ranks third (40 citations), substantially higher than its position in other question tiers—indicating that it carries more weight for implementation-oriented questions.

This means that the question “Which platform should I publish on?” must be broken down one level further: which type of question do you want to cite you?


8. Limitations of This Dataset

These are stated at the end because they determine how broadly all the conclusions above can be applied:

  1. A single category and a single question set. All 45 questions concern the single topic of GEO. If the category changes, the conclusions probably do not apply.
  2. The window is only two weeks. The citation pool changes—we observed that the number of hits for questions in the same tier can fall from 3 to 0 between two adjacent windows. No single-window reading can serve as a baseline; at least three points are needed before discussing a trend.
  3. Citation counts are based on the number of URL appearances. Repeated appearances of the same URL are counted repeatedly and do not equal “the number of different articles.”
  4. There are 11 engines, not 12. One did not run and one completed only about 60% of the scans; the latter’s preference data is weaker.
  5. This is an observation of “whom the engines cited,” not a method for “how to get cited.” The former is a necessary condition, not a sufficient one.
Y
About the author
YinJen GEO Team

YinJen's Generative Engine Optimization (GEO) research & field team — we track how content gets cited and surfaced across ChatGPT, Claude, Gemini, Perplexity, Doubao, DeepSeek and other major AI engines. This series is first-hand field notes.

See how YinJen does GEO →