← Back to blog/Blog·September 30, 2026·48 min

China AI Engine Citation Sources Monthly, Issue 1 (September 2026): 13,203 Citations from 2,225 Domains Across Three Rounds, with a Reproducible Method

In September 2026 we put the same 45 Chinese-language GEO questions to AI engines in three weekly rounds (the product covers 12 AI engines; 11 were actually run in this issue, 9 of them within the reliable range) and counted the citations in their answers by domain: 13,203 citations from 2,225 domains. The 10 most-cited domains account for 24.9% and the top 100 for 63.8%; 1,313 domains appeared only once across the three rounds. This article gives the overall top 20, how the head moved across the three rounds, the head of each of the five question layers, a share-based comparison with August’s issue-5, and a method and script you can follow to check the figures that come from the public file. The raw dataset, issue-11, is public on GitHub. These are measurements on a single merchant’s question set, not an industry average.

Y
YinJen GEO Team
Generative Engine Optimization · YinJen

This is Issue 1 of the *China AI Engine Citation Sources Monthly*. Every week we put the same set of 45 Chinese-language GEO questions to a group of AI engines and store the citation links in their answers exactly as given. This issue merges the three September weekly monitoring rounds that had closed when the data was pulled (29 September) and counts them by domain: 13,203 citations from 2,225 distinct domains.

Four readings first; every section below gives the source and the method:

  • The head is not concentrated. The 10 most-cited domains account for 24.9% of all citations, and the top 100 for 63.8%. Beyond the top 100 there are another 2,125 domains, with an average of just 2.25 citations each.
  • The long tail is long, and it changes every round. Of the 2,225 domains, 1,313 appeared only once across the three rounds. Only 356 appeared in all three rounds, yet they contributed 77.5% of the citations.
  • A single round can mislead. zhuanlan.zhihu.com, in the top 10, received 61, 34 and 126 citations in the three rounds; bilibili.com, in 25th place, received 101, 2 and 2.
  • Different questions, different heads. In the "Technical & integration" layer, the top domain is developer.aliyun.com; in the "Regional" layer it is cnblogs.com; in the other three layers it is baijiahao.baidu.com or iesdouyin.com.

The raw dataset is public in the chinese-ai-engine-sources repository under the GitHub organization ZhiMaHang, as data/2026-09-18-issue-11.json (issue-11 below). Every figure in this article can either be computed directly from that file or is marked in the text as coming from the underlying table of the same snapshot.


1. Scope: read this section before the numbers

ItemThis issue
Data sourceYinJen's self-monitoring records (GEO monitoring we run on our own brand). These are measurements on a single merchant's question set, not an industry average
Question set45 Chinese-language GEO monitoring questions in five layers: Core category 21, Long-tail Q&A 7, Technical & integration 7, Regional 4, Competitor comparison 6. The questions are public in the same repository as prompts/zh-geo-45.md and match the questions actually being run word for word
EnginesThe product covers 12 AI engines; 11 were actually run in this issue, 9 of them within the reliable range (see the table below)
RoundsThe three September weekly monitoring rounds that had closed when the data was pulled, i.e. rounds 9, 10 and 11 of our weekly monitoring. Split by round ID, not by date window
Windows (Beijing time)Round 9: 09-01 14:13 ~ 09-05 12:16; Round 10: 09-09 09:53 ~ 09-14 15:52; Round 11: 09-17 11:40 ~ 09-18 11:48
Not includedRound 12 (from 09-24): re-runs were still unfinished when the data was pulled, so it moves to Issue 2
Unit1 citation = 1 URL given in an answer. The same URL appearing under different questions, engines or rounds counts once each; repeats within the same answer are not counted
DomainsTaken as-is, subdomains not merged: blog.csdn.net, m.blog.csdn.net and csdn.net each count separately
OrderingBy count, descending; ties broken alphabetically by domain
SnapshotPulled at 2026-09-29 16:08 Beijing time. Monitoring records can be rewritten in place by re-runs, so a later recomputation may differ slightly
DesensitizationBefore publication, 25 domains (42 citations) were removed under fixed rules; none of them is in the top 100. All totals in this article are post-removal figures; see Section 8

Two numbering systems—don't mix them up. "Issue 1" is this article series' issue number; "issue-11" is the dataset number, which takes the round number of the last weekly monitoring round included. The issue-5 dataset published in August is single-round data from round 5; issue-6 to issue-10 were not published separately, and rounds 9 and 10 have been merged into this issue.

Engine coverage (45 questions per round, so each engine should run 135 times across the three rounds; "valid answers" means answers with status ok)

Engine (key in the dataset)Valid answersCitations this issueStatus
Ernie (wenxin)135/1352,652Reliable range
Yuanbao (hunyuan)134/1352,354Reliable range (1 failure in round 11)
Doubao (doubao)135/1352,016Reliable range
DeepSeek (deepseek)135/1351,816Reliable range
GLM (glm)135/1351,345Reliable range
StepFun (stepfun)135/135996Reliable range
Qwen (qwen)135/135763Reliable range
Perplexity (perplexity)135/135511Reliable range
Gemini (gemini)134/135200Reliable range (1 failure in round 11)
ChatGPT (chatgpt)92/135284⚠️ Outside the reliable range (only 2/45 in round 11)
Kimi (kimi)81/135266⚠️ Outside the reliable range (36/45 in round 10; 0/45 in round 11 because the login session had expired)
Claude (claude)0/1350Deliberately not run in any of the three rounds; not counted as run
  • The reliable-range criterion is our own: a combined valid rate of at least 95% across the three rounds, with no single round below 90%.
  • Citations from ChatGPT and Kimi are still counted in the totals and in every list; they are flagged, not removed. Gaps are reported as they are—not filled in, not smoothed over.
  • The citations of the 11 engines add up to 13,203, matching the total (the script in Section 7 checks this).

2. Overview: the head takes a quarter, the long tail more than a third

The three rounds received 4,555, 4,372 and 4,276 citations respectively. In round 11, ChatGPT and Kimi barely ran, and that factor is mixed into the round-to-round differences in the totals.

CutDomainsCitationsShare of all citations
No. 1 (baijiahao.baidu.com)16154.66%
Top 10103,28724.9%
Top 20205,08438.5%
Top 1001008,42163.8%
Beyond the top 1002,1254,78236.2%
Of which, appearing only once across the three rounds1,3131,3139.9%
  • 25 domains were cited 100 times or more; the threshold for the top 100 is 17 citations (No. 100 has 17; in the underlying table No. 101 has 16, so there is no tie at the cut).
  • The first five rows of the table can be computed directly from issue-11. The last row, and the per-round totals at the start of this section, come from the underlying table of the same snapshot; the public JSON does not contain them, so we have also written them into the issue section of the repository README.

3. Overall top 20

#DomainCitationsShare
1baijiahao.baidu.com6154.66%
2iesdouyin.com4043.06%
3blog.csdn.net3572.70%
4cnblogs.com3522.67%
5ithome.com3452.61%
6developer.volcengine.com2772.10%
7developer.aliyun.com2662.01%
8sohu.com2301.74%
9zhuanlan.zhihu.com2211.67%
10cloud.tencent.com2201.67%
11jiemian.com2141.62%
12page.sm.cn2061.56%
13mp.weixin.qq.com1971.49%
14baike.baidu.com1921.45%
15m.toutiao.com1921.45%
16github.com1691.28%
17cet.com.cn1671.26%
18ima.qq.com1651.25%
19m.blog.csdn.net1521.15%
20chinadevelopment.com.cn1431.08%

The full top 100 is in top_domains in issue-11. Keep two things in mind when reading this table:

  1. It counts domains, not articles. A domain in the table only means the engines cited some pages under it on these questions; it says nothing about who wrote any of them.
  2. A count is only a count. It does not indicate a site's content quality or credibility, and we make no judgement about any site in the table.

4. The three rounds side by side: how the top 10 moved

Per-round counts come from the underlying table of the same snapshot (the public JSON only has the three-round totals; this table is also written into the repository README).

#DomainRound 9Round 10Round 11
1baijiahao.baidu.com154253208
2iesdouyin.com20411684
3blog.csdn.net11814792
4cnblogs.com141111100
5ithome.com12012798
6developer.volcengine.com1089574
7developer.aliyun.com9173102
8sohu.com4869113
9zhuanlan.zhihu.com6134126
10cloud.tencent.com816970
  • Across the three per-round top-10 lists, only 6 domains are in all three: baijiahao.baidu.com, iesdouyin.com, blog.csdn.net, cnblogs.com, ithome.com and developer.aliyun.com.
  • 99 of the top 100 appeared in all three rounds; the exception is quanmin.baidu.com, whose 36 citations all came in round 9.
  • Across all 2,225 domains: the 356 that appeared in all three rounds (16.0% of domains) contributed 77.5% of the citations; 1,461 (65.7% of domains) appeared in only one round.

How to read it: the head is relatively stable, while the long tail turns over heavily from round to round. A single round's reading tells you what happened in that round; it is not a baseline. That is why we moved this list to a monthly merge of three rounds. Also, the ChatGPT and Kimi gaps in round 11 pull that round's numbers down overall, so a domain rising or falling in round 11 cannot simply be read as it "heating up" or "cooling down".


5. By question layer: different questions, different heads

The five layers of the question set correspond to five ways of asking. The top 5 of each layer (the full top 15 for each layer is in top_domains_by_prompt_layer in issue-11):

Layer (questions)Top 5 domains (citations)
Core category (21)baijiahao.baidu.com 290 · iesdouyin.com 192 · blog.csdn.net 191 · ithome.com 181 · developer.volcengine.com 156
Long-tail Q&A (7)baijiahao.baidu.com 120 · cnblogs.com 55 · developer.aliyun.com 54 · sohu.com 53 · blog.csdn.net 52 (tied with No. 6, cloud.tencent.com)
Technical & integration (7)developer.aliyun.com 105 · baijiahao.baidu.com 101 · blog.csdn.net 60 · ima.qq.com 52 · cloud.tencent.com 48
Regional (4)cnblogs.com 78 · baike.baidu.com 58 · iesdouyin.com 58 · baijiahao.baidu.com 53 · jiemian.com 44
Competitor comparison (6)iesdouyin.com 128 · ithome.com 95 · geo.newrank.cn 74 · a.newrank.cn 67 · mp.weixin.qq.com 52
  • Only 3 domains are in the top 15 of all five layers: baijiahao.baidu.com, cnblogs.com and developer.volcengine.com.
  • Number of domains in each layer's top 15 that appear in no other layer's top 15: Competitor comparison 5, Regional 4, Technical & integration 2, Core category 1 (github.com), Long-tail Q&A 0. The two layers that ask about a specific product or a specific region have more domains of their own than the other three.
  • The layers have different numbers of questions (21 versus 4), so compare layers by structure only, not by absolute counts.
  • In the Long-tail Q&A, Competitor comparison and Regional layers, No. 15 and No. 16 are tied and the list is cut alphabetically by domain; this is also stated in the notes field of issue-11.

6. Against August's issue-5: compare shares and positions only

issue-5 is single-round data from round 5 in August (4,636 citations). The questions are identical in both issues, but one is a single round and the other a merge of three, and engine coverage differs, so absolute numbers cannot be compared—only shares and positions.

DomainSept. positionSept. shareAug. positionAug. share
baijiahao.baidu.com14.66%43.28%
iesdouyin.com23.06%33.58%
blog.csdn.net32.70%24.38%
cnblogs.com42.67%15.57%
ithome.com52.61%53.04%
developer.volcengine.com62.10%81.92%
developer.aliyun.com72.01%210.93%
sohu.com81.74%230.86%
zhuanlan.zhihu.com91.67%141.32%
cloud.tencent.com101.67%131.36%
  • 6 domains are in both issues' top 10, and 68 in both issues' top 100.
  • The combined share of the top 10 fell from 30.4% to 24.9%. This cannot be read as "the head getting weaker": merging three rounds accumulates long-tail domains that appear in only one round, which dilutes the head's share by construction; and the two issues are missing different engines (in August, Gemini ran only 28/45; in September, ChatGPT and Kimi did not run in full).
  • Changes in position are likewise just the difference between two observations, not a trend. It takes several more issues on the same basis before a direction can be discussed.

7. How to reproduce

The script below reads only the issue-11 file; it checks the total and computes the figures in Sections 2, 3 and 5 of this article:

import json

d = json.load(open("data/2026-09-18-issue-11.json", encoding="utf-8"))
total = d["_meta"]["citations_total"]                     # 13203
assert sum(d["citations_by_engine"].values()) == total    # (1) per-engine citations add up to the total

for n in (1, 10, 20, 100):                                # (2) sum and share of the overall top N
    s = sum(x["citations"] for x in d["top_domains"][:n])
    print(f"top {n}: {s} citations, {s / total:.2%}")

rest_domains = d["_meta"]["unique_domains"] - 100         # (3) domains and citations beyond the top 100
rest_cites = total - sum(x["citations"] for x in d["top_domains"])
print(rest_domains, rest_cites, round(rest_cites / rest_domains, 2))

pos = {x["domain"]: (i + 1, x["citations"]) for i, x in enumerate(d["top_domains"])}
print(pos.get("zhuanlan.zhihu.com"))                      # (4) position and count of a given domain

print(d["top_domains_by_prompt_layer"]["tech"][:5])       # (5) head of a given question layer

To compare with issue-5, apply the same calculations to data/2026-08-05-issue-5.json.

How the data was pulled: citations are taken from the citation list stored with each valid answer; the filter is "this question set, not deleted, status ok, round ID among the three rounds above"; citations are grouped and counted by the domain as recorded, then removed according to the rules in Section 8. Database rows can be rewritten in place by re-runs, so if a recomputation does not match, check the snapshot time first, then the scope.


8. Desensitization: what we did before publishing

Before publishing the dataset, we desensitized it under four fixed rules. The summary reads (translated from the Chinese original):

Before this issue (issue-11, merged monthly for 2026-09) was published, it was desensitized under four fixed rules: ① remove brand domains and industry domains that might point to third parties (merchants we serve); ② verify that every published figure comes only from the monitoring records of this question set, with no data from any other merchant; ③ publish only domain-level aggregates (the overall list, per-engine totals, the per-layer lists for the question set), not slices at the level of individual questions or keywords; ④ keep an itemized self-check record. In this issue, 25 domains and 42 citations were removed, all of them outside the top 100. All published totals are post-removal figures.

We do not publish which domains were removed: listing them would itself point to third parties. What we can publish is the impact: before and after removal, the top 100 and the top 15 of each of the five question layers are unchanged, entry for entry.


9. What this list is—and isn't—useful for in GEO

What you can use it for:

  1. Seeing where engines tend to draw material from on questions in this category. When planning where to publish, it is a reference surface—but only a reference surface.
  2. Looking by question layer, not just at the overall list. When users ask "how do I integrate this" versus "is there one in my city", the sources engines cite differ a lot.
  3. Looking at persistence. Domains that appear in all three rounds deserve a place on a watch list more than domains that show up in just one round.

What you can't use it for:

  1. This article makes no promise about indexing, rankings or any other outcome. The counts in the tables are observations from three past rounds and do not predict whether any site or any piece of content will be cited in future.
  2. Extrapolating to other categories. The question set covers only the GEO category; in another category the conclusions will most likely differ.
  3. Treating a domain as your own asset. A platform being in the table does not mean your articles on that platform were cited. To tell whether your own content was cited, you must match at the exact URL level; scripts/domain_vs_article.py in the repository was written for exactly this.
  4. Judging quality by count. A count only says something appeared often, not that its content is better.

10. Next issue

Issue 2 starts from round 12. The rule is fixed: each issue takes all weekly monitoring rounds that come after the previous issue and have closed by the time the data is pulled, listing each round; any round that has not closed at that point is left for the next issue. From this issue on, the repository is updated monthly.


Sources and scope: All figures come from rounds 9, 10 and 11 of YinJen's weekly self-monitoring (Beijing time 2026-09-01 ~ 09-18), pulled at 2026-09-29 16:08 Beijing time; 45 Chinese-language GEO questions; the product covers 12 AI engines, 11 were actually run in this issue, 9 of them within the reliable range; measurements on a single merchant's question set, not an industry average. For the dataset and the question set, see GitHub: ZhiMaHang/chinese-ai-engine-sources (issue-11).

This article was written by Suzhou ZhiMaHang Technology Co., Ltd. (the YinJen team). The raw datasets behind all data in this article are public under the GitHub organization ZhiMaHang (verified domain ownership: zhimahang.com). Website: zhimahang.com

Data in this article comes from production-environment measurements by a GEO monitoring tool built in-house by the author's company. The text was drafted with AI assistance and reviewed by the author before publication.


Further reading

Y
About the author
YinJen GEO Team

YinJen's Generative Engine Optimization (GEO) research & field team — we track how content gets cited and surfaced across ChatGPT, Claude, Gemini, Perplexity, Doubao, DeepSeek and other major AI engines. This series is first-hand field notes.

See how YinJen does GEO →