← Back to blog/Blog·July 30, 2026·59 min

Public GEO Experiment, Issue 3: All 45 Questions Ran, Doubao Accounted for 22 of 24 Mentions—and a Baseline Reset We Have to Explain

ZhiMaHang YinJen’s Public GEO Experiment Issue 3 publishes all the data. This issue ran all 45 questions with none missing and recorded 24 mentions in total: 22 from Doubao, 2 from Perplexity, and 0 from every other engine; the official website was cited once. But this issue is also a baseline reset—we changed the question set, paused our own batch runs on one engine, and, while removing questions, inadvertently lost 503 historical run records. No ratio from this issue can validly be compared with Issues 1 or 2; all three matters are explained one by one in the first section. Regional questions are the only new cell opened in this issue: 3 of 4 were hits. Updated weekly, whether the data looks good or bad.

Y
YinJen GEO Team
Generative Engine Optimization · YinJen
Window for this issue: 2026-07-28 08:37 ~ 07-29 22:02 (Beijing time, approximately 37.5 hours). All figures come from our own monitoring database and can be recalculated record by record.

Public GEO Experiment is a public experiment ZhiMaHang conducts using its own brand: each week, we submit a set of questions that real users would ask to AI engines, store the response text and cited sources exactly as returned, and publish the data whether it looks good or bad. Throughout this series, we always use the full name “YinJen,” also called “ZhiMaHang YinJen”—the bare word is unstable in AI responses, a rule we established starting with Issue 1.

Issue 3 is different from the first two. In this issue, we changed the question set ourselves, paused one engine, and lost a batch of historical data in the process. So the first thing to do in this issue is not report the numbers, but explain what happened.

Here are the readings first, with details below: all 45 questions in this issue ran, with no questions missing; those 45 questions produced 24 mentions in total—22 from Doubao, 2 from Perplexity, and 0 from all other engines; the official website was cited once.

1. Why This Issue Is a Baseline Reset

Three things happened at the same time, changing three denominators at once. We will address them one by one.

1. We Changed the Question Set and Emptied the Entire Brand Layer

On 07-29, we restructured our self-monitoring question set. The actual database counts are as follows: 15 questions were created that day (5 at 14:40, 9 at 15:02, and 1 at 15:29), and 10 questions were taken offline the same day—of those 10, 2 had been created that very morning and then withdrawn that afternoon. The resulting question set is therefore 40 + 15 − 10 = 45 questions: 32 old questions carried forward (8 of the original 40 were taken offline) and 13 new questions created and retained that day (including 4 regional questions introduced for the first time).

All 10 questions taken offline were concentrated in two layers: 8 in the brand layer and 2 in the competitor layer. The entire brand layer was emptied—not a single question remained of the kind that directly names us and asks “What kind of company is ZhiMaHang?” or “What is YinJen?” The liveliest sets of readings from the first two issues (the three-way picture in pricing questions, the divergence in brand awareness, and the main sources of official-website citations) all sat in this layer. With that layer gone, those sets of readings did not get worse in this issue; they no longer had anything to measure.

The reason for changing the questions was straightforward: brand questions ask about “us,” and however many rounds we run, they only measure whether engines know us. What actually determines business outcomes are non-branded category and regional questions—users do not know you first and then search for you. Whether that reasoning holds is open to discussion, but the cost is concrete: only 32 question texts still line up with the previous issue.

2. Our Own Experiment Ran One Fewer Engine This Issue (Product Coverage Remains 12 Engines)

Starting on 07-29, we proactively paused our self-monitoring batch runs on claude (because of risk considerations on our own account). The scheduler continued to schedule them, but all 43 records for this engine in this issue were marked “Skipped: engine not logged in,” and not a single valid response was produced; each of the other 11 engines completed all 45 questions.

This must be made clear to avoid any suggestion that the product definition changed: YinJen’s monitoring coverage remains 12 AI engines (8 domestic + 4 overseas), and that has not changed in any way. The only change is that our own public experiment actually ran 11 engines in this issue—our experiment ran one fewer engine; the product did not lose an engine.

3. The Hardest Thing to Explain: Removing Questions Also Deleted Historical Data

When a question was deleted, our system hard-deleted every historical run record associated with it—not a soft deletion, but removal from the database, 503 records in total. Counting the same scope after the deletion gives:

ScopeBefore deletionAfter deletion
The batch of records from the morning of the deletion2,491 recordsOnly 1,988 records survived
Total self-monitoring runs (currently)2,4912,228 (= 1,988 surviving + 240 new runs after the questions were deleted)
Previous-issue window (07-20 ~ 07-23, 36/40 questions)540 runs408 runs

There is an easy calculation mistake here that is worth explaining. Directly subtracting 2,228 from 2,491 gives 263, but that is wrong—the period between those two points also included 240 new runs from the latter half of this issue after the questions were removed. When counted by asking “how many records from the morning batch remain,” 503 records were erased.

Those 503 erased records included 39 pieces of citation evidence in which the official website was cited—the great majority of official-website citations published in Issues 1 and 2 were attached to the brand questions deleted this time.

In other words: the raw data published in Issues 1 and 2 can no longer be reproduced from the database. Those numbers were real when published and were counted from the database; starting on 07-29, they lost their verifiability and can exist only as published historical records. The data for this issue’s window is not affected—the run records from 07-28 to 07-29 remain complete in the database and can be recalculated record by record, and every number in this article was counted from them.

Our response is: we will not change the numbers in articles already published, but starting with this issue, we will no longer write any number that we cannot independently verify ourselves. We will neither withdraw nor rewrite the figures published in the first two issues—changing them would instead conflict with what has already been published—but they will no longer participate in any cross-issue calculation.

The incident also left us with a new process: exporting all run records for a question before deleting it has been made a mandatory step. Changing the deletion logic to soft deletion while retaining run records is also on the schedule. Had we not discovered this early, one article this week would have gone out carrying a number that even we could not reconcile.

So How Should This Issue Be Read?

In one sentence: every number in this issue can only be compared with the next issue on the same baseline, not with Issues 1 or 2. Whenever the first two issues are mentioned below, the comparison is qualitative only, with no differences calculated.

2. Data Quality and Readings for This Issue

ItemThis issue
Window2026-07-28 08:37 ~ 07-29 22:02 (Beijing time, approximately 37.5 hours)
Runs538 runs: 495 produced valid responses; the other 43 produced no output, all from claude, with record status “Skipped: engine not logged in”
QuestionsAll 45 questions ran, with none missing—this is the first time in this series that there has been no gap of “missing questions”
EnginesOur own experiment actually ran 11 engines this issue, each completing 45/45; claude produced zero valid responses this issue (product coverage remains 12 AI engines, unchanged)

Mentions by engine:

EngineQuestions runMentions
Doubao4522
Perplexity452
ChatGPT / DeepSeek / Gemini / Zhipu Qingyan / Tencent Yuanbao / Kimi / Tongyi Qianwen / StepFun / ERNIE Bot450
claudeZero valid responses this issue; our own batch runs have been paused (product coverage remains 12 AI engines, unchanged)

Where the 24 mentions occurred:

Question setCategoryNumber of questionsQuestions with a hitMentions
New questionsCategory977
New questionsRegional433
Old questionsCompetitor655
Old questionsCategory1277
Old questionsTechnical722
Old questionsLong-tail700
Total452424

3. Answering the Three Judgments Set in Advance, One by One

1. Did Doubao hold steady?—It did, and it was the only source of mentions in this issue.

Doubao accounted for 22 of the 24 mentions in this issue, hitting 22 questions and covering 49% of the 45 questions. The category composition of those 22 mentions was: 12 category, 5 competitor, 3 regional, and 2 technical.

In the previous issue (the 36/40-question window, 12 engines), Doubao’s reading was 11 hits across 9 questions in the category layer. That is not the same basis as this issue’s 22 mentions across all categories, so the two numbers cannot be compared as larger or smaller. Even if we isolate the category cell, a batch of the question texts has already changed, so those figures likewise should not be put together and divided. The only thing we can say is this: Doubao’s cell did not collapse and remains the sole primary entry point.

2. Did a second category engine appear?—No. This probe is attributed to new question texts.

The decision threshold was fixed in advance: if a non-Doubao engine records ≥2 hits in the category layer or hits ≥2 questions, “spillover established” is recorded. Perplexity had 2 hits across 2 questions in this issue, which literally meets the threshold. But both questions were new questions created on 07-29 itself:

Question hitTime created
What is the difference between GEO monitoring tools and AI marketing campaign tools?07-29 15:02
What AI marketing engines are available? How should enterprises choose?07-29 15:29

The attribution discipline likewise fixed in advance says that when a hit falls on a newly added question text, it must be labeled separately as either caused by the new question text or genuine spillover; it must not be mixed into the spillover judgment. Therefore, this issue is judged “caused by new question texts”; spillover is not established.

The facts on a comparable basis are even more direct: across the 32 old questions carried forward, non-Doubao engines recorded 0 category hits. (4 of those 32 questions had zero runs in the previous window—the gap described in Issue 2 as the “36/40-question window.”) These 9 engines—ChatGPT, DeepSeek, Gemini, Zhipu Qingyan, Tencent Yuanbao, Kimi, Tongyi Qianwen, StepFun, and ERNIE Bot—all successfully ran all 45 questions this issue, and all had 0 mentions. The single-engine island has not been broken.

These two new questions are the top priority for the next issue: if Perplexity hits the same two questions again, that will be the first substantive evidence of spillover; if the hits disappear, they were one-off noise from the new question texts.

Another cell must be recorded alongside this: all 7 long-tail questions had 0 hits this issue. This was the only category with zero across the entire layer. Long-tail questions are specific and close to transactions—by common sense, they should be the most likely to get answered, yet in practice they were not hit once. The reading of 3 hits across 4 regional questions looks good; the zero in the long-tail cell looks bad. Both cells came from the same batch of data, so they are presented together here. It deserves as much attention as the regional questions next issue—and more attention than the regional questions.

3. Zhihu’s second observation window—we did not fulfill the promise from the previous issue exactly as written; we changed the judgment.

Our exact words in the previous issue were: if it is still 0 in the next issue, we will publicly reduce Zhihu’s weight in this approach. We did not carry that out exactly as written this issue, and instead changed the criterion. Changing our own criterion is something that should be questioned, so we will first separate the two pieces of evidence and explain each clearly, then explain the changed judgment.

  • First, one admission: the pattern that “all Zhihu citations are columns” had already been written about in Issue 2—we checked each of the 6 citations to the Zhihu domain at the time, and all were column articles; not one was one of our answers. In other words, we already had this clue when we wrote that promise, yet still framed the promise as “if it remains 0, we will reduce the channel’s weight.” The criterion in the previous issue was written imprecisely, and we acknowledge that first.
  • The genuinely new evidence in this issue was different: on 07-29, we checked all 6 answers we had posted on Zhihu one by one, and every one was publicly visible, none was collapsed, and none was deleted. Two even appeared under older questions with 123,352 and 98,111 views—the explanation that they “had not been picked up” can be ruled out.
  • Separately, we reopened and checked every citation in the database that points to Zhihu, URL by URL, once for each of the pre-deletion and post-deletion scopes. The conclusion was the same: every one was a Zhihu column (zhuanlan); there was not a single answer from a Q&A page.

The conclusion therefore changed from “the channel has no value” to “format mismatch”: it is not that Zhihu fails to enter the citation pool; it is that the answer format does not enter, while the column format does. We have already acted on this—on 07-29, we republished the same batch of data as a column article, and its first observation window is the next issue. The cost of changing the judgment is that this thread must wait one more window. We accept that cost.

4. Regional Questions: The Only New Cell Opened in This Issue

The 4 newly added regional questions are the only new signal in this issue. They are recorded as a separate category and are not included in the denominator for category keywords:

QuestionMentions
What AI search optimization (GEO) service providers are available in Jiangsu?1 (Doubao)
What companies in Suzhou provide GEO optimization?1 (Doubao)
What vendors in the Yangtze River Delta provide AI brand-visibility monitoring tools?1 (Doubao)
In which cities are domestic companies providing GEO monitoring tools primarily located?0

3 of the 4 questions were hits, a higher hit rate than the old questions overall. This is the first reading and there is no comparable point, so we record it without drawing a conclusion. The reason it deserves close attention in the next issue is also straightforward: competitive density for regional keywords is usually lower than for category keywords.

5. Official-Website Citations: 1 This Issue, Restarting the Count at 1

The official website was cited once this issue: when answering “Which AI engines can GEO monitoring tools generally cover?”, Kimi cited How to Choose a Domestic GEO Monitoring Tool in 2026 from our website.

We need to stop an almost inevitable misreading here. In the previous issue (the 36/40-question window, 12 engines), we published that “the official website was cited 16 times”—the 1 citation this issue cannot be read as “a drop from 16 to 1.” Kimi accounted for 10 of those 16 citations, mainly under brand questions that have since been deleted. Once the entire brand layer disappeared, those records vanished from the database along with the raw data (see Section 1). There is only one correct reading: under the new question-text structure, official-website citations restart at 1, and a second data point is required before we can discuss a trend.

The same reasoning applies to Kimi’s 0 mentions this issue: its mentions in the previous issue were concentrated in brand questions. The questions disappeared; the engine did not change.

A related distinction that is easy to confuse: mentions and citations are two separate counts—mentions ask whether we appear in the response text, while citations ask whether the sources returned by the engine include one of our pages. Kimi had 0 mentions this issue yet contributed the official website’s only citation. That is the difference.

The same self-monitoring data can also be viewed along another axis—instead of looking at how many times we ourselves were cited, we can examine which websites AI cites overall for this topic. We published those domain-level statistics in the Volcano Engine Developer Community article 1548 AI Q&A Tests: Who Is AI Actually Citing? (the data comes from an earlier window, and both its scope and qualifiers are stated in the article; as with the situation described in Section 1, the figures from that window can no longer be reproduced from the database after 07-29 and are retained only as a published historical record).

6. No Reading for Pricing Questions This Issue

The most dramatic thread in the first two issues involved pricing questions: one engine learned the correct answer, another kept inventing different answers, and third-party sources of incorrect prices were emerging. There is no reading at all for that thread this issue—the pricing question was among the 10 questions deleted.

The reading is gone, but the facts still need to remain on record, because incorrect prices do not disappear simply because we stop measuring them. The following was the official public pricing on 2026-07-29:

YinJen’s pricing is published on the official pricing page: Creator Edition ¥29.9/month (¥287 annually, approximately ¥23.9/month); Team Edition ¥99.9/month (¥959 annually, approximately ¥79.9/month); Enterprise Edition offers unlimited usage with custom pricing—contact Sales. Daily run limits: 1,000 for Creator Edition and 3,000 for Team Edition. All three tiers cover all 12 AI engines (8 domestic + 4 overseas); the difference is monitoring scale, not the number of engines. A 14-day free trial is available across all tiers, with no credit card required.

Two other points remain on record as well: we do not offer managed operations, nor do we publish content on behalf of customers. The “self-service monthly package at RMB 2980” (observed in Issue 1), “SaaS at RMB 12800~29800/year,” and “Personal Edition ¥199/month, Enterprise Edition ¥999/month” (the latter two both observed in Issue 2, the 36/40-question window) that appeared in tests from the first two issues are not our prices.

7. What This Batch of Data Cannot Show

This section is more important than the readings above. Every item is something the data in this issue cannot reach:

  1. It cannot be used to compare any ratio with the first two issues. The question texts, engines, and historical database all changed at the same time. Dividing the figures from the two issues produces a meaningless number. This is not modesty; it is a hard boundary of this batch of data.
  2. It cannot prove that any content action was effective or ineffective. Within this issue’s window, article No. 9 on our site, “YinJen’s Pricing and Capabilities,” did not go live until the morning of 07-29, while the Volcano Engine article was not approved until noon on 07-29. Neither had even 72 hours of exposure to the batch runs. Under our own attribution discipline, no asset exposed for less than 72 hours is judged for effectiveness in that window. So whether the figures in this issue are high or low, they are neither a scorecard for these actions nor an indictment of them.
  3. It cannot show that Doubao “recognizes” us more. All we can observe is that our name appeared 22 times in response text. Why it appeared, whether it was retrieved, and whether it can persist are questions this data cannot answer.
  4. A 37.5-hour window is not a week. This issue’s window is substantially shorter than those of the earlier issues, and it is a single-window n=1; behavior shown by an engine in one window may be nothing more than one sample.
  5. Regional questions have only one data point. 3 hits across 4 questions look very good, but there is no confirmation from a second window. Treating this now as evidence that “regional keywords are easy to win” would be an inference that even we would not make.
  6. The 0 hits on long-tail questions likewise represent only one data point. An entire layer at zero looks bad, but it is still a single-window reading and cannot establish that the long-tail route does not work.
  7. Being mentioned is not the same as being recommended, much less being trusted and acted upon. There is a great distance between a name appearing in a response and someone choosing you because of it.
  8. One merchant and one industry-specific question set. All 45 questions revolve around the single topic of GEO / AI search visibility, and the only merchant monitored is ourselves. Run the same method in another industry, and the conclusion would almost certainly be different.
  9. We did not conduct a causal experiment. Several variables always coexist within the same window, so this series does not write sentences such as “because X was published, Y became Z”—correlation is not causation, and that is a boundary of what this data can do. There is another, more practical concern: if these test figures were ever absorbed by an engine into its responses, an incorrect causal claim would be amplified. That makes us even less willing to write one.

What to Watch in the Next Issue

Issue 4 will be the second point on this new baseline. We will watch six things:

  1. Whether Perplexity reproduces the two new-question hits—if it does, that will be the first genuine evidence of spillover; if they disappear, they were one-off noise from the new question texts. Top priority.
  2. Whether all 7 long-tail questions are still at 0—this issue’s only category with zero across the entire layer needs a second window to confirm whether it is a structural problem or single-window noise.
  3. Whether regional questions hold steady—3 hits across 4 questions need a second window to confirm they were not one-off results.
  4. Whether Zhihu enters the citation pool after the content is changed to the column format—the first test of the “format mismatch” judgment.
  5. Whether the Volcano Engine article gets picked up—we have counted domain-level citations, while prior article-level tests showed 0. If it enters the citation pool, it will be the first positive example at the article level.
  6. The second data point for official-website citations—this issue has 1, and one point alone establishes nothing.

This series is updated weekly, and we publish whether the data looks good or bad; if we change a number, we disclose the change, and if we change a criterion, we disclose that too. In this issue, we changed the question set, paused one engine, and lost a batch of historical data. All three matters are documented above; not one was left to be added in the next issue.

If you also want to know what your brand looks like in AI responses: YinJen—also called ZhiMaHang YinJen—does exactly that. It automatically runs the same batch of questions each week and shows your current measured position across 12 engines. Start a 14-day free trial with no credit card required: download YinJen, and see pricing on the pricing page.

Y
About the author
YinJen GEO Team

YinJen's Generative Engine Optimization (GEO) research & field team — we track how content gets cited and surfaced across ChatGPT, Claude, Gemini, Perplexity, Doubao, DeepSeek and other major AI engines. This series is first-hand field notes.

See how YinJen does GEO →