On-site canonical version|slug: public-geo-experiment-issue-4 Time zones: Monitoring data is always in Beijing time (stored in UTC and converted); publication and internal-action dates are always in PDT. Every date is labeled with its time zone. Data window: Completed monitoring rounds R4~R7. The final round, R7, was cy_5104f371776e, Beijing time 2026-08-18 08:38 ~ 08-20 17:18. Not a single number from any round still in progress is included—not even as a “trend observation.” All figures come from our own monitoring database; see Section 2 for the methodology. Numbering: “Issue N” refers only to the issue number of an article in this series (this is Issue 4). Monitoring rounds are always written as R4~R7. The two numbering systems are not interchangeable.
Abstract: All four rounds from R4~R7 had the same denominator of 495 valid samples (= 11 engines that produced valid samples in the round × 45 questions), and the question sets were verified to be identical. Mentions of our own brand, “YinJen,” fell 24 → 18 → 4 → 0. In the same batch of responses, competitors kept being named, and their own official sites kept appearing in the citation lists: zhituishidai.com recorded 1 / 7 / 3 / 8 across the four rounds, aidso.com recorded 3 / 17 / 26 / 11, while our zhimahang.com recorded 1 / 2 / 2 / 0. We are the ones that fell to zero. ⚠️ These figures show only whose names were written into the answers and which domains were written into the citation lists for these 45 questions and 11 valid engines. They do not indicate market share or relative quality. This issue also fulfills a promise left outstanding from Issue 2 and publicly corrects two errors we discovered ourselves.
“Public GEO Experiment” is ZhiMaHang’s open experiment using its own brand: we send questions that real users would ask to AI engines in rounds, store the answer text and citation sources exactly as returned, and publish the results whether they look good or bad. First, two admissions: first, we stopped updating—the previous issue was published on 2026-07-29 (PDT), and 27 days had elapsed by the publication date of 2026-08-25 (PDT), even though the end of that article said “one update every week”; second, the promise made in Issue 2 still had not been fulfilled publicly. Section 5 explains it.
1. Four-Round Side-by-Side Table
Item (the denominator for every round is 495 status='ok' responses, with rounds separated by cycle_id) | R4 | R5 | R6 | R7 |
|---|
cycle_id | cy_b4e0… | cy_f0ac… | cy_2c2f… | cy_5104… |
| Round capacity (rows dispatched) | 540 | 540 | 540 | 540 |
Valid samples (status='ok') | 495 | 495 | 495 | 495 |
| Body-text mentions · YinJen (union of four aliases) | 24 | 18 | 4 | 0 |
Body-text mentions · AIDSO (Latin string, LIKE BINARY) | 76 | 84 | 87 | 63 |
| Body-text mentions · AIDSO (Chinese “爱搜”) | 76 | 85 | 85 | 71 |
| Body-text mentions · Zhituishidai (clean union methodology) | 54 | 46 | 39 | 34 |
zhimahang.com appears in citation list (ours) | 1 | 2 | 2 | 0 |
zhituishidai.com appears in citation list | 1 | 7 | 3 | 8 |
aidso.com appears in citation list | 3 | 17 | 26 | 11 |
All four rounds have the same structure: each round dispatched 12 engines × 45 questions = 540 rows. All 45 rows for claude were not run in every round (see the next section), so each round had 495 = 11 valid engines × 45 questions.
We verified the prerequisite for horizontal comparison: the prompt_id intersections between R4, R5, and R6 and R7 were each 45/45. Equal denominators only mean equal numbers of valid rows; identical question sets are what make the four columns comparable. But the four “body-text mention” rows cannot be compared with one another: our row uses the union of four aliases (引见 / 智码航 / yinjen / zhimahang, including Chinese); the two AIDSO rows separately use the Latin string alone and the Chinese name alone; and Zhituishidai uses a three-string union. The matching rules differ, so each row can only be compared vertically against itself. The bottom three “citation list” rows use the same yardstick (hostname matching) and can be compared with one another.
In R7, AIDSO appeared in 24/45 questions and across 11/11 valid engines (methodology = Latin-string LIKE BINARY '%AIDSO%', from the same source as the 63 in the table; question coverage was not retrieved for the Chinese-name methodology), while we recorded 0 for the same questions and the same engines. The comforting excuse that “AI does not recommend specific products in this category” does not hold up.
2. Methodology: How We Counted, Two Corrections, and Three Blind Spots
1. Separate rounds by `cycle_id`, not by date. The four time windows overlap (the final R6 record landed at Beijing time 08-20 15:45, later than the start of R7). Separating by date would mix different rounds into the same cell.
2. There were 495 valid samples, not 540. 540 is the round capacity. All 45 claude rows in all four rounds had status='skipped' and message='engine not logged in', were stopped by the login precheck, and had empty body text—the engine was never asked at all, rather than having been asked and failing to answer—so they are not included in the denominator. Attribution: the collection gap was caused by an expired login session on our account and was unrelated to the engine’s service itself. It skipped every row for four consecutive rounds beginning on Beijing time 2026-07-28, and none of this article’s conclusions cover it.
Keep the product methodology and experiment methodology separate: YinJen monitors 12 AI engines, and all three pricing tiers cover all 12; they differ only in monitoring scale. This public experiment, however, had only 11 engines produce valid samples across four consecutive rounds, so 495 = 11 × 45. Every reference to “11 engines” in this article means the latter.
3. How mentions were counted, plus one boundary. We first retrieved the mention flag, then bypassed it and scanned the full text of all 495 responses for four aliases. The two methods aligned exactly in all four rounds (24 / 18 / 4 / 0); brand_aliases had an updated_at of Beijing time 07-12, before R4. For our own brand, we reported only the union methodology rather than separate Latin-only and Chinese-only methodologies. That is a methodological asymmetry, and we will add both next issue. Boundary: “mentioned in the body text” is taken from `response_text`, while “entered the citation list” is taken from `citations`. They are counted and reported separately throughout.
4. Correction one: one contaminated methodology. The Zhituishidai cell originally counted response_text LIKE '%GENO%', but the database’s default collation is case-insensitive, so three objects were mixed into the same cell: GENO, GenOptima, and Dageno / dageno.ai.
`Dageno / dageno.ai` and Zhituishidai / GenOptima are separate entities. Verification method: zhituishidai.com was checked through both DoH resolution and its page title. The title was “Zhituishidai GenOptima — Brand Consensus Operating System for the AI Search Era,” checked on 2026-08-25 (PDT)—this is also the basis for treating GenOptima and “Zhituishidai” as names for the same entity. The entity information for dageno.ai did not match. Corroborating evidence: in R7, among the 34 responses matched by the clean methodology, the number citing dageno.* was 0; among the same round’s 495 responses, 7 cited it, and the contexts did not overlap at all. We make no other attribute judgment beyond “separate entities” (the draft modifier “overseas GEO tool vendor” had no verifiable basis and was removed).
Recounted separately (R4~R7, each with a denominator of 495): GENO 43 / 32 / 28 / 23, GenOptima 15 / 25 / 24 / 15, and pure false matches 6 / 4 / 1 / 6. Clean methodology = GENO ∪ GenOptima ∪ Chinese “智推时代,” with LIKE BINARY used for case sensitivity = 54 / 46 / 39 / 34. The contaminated methodology yielded 53 / 45 / 35 / 32. The trend is the same but the figures differ, so we are publishing the clean figures. An internal draft used the contaminated methodology and treated dageno.ai as Zhituishidai’s own domain in its analysis. That was wrong, and the draft was discarded before publication. We therefore set this rule: all secondary statistics previously produced with the contaminated methodology must be rerun. The table in Section 4 is the rerun version.
5. The two AIDSO rows: neither spelling contains the other. First, the basis: aidso.com was checked through both DoH resolution and its page title. The title was “AIDSO爱搜—GEO Search Ranking Optimization Platform,” and og:site_name was “AIDSO爱搜,” checked on 2026-08-25 (PDT). On that basis, the two strings were treated as names for the same entity, but they are still reported in separate rows—R6 had 87 Latin-string matches > 85 Chinese-name matches, while R7 reversed to 63 < 71, so neither set contains the other. R7’s measured set difference (the same denominator of 495 ok responses, separated by cycle_id): 62 matched both, 9 matched only the Chinese name, 1 matched only the Latin string, and the union was 72. ⚠️ 71−63=8 is the net difference; reading it as “8 responses used only the Chinese name” is wrong. For Zhituishidai in the same round: intersection 23, Chinese-only 8, Latin-only 3, and union 34, confirming the cell above.
6. Correction two: the old conclusion that “competitors’ own domains were never cited” is withdrawn. The error began with checking the wrong domain: we checked a parked domain that no longer corresponded to that vendor, got 0, and treated it as the conclusion. Rechecking with the verified official domains reversed the conclusion—zhituishidai.com appeared in the citation lists of 1 / 7 / 3 / 8 responses across the four rounds, and aidso.com appeared in 3 / 17 / 26 / 11. Of those 8 R7 responses, 3 also matched the clean methodology in their body text (kimi 1, hunyuan 2), while the other 5 cited the domain without naming the brand in the body text (perplexity 3, hunyuan 2). So the real comparison in this issue is not “nobody relies on their own site,” but rather: competitors’ own sites are being cited, while our cell fell to zero. This and Item 4 are two sides of the same type of error: check the wrong object once, get 0, and draw a conclusion.
7. Three blind spots.
- In R7, 124/540 rows were rerun and overwritten in place (their statuses and body text were rewritten in situ), erasing the original failure records, so the round’s true collection failure rate cannot be measured.
- Of the 495
ok rows in R7, 7 were not actually answers: 2 were engine-side error messages recorded as ok, and 5 were refusals. - For R7, all 45 gemini responses had empty `citations` arrays (a citation-collection fault; the body text was normal). Together with claude not running, R7’s citation statistics (the R7 column of the final three rows in the Section 1 table and the source table in Section 4) are based on citation data from only 10 engines; R4~R6 citation statistics had citation output from 11 engines (claude did not run). ⚠️ These three numbers are not the same thing: 12 is the product’s monitoring-coverage methodology, 11 is the number of engines producing valid samples in this experiment, and 10 is the number of engines that actually produced citation output in R7.
3. How We Fell to Zero
First, rule out three “false zeros.” It was not a failure to produce answers: none of R7’s 495 responses was empty, and the average length was approximately 1796 characters (R6: 1806). The detector was not broken: the four-alias scan that bypassed the flag found 24/18/4 matches in R4/R5/R6, aligning exactly with the flag, and only fell to zero in R7. Those questions were not omitted: we made a review queue from the 19 questions that had produced matches in R5 and R6 (R5 doubao 15, R6 kimi 2, R6 Perplexity 2), and every one was asked again this round, every one had `status='ok'`, and every one no longer mentioned us. These are measured results, not an inference from a sample.
So the 0 is real: the engines were asked, they answered, but their answers did not include us.
Why? We have only one inference, and it must be labeled as such. Of the 13 articles on our site, only 4 can currently be verified as having entered the citation pool:
The methodology is fixed as follows: the objects counted are URL rows in citations whose hostnames match our site’s domain (domain-level matching, not article-level body-text matching). All 7 rows had a status of ok and occurred across 6 responses (1 R5 response cited 2 articles). Separated by cycle_id, the R4~R7 response counts were 1 / 2 / 2 / 0 and the URL-row counts were 1 / 3 / 2 / 0. 1 additional URL row came from an ad hoc run before cycle-based monitoring (its cycle_id was NULL) and is not included in the four rounds. Engine attribution was checked row by row: all 6 rows within the four rounds were from kimi, and the 1 row outside the rounds was from chatgpt.
The other three formats—experiment logs (the first three issues in this series), customer cases, and instrumentation methodology—have never entered the citation pool to date. But a boundary must be added: two of the first three issues still have not been indexed by Bing, and indexing is a necessary but not sufficient condition for being cited. This observation therefore mixes two causes—“format” and “never entered the index at all”—that cannot be disentangled. We are also stating our publishing capacity honestly: during the three weeks from 08-03 ~ 08-24 (PDT, inclusive), we published only 1 article; every other day had 0.
Inference: this looks more like a content-format mismatch than a channel problem. But it is only an inference—the citations are concentrated in one engine, the sample size is in the single digits, and the question set covers only one merchant and one category, so it cannot support a causal conclusion. Moreover, our mentions falling to zero and citations of our own site falling to zero happened at the same time and point in the same direction, but they cannot corroborate each other: without a mention, a citation to us is already unlikely. They share the same source and are not two independent pieces of evidence.
4. Distribution of Cited Sources: Observations Only
For the 34 R7 responses that matched Zhituishidai using the clean methodology, the distribution of citation sources after grouping by hostname was as follows (rerun results, replacing the contaminated-methodology table in the draft):
| Hostname | Number of responses in which it appeared |
|---|
| ithome.com | 14 |
| cnblogs.com | 11 |
| baike.baidu.com | 7 |
| blog.csdn.net / cet.com.cn / finance.sina.com.cn / jiemian.com / tech.ifeng.com | 5 each |
| 163.com / baijiahao.baidu.com / csdn.net / developer.volcengine.com / m.jiemian.com | 4 each |
Table footnote (required reading): n = 34 (the number of R7 responses matching the clean methodology). Of those, 31 had ≥1 citation and 3 had none; after deduplication, there were 232 unique “response × hostname” pairs and 0 parse failures. Grouping rule: within one response, the same hostname is counted only 1 time; counts are summed across responses. Hostname normalization = lowercase + remove protocol + remove path/query/fragment + remove port + remove leading www.; subdomains are not merged. A difference of 1~2 occurrences has no distinguishing power. This table records only “which domains appeared”; it is not used for ranking, much less for placement decisions. 5 domains tie at n=4, and 13 domains are listed in full for completeness.
Observation 1: all 13 are public third-party platforms, but do not read that as “the citation lists contain only third parties.” Both vendors whose official domains were verified had their own sites appear in citation lists (see the final two rows of the Section 1 table). Specifically for the 34-response slice in this table, zhituishidai.com appeared in 3 responses, below the n≥4 threshold and therefore not visible here. Third parties are still the majority (ithome 14 and cnblogs 11 versus the vendor’s own site at 3), so the judgment that “third-party content is worth placing” still stands—but not because “they do not use their own sites.” Their own sites have been cited all along.
Observation 2: we own or previously owned assets on 3 of these platforms (CNBlogs, CSDN, and Volcano Engine Community), while our off-site assets (9 articles across 7 platforms) have 0 cumulative article-level matches (R7 was not retrieved separately). Same platforms, same engines; what was cited was not ours.
What we cannot say must also be explicit: these data show only which domains were written into the citation lists of answers for these 45 questions and 11 valid engines. They cannot tell us “what those companies did.” We do not evaluate their practices, speculate about tactics, or infer motives. This distribution also comes from only 10 engines.
5. Fulfilling the Promise Left Outstanding from Issue 2
The exact line from Issue 2 (published 2026-07-22, PDT) was: “If it is still 0 next issue, we will publicly reduce Zhihu’s weight in this playbook.”
What happened afterward: Issue 3 (07-29, PDT) did not execute it as written. The criterion was changed from “the channel has no value” to “the content format is mismatched,” and the promise was pushed to the next observation window. On 2026-08-11 (PDT), we internally ranked platforms using combined R4~R6 citation data and deprioritized Zhihu Columns on that basis; then we stopped updating. We made the internal change but never said so publicly. That is a genuine breach of a promise, not a “delay” or a “methodology adjustment.” We are rectifying it now.
Basis for the deprioritization. ⚠️ The ranking was based on R4~R6 data as retrieved on 2026-08-11 (PDT), when R6 had not yet finished running—it was not a complete rerun of all three rounds. At that point, the sample contained 821 runs with citations, 7,317 citations, and citation output from 11 engines (claude did not run). Before publication, we rechecked the complete R4~R6 data once (1,108 runs with citations and 10,230 citations). The platform-ranking conclusion did not change, but to represent the decision basis at the time truthfully, we retain the 08-11 snapshot values here and label them as a snapshot. We grouped domains into “platforms,” limited the count to platforms where we can contribute, and used number of engines covered rather than total citation volume as the primary sort key—a source cited by 9 engines is a stronger signal than one cited by only 1. Results: Zhihu Columns covered 2/11 engines (84 citations), and WeChat Official Accounts also covered 2/11 (111); Volcano Engine Community covered 10/11, CNBlogs 9/11, CSDN 9/11, and Tencent Cloud Community 9/11. The Zhihu Column article we published ourselves never appeared in these citation data.
The conclusion goes only as far as the evidence supports. These two channels had the lowest engine coverage (2/11 each), so we lowered their priority in our GEO ledger. No cost-side data was retrieved, so we make no “cost-effectiveness” judgment (the draft said “poor cost-effectiveness,” which judged an unmeasured dimension and was changed). The action is deprioritization, not discontinuation: they may serve roles such as branding and private-traffic conversion, and should be evaluated for those roles rather than mixed into the GEO ledger.
Boundary of this evidence itself: “which platforms engines have cited” is a necessary condition, not a sufficient condition—we have articles on CSDN and Volcano Engine Community, yet their article-level match count is still 0. Choosing the right platform is only the first step; being indexed and being selected are the next two. These data also come from one merchant’s 45 GEO-category questions; the conclusion would most likely differ in another category.
Of the six observation items written down in Issue 3, this issue answers four (the two new Perplexity questions fell to zero in R7; Zhihu Columns did not appear in R4~R6; citations of our own site were 1 / 2 / 2 / 0; the Volcano Engine article was not retrieved for R7 and was 0 in each of the previous three rounds). Data for the long-tail and regional questions were not retrieved; those remain outstanding.
6. Next Steps
We are not promising results; we are stating only what we plan to change.
- Redirect content toward proven formats without increasing output. Invest in comparison, methodology, and definition formats, for one reason only: they are the only formats currently verifiable as having entered the citation pool. This does not create an expectation that “publishing it will get it cited”—the sample is in the single digits and is confounded by the “not indexed” factor.
- Make citations of our own site a permanent standalone observation item. The sharpest cell in this issue is that our
zhimahang.com fell from 1/2/2 to 0 while the other two domains in the same table did not. Observe only; set no target value. - Add vendor-name and domain verification to every round’s routine checks. Both corrections in this issue arose because that step was missing. Both errors followed the same pattern: check the wrong object once, get 0, and draw a conclusion.
- Either restore the claude login path or label it as unmeasured for the long term. Writing vaguely that “12 engines were run once” after four consecutive rounds with zero samples would be dishonest. We must also state clearly that this was our login-session problem, not an engine-service problem.
What these data cannot establish: they cannot prove that any content action was effective or ineffective, nor can they prove that one company’s product is better or has greater market share. Being mentioned does not mean being recommended, much less that someone chose you because of it. Run the same method in another industry and the conclusion will almost certainly differ. We did not run a causal experiment, so this series does not make causal statements. The rules remain unchanged: if the numbers change, disclose the change; if the criteria change, disclose the change; if a promise remains outstanding, admit it publicly. We recorded 0 ourselves in this issue, and we are publishing it as is.
Want to know what your own brand looks like in AI answers? YinJen (also called “ZhiMaHang YinJen”) does exactly that—continuous monitoring across 12 AI engines, automatically running the same batch of questions round by round and storing the answer text and citation sources exactly as returned. The number of engines that actually produce valid samples in each round fluctuates with login status; it was 11 for this issue, and the product identifies that count. Start a 14-day free trial with no credit card required, download, and see the pricing page.