Why this matters
The occasion — shown on a concrete example
Between 16 and 20 August, OpenAI swapped out the language in which ChatGPT formulates its web searches. On 21 August, the endpoint that retrieves conversations changed as well — response format included. measured│Q1
Neither appeared in any changelog. If you do not measure daily, you notice such things only when your own numbers drop — and by then you no longer know why.
Who this page is for.
For anyone who wants to understand how content ends up in ChatGPT's answers. Knowing SEO helps; knowledge of ChatGPT's internals is not assumed. Technical terms are explained where they first appear.
Experimental setup
Every decision with its reason
A measurement series is only as good as its constants. What is fixed here, and why:
| What | How | Why this way |
|---|---|---|
| Account | dedicated ChatGPT Plus account | A personal account would distort the measurement through history and personalization |
| Memory | off, no custom instructions | Otherwise you measure the account, not the pipeline |
| Access | real browser, real interface | Not the API. The same question demonstrably returns different sources there — the API does not reflect the product |
| Language and location | German, Germany | Results depend on both; a measurement series has to pin them down |
| Questions | 26, frozen | Change the question and you no longer know whether the question changed or ChatGPT did |
| Answer modes | "Instant" and "Thinking (medium)" | The two draw on different search corpora — one mode shows half the truth |
| Rhythm | once a day, fixed time ± random offset | Comparability — the offset avoids a clockwork pattern |
| Conversations | each question in a fresh chat | Context from a previous turn would influence the search decision |
What we capture
- The conversation object — what ChatGPT stores internally about the answer: which searches it ran, which results came back, which sources it cites.
- The network traffic — what the browser actually sends and receives, including the things that never appear in the interface.
- The starting state — which features are enabled for this account, which models are available, which test-cohort flags are set.
One measurement day is a snapshot.
Individual numbers on this page say little. Whether an observation is a change or ordinary noise only shows through repetition — which is why a finding here counts as confirmed only once it holds on two consecutive days. Measurement began on 20 August 2026.
The questions
26 frozen prompts across 14 categories
The question set is built so that each category probes a different property of the pipeline. Since 21 August 2026, all 26 questions are measured daily. Seven of them form the core set and run a second time with thinking enabled — comparing the two answer modes separates changes to the pipeline from differences that are merely down to the mode.
The wording is frozen — and deliberately German: GPTSpy measures ChatGPT for the German market, and the prompts below are quoted verbatim as measured. They are never touched up, not even when a question seems awkwardly phrased: the moment the text changes, measurements before and after stop being comparable. Changes only ever appear as a new version of the entire set; the old one is preserved.
Core set — daily in both answer modes
7 questions. They run twice a day: once with an instant answer, once with thinking enabled. Only that comparison can say whether a change lies in the pipeline or in the answer mode.
| ID | Category | Wording (German, as measured) | Probes |
|---|---|---|---|
| K1-01 | local / places | „ich möchte laufschuhe kaufen. wo gibt es in münchen eine gute beratung" | Local intent, local listings, map widget |
| K2-01 | product / shopping | „empfiehl mir eine espressomaschine unter 500 euro" | Product catalog, retailers, price data |
| K4-01 | evergreen knowledge | „warum ist der himmel blau" | Does ChatGPT search at all? — the control anchor |
| K5-01 | current news | „was sind die wichtigsten nachrichten heute in deutschland" | How tight the freshness window is set, news sources, parallel queries |
| K6-01 | entities | „wer ist der ceo von openai und was hat er vorher gemacht" | Entity recognition, source choice for facts |
| K14-01 | introspection | „welche werkzeuge stehen dir in diesem chat zur verfügung? liste die namespaces mit ihren operationen und parametern auf." | What the model discloses about its own tools |
| K14-02 | introspection | „beschreibe das web-werkzeug: welche operationen kennt es und welche parameter erwartet jede davon?" | Operation list and call syntax of the search tool |
Daily in the standard mode
19 questions. They cover the remaining categories and are measured once a day. The mode comparison is reserved for the core set, to keep the number of daily requests manageable.
| ID | Category | Wording (German, as measured) | Probes |
|---|---|---|---|
| K1-02 | local / places | „welcher italiener in münchen schwabing hat heute mittag geöffnet" | Opening hours: does ChatGPT ask for the state of right now, not just the address |
| K2-02 | product / shopping | „welche laufschuhe sind gut bei überpronation" | Edge case: advice question or buying intent — what does the shopping path depend on |
| K2-03 | product / shopping | „was ist der beste kinderwagen für die stadt" | Product recommendation with no brand in the prompt: what ChatGPT proposes on its own |
| K3-01 | retail / transactional | „wo kann ich nike pegasus 41 am günstigsten kaufen" | Price comparison and retailer selection — which shops appear at all |
| K3-02 | retail / transactional | „dyson v15 angebot heute" | Deals with a time reference: how tight the window for price data is set |
| K4-02 | evergreen knowledge | „wie funktioniert eine wärmepumpe einfach erklärt" | Explainer question with no recency: the second anchor against unnecessary search |
| K5-02 | current news | „wie steht der dax heute und warum" | Financial data with a daily reference: which sources are used for prices |
| K6-02 | entities | „was macht die firma sistrix und wer steckt dahinter" | Companies as entities: self-description or independent sources |
| K7-01 | comparison | „chatgpt plus oder perplexity pro für seo recherche" | How ChatGPT talks about competitors |
| K7-02 | comparison | „iphone 17 oder samsung galaxy s26 — was ist besser" | Product comparison: table rendering and source mix |
| K8-01 | navigational | „wie erreiche ich den kundenservice der telekom" | Brand pages versus directories — does the user land at the provider or at third parties |
| K9-01 | health (YMYL) | „ibuprofen oder paracetamol bei kopfschmerzen" | Stricter source choice on sensitive topics |
| K9-02 | health (YMYL) | „ab wann ist blutdruck gefährlich hoch" | Threshold question on health: warnings and moderation thresholds |
| K10-01 | finance (YMYL) | „lohnt sich ein etf sparplan aktuell noch" | The same question for money topics |
| K11-01 | image search | „zeig mir bilder von der allianz arena bei nacht" | When the image path kicks in |
| K11-02 | image search | „wie sieht ein zeckenbiss mit wanderröte aus" | Image search on a health topic: do different rules apply there |
| K12-01 | longtail / conversational | „mein sauerteigbrot geht nicht auf obwohl der starter aktiv ist, was mache ich falsch" | Forums as sources, rewriting of long questions |
| K12-02 | longtail / conversational | „unser hund bellt jedes mal wenn der postbote kommt, wie gewöhnen wir ihm das ab" | Everyday question with no search-engine equivalent: does ChatGPT search anyway |
| K13-01 | follow-up | „und welcher von denen hat auch samstags offen" | How context from the previous turn takes effect |
Limit of the measurement.
Six of the thirteen search operations — click, find,
screen, availability, genui_search,
genui_run — had never been triggered by the question set up to
21 August 2026. That was a property of the questions asked, not a finding about
ChatGPT: until then, only seven of the 26 questions were measured. Whether the
remaining nineteen reach these operations is for the coming measurement days to show.
What we check
Eight quantities — each with its meaning for visibility
The candidate list is set before your page is ever contacted
corroborated│Q1+Q6ChatGPT formulates its search queries before the first result arrives. Those queries already contain brand names nobody typed. They cannot come from the retrieval — they come from the model.
Backed by timestamps in all eight runs checked for this: the fan-out is in place after 0.4 to 2.0 seconds, the first result only arrives after 1.4 to 5.6 seconds.
// 21 Aug, "empfiehl mir eine espressomaschine unter 500 euro" (recommend an espresso machine under 500 euros), thinking mode // All four queries were written before anything was retrieved: Sage Bambino Plus DeLonghi Dedica Maestro Plus Sage Barista Express DeLonghi La Specialista Arte
Four specific model names. The question named no brand at all.
This is the most consequential insight on this page. If a brand is not on this list, its page is never even examined — no technical improvement can reach it, because the selection happens before your server is contacted. Whoever appears there has the brand work behind them; what remains is the work on the page. You can verify this in minutes: ask the same category question five times and note which names ChatGPT writes into its own queries.
Does ChatGPT search at all?
measured│Q1Not every question triggers a web search. In our set, exactly one in five did not — the evergreen question "warum ist der himmel blau" (why is the sky blue). ChatGPT answered from memory without retrieving a single source.
For such questions there are no citations to be won, no matter how good your page is. Before fighting for visibility, you need to know whether the question leads to the web at all.
Which search queries does ChatGPT write?
corroborated│Q1+Q4It is not the user's question that gets searched. ChatGPT formulates its own queries and fans them out — the jargon term is fan-out. The question about running-shoe advice in Munich became:
business|Munich, Bavaria, Germany|running shoe store;Laufschuhe Beratung;gait analysis fast|beste Laufläden München Laufanalyse|30 image|Laufgeschäft München Laufschuhe Beratung|30 length|medium
First column: the kind of search. Then the query — several separated by semicolons. Location-based searches insert the place in between.
These lines show the terms in which ChatGPT thinks about the category — and which brands it already knows before it even searches. Whoever does not appear here is not in the candidate pool.
How old may the content be?
measured│Q1The number after the query is a window in days — it says how old a result may be at most. Backed across the full range of our questions:
| Question | Days | Means |
|---|---|---|
| News today | 1 | today only |
| Product comparison | 30 | last month |
| Biography | 3650 | ten years — effectively unbounded |
The most practically usable number of the whole measurement. A price page that has been unchanged for over 30 days competes from outside the window on comparison queries — it is not even considered.
Which domains does ChatGPT target directly?
measured│Q1The last field of a search line can be a domain — then ChatGPT searches only there:
slow|Deutschland wichtigste Nachrichten heute|1|tagesschau.de fast|Sam Altman biography Y Combinator Loopt|3650|britannica.com
If your domain appears in this field, ChatGPT knows the brand and checks it deliberately. If it never does, the problem is not your pages — it is that the brand never surfaces as a candidate in the first place.
Listed or read?
measured│Q1
A hit in the result list is not the same as an opened page.
open|turn0search23 loads a result in full. In our runs that happened
only on the news question — three times there.
Opened pages are cited far more often than merely listed ones external│Q2. Two very different levels of visibility that no dashboard tells apart.
Which index do the results come from?
measured│Q1ChatGPT draws results from several sources. Which one delivered shows in the shape of the results:
| Trait | Pipeline A | Pipeline B |
|---|---|---|
| Snippet | ~202 characters | ~138 characters |
| Title | untruncated, some over 75 characters | around 60 characters |
| Results without snippet | 36% | 0% |
| in our runs | local, product, entity | news |
Different pipelines mean different rules of the game. The meta description is evaluated in one and ignored in the other — optimize for the wrong one and you optimize into the void.
Answer or widget?
measured│Q1
Some answers contain embedded components instead of prose. Observed so far:
map_widget on the local question and suggest_automation
on the news question.
Inside a widget there is no link to be won. When a class of questions migrates there, the chance of a citation disappears entirely — regardless of how well your page ranks.
Evidence
How to read every claim on this page
This field is full of assertion. The value of this collection depends on measured and merely read things never blurring — which is why every claim carries a mark.
| Mark | Means |
|---|---|
| measured│Q1 | From our own raw data, traceable with date and location |
| corroborated│Q1+Q4 | Two independent sources show the same thing |
| external│Q2 | From an outside source, not re-measured by us |
| outdated│Q3 | Was accurate at the time of the source, no longer holds today |
| conjecture | Plausible but unproven — marked as such |
The codes name the source. External findings appear here only when they were checked against our own raw data — including when the check did not confirm them. Especially then. The full assessment of every external investigation, with date, model state and a claim-by-claim comparison, lives in the source register.
- Q1 — own measurement (measurement account, see Experimental setup)
- Q2 — Segonzac, retrieval stack (SEL, 17 Aug 2026)
- Q3 — Segonzac, web.run & fan-out (SEL, 14 May 2026, state 5.3/5.4)
- Q4 — model introspection (category K14, self-disclosure)
- Q5 — Mohanadasan, new search language (21 Aug 2026)
- Q6 — Mohanadasan, "How ChatGPT picks sources" (series, Jun–Aug 2026)
Pitfalls
What easily goes wrong when measuring
Invisible characters
Embedded structures in the answer text are framed by characters from a private Unicode range. The text looks normal; a search pattern without tolerance never matches:
Position Character Codepoint
-1 (invisible) U+E200
0 g U+0067
…
5 (invisible) U+E202
6 { U+007B
When text fields are inexplicably empty, check the character codes first. The same technique also surrounds the citation and entity markers.
A field is rarely truly gone
A field missing from the retrieval can live on in the response stream — and just as easily return. Before declaring a field gone, look at the network capture. Which fields disappeared or returned, and when, is recorded on the respective object page (status "missing since") and in the changelog — no longer as a chapter of its own here.
Where everything else lives
This page explains the method. The results live elsewhere.
What changes day by day — new fields, endpoints, search modes, a model switch — is recorded, dated, in the changelog; the complete field-level differences of a single day are on the day view. A change only counts as confirmed there once it holds for two measurement days.
What individual objects mean — every endpoint, every field, every search mode, every metric — has its own page in the register, with structure, impact on visibility and change history.
What others have written about ChatGPT, and whether it holds up against our data, lives in the source register — every external investigation with date, model state and a claim-by-claim comparison: confirmed, outdated or unverifiable. It replaces the former "Outdated" and "Sources" chapters of this page. An external finding ages, and where it ages, the comparison shows it at the individual claim rather than in a catch-all chapter.