GPTSpy · Knowledge base · As of 21 August 2026

How ChatGPT searches and builds its answers

ChatGPT answers questions by searching the web often — but not always. How it searches — which queries it writes, which pages it opens, how old content may be — changes constantly. This page is our attempt to document it.

We ask the same questions every day and record what changes. This page holds what can be backed by evidence.

Measuring since20 Aug 2026
Questions in the set26
Runs per day33

Why this matters

The occasion — shown on a concrete example

Between 16 and 20 August, OpenAI swapped out the language in which ChatGPT formulates its web searches. On 21 August, the endpoint that retrieves conversations changed as well — response format included. measured│Q1

Neither appeared in any changelog. If you do not measure daily, you notice such things only when your own numbers drop — and by then you no longer know why.

Who this page is for.

For anyone who wants to understand how content ends up in ChatGPT's answers. Knowing SEO helps; knowledge of ChatGPT's internals is not assumed. Technical terms are explained where they first appear.

Every claim carries its evidence: measured corroborated external outdated conjecture — what the marks mean

Experimental setup

Every decision with its reason

A measurement series is only as good as its constants. What is fixed here, and why:

WhatHowWhy this way
Accountdedicated ChatGPT Plus account A personal account would distort the measurement through history and personalization
Memoryoff, no custom instructions Otherwise you measure the account, not the pipeline
Accessreal browser, real interface Not the API. The same question demonstrably returns different sources there — the API does not reflect the product
Language and locationGerman, Germany Results depend on both; a measurement series has to pin them down
Questions26, frozen Change the question and you no longer know whether the question changed or ChatGPT did
Answer modes"Instant" and "Thinking (medium)" The two draw on different search corpora — one mode shows half the truth
Rhythmonce a day, fixed time ± random offset Comparability — the offset avoids a clockwork pattern
Conversationseach question in a fresh chat Context from a previous turn would influence the search decision

What we capture

  • The conversation object — what ChatGPT stores internally about the answer: which searches it ran, which results came back, which sources it cites.
  • The network traffic — what the browser actually sends and receives, including the things that never appear in the interface.
  • The starting state — which features are enabled for this account, which models are available, which test-cohort flags are set.

One measurement day is a snapshot.

Individual numbers on this page say little. Whether an observation is a change or ordinary noise only shows through repetition — which is why a finding here counts as confirmed only once it holds on two consecutive days. Measurement began on 20 August 2026.

The questions

26 frozen prompts across 14 categories

The question set is built so that each category probes a different property of the pipeline. Since 21 August 2026, all 26 questions are measured daily. Seven of them form the core set and run a second time with thinking enabled — comparing the two answer modes separates changes to the pipeline from differences that are merely down to the mode.

The wording is frozen — and deliberately German: GPTSpy measures ChatGPT for the German market, and the prompts below are quoted verbatim as measured. They are never touched up, not even when a question seems awkwardly phrased: the moment the text changes, measurements before and after stop being comparable. Changes only ever appear as a new version of the entire set; the old one is preserved.

Core set — daily in both answer modes

7 questions. They run twice a day: once with an instant answer, once with thinking enabled. Only that comparison can say whether a change lies in the pipeline or in the answer mode.

IDCategoryWording (German, as measured)Probes
K1-01local / places„ich möchte laufschuhe kaufen. wo gibt es in münchen eine gute beratung"Local intent, local listings, map widget
K2-01product / shopping„empfiehl mir eine espressomaschine unter 500 euro"Product catalog, retailers, price data
K4-01evergreen knowledge„warum ist der himmel blau"Does ChatGPT search at all? — the control anchor
K5-01current news„was sind die wichtigsten nachrichten heute in deutschland"How tight the freshness window is set, news sources, parallel queries
K6-01entities„wer ist der ceo von openai und was hat er vorher gemacht"Entity recognition, source choice for facts
K14-01introspection„welche werkzeuge stehen dir in diesem chat zur verfügung? liste die namespaces mit ihren operationen und parametern auf."What the model discloses about its own tools
K14-02introspection„beschreibe das web-werkzeug: welche operationen kennt es und welche parameter erwartet jede davon?"Operation list and call syntax of the search tool

Daily in the standard mode

19 questions. They cover the remaining categories and are measured once a day. The mode comparison is reserved for the core set, to keep the number of daily requests manageable.

IDCategoryWording (German, as measured)Probes
K1-02local / places„welcher italiener in münchen schwabing hat heute mittag geöffnet"Opening hours: does ChatGPT ask for the state of right now, not just the address
K2-02product / shopping„welche laufschuhe sind gut bei überpronation"Edge case: advice question or buying intent — what does the shopping path depend on
K2-03product / shopping„was ist der beste kinderwagen für die stadt"Product recommendation with no brand in the prompt: what ChatGPT proposes on its own
K3-01retail / transactional„wo kann ich nike pegasus 41 am günstigsten kaufen"Price comparison and retailer selection — which shops appear at all
K3-02retail / transactional„dyson v15 angebot heute"Deals with a time reference: how tight the window for price data is set
K4-02evergreen knowledge„wie funktioniert eine wärmepumpe einfach erklärt"Explainer question with no recency: the second anchor against unnecessary search
K5-02current news„wie steht der dax heute und warum"Financial data with a daily reference: which sources are used for prices
K6-02entities„was macht die firma sistrix und wer steckt dahinter"Companies as entities: self-description or independent sources
K7-01comparison„chatgpt plus oder perplexity pro für seo recherche"How ChatGPT talks about competitors
K7-02comparison„iphone 17 oder samsung galaxy s26 — was ist besser"Product comparison: table rendering and source mix
K8-01navigational„wie erreiche ich den kundenservice der telekom"Brand pages versus directories — does the user land at the provider or at third parties
K9-01health (YMYL)„ibuprofen oder paracetamol bei kopfschmerzen"Stricter source choice on sensitive topics
K9-02health (YMYL)„ab wann ist blutdruck gefährlich hoch"Threshold question on health: warnings and moderation thresholds
K10-01finance (YMYL)„lohnt sich ein etf sparplan aktuell noch"The same question for money topics
K11-01image search„zeig mir bilder von der allianz arena bei nacht"When the image path kicks in
K11-02image search„wie sieht ein zeckenbiss mit wanderröte aus"Image search on a health topic: do different rules apply there
K12-01longtail / conversational„mein sauerteigbrot geht nicht auf obwohl der starter aktiv ist, was mache ich falsch"Forums as sources, rewriting of long questions
K12-02longtail / conversational„unser hund bellt jedes mal wenn der postbote kommt, wie gewöhnen wir ihm das ab"Everyday question with no search-engine equivalent: does ChatGPT search anyway
K13-01follow-up„und welcher von denen hat auch samstags offen"How context from the previous turn takes effect

Limit of the measurement.

Six of the thirteen search operations — click, find, screen, availability, genui_search, genui_run — had never been triggered by the question set up to 21 August 2026. That was a property of the questions asked, not a finding about ChatGPT: until then, only seven of the 26 questions were measured. Whether the remaining nineteen reach these operations is for the coming measurement days to show.

What we check

Eight quantities — each with its meaning for visibility

The candidate list is set before your page is ever contacted

corroborated│Q1+Q6

ChatGPT formulates its search queries before the first result arrives. Those queries already contain brand names nobody typed. They cannot come from the retrieval — they come from the model.

Backed by timestamps in all eight runs checked for this: the fan-out is in place after 0.4 to 2.0 seconds, the first result only arrives after 1.4 to 5.6 seconds.

// 21 Aug, "empfiehl mir eine espressomaschine unter 500 euro" (recommend an espresso machine under 500 euros), thinking mode
// All four queries were written before anything was retrieved:
Sage Bambino Plus
DeLonghi Dedica Maestro Plus
Sage Barista Express
DeLonghi La Specialista Arte

Four specific model names. The question named no brand at all.

Why this matters

This is the most consequential insight on this page. If a brand is not on this list, its page is never even examined — no technical improvement can reach it, because the selection happens before your server is contacted. Whoever appears there has the brand work behind them; what remains is the work on the page. You can verify this in minutes: ask the same category question five times and note which names ChatGPT writes into its own queries.

Does ChatGPT search at all?

measured│Q1

Not every question triggers a web search. In our set, exactly one in five did not — the evergreen question "warum ist der himmel blau" (why is the sky blue). ChatGPT answered from memory without retrieving a single source.

Why this matters

For such questions there are no citations to be won, no matter how good your page is. Before fighting for visibility, you need to know whether the question leads to the web at all.

Which search queries does ChatGPT write?

corroborated│Q1+Q4

It is not the user's question that gets searched. ChatGPT formulates its own queries and fans them out — the jargon term is fan-out. The question about running-shoe advice in Munich became:

business|Munich, Bavaria, Germany|running shoe store;Laufschuhe Beratung;gait analysis
fast|beste Laufläden München Laufanalyse|30
image|Laufgeschäft München Laufschuhe Beratung|30
length|medium

First column: the kind of search. Then the query — several separated by semicolons. Location-based searches insert the place in between.

Why this matters

These lines show the terms in which ChatGPT thinks about the category — and which brands it already knows before it even searches. Whoever does not appear here is not in the candidate pool.

How old may the content be?

measured│Q1

The number after the query is a window in days — it says how old a result may be at most. Backed across the full range of our questions:

QuestionDaysMeans
News today1today only
Product comparison30last month
Biography3650ten years — effectively unbounded
Why this matters

The most practically usable number of the whole measurement. A price page that has been unchanged for over 30 days competes from outside the window on comparison queries — it is not even considered.

Which domains does ChatGPT target directly?

measured│Q1

The last field of a search line can be a domain — then ChatGPT searches only there:

slow|Deutschland wichtigste Nachrichten heute|1|tagesschau.de
fast|Sam Altman biography Y Combinator Loopt|3650|britannica.com
Why this matters

If your domain appears in this field, ChatGPT knows the brand and checks it deliberately. If it never does, the problem is not your pages — it is that the brand never surfaces as a candidate in the first place.

Listed or read?

measured│Q1

A hit in the result list is not the same as an opened page. open|turn0search23 loads a result in full. In our runs that happened only on the news question — three times there.

Why this matters

Opened pages are cited far more often than merely listed ones external│Q2. Two very different levels of visibility that no dashboard tells apart.

Which index do the results come from?

measured│Q1

ChatGPT draws results from several sources. Which one delivered shows in the shape of the results:

TraitPipeline APipeline B
Snippet~202 characters~138 characters
Titleuntruncated, some over 75 charactersaround 60 characters
Results without snippet36%0%
in our runslocal, product, entitynews
Why this matters

Different pipelines mean different rules of the game. The meta description is evaluated in one and ignored in the other — optimize for the wrong one and you optimize into the void.

Answer or widget?

measured│Q1

Some answers contain embedded components instead of prose. Observed so far: map_widget on the local question and suggest_automation on the news question.

Why this matters

Inside a widget there is no link to be won. When a class of questions migrates there, the chance of a citation disappears entirely — regardless of how well your page ranks.

Evidence

How to read every claim on this page

This field is full of assertion. The value of this collection depends on measured and merely read things never blurring — which is why every claim carries a mark.

MarkMeans
measured│Q1 From our own raw data, traceable with date and location
corroborated│Q1+Q4 Two independent sources show the same thing
external│Q2 From an outside source, not re-measured by us
outdated│Q3 Was accurate at the time of the source, no longer holds today
conjecture Plausible but unproven — marked as such

The codes name the source. External findings appear here only when they were checked against our own raw data — including when the check did not confirm them. Especially then. The full assessment of every external investigation, with date, model state and a claim-by-claim comparison, lives in the source register.

Pitfalls

What easily goes wrong when measuring

Invisible characters

Embedded structures in the answer text are framed by characters from a private Unicode range. The text looks normal; a search pattern without tolerance never matches:

Position  Character  Codepoint
      -1  (invisible)  U+E200
       0  g          U+0067
     …
       5  (invisible)  U+E202
       6  {          U+007B

When text fields are inexplicably empty, check the character codes first. The same technique also surrounds the citation and entity markers.

A field is rarely truly gone

A field missing from the retrieval can live on in the response stream — and just as easily return. Before declaring a field gone, look at the network capture. Which fields disappeared or returned, and when, is recorded on the respective object page (status "missing since") and in the changelog — no longer as a chapter of its own here.

Where everything else lives

This page explains the method. The results live elsewhere.

What changes day by day — new fields, endpoints, search modes, a model switch — is recorded, dated, in the changelog; the complete field-level differences of a single day are on the day view. A change only counts as confirmed there once it holds for two measurement days.

What individual objects mean — every endpoint, every field, every search mode, every metric — has its own page in the register, with structure, impact on visibility and change history.

What others have written about ChatGPT, and whether it holds up against our data, lives in the source register — every external investigation with date, model state and a claim-by-claim comparison: confirmed, outdated or unverifiable. It replaces the former "Outdated" and "Sources" chapters of this page. An external finding ages, and where it ages, the comparison shows it at the individual claim rather than in a catch-all chapter.