What of AEO really works
The honest answer is uncomfortable: in the only controlled comparison so far, 3 of 54 method-domain combinations were significant — and for question-answer tasks, not a single one. Anyone promising you a large effect from text optimisation has not read the studies.
The answer in one sentence
It is not the form of a text that decides, but what verifiably stands in it. Format cosmetics do not work; figures with a source, clean definitions and real comparisons do.
The controlled comparison
C-SEO Bench (NeurIPS 2025) tested common optimisation methods against control groups — the difference from the usual observational studies is that here an actual comparison was made instead of counting correlations.
What still stood out in the observation
A large observational measurement with 21,143 citations and 18,151 fetched pages found clear differences — not by format, but by evidence genre. That is correlation, not causation; the authors themselves call their value an “observational proxy”.
- +76,9 %CodeEspecially with technical topics.
- +61,6 %Figures and statisticsWith a named source — without one it is a claim.
- +57,3 %DefinitionsThe sentence “X is …” is the unit most readily taken over.
- +55,3 %ComparisonsWhoever compares is facing a decision.
- +41,2 %InstructionsSteps in a comprehensible order.
- −5,74 %Question-answer formatThe only negative value in the measurement.
Why we do not build FAQ pages
Das Frage-Antwort-Format war der einzige Typ mit negativem Wert. Trotzdem ist es der meistverkaufte Baustein im GEO-Markt — wir haben Angebote gesehen, die 195 FAQ-Seiten je Kunde vorsehen. Die Autoren der Messung schreiben wörtlich: „surface format alone is not sufficient signal“.
Instead we build concept pages: the answer in one sentence, a precise definition, a delimitation, an example and figures with a source. Those are exactly the genres with the highest measured values — and at the same time what a person wants to read.
What follows from this
If text optimisation barely works, the leverage lies elsewhere: in whether a page may be fetched at all and whether it is cited by the sources AI systems actually draw on.
Stand: 1. September 2026 · C-SEO Bench (NeurIPS 2025) · Beobachtungsmessung arxiv.org/html/2604.25707v2