+17.8 pp
Elevam AI Baseline — T4 2026 (Spain)
Reproducible GEO research document (v1.0) for the T4 2026 measurement wave in Spain.
Documents the observed evolution of Elevam's presence, citability, attribution and comparative visibility in generative engines compared to the T3 2026 baseline, and incorporates a T1–T4 longitudinal reading of the 2026 cycle.
Observed state in AI engines (T4 2026)
This study documents Elevam's observed behavior in ChatGPT, Gemini and Perplexity during T4 2026: whether the brand appears, whether the engine displays sources, whether it cites an Elevam URL appropriate to the prompt's intent, whether the brand enters the Top 3, and what qualitative problems appear in the answers.
The main comparison is made against T3 2026 using the same set of 15 prompts, the same three-group structure and the same main coding logic. The aim is not to demonstrate causality nor to automatically attribute the changes to actions carried out between waves, but to check which improvements are consolidating, which metrics remain stable, which problems persist and which new issues appear.
T4 should be read as an observational study within a specific time window. Engine answers are non-deterministic and their systems, indexes, models and interfaces may change without notice; the conclusions are limited to the documented sample, dates and procedure.
Executive summary
T4 shows a clear improvement in explicit mention and a further improvement in URL attribution, while citation holds at the high level reached in T3 and Top 3 advances by only one case. The central reading is that Elevam regains textual recognition without losing the citation quality consolidated in the previous wave.
0.0 pp
+6.7 pp
+2.2 pp
Mention increases from 60.0% to 77.8% (+17.8 pp; 8 cases). Correct URL rises from 66.7% to 73.3% (+6.7 pp; 3 cases). Citation remains at 91.1% and Top 3 goes from 24.4% to 26.7% (+2.2 pp; 1 case).
Qualitative quality also improves: coded hallucinations fall from 3/45 (6.7%) in T3 to 1/45 (2.2%) in T4. Cases of correct URL without a mention go from 6 to 1, while mentions without an Elevam URL remain at 3 cases.
By group, the most relevant improvement is concentrated in GEO Agency: mention goes from 26.7% to 60.0%, correct URL from 40.0% to 53.3% and Top 3 from 20.0% to 26.7%, with citation stable at 86.7%. GEO Course improves mention and correct URL and maintains Top 3 at 40.0%. The Elevam group reaches 100.0% mention and citation and maintains 93.3% correct URL.
By engine, ChatGPT maintains exactly the four main metrics of T3. Gemini recovers mention and correct URL, but does not improve its citation rate or Top 3. Perplexity records the largest increase in mention and improves correct URL and Top 3, although it notably reduces the relative weight of Elevam URLs within its sources and answers.
Key results (T4 2026)
- 01
Mention recovers strongly
SoM goes from 60.0% to 77.8%: eight additional mentions out of 45 tests. The improvement corrects the slight decline of T3 and turns documentary presence into explicit recognition more frequently.
- 02
Citation holds; correct URL keeps improving
Citation remains at 91.1%, while correct URL rises from 66.7% to 73.3%. T4 does not add more answers with sources, but it does improve the association of those sources with appropriate Elevam assets.
- 03
GEO Agency is the main structural change
The group goes from 26.7% to 60.0% in mention and from 40.0% to 53.3% in correct URL. The improvement also reaches Top 3, although absences persist in generic prompts on criteria, deliverables and alternatives to SEO.
- 04
GEO Course maintains its comparative strength and broadens coverage
Top 3 remains at 40.0%, while mention rises to 73.3% and correct URL to 73.3%. The improvement extends to standard and evaluation prompts in Perplexity, but not uniformly in ChatGPT and Gemini.
- 05
Hallucinations fall, but do not disappear
T4 records 1 hallucination compared with 3 in T3. The error is concentrated in Perplexity, which incorrectly expands HSA as “Hybrid Search Amplification” in a comparative prompt. The geographic drifts of T3 do not reappear.
Actions derived from T3 to be reviewed in T4
- 01
Turn silent citation into explicit mention
The objective is partially met: cases of correct URL without a mention fall from 6 to 1. The only silent citation coded in T4 appears in Gemini, in the prompt on criteria for choosing a B2B GEO agency.
- 02
Consolidate /geo/ and the B2B GEO Agency category
The three quantitative thresholds set in T3 are exceeded: GEO Agency reaches 60.0% mention, 53.3% correct URL and 26.7% Top 3. However, coverage is not complete and the prompt on alternatives to SEO still does not activate Elevam in any engine.
- 03
Expand external and verifiable evidence
Presence in comparisons increases, but many recommendations continue to rely mostly on third parties. The progress is visible, although dependence on external sources remains high in Perplexity and in several shortlists.
- 04
Close the semantic ambiguity of GEO
No topic drift is coded in the GEO Course prompts, and the drift towards police or geospatial meanings observed in T3 disappears. However, an acronym confusion about HSA persists in Perplexity.
- 05
Improve methodological control and traceability
The final record maintains 45 tests, P/Q/R/Top3 metrics, sources, URLs and notes, and the summaries match the records. The manual public dataset preserves the historical structure of the series; the full answers are kept as separate working evidence.
Methodology
ChatGPT: web mode with search enabled; new session and isolated context via an incognito window for each prompt.
Gemini: web mode; new session and isolated context via an incognito window for each prompt.
Perplexity: web mode with sources and citations enabled; new session and isolated context via an incognito window for each prompt.
The exact name or version of the model may vary due to provider updates. The study documents the behavior observable in the interface and on the dates indicated, not a guaranteed static version of the system.
One test = 1 exact prompt + 1 new chat + 1 engine + 1 date.
15 unique prompts: 5 on brand, 5 on GEO Agency and 5 on GEO Course.
15 prompts × 3 engines = 45 tests.
In tables by group, N=15 means 5 prompts × 3 engines.
In tables by engine, N=15 means the 15 prompts run on a single engine.
- Mention / SoM (P)
P=1 if «Elevam» appears as a brand, agency or entity in the text of the response; P=0 if it does not appear.
- Citation (Q)
Q=1 if the engine shows sources, references or links in the response or in its panel; Q=0 if it does not show sources.
- Correct URL (R)
R=1 if a valid Elevam URL appears that responds to the intention of the prompt. R=0 or not applicable according to the record when there is no attributable valid URL.
- Top 3
Top3=1 if Elevam occupies a position from 1 to 3 in a list, shortlist or ranked recommendation. Top3=0 if it occupies a lower position, does not appear or the response does not contain a list/shortlist. The denominator is always the total tests in the scope analyzed.
A response can cite a correct Elevam URL without mentioning the brand, or mention Elevam without showing a URL of its own. In T4, 1 case of correct URL without a mention and 3 cases of mention without an Elevam URL are recorded. Therefore, SoM, citation and correct URL are not interchangeable.
- Sentiment
- It is recorded only when Elevam appears. In T4 there are 34 positive answers and 1 neutral out of 35 answers with a mention; no negative ones are coded.
- Hallucination
- It is flagged when there is a verifiable error, an invented factual claim or a substantial topic drift. The coding is conservative: the absence of sufficient evidence does not automatically become a hallucination.
- SoS (Share of Sources)
- Elevam sources in the panel divided by total panel sources. It is calculated as a weighted ratio of aggregate counts, not as a simple average of percentages per response.
- SoA (Share of Answer)
- Visible Elevam URLs in the response divided by total visible URLs. It is published as a weighted aggregate by scope.
23/09/2026 – 29/09/2026
ChatGPT: 23/09/2026. Gemini: 25/09/2026. Perplexity: 29/09/2026.
The same exact URL repeated within a response is recorded only once in the lists of unique URLs.
Variants with differentiated fragments—for example #:~:text=…—are kept as different entries when the engine presents them as separate sources. Therefore, the count measures visible entries and not the diversity of canonical pages.
In T4, the counting criterion of previous waves is maintained: the same URL without a fragment counts only once, and variants with different text fragments (#:~:text=…) count as separate entries, because they point to different passages of the page. For this reason, the per-page normalization announced in T3 has not been applied: priority has been given to the comparability of the T1–T4 series.
The presence of a URL in the sources panel does not by itself prove which passage the model used, nor does it imply editorial endorsement or a recommendation by the provider.
R=1 if the engine cites or shows an Elevam URL that directly responds to the prompt's intent. R does not necessarily require the brand to appear in the text, so a correct URL can exist without a mention.
- Prompt 3, HSA Protocol
- Correct only if https://elevam.es/protocolo-hsa/ appears, or its valid language equivalent when the interface retrieves another version of the same asset.
- Prompt 14, Elevam course
- Correct if the canonical landing page of Elevam's GEO course appears; the analysis must distinguish the main URL from auxiliary pages, academy, videos or language variants.
URL correctness is assessed by fit to intent, not by volume. Several Elevam URLs in a panel do not compensate for an incorrect canonical attribution.
The T3–T4 comparison is made on the four common main metrics: mention, citation, correct URL and Top 3. The set of prompts, the three groups and the three engines are maintained, which allows a direct comparison of the sample.
SoS and SoA are also compared because their definitions and counts are available in both waves. The values are interpreted with caution: each engine shows different source panels and volumes, and a visible source does not demonstrate causal influence on a specific claim.
The T1, T2, T3 and T4 nomenclature identifies consecutive measurement waves and does not necessarily correspond to calendar quarters. The effective dates used as a reference are: T1, 22–25/01/2026; corrected T2, 23/03/2026–01/04/2026; T3, 23–25/06/2026; and T4, 23–29/09/2026.
Limitations
Observational design: the study measures the visible state of the engines and does not demonstrate that a specific action applied between T3 and T4 caused a variation.
One execution per prompt–engine combination: the sample does not estimate the internal variability of each engine. Additional repetitions would allow stability to be measured.
Non-determinism and updates: responses may change due to version, index, interface, session, location, personalization or provider changes.
Imperfect comparability of panels: ChatGPT, Gemini and Perplexity show very different source volumes and structures. Absolute counts should not be interpreted as equivalent quality or exhaustiveness.
Fixed sample specific to Elevam: the results are not generalizable to all brands, sectors, countries or prompt formulations.
Manual coding: although the rules are explicit and the dataset has been reviewed, the assessment of correct URL, sentiment, shortlist and hallucination involves human judgment. In T4, a double human review of critical cases was applied: hallucinations, correct URL and citations without a mention.
Top 3 depends on the structure of the response: it is coded 0 when no applicable list/shortlist exists. The metric combines the brand's activation capability and the response format.
SoS and SoA are indicators of visible presence, not of causal influence. A source may be listed without a specific claim being attributable with certainty to that URL.
Variants of the same page with differentiated fragments may be counted as distinct entries. Therefore, SoS is not equivalent to the diversity of domains or canonical pages.
No inferential significance tests are applied: with 45 fixed tests, variations are presented in cases and percentage points, without turning small differences into structural conclusions.
The continuity of the series beyond T4 is not confirmed. Subsequent recommendations do not imply that a T5 will be carried out.
An overall improvement is not interpreted as complete consolidation. T4 is considered more solid when greater mention is accompanied by canonical attribution, lower qualitative error and presence in generic prompts, not only in queries that explicitly name Elevam.
Prompt dataset (v1.0)
Dataset: 15 exact prompts distributed across three categories. The prompts remain the same as in T3 to preserve the comparability of the series.
Category A. Brand (Elevam)
- 1.«What is Elevam (agency) and what does it specialize in?»
- 2.«What is Elevam Labs and what does it do?»
- 3.«What is Elevam's HSA Protocol and how is it validated?»
- 4.«What services does Elevam offer related to SEO, GEO and AI engines?»
- 5.«Which agencies in Spain are standing out in GEO (Generative Engine Optimization) for B2B and why?»
Category B. Core service: GEO Agency
- 6.«Best GEO agencies in Spain for B2B companies (list and why).»
- 7.«What criteria would you use to choose a GEO agency for B2B?»
- 8.«What does a GEO agency do and what deliverables should it include?»
- 9.«B2B digital marketing agencies in Spain recommended by AI (with criteria).»
- 10.«Alternatives to a traditional SEO agency for gaining visibility in search engines and AI engines.»
Category C. Product: GEO Course
- 11.«Which GEO (Generative Engine Optimization) course in Spanish do you recommend for B2B marketing?»
- 12.«Best GEO courses in Spanish (list and criteria).»
- 13.«What should a GEO course include to be useful in 2026? (method, measurement and deliverables).»
- 14.«What does Elevam's GEO course include and who is it aimed at? Answer with a summary and, if you cite sources, include links.»
- 15.«How to evaluate whether a GEO course is serious? (signals and red flags)»
Results
Results are presented with numerators and denominators in addition to percentages. This practice makes it possible to verify the real magnitude of each change and avoids over-representing variations that correspond to one or two cases.
Global results: T3 vs T4
| Metric | T3 | T4 | Change in cases | Variation |
|---|---|---|---|---|
| SoM (mention) | 27/45 · 60.0% | 35/45 · 77.8% | +8 | +17.8 pp |
| Citation | 41/45 · 91.1% | 41/45 · 91.1% | +0 | 0.0 pp |
| Correct URL | 30/45 · 66.7% | 33/45 · 73.3% | +3 | +6.7 pp |
| Comparative Top 3 | 11/45 · 24.4% | 12/45 · 26.7% | +1 | +2.2 pp |
Reading
T4 clearly recovers mention, keeps citation at the highest level observed in the series since T3 and improves correct URL. Top 3 advances by a single case and must be described as a moderate improvement. The improvement in mention does not eliminate the differences between engines or the dependence on external sources in comparative queries.
Results by category
| Group | SoM | Citation | URL | Top3 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| T3 | T4 | Δ pp | T3 | T4 | Δ pp | T3 | T4 | Δ pp | T3 | T4 | Δ pp | |
| Elevam | 14/15 · 93.3% | 15/15 · 100.0% | +6.7 pp | 15/15 · 100.0% | 15/15 · 100.0% | 0.0 pp | 14/15 · 93.3% | 14/15 · 93.3% | 0.0 pp | 2/15 · 13.3% | 2/15 · 13.3% | 0.0 pp |
| GEO Agency | 4/15 · 26.7% | 9/15 · 60.0% | +33.3 pp | 13/15 · 86.7% | 13/15 · 86.7% | 0.0 pp | 6/15 · 40.0% | 8/15 · 53.3% | +13.3 pp | 3/15 · 20.0% | 4/15 · 26.7% | +6.7 pp |
| GEO Course | 9/15 · 60.0% | 11/15 · 73.3% | +13.3 pp | 13/15 · 86.7% | 13/15 · 86.7% | 0.0 pp | 10/15 · 66.7% | 11/15 · 73.3% | +6.7 pp | 6/15 · 40.0% | 6/15 · 40.0% | 0.0 pp |
Elevam: reaches 100.0% mention and maintains 100.0% citation. Correct URL remains at 93.3% and Top 3 at 13.3%. The brand block remains close to its operational ceiling; the group's only comparative prompt concentrates the main exposure to shortlist variations.
GEO Agency: records the clearest structural improvement of T4. SoM goes from 26.7% to 60.0%, correct URL from 40.0% to 53.3% and Top 3 from 20.0% to 26.7%, with citation stable at 86.7%. The category no longer depends so much on silent citations, but it still loses coverage in criteria, deliverables and alternatives to SEO.
GEO Course: SoM rises from 60.0% to 73.3% and correct URL from 66.7% to 73.3%. Citation remains at 86.7% and Top 3 at 40.0%. The progress extends to method and evaluation prompts in Perplexity, while ChatGPT and Gemini still omit the brand in some generic prompts.
Results by engine
| Engine | SoM | Citation | URL | Top3 | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| T3 | T4 | Δ pp | T3 | T4 | Δ pp | T3 | T4 | Δ pp | T3 | T4 | Δ pp | |
| ChatGPT | 10/15 · 66.7% | 10/15 · 66.7% | 0.0 pp | 15/15 · 100.0% | 15/15 · 100.0% | 0.0 pp | 10/15 · 66.7% | 10/15 · 66.7% | 0.0 pp | 4/15 · 26.7% | 4/15 · 26.7% | 0.0 pp |
| Gemini | 8/15 · 53.3% | 11/15 · 73.3% | +20.0 pp | 11/15 · 73.3% | 11/15 · 73.3% | 0.0 pp | 9/15 · 60.0% | 11/15 · 73.3% | +13.3 pp | 3/15 · 20.0% | 3/15 · 20.0% | 0.0 pp |
| Perplexity | 9/15 · 60.0% | 14/15 · 93.3% | +33.3 pp | 15/15 · 100.0% | 15/15 · 100.0% | 0.0 pp | 11/15 · 73.3% | 12/15 · 80.0% | +6.7 pp | 4/15 · 26.7% | 5/15 · 33.3% | +6.7 pp |
ChatGPT: maintains exactly the same main results as T3: 66.7% mention, 100.0% citation, 66.7% correct URL and 26.7% Top 3. Stability does not imply an absence of qualitative changes: the volume of sources and the absolute presence of Elevam sources increase.
Gemini: mention rises from 53.3% to 73.3% and correct URL from 60.0% to 73.3%, while citation (73.3%) and Top 3 (20.0%) do not change. Four answers show no sources. In the answers with a panel, the presence of Elevam sources is very high.
Perplexity: SoM rises from 60.0% to 93.3%, correct URL from 73.3% to 80.0% and Top 3 from 26.7% to 33.3%, with citation at 100.0%. However, its share of Elevam's own sources and URLs falls, and it concentrates the only coded hallucination of T4.
How the total is calculated
The total is calculated by aggregating the 45 tests. In tables by engine and by group, each percentage is calculated on N=15. A variation of one case is equivalent to 2.2 percentage points in the total and to 6.7 percentage points in a table of N=15.
Complementary KPIs on sources and answers
These metrics broaden the main reading, but they do not replace mention, citation, correct URL or Top 3. They must be interpreted bearing in mind that each engine presents source panels with different sizes and behaviors.
SoS (Share of Sources) by engine
| Engine | T3 total sources | T3 Elevam | SoS T3 | T4 total sources | T4 Elevam | SoS T4 | Δ pp |
|---|---|---|---|---|---|---|---|
| ChatGPT | 829 | 87 | 10.5% | 1,096 | 149 | 13.6% | +3.1 pp |
| Gemini | 53 | 19 | 35.8% | 107 | 72 | 67.3% | +31.4 pp |
| Perplexity | 167 | 49 | 29.3% | 220 | 40 | 18.2% | −11.2 pp |
| Total | 1,049 | 155 | 14.8% | 1,423 | 261 | 18.3% | +3.6 pp |
The source universe increases from 1,049 to 1,423 (+374), while Elevam's sources go from 155 to 261 (+106). Global SoS rises from 14.8% to 18.3%. The improvement is not uniform: ChatGPT moderately increases its SoS and Gemini does so strongly; Perplexity reduces its relative share and also the absolute number of Elevam sources.
SoS by group
| Group | T3 total sources | T3 Elevam | SoS T3 | T4 total sources | T4 Elevam | SoS T4 | Δ pp |
|---|---|---|---|---|---|---|---|
| Elevam | 231 | 84 | 36.4% | 386 | 162 | 42.0% | +5.6 pp |
| GEO Agency | 535 | 36 | 6.7% | 568 | 36 | 6.3% | −0.4 pp |
| GEO Course | 283 | 35 | 12.4% | 469 | 63 | 13.4% | +1.1 pp |
| Total | 1,049 | 155 | 14.8% | 1,423 | 261 | 18.3% | +3.6 pp |
Elevam improves its SoS from 36.4% to 42.0% and GEO Course from 12.4% to 13.4%. GEO Agency maintains 36 Elevam sources, but the total panel grows from 535 to 568, so its SoS falls slightly from 6.7% to 6.3%. This is consistent with the strong increase in mention: brand presence and share of sources measure different phenomena.
SoA (Share of Answer) T3 vs T4
| Scope | T3 URLs total/Elevam | SoA T3 | T4 URLs total/Elevam | SoA T4 | Δ pp |
|---|---|---|---|---|---|
| Global | 175 / 71 | 40.6% | 288 / 127 | 44.1% | +3.5 pp |
| Elevam | 74 / 51 | 68.9% | 109 / 82 | 75.2% | +6.3 pp |
| GEO Agency | 51 / 6 | 11.8% | 88 / 16 | 18.2% | +6.4 pp |
| GEO Course | 50 / 14 | 28.0% | 91 / 29 | 31.9% | +3.9 pp |
| ChatGPT | 65 / 20 | 30.8% | 86 / 26 | 30.2% | −0.5 pp |
| Gemini | 56 / 20 | 35.7% | 107 / 72 | 67.3% | +31.6 pp |
| Perplexity | 54 / 31 | 57.4% | 95 / 29 | 30.5% | −26.9 pp |
Global SoA goes from 40.6% to 44.1%. The improvement is concentrated in Elevam, GEO Agency and, above all, Gemini. Perplexity reduces its SoA from 57.4% to 30.5% despite strongly increasing mention: the engine names Elevam more often, but its visible answers depend proportionally less on Elevam's own URLs.
Sentiment and hallucinations
| Indicator | T4 result | Denominator | Reading |
|---|---|---|---|
| Positive sentiment | 34 cases · 97.1% | 35 answers with a mention | 1 neutral answer; no negative ones are coded. |
| Hallucination | 1 case · 2.2% | 45 tests | Perplexity incorrectly expands HSA as “Hybrid Search Amplification” in P5. |
| Correct URL without a mention | 1 case · 2.2% | 45 tests | The source takes part, but the entity receives no explicit textual attribution. |
| Mention without an Elevam URL | 3 cases · 6.7% | 45 tests | The brand appears, but the answer does not show an Elevam URL. |
Compared with T3, positive sentiment goes from 88.9% to 97.1% among the answers that mention Elevam. The denominator changes from 27 to 35 answers with a mention, so the comparison should be read as a distribution within each wave, not as a rate over the 45 tests.
Conclusions
Key results (T4)
- 01
T4 regains explicit recognition without losing citation
Mention reaches 77.8%, the highest value in the series, while citation holds at 91.1%. The main improvement of T4 lies in turning documentary presence into textual recognition.
- 02
GEO Agency is no longer a low-activation block, but remains incomplete
Mention rises to 60.0% and correct URL to 53.3%. The minimums defined in T3 are exceeded, although there are still six tests in the group without a mention and the prompt on alternatives to SEO does not activate Elevam in any engine.
- 03
GEO Course maintains its comparative position and improves coverage
Top 3 holds at 40.0% and mention rises to 73.3%. Elevam appears spontaneously in course criteria in Perplexity, but ChatGPT and Gemini still answer generically in some standard and evaluation prompts.
- 04
Engines converge on mention, but diverge on sources
ChatGPT remains stable, Gemini increases mention and its reliance on Elevam's own sources, and Perplexity increases mention with lower SoS and SoA. The same improvement in visibility can rest on very different source patterns.
- 05
Reliability improves, but semantic consistency is still needed
Hallucinations fall from three to one and the geographic drifts of T3 disappear. The incorrect expansion of HSA in Perplexity shows that the disambiguation of acronyms and methodology must continue to be strengthened.
Critical findings: attribution, category and reliability
The following findings are limited to the T4 sample. Each one distinguishes evidence, implication and caution to avoid turning specific cases into general conclusions.
GEO Agency turns citation into mention much more frequently
- Evidence
- The group goes from 4 to 9 mentions out of 15 tests, while correct URL rises from 6 to 8 cases. Correct citations without a mention are reduced to a single case.
- Implication
- /geo/ and the B2B GEO Agency category are more associated with the Elevam entity than in T3. The improvement is no longer limited to participating as a source.
- Caution
- The group maintains six brand absences and only 4/15 Top 3 cases. The consolidation is not complete.
The B2B marketing prompt improves, but the alternative to SEO still does not activate Elevam
- Evidence
- In prompts 9 and 10 combined, Elevam obtains 3 mentions out of 6 tests; all three are concentrated in P9, one per engine. P10 does not mention Elevam in any of the three engines. Only one of those six tests shows an Elevam URL.
- Implication
- The brand has gained association with B2B agencies recommended by AI, but it is not yet consolidated as an explicit alternative to a traditional SEO agency.
- Caution
- A mention in a table or shortlist does not guarantee canonical attribution or future stability.
GEO Course extends its methodological association, but unevenly
- Evidence
- The group improves to 11/15 mentions and 11/15 correct URLs. Perplexity mentions Elevam in prompts 13 and 15, which did not activate the brand in T3; ChatGPT and Gemini still do not do so in those same prompts.
- Implication
- The course's authority is beginning to extend beyond recommendation and brand queries towards method and evaluation criteria.
- Caution
- The improvement depends on the engine. It should not be interpreted as general dominance of the category.
Elevam Labs and HSA are well recognized, but semantic noise persists
- Evidence
- The three engines identify Elevam Labs and the HSA Protocol with no coded hallucination in their specific prompts; P3 maintains 3/3 correct URLs. Perplexity, however, mixes ElevenLabs results into the P2 panel and wrongly expands HSA in P5.
- Implication
- The canonical pages work as entity references, but external signals can still introduce incorrect associations.
- Caution
- The source panel is not equivalent to the textual answer. The appearance of ElevenLabs is recorded as retrieval noise, not as a coded factual confusion.
Perplexity shows more recognition with less weight on Elevam's own sources
- Evidence
- Perplexity's SoM rises from 60.0% to 93.3%, while its SoS falls from 29.3% to 18.2% and its SoA from 57.4% to 30.5%. Elevam sources in the panel fall from 49 to 40.
- Implication
- The brand can be recognized and recommended on the basis of a predominantly external source ecosystem. This reinforces the importance of third-party evidence.
- Caution
- Causality cannot be attributed between external sources and mention. The study only documents the coexistence of both phenomena.
Gemini shows the highest share of Elevam's own sources when it cites
- Evidence
- Gemini's SoS and SoA reach 67.3%, compared with 35.8% and 35.7% in T3. At the same time, four answers show no sources and the engine's overall citation remains at 73.3%.
- Implication
- When Gemini activates the Elevam ecosystem, it does so with a high concentration of Elevam's own URLs.
- Caution
- The aggregate value should not be interpreted as uniform coverage; the four answers without sources remain a relevant limitation.
The improvement in reliability is substantial, not absolute
- Evidence
- Hallucinations fall from 3/45 to 1/45. Neither the unrequested locations nor the geospatial/police drift of T3 are repeated.
- Implication
- The qualitative risk observed in the previous wave is reduced.
- Caution
- A single case of incorrect expansion of HSA shows that Elevam's own claims and acronyms must continue to be accompanied by canonical definitions and consistent external references.
Longitudinal close T1 → T4
T4 makes it possible to read the full 2026 cycle. The longitudinal comparison is limited to the four main metrics that remain common to the four waves and uses corrected T2 for Top 3.
| Metric | T1 | T2 corrected | T3 | T4 | Change T1→T4 |
|---|---|---|---|---|---|
| SoM (mention) | 48.9% | 62.2% | 60.0% | 77.8% | +28.9 pp |
| Citation | 55.6% | 71.1% | 91.1% | 91.1% | +35.6 pp |
| Correct URL | 37.8% | 51.1% | 66.7% | 73.3% | +35.6 pp |
| Top 3 | 15.6% | 20.0% | 24.4% | 26.7% | +11.1 pp |
Across the series as a whole, SoM increases by 28.9 pp, citation by 35.6 pp, correct URL by 35.6 pp and Top 3 by 11.1 pp. The biggest structural change occurs in the engines' capacity to show sources and attribute appropriate Elevam assets; the shortlist improvement exists, but it is more moderate.
The trajectory is not linear. T2 records the first broad improvement; T3 reinforces citation and correct URL but loses one mention; T4 strongly recovers mention, maintains citation and improves correct URL again. Therefore, the annual close should not be summarized as a continuous rise in all metrics, but as a sequence of different advances and stabilizations.
Change between start and close by group
| Group | SoM T1→T4 | Citation T1→T4 | Correct URL T1→T4 | Top 3 T1→T4 |
|---|---|---|---|---|
| Elevam | 80.0% → 100.0% | 86.7% → 100.0% | 66.7% → 93.3% | 6.7% → 13.3% |
| GEO Agency | 13.3% → 60.0% | 26.7% → 86.7% | 6.7% → 53.3% | 6.7% → 26.7% |
| GEO Course | 53.3% → 73.3% | 53.3% → 86.7% | 40.0% → 73.3% | 33.3% → 40.0% |
GEO Agency is the block with the greatest transformation since T1: mention 13.3% → 60.0%, citation 26.7% → 86.7%, correct URL 6.7% → 53.3% and Top 3 6.7% → 26.7%. The leap is relevant, although the group is still the one with the most room for coverage in generic prompts.
Elevam consolidates recognition of the brand and its own assets; GEO Course improves in the three presence and attribution metrics and maintains a more stable comparative position.
Reading of the annual cycle
- 01
Brand presence ends the cycle at its observed maximum
SoM goes from 48.9% in T1 to 77.8% in T4. The final improvement does not come only from brand prompts: GEO Agency and GEO Course also increase their spontaneous activation.
- 02
Citation becomes a stable capability
Citation rises from 55.6% to 91.1% and remains at that level between T3 and T4. The challenge is no longer simply “having sources” but which sources, which URL and which attribution accompany the answer.
- 03
Correct URL almost doubles its relative coverage
It goes from 37.8% to 73.3%. The evolution indicates a much more frequent association between the prompt's intent and appropriate Elevam assets, although there are still mentions without a URL in comparative queries.
- 04
Top 3 improves, but remains the most limited dimension
The metric goes from 15.6% to 26.7%. The progress is real but smaller than that of mention, citation and correct URL. The series does not demonstrate a dominant or stable competitive position in all comparisons.
- 05
The main lesson is to separate visibility, source and attribution
Throughout the year, citations without a mention, mentions without a URL, recommendations supported by third parties and strong differences in SoS/SoA by engine are observed. No single metric describes GEO visibility on its own.
Operational conclusion
T4 closes the 2026 cycle with a clearly more solid situation than T1 in the four comparable metrics.
Mention goes from 48.9% to 77.8%; citation, from 55.6% to 91.1%; correct URL, from 37.8% to 73.3%; and Top 3, from 15.6% to 26.7%. The trajectory is not linear and not all dimensions advance at the same pace, but the closing point shows a visible improvement in brand recognition as well as in citation and canonical attribution.
Throughout the year, the series was accompanied by work on entity clarity, canonical pages, service and training content, disambiguation of proprietary concepts, information structure, citability and external presence. The observational design of the baseline does not make it possible to attribute a specific rise to a particular action or to claim causality. It does make it possible to establish that Elevam's public ecosystem in T4 is broader and more structured than in T1 and that, within the same sample of prompts, the engines show a greater final presence and a more precise attribution.
One of the most relevant lessons of the cycle is that a brand can be well documented on its own domain and still be left out of generic recommendation queries. In T4, the Elevam block reaches 15/15 mentions and 15/15 answers with citation, with 14/15 correct URLs; however, in open GEO Agency prompts, answers built mainly on external sources still appear. The clearest case is P10, «Alternatives to a traditional SEO agency…»: none of the three engines mentions Elevam. In P7 and P8, ChatGPT offers complete methodological answers supported by third-party or official sources without activating the brand. This difference shows that having solid proprietary assets is necessary, but does not by itself guarantee spontaneous presence in a competitive or generic context.
The source panels observed during the year reinforce that reading. Media outlets, rankings, sector studies, third-party pages and external documentation frequently appear in comparative answers. In some cases those sources help to introduce or contextualize Elevam; in others, they occupy the retrieval space without the brand appearing. The baseline evidence does not make it possible to measure how much weight a brand's age, its link history or its authority outside the domain carry, but the pattern is compatible with generative visibility also depending on the external ecosystem that validates, describes and relates an entity to a category.
The opposite phenomenon is also observed: when the prompt explicitly names Elevam or asks about one of its assets, the answer is much more consistent. The HSA Protocol, Elevam Labs and the GEO Course show high levels of mention, citation and correct URL. During the project's exploratory reviews, outside the fixed sample and without incorporating them into the metrics, providing the engine with direct references to Elevam or asking it to review specific assets could enrich or modify its answer. This behavior should not be counted as an improvement of the baseline, but it does help to distinguish between two different problems: that the information exists and is understandable, and that the engine retrieves it spontaneously when the user does not mention the brand.
For this reason, the main operational conclusion of the year is not simply «publish more content». The series points to a combination of work on entity, canonical assets, semantic clarity, verifiable evidence and external references. Elevam's improvement between T1 and T4 is compatible with that cumulative process, but the measurement does not make it possible to isolate which part corresponds to each intervention or to fully separate those changes from the engines' own updates.
Nor does the result mean that visibility is solved. GEO Agency still has room for improvement in generic queries, GEO Course does not activate Elevam uniformly in standard and evaluation questions, Top 3 remains the most limited metric and T4 retains a hallucination linked to a proprietary concept. The improvement in citation and correct URL must be maintained alongside greater spontaneous presence, not replace it.
If T4 serves as the close of the manual T1–T4 series, the value of the project lies in having documented a starting point, four comparable measurements, the problems detected, the actions applied and a verifiable final state. It is not necessary to extend the same series out of inertia for the work to remain useful. T4 can stand as the 2026 closing baseline and serve as a reference for future audits or for new, more specific lines of research on generative search, authority, external sources, AI Overviews, AI Mode, sector benchmarks and client evidence.
Actions derived from the baseline (T4 2026)
The actions are linked exclusively to evidence observed in T4. Given that the continuity of the series is not confirmed, they are set out as maintenance recommendations and as verifiable objectives only if a decision is made to carry out a new wave.
1. /geo/ and B2B GEO Agency
Consolidate /geo/ as the canonical reference for commercial, comparative, criteria and deliverables intent.
Create or strengthen extractable blocks on how to choose a B2B GEO agency, deliverables, baseline, sources, authority, roadmap and re-measurement.
Work specifically on the prompt about alternatives to a traditional SEO agency: in T4 it obtains 0/3 mentions of Elevam.
Increase external mentions that link to /geo/ and explicitly describe the relationship between GEO, B2B, pipeline, CRM and measurement when verifiable evidence exists.
2. Broad B2B marketing category and alternatives to SEO
Strengthen content that neutrally compares traditional SEO agency, GEO/AEO, consulting, in-house team, Digital PR and hybrid models.
Document B2B cases and evidence on demand, pipeline, CRM, automation and revenue only when verifiable data exist.
Ensure that the B2B recommendations that already mention Elevam include a canonical Elevam URL; in P9 there are three mentions but only one test shows an Elevam URL.
3. /curso-de-geo/ and category authority
Keep /curso-de-geo/ as the main reference and synchronize data with Academy and the language versions.
Publish or consolidate a neutral, citable guide on what a GEO course should include and how to assess whether it is serious, with a checklist, method, measurement, deliverables and red flags.
Increase reviews, student cases and independent references that validate the course beyond the information declared by Elevam.
Extend to ChatGPT and Gemini the association that Perplexity already shows in P13 and P15.
4. Elevam Labs, HSA and entity consistency
Maintain the explicit disambiguation of Elevam Labs against ElevenLabs and strengthen direct external references to /elevam-labs/.
Consolidate /protocolo-hsa/ as the canonical definition of Human–Search–AI and review external mentions that may incorrectly expand the acronym.
Link HSA to reproducible examples, baselines and interpretation limits to reduce reinterpretations by the engine.
Review the consistency of naming and corporate data: in T4, a one-off typo «levam Labs» appears in a ChatGPT answer.
5. External evidence and canonical attribution
Prioritize external references that not only mention Elevam, but also link to the appropriate URL for the intent.
Review third-party comparisons, rankings and studies to avoid inaccurate claims, acronyms or descriptions.
Separate, in public communication, Elevam's own data, observed results, third-party validations and editorial descriptions.
In Perplexity, especially strengthen the connection between mention and Elevam's own URL: the engine increases SoM, but reduces SoS and SoA.
6. Methodological archive and close of the series
Preserve the T1–T4 datasets, public documents, working answers and per-test evidence as the project's reproducible archive.
Explicitly document the T2 Top 3 correction to prevent old figures from reappearing in future comparisons.
Maintain a T1–T4 master evolution table with stable numerators, denominators and definitions.
If T5 is not carried out, use T4 as the closing baseline for future one-off audits without presenting the series as an indefinite continuous measurement.
Improvements and aspects that will need to be reviewed in T5
T5 is not confirmed. This section does not propose continuing the series out of obligation: it sets out a simple set of checks that would make it possible to know, if the same sample is repeated, whether the progress of T4 holds, improves or regresses.
| What to check | Situation in T4 | What to validate if T5 is carried out | Why it matters |
|---|---|---|---|
| Global mention | 35/45 · 77.8% | Maintain at least 35/45 mentions. | It would confirm that the brand recovery of T4 was not a one-off. |
| Global citation | 41/45 · 91.1% | Maintain at least 41/45 answers with sources. | It makes it possible to check that the improvement in mention is not achieved at the cost of documentary support. |
| Correct URL | 33/45 · 73.3% | Maintain 33/45 or improve the number of correct attributions. | It measures whether the brand continues to be associated with the appropriate asset for each intent. |
| Top 3 | 12/45 · 26.7% | Maintain 12/45 and check whether at least one additional case is gained. | It is the competitive dimension that has advanced least during the year. |
| Hallucinations | 1/45 · 2.2% | Reduce to 0/45 without relaxing the review criterion. | It checks the reliability of proprietary concepts, acronyms and attributions. |
| GEO Agency | 9/15 mentions; 8/15 correct URL; P10 = 0/3 mentions | Maintain the group's levels and achieve at least 1 mention in P10. | P10 still represents the main absence in a relevant generic intent. |
| GEO Course · P13/P15 | 2/6 Elevam mentions | Reach at least 3/6 mentions. | It would extend the course's association from recommendations to standard and evaluation prompts. |
How to read these criteria
The values in the table are neither commercial objectives nor a prediction. They are control points: if T5 were carried out with the same prompts and rules, they would make it possible to distinguish simply between an improvement that is consolidating, a metric that is stabilizing and a regression. If there is no T5, the table remains useful as a reference for a future one-off audit.
If T4 closes the series
If it is decided to end the manual series here, T4 can be considered the closing point of the 2026 cycle. The T1–T4 datasets, the metric definitions, the T2 Top 3 methodological correction, the recorded sources and URLs and the traceability of changes must be preserved. This makes it possible to use the series as a historical reference without presenting it as an indefinite periodic measurement.
Further lines of research
The experience accumulated during T1–T4 makes it possible to propose more specific studies, with questions different from those of this baseline:
AI Overviews and AI Mode: measure when a brand appears, which sources Google uses and how attribution changes between organic results, AI Overviews and AI Mode.
Historical authority versus specialization: compare entities with a great deal of accumulated authority against more recent specialists to observe what carries more weight in generic and niche queries.
External sources and brand retrieval: study which types of media, rankings, directories, studies or third-party links coincide with a higher probability of mention and citation.
Testimonials, cases and client evidence: analyze whether publishing verifiable cases, reviews and testimonials changes the brand's presence in recommendation, trust or selection prompts.
Sector benchmarks: apply a comparable methodology to several companies in the same sector to better separate the effect of brand, category, authority and specialization.
Response stability: repeat the same sample several times per engine to measure variability and distinguish structural changes from fluctuations inherent to non-deterministic systems.
Study reference
Elevam (2026). Elevam AI Baseline — T4 2026 (Spain) (v1.0). GEO research, Elevam Labs.
Main comparative baseline: Elevam AI Baseline — T3 2026 (Spain).
Longitudinal series: T1 2026 → T2 2026 corrected v1.1 → T3 2026 → T4 2026.
T2 note: For Top 3, only the corrected value of 20.0% is used; the old figure prior to the methodological correction is not used.
The public Excel file contains the record of the 45 tests and the aggregates by group and by engine used in this document.
Dataset license: Creative Commons Attribution 4.0 International (CC BY 4.0). You may share and adapt the data, including for commercial purposes, crediting Elevam and linking to the original study. View the license terms