Prepared in application for AI Google Ads Digital Practitioner · TELUS

A Rubric-Based Framework for Search Advertising Quality Evaluation

Theoretical foundations and three applied case studies
Guatemala City, Guatemala · July 2026
This work develops a rubric for evaluating search-ad relevance and applies it across three studies: scoring nine live advertisements, architecting a Google Ads campaign, and classifying a thirty-keyword universe by intent. Every judgment is recorded as a named, scored dimension with a written justification the same standard a calibrated, multi-rater evaluation program requires of its raters.
9
Advertisements scored
3
Search queries evaluated
24
Keywords architected
30
Keywords classified
5
Rubric dimensions
§ 01 — Overview

What this work does

Search advertising sits at the intersection of two evaluation problems. One is technical: does an advertisement match the literal terms of a query. The other is interpretive: does it satisfy the need that motivated the search in the first place. Google's own Quality Score methodology treats this as a measurable diagnostic rather than a matter of opinion, built from expected click-through rate, ad relevance, and landing page experience, each rated relative to comparable advertisers.

As AI systems take on more of the work of ranking and serving advertisements, human evaluators are needed to judge, calibrate, and correct those systems' outputs against real search intent. That is the applied problem addressed by roles such as the AI Google Ads Digital Practitioner position at TELUS, which requires assessing ad relevance against search behavior, keywords, and intent, and producing feedback that can improve AI-powered advertising systems.

That responsibility is not well served by holistic, unstated judgment. It requires a rubric: a small number of named dimensions, each scored independently, each defensible to a second reviewer working from the same evidence. This work builds that rubric and demonstrates it across three applied case studies.

§ 02 — Theoretical Framework

Five concepts, one evaluation standard

Every scoring decision in the three case studies traces back to one of five established concepts, rather than to an unstated impression.

Search Intent Classification
Broder, 2002 · Jansen et al., 2008
Queries fall into three foundational categories (informational, navigational, transactional) later extended into a fourth, practitioner-defined category: commercial investigation, where a user is comparing options ahead of a purchase decision without yet being ready to transact.
Ad Relevance & Quality Score
Google, n.d.-b
A 1–10 diagnostic, not a ranking input, built from three components (expected click-through rate, ad relevance, and landing page experience) each benchmarked against other advertisers competing for the same query.
Keyword Match-Type Architecture
Google, n.d.-a
Exact, phrase, and broad match govern how closely a query must correspond to a keyword before an ad becomes eligible. Because broader types subsume narrower ones, match-type selection is fundamentally a precision–reach trade-off.
Inter-Rater Reliability
Cohen, 1960
Judgments made under a shared rubric must agree beyond what chance alone would produce. A rubric with named, independently scored dimensions (rather than a single holistic verdict) is what makes that check possible in the first place.
Decision Intelligence
Kozyrkov, 2019
The discipline of turning information into better action at scale. Each relevance verdict, structure choice, or keyword classification is treated as a decision made under uncertainty documented well enough to be reviewed, challenged, and improved.
Methodology. Three independent studies were conducted, unified by one evidentiary standard: every judgment is expressed as a named, scored dimension, accompanied by a written justification sufficient for a second reviewer to reproduce or challenge it.
§ 03 — Case Study I

Ad Quality Evaluation

Nine live advertisements evaluated across three real search queries, using a five-dimension rubric scored on the same 1–10 scale as Google's Quality Score diagnostic.

DimensionWhat it measures
Keyword AlignmentCorrespondence between the query's literal terms and the advertisement's copy
Intent MatchWhether the advertisement addresses the underlying reason for the search, not just its literal wording
Ad RelevanceClarity, specificity, and quality of the advertisement copy itself
Landing Page FitCorrespondence between the advertisement's promise and the page it links to
Trust SignalsCredibility markers (social proof, domain authority, specificity) supporting the advertisement's claims
Table 1. Five-dimension ad relevance rubric. Each dimension is scored 1–10.

Query 1 — "affordable CRM software for small business"

Commercial investigation intent · Three advertisements evaluated
HubSpot CRM
9.4
Salesforce
5.8
bestbusinesstools.net
3.0
High relevanceMedium relevanceLow relevance

HubSpot's advertisement addressed both query modifiers (affordable and small business) directly in its headline through a free-tier offer, reinforced by a social-proof claim specific to the small-business segment. Salesforce scored well on trust signals through brand recognition, but its enterprise-audience framing ("Enterprise-level features," "Top Companies") contradicted the price-sensitive intent behind the query. The comparison-site advertisement never mentioned the product category and offered minimal trust signals.

Query 2 — "online Spanish courses for adults"

Transactional intent · Two advertisements evaluated
Babbel
9.3
spanishschool.edu
6.1

Babbel addressed both the audience (adults) and modality (online) while reframing the adult learner's core constraint (limited time) as a feature ("15 Min/Day"). The second advertisement was competently written but omitted both terms, leading to a specific, documented recommendation: add adults and online to the headline.

Figure 1. Overall relevance scores by evaluated advertisement, reflecting the mean of five rubric dimensions per advertisement, rounded to one decimal place.

Across both queries, the pattern held: the lowest-scoring advertisements were not poorly written in isolation they were well written for a different audience or a different intent than the one the query expressed. That distinction between copy quality and intent match is exactly what a single holistic relevance judgment fails to separate, and exactly what a five-dimension rubric is designed to expose.

§ 04 — Case Study II

Search Campaign Blueprint

A complete Google Ads account architecture for LexGuatemala, an immigration and corporate law firm, organized around one key performance indicator: cost per consultation lead.

The architecture comprises four ad groups, each aligned to a distinct search intent, together totaling 24 keywords and 8 responsive search ad assets.

Ad GroupSearch IntentRepresentative KeywordMatch Type
AG-01 · Visa & Immigration ConsultationTransactional[abogado de visas Guatemala]Exact
AG-02 · Corporate & Business LawCommercial Investigation[abogado corporativo Guatemala]Exact
AG-03 · US Diaspora (English)Transactional[Guatemala immigration lawyer]Exact
AG-04 · Brand & Competitor DefenseNavigational[LexGuatemala]Exact
Table 2. Ad group architecture, LexGuatemala campaign. All four ad groups maintain independent negative-keyword lists to prevent cross-group query overlap.

AG-03 replicates the transactional strategy of AG-01 for an English-speaking diaspora audience, with independently written copy rather than direct translation (for example, "Guatemala Immigration Lawyer · Online Consultations · Trusted Legal Experts · Get Help Now") and its own negative-keyword list, built to exclude U.S.-domestic immigration queries that would otherwise overlap structurally with the target set. This bilingual design is direct, demonstrated evidence of working proficiency in both English and Spanish within a single applied campaign.

Quality Score diagnostic — Ad Group 01

Ad Relevance
Below averageAverageAbove average
Expected CTR
Landing Page Experience
Figure 2. Quality Score component diagnostic, Ad Group 01. Ratings are relative to comparable advertisers competing for the same queries over a rolling 90-day window.

Expected click-through rate and ad relevance both rated above average; landing page experience rated only average the diagnosed constraint on the group's overall score. The corresponding recommendation was landing-page-specific: incorporate the primary keyword into the page's main heading, reduce load time, and surface a trust signal above the fold, rather than adjusting advertisement copy that was already performing above benchmark.

§ 05 — Case Study III

Keyword Intent Mapping

A thirty-keyword universe for an online fitness-coaching service in the Guatemalan market, classified using a four-part intent taxonomy and prioritized by proximity to conversion rather than by raw search volume.

Intent CategoryDefining CharacteristicBudget Prioritization
InformationalSeeking knowledge; not yet evaluating providersLow — nurture rather than convert
Commercial InvestigationComparing providers ahead of a decisionMedium — converts within 2–3 days of first contact
TransactionalReady to actHigh — prioritized for budget allocation
NavigationalSearching for a known brand by nameProtect, not conquest
Table 3. Intent classification framework applied to the keyword universe.

Five keywords were classified as top-priority transactional queries for example "entrenador personal online Guatemala" and "rutina personalizada online precio," exact- and phrase-matched respectively while informational queries such as "cómo empezar a hacer ejercicio en casa" were retained for reach but assigned the lowest bidding priority.

Share of search volume by intent stage

Informational
100%
Commercial
72%
Transactional
44%
Navigational
18%
Figure 3. Values represent the proportion of total classified search volume retained at each successive intent stage.

Volume falls from 100% at the informational stage to 18% at the navigational stage, with the steepest single decline occurring between commercial investigation and transactional. The budget implication is direct: the highest-value inflection point in the funnel is not where search volume is largest, but where it drops fastest.

§ 06 — Alignment

How this maps to the role

The three case studies were conducted independently, for three different markets, but share one methodological commitment: every evaluative judgment is expressed as a named, scored, and justified decision rather than an unstated impression.

RequirementEvidenceSource
Evaluate ad relevance against search queries, keywords, and intentNine advertisements scored across three queries on one five-dimension rubric; two advertisements on the same query scored 9.4 and 3.0, with the gap fully documentedCase Study I
Search intentFour-part intent framework applied to a 30-keyword universe with priority rankingCase Study III
Keyword targeting and match types24 keywords architected across Exact, Phrase, and Broad match in a live campaign; 30 more classified by intent and match typeCase Studies II–III
Quality ScorePer-ad-group diagnostic (expected CTR, ad relevance, landing page experience) with a specific optimization action attachedCase Study II
Landing page experienceScored as one of five rubric dimensions on every evaluated advertisement; rated a second time within the Quality Score diagnosticCase Studies I–II
Campaign optimization; conversion-focused advertisingFull four-ad-group architecture built around one stated KPI cost per consultation leadCase Study II
Apply guidelines consistently; attention to detailIdentical rubric applied without exception across all nine evaluated advertisementsCase Study I
Provide clear, actionable feedbackEvery evaluation closes with a specific, written recommendation, down to the exact headline revision proposedCase Study I
Table 4. Alignment between position requirements and demonstrated evidence.

A note on limitations. All three evaluations were conducted by a single rater. Cohen's coefficient of agreement (and the broader discipline of inter-rater reliability it founded) exists precisely to quantify whether a second, independent evaluator would reach the same verdict from the same evidence, which single-rater work cannot itself demonstrate. The rubric in Case Study I and the taxonomy in Case Study III were both built with named, discrete dimensions specifically so a second rater could apply them and a coefficient of agreement could, in principle, be computed. Calibration in a multi-rater setting is the natural next application of a framework built this way, not a departure from it.

§ 07 — Background

Professional background

Research-driven digital marketing background combining Information Technology, Aeronautical Administration, and doctoral research in high-performance management.

Digital Support
Maxpower / Nutrisano
2019 – 2025
  • Conducted market and customer research to identify audience needs and develop buyer personas
  • Designed content strategies for social media and email marketing campaigns; built landing pages and lead magnets for lead generation
  • Applied SEO, content marketing, and conversion-focused communication principles to improve engagement
  • Built small web-based projects with HTML, CSS, and JavaScript to support workflow and content organization
  • Translated technical and scientific information into clear, audience-focused content across web, email, and digital channels
Community Programs Coordinator
Unidad para la Prevención Comunitaria de la Violencia
2013 – 2019
  • Coordinated community programs and multi-stakeholder initiatives, ensuring structured execution and communication
  • Tracked KPIs, timelines, and program execution against defined objectives and deliverables
  • Supported analysis of program performance and continuous improvement
PhD Candidate, High-Performance Management
Universidad Galileo
2023 – Present
  • Research on circadian rhythms, human performance, productivity, and decision-making
  • Literature review, synthesis, conceptual analysis, and structured research documentation
DJPR9.0 — Decision Journey Process Reduction
Applied Intelligence Portfolio
2026
  • Built the evaluation framework presented on this page: search intent analysis, ad relevance scoring, and landing page experience assessment across three documented case studies
Education
PhD Candidate, High-Performance Management
Universidad Galileo
2023 – Present
MSc, Traffic & Digital Marketing (SEO, SEM & Content)
IEBS / UCAM
2023 – Present
Google Ads Search Certification
Google Skillshop
In progress · Expected 2026
MSc, Quality Management
Universidad Galileo
2021 – 2022
Postgraduate Degree, Quality Planning & Assurance
Universidad Galileo
2020 – 2021
BSc, Information Technology & Aeronautical Administration
Universidad Galileo
2016 – 2020
Skills
Digital Marketing
SEO, SEM & content strategy · Landing page optimization · Email marketing · Audience segmentation
Analytics & Performance
Engagement tracking · KPI monitoring · Data-informed decision support · Trend analysis
Communication Systems
Content workflow design · Documentation systems · Cross-functional coordination
Tools
Microsoft Office · Google Workspace · WordPress / Notion · HTML/CSS/JS · Email marketing systems
Languages
SpanishNative
EnglishB2 · Technical & Academic
§ 08 — References

Sources cited

01
Broder, A. (2002). A taxonomy of web search. ACM SIGIR Forum, 36(2), 3–10. doi.org/10.1145/792550.792552
02
Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. doi.org/10.1177/001316446002000104
03
Google. (n.d.-a). About keyword matching options. Google Ads Help. Retrieved July 27, 2026, from support.google.com/google-ads/answer/7478529
04
Google. (n.d.-b). About Quality Score for Search campaigns. Google Ads Help. Retrieved July 27, 2026, from support.google.com/google-ads/answer/6167118
05
Jansen, B. J., Booth, D. L., & Spink, A. (2008). Determining the informational, navigational, and transactional intent of Web queries. Information Processing & Management, 44(3), 1251–1266. doi.org/10.1016/j.ipm.2007.07.015
06
Kozyrkov, C. (2019, August 2). Introduction to decision intelligence: A new discipline for leadership in the AI era. Medium. medium.com/.../introduction-to-decision-intelligence