Search advertising sits at the intersection of two evaluation problems. One is technical: does an advertisement match the literal terms of a query. The other is interpretive: does it satisfy the need that motivated the search in the first place. Google's own Quality Score methodology treats this as a measurable diagnostic rather than a matter of opinion, built from expected click-through rate, ad relevance, and landing page experience, each rated relative to comparable advertisers.
As AI systems take on more of the work of ranking and serving advertisements, human evaluators are needed to judge, calibrate, and correct those systems' outputs against real search intent. That is the applied problem addressed by roles such as the AI Google Ads Digital Practitioner position at TELUS, which requires assessing ad relevance against search behavior, keywords, and intent, and producing feedback that can improve AI-powered advertising systems.
That responsibility is not well served by holistic, unstated judgment. It requires a rubric: a small number of named dimensions, each scored independently, each defensible to a second reviewer working from the same evidence. This work builds that rubric and demonstrates it across three applied case studies.
Every scoring decision in the three case studies traces back to one of five established concepts, rather than to an unstated impression.
Nine live advertisements evaluated across three real search queries, using a five-dimension rubric scored on the same 1–10 scale as Google's Quality Score diagnostic.
| Dimension | What it measures |
|---|---|
| Keyword Alignment | Correspondence between the query's literal terms and the advertisement's copy |
| Intent Match | Whether the advertisement addresses the underlying reason for the search, not just its literal wording |
| Ad Relevance | Clarity, specificity, and quality of the advertisement copy itself |
| Landing Page Fit | Correspondence between the advertisement's promise and the page it links to |
| Trust Signals | Credibility markers (social proof, domain authority, specificity) supporting the advertisement's claims |
HubSpot's advertisement addressed both query modifiers (affordable and small business) directly in its headline through a free-tier offer, reinforced by a social-proof claim specific to the small-business segment. Salesforce scored well on trust signals through brand recognition, but its enterprise-audience framing ("Enterprise-level features," "Top Companies") contradicted the price-sensitive intent behind the query. The comparison-site advertisement never mentioned the product category and offered minimal trust signals.
Babbel addressed both the audience (adults) and modality (online) while reframing the adult learner's core constraint (limited time) as a feature ("15 Min/Day"). The second advertisement was competently written but omitted both terms, leading to a specific, documented recommendation: add adults and online to the headline.
Across both queries, the pattern held: the lowest-scoring advertisements were not poorly written in isolation they were well written for a different audience or a different intent than the one the query expressed. That distinction between copy quality and intent match is exactly what a single holistic relevance judgment fails to separate, and exactly what a five-dimension rubric is designed to expose.
A complete Google Ads account architecture for LexGuatemala, an immigration and corporate law firm, organized around one key performance indicator: cost per consultation lead.
The architecture comprises four ad groups, each aligned to a distinct search intent, together totaling 24 keywords and 8 responsive search ad assets.
| Ad Group | Search Intent | Representative Keyword | Match Type |
|---|---|---|---|
| AG-01 · Visa & Immigration Consultation | Transactional | [abogado de visas Guatemala] | Exact |
| AG-02 · Corporate & Business Law | Commercial Investigation | [abogado corporativo Guatemala] | Exact |
| AG-03 · US Diaspora (English) | Transactional | [Guatemala immigration lawyer] | Exact |
| AG-04 · Brand & Competitor Defense | Navigational | [LexGuatemala] | Exact |
AG-03 replicates the transactional strategy of AG-01 for an English-speaking diaspora audience, with independently written copy rather than direct translation (for example, "Guatemala Immigration Lawyer · Online Consultations · Trusted Legal Experts · Get Help Now") and its own negative-keyword list, built to exclude U.S.-domestic immigration queries that would otherwise overlap structurally with the target set. This bilingual design is direct, demonstrated evidence of working proficiency in both English and Spanish within a single applied campaign.
Expected click-through rate and ad relevance both rated above average; landing page experience rated only average the diagnosed constraint on the group's overall score. The corresponding recommendation was landing-page-specific: incorporate the primary keyword into the page's main heading, reduce load time, and surface a trust signal above the fold, rather than adjusting advertisement copy that was already performing above benchmark.
A thirty-keyword universe for an online fitness-coaching service in the Guatemalan market, classified using a four-part intent taxonomy and prioritized by proximity to conversion rather than by raw search volume.
| Intent Category | Defining Characteristic | Budget Prioritization |
|---|---|---|
| Informational | Seeking knowledge; not yet evaluating providers | Low — nurture rather than convert |
| Commercial Investigation | Comparing providers ahead of a decision | Medium — converts within 2–3 days of first contact |
| Transactional | Ready to act | High — prioritized for budget allocation |
| Navigational | Searching for a known brand by name | Protect, not conquest |
Five keywords were classified as top-priority transactional queries for example "entrenador personal online Guatemala" and "rutina personalizada online precio," exact- and phrase-matched respectively while informational queries such as "cómo empezar a hacer ejercicio en casa" were retained for reach but assigned the lowest bidding priority.
Volume falls from 100% at the informational stage to 18% at the navigational stage, with the steepest single decline occurring between commercial investigation and transactional. The budget implication is direct: the highest-value inflection point in the funnel is not where search volume is largest, but where it drops fastest.
The three case studies were conducted independently, for three different markets, but share one methodological commitment: every evaluative judgment is expressed as a named, scored, and justified decision rather than an unstated impression.
| Requirement | Evidence | Source |
|---|---|---|
| Evaluate ad relevance against search queries, keywords, and intent | Nine advertisements scored across three queries on one five-dimension rubric; two advertisements on the same query scored 9.4 and 3.0, with the gap fully documented | Case Study I |
| Search intent | Four-part intent framework applied to a 30-keyword universe with priority ranking | Case Study III |
| Keyword targeting and match types | 24 keywords architected across Exact, Phrase, and Broad match in a live campaign; 30 more classified by intent and match type | Case Studies II–III |
| Quality Score | Per-ad-group diagnostic (expected CTR, ad relevance, landing page experience) with a specific optimization action attached | Case Study II |
| Landing page experience | Scored as one of five rubric dimensions on every evaluated advertisement; rated a second time within the Quality Score diagnostic | Case Studies I–II |
| Campaign optimization; conversion-focused advertising | Full four-ad-group architecture built around one stated KPI cost per consultation lead | Case Study II |
| Apply guidelines consistently; attention to detail | Identical rubric applied without exception across all nine evaluated advertisements | Case Study I |
| Provide clear, actionable feedback | Every evaluation closes with a specific, written recommendation, down to the exact headline revision proposed | Case Study I |
A note on limitations. All three evaluations were conducted by a single rater. Cohen's coefficient of agreement (and the broader discipline of inter-rater reliability it founded) exists precisely to quantify whether a second, independent evaluator would reach the same verdict from the same evidence, which single-rater work cannot itself demonstrate. The rubric in Case Study I and the taxonomy in Case Study III were both built with named, discrete dimensions specifically so a second rater could apply them and a coefficient of agreement could, in principle, be computed. Calibration in a multi-rater setting is the natural next application of a framework built this way, not a departure from it.
Research-driven digital marketing background combining Information Technology, Aeronautical Administration, and doctoral research in high-performance management.