Evaluate agencies, consultancies, and other marketing partners with a structured, evidence-based process using the AAMA Agency Evaluation Scorecard. This editable Excel workbook helps organizations compare strategic capability, creative judgment, media expertise, research, measurement, team quality, project management, commercial value, and organizational fit without allowing presentation polish or personal chemistry to dominate the selection.
Agency selection involves more than choosing the team with the most impressive pitch. A strong evaluation process should establish the criteria before proposals are scored, assign greater weight to the factors that matter most to the assignment, document the evidence behind each rating, identify mandatory requirements, and preserve the reasoning behind the final decision.
Download the Agency Evaluation Scorecard
The AAMA Agency Evaluation Scorecard is provided as an editable Microsoft Excel workbook with weighted calculations, scoring controls, minimum-requirement checks, category rollups, rankings, and finalist-review tools already built in. Teams can customize the criteria and weights before evaluation begins, then score up to four agencies using the included structure.
The workbook is intentionally neutral and can be adapted to advertising agencies, media agencies, creative agencies, digital agencies, research firms, public-relations agencies, marketing consultancies, production partners, freelancers, or other external professional-service providers.
What the Scorecard Helps You Evaluate
The workbook is designed to make agency evaluation more consistent and transparent. It separates the criteria used to judge capability from the evidence used to support the judgment and from the additional due diligence that may affect the final selection.
The default model uses 15 criteria totaling 100 weighted points. Organizations can modify those criteria, descriptions, weights, and minimum requirements to reflect the actual assignment rather than using a generic selection process for every engagement.
What’s Included
The workbook contains five working sheets:
- Instructions
- Criteria & Weights
- Evaluation
- Agency Summary
- Finalist Review
Together, these sheets create a complete evaluation process from criteria design through scoring, comparison, due diligence, and final recommendation.
Customize the Criteria Before Scoring
The Criteria & Weights sheet should be reviewed before proposals, presentations, or finalist meetings are scored. This helps prevent teams from changing the rules after they have already developed preferences for particular agencies.
Each criterion includes a category, definition, weight, minimum-requirement designation, and examples of evidence that should be reviewed. The default weighting totals 100%, and the workbook highlights whether the weighting remains properly normalized after changes are made.
Default Evaluation Criteria
The scorecard begins with 15 professional evaluation criteria covering the areas most likely to affect agency performance.
These include:
- Strategic Understanding
- Strategic Approach
- Relevant Experience
- Creative Quality & Judgment
- Media / Channel Capability
- Research & Audience Insight
- Measurement & Analytics
- Proposed Team
- Working Relationship & Communication
- Project Management & Delivery
- Technology & Operational Fit
- Commercial Value
- Transparency & Scope Clarity
- References & Evidence
- Organizational / Cultural Fit
These defaults provide a practical starting point rather than a required universal standard. An organization selecting a research firm, media agency, production company, or specialized technology partner may reasonably use a different weighting structure.
Weight What Actually Matters
Not every criterion should influence the decision equally. A strategically complex assignment may place substantial weight on strategic understanding, research, and measurement, while a production-heavy engagement may require more emphasis on execution capability, technical requirements, staffing, and delivery.
Weighting forces the organization to make those priorities visible before scoring. It also reduces the risk that a relatively minor strength, such as an entertaining presentation or attractive speculative creative, unintentionally outweighs a more important weakness.
Keep the Total Weight at 100%
The workbook automatically totals the criterion weights. The total should equal 100% before formal evaluation begins.
If the weighting does not equal 100%, the evaluation is no longer normalized to the intended 100-point scale. Teams should resolve the allocation before scoring rather than adjusting the mathematics after agency results are known.
Use the 1-to-5 Scoring Scale
Each criterion is scored using a five-point scale.
1: Materially Weak indicates that the response does not meet the requirement or creates a substantial concern.
2: Below Expectations indicates that the agency partially addresses the requirement but significant gaps remain.
3: Meets Expectations represents an acceptable and credible response supported by sufficient evidence.
4: Strong indicates meaningful strengths beyond the basic requirement.
5: Exceptional represents unusually strong capability, fit, evidence, or value relative to the assignment.
A score of 3 should represent genuine adequacy rather than failure. If evaluators treat every competent response as a 4 or 5, the scale loses much of its ability to differentiate agencies.
Calculate Weighted Scores
The workbook converts each criterion score into weighted points using its assigned percentage.
A criterion weighted at 10% and scored 5 out of 5 contributes 10 points to the final 100-point score. The same criterion scored 3 contributes 6 points.
This allows strategically important criteria to influence the final result more than lower-priority factors while preserving one common comparison scale.
Require Evidence for the Score
A numeric score becomes more useful when another reviewer can understand why it was assigned. The Evaluation sheet therefore includes a dedicated Evidence / Rationale field beside every rating.
Useful evidence may include proposals, strategic recommendations, working sessions, case studies, references, staffing plans, media plans, research examples, dashboards, project-management materials, budgets, demonstrations, or direct answers from the agency team.
The objective is not to produce excessive documentation. It is to prevent important ratings from becoming unsupported statements such as “we liked them” or “they seemed smart.”
Separate Reputation From Evidence
Agency reputation can be relevant, but it should not substitute for evaluating the actual team and proposal. A well-known agency may assign a relatively junior team to the account, while a less familiar firm may provide stronger senior involvement and more relevant experience.
The scorecard focuses the evaluation on what the agency proposes to do, who will perform the work, what evidence supports its capabilities, and how well those capabilities fit the assignment.
Evaluate Strategic Understanding
Strategic Understanding evaluates whether the agency understands the business problem, audience, market context, and assignment rather than simply responding to the tactics listed in the request.
A strong agency should be able to explain what problem it believes the work needs to solve and how its proposed approach connects to that problem. Repeating terminology from the request without demonstrating deeper understanding should not receive the same score as a thoughtful strategic interpretation.
Evaluate the Strategic Approach
The Strategic Approach criterion examines whether the proposed method connects objectives, audience insights, positioning, channels, creative work, execution, and measurement into one coherent plan.
An agency may have strong individual capabilities but still present a fragmented approach in which media, creative, research, and reporting operate independently. The evaluation should consider whether the proposed system makes strategic sense as a whole.
Consider Relevant Experience Carefully
Relevant experience can reduce execution risk, but organizations should avoid treating industry familiarity as an automatic requirement when the underlying problem may be transferable across categories.
The scorecard therefore evaluates experience with similar problems, audiences, channels, organizational complexity, or market conditions rather than asking only whether the agency has worked with a nearly identical company.
An agency with extensive category experience can still be weak if its examples are outdated, superficial, or dependent on a completely different team.
Evaluate Creative Judgment, Not Style Alone
Creative evaluation should consider whether the agency can solve communication problems effectively, not simply whether reviewers personally like the visual style of the portfolio.
A strong creative partner should demonstrate judgment about audience, message, channel, context, accessibility, production, and brand strategy. Different assignments may require different aesthetics, so a portfolio should be evaluated for thinking and craft rather than resemblance to one preferred look.
Evaluate Media & Channel Capability
Media expertise should be assessed according to the actual channels required by the assignment.
The agency should demonstrate appropriate depth in planning, buying, optimization, trafficking, audience strategy, platform requirements, measurement, and coordination where those functions are relevant. Teams should also determine which capabilities are performed directly and which depend on subcontractors, affiliated companies, or other outside partners.
Evaluate Research & Audience Insight
Research capability matters when the assignment requires the agency to understand customers, markets, competitors, behavior, or changing conditions.
The agency should demonstrate how it distinguishes evidence from assumption and which methods it uses when existing information is insufficient. A polished persona or audience presentation should not receive a high score simply because it appears detailed if the underlying conclusions cannot be traced to credible evidence.
Evaluate Measurement & Analytics
The Measurement & Analytics criterion examines whether the agency can define success, calculate important metrics consistently, understand attribution limitations, work with available data sources, and connect campaign performance with meaningful organizational outcomes.
This criterion is designated as a default minimum requirement because weak measurement can make the work difficult to evaluate even when the creative or media execution appears strong.
Organizations can change that designation when the nature of the engagement makes another standard more appropriate.
Evaluate the Proposed Team
Agency selection should consider the people who will actually perform and manage the work.
The Proposed Team criterion evaluates relevant experience, role clarity, senior involvement, staffing levels, availability, continuity, and whether the team appears appropriate for the complexity of the assignment.
Pitch-team seniority should not be confused with account-team seniority. Organizations should confirm which individuals are committed to the engagement and how much of their time will realistically be available.
Evaluate Working Relationship Without Turning It Into a Popularity Contest
Communication and collaboration matter in a professional-services engagement. Teams need to exchange information, challenge assumptions, resolve problems, and make decisions efficiently.
The scorecard includes Working Relationship & Communication while intentionally assigning it less weight than several strategic and operational capabilities. This helps organizations consider practical compatibility without allowing subjective chemistry to overwhelm stronger evidence.
Evaluate Project Management & Delivery
Creative and strategic capability loses value when the agency cannot manage schedules, dependencies, approvals, scope, quality control, or production reliably.
The Project Management & Delivery criterion examines the agency’s operating process and its ability to manage the work from kickoff through implementation. Useful evidence can include project plans, responsibility structures, QA processes, escalation rules, scope-management procedures, or references from previous clients.
Evaluate Technology & Operational Fit
An agency may need to operate within existing advertising platforms, analytics systems, CRM environments, content-management systems, security requirements, procurement processes, or other organizational infrastructure.
The Technology & Operational Fit criterion helps teams determine whether the proposed partner can work within those conditions without introducing unnecessary risk or complexity.
This can be especially important for larger organizations, regulated environments, complex ecommerce operations, or engagements involving customer data.
Evaluate Commercial Value
Commercial evaluation should consider more than which agency submits the lowest fee.
The Commercial Value criterion asks whether the proposed investment, resource allocation, staffing, and cost structure are reasonable for the work and expected outcomes. A higher-priced agency can represent better value when it provides stronger expertise, senior involvement, lower execution risk, or greater capability.
The evaluation should also consider whether important third-party expenses are included or excluded from the stated price.
Evaluate Scope Transparency
A proposal becomes difficult to compare when important assumptions remain hidden.
The Transparency & Scope Clarity criterion examines whether the agency clearly identifies inclusions, exclusions, revision limits, external costs, ownership, responsibilities, dependencies, and circumstances that could change price or schedule.
This criterion is designated as a default minimum requirement because unclear scope can create significant problems after selection even when the proposed strategy is strong.
Check References & Evidence
Agency claims should be validated where the importance of the engagement warrants it.
Reference checks can help determine whether the proposed team communicates effectively, meets commitments, handles problems responsibly, maintains senior involvement, manages budgets appropriately, and behaves similarly after winning the business.
The scorecard treats references as one evidence source rather than a ceremonial final step.
Consider Organizational Fit
Organizations and agencies can struggle even when both are individually capable if their operating styles are fundamentally incompatible.
The Organizational / Cultural Fit criterion considers pace, decision style, working norms, communication expectations, formality, and other practical characteristics that may affect the relationship. This criterion should be used carefully and supported by observable working behavior rather than vague judgments about whether a team feels like “our kind of people.”
Use Minimum Requirements
Some criteria may represent requirements the selected agency must satisfy regardless of its overall weighted score.
The default scorecard identifies Strategic Understanding, Strategic Approach, Measurement & Analytics, Proposed Team, Project Management & Delivery, and Transparency & Scope Clarity as minimum requirements. Organizations can change these designations before evaluation begins.
The Evaluation sheet includes a separate Meets Minimum field so a serious requirement failure does not disappear inside a strong overall average.
Review Gate Failures Separately
The Agency Summary automatically counts failures among criteria marked as minimum requirements.
An agency with one or more failures receives a REVIEW gate status even if its weighted score is high. This does not automatically mean the agency must be rejected, but it makes the issue visible before a selection decision is made.
A high aggregate score should not be used to average away a critical weakness.
Compare Agencies on a 100-Point Scale
The Agency Summary automatically calculates the total weighted score for each agency and ranks the four default candidates.
The score is intended to make structured comparison easier, not to create the illusion that agency selection can be reduced to mathematics alone. A difference of 82.4 versus 81.8 should not be treated as meaningful unless the underlying evidence supports a meaningful distinction.
The ranking should always be reviewed alongside minimum requirements, references, pricing, risks, team availability, and other due diligence.
Review Performance by Category
The Agency Summary also groups weighted scores by broader capability category.
These categories include:
- Business & Strategy
- Capability
- Creative
- Media & Distribution
- Research
- Measurement
- Team
- Operations
- Commercial
- Validation
- Fit
Category comparison can reveal differences that the overall total hides. Two agencies may produce similar total scores while one is considerably stronger strategically and the other is stronger operationally.
Use the Finalist Review
The Finalist Review separates formal scoring from additional due diligence that may affect the award decision.
Teams can document:
- Proposed fee or budget
- Pricing structure
- Reference findings
- Conflicts or independence concerns
- Procurement or security issues
- Material risks
- Finalist status
This avoids forcing every selection issue into the same numeric model.
Review Pricing Outside the Score Too
Commercial Value belongs in the weighted evaluation, but the actual proposed fee and pricing structure should still remain visible during final selection.
Two agencies may receive similar commercial scores while proposing materially different investments, staffing models, payment structures, or third-party costs. The Finalist Review gives decision-makers a separate place to consider those facts alongside the capability scores.
Document Conflicts & Independence
Some engagements may require consideration of competitive conflicts, client conflicts, media relationships, vendor incentives, ownership relationships, commissions, referral structures, or other issues that could affect independence.
These issues may not fit cleanly into a performance score. The Finalist Review therefore provides a dedicated place to document them before award.
Organizations should apply their own conflict and procurement policies where applicable.
Complete Reference Checks
The workbook includes a due-diligence log for documenting specific reference questions, findings, and their implications.
Reference conversations become more useful when they ask about matters that can influence the current assignment. These may include senior-team continuity, scope control, budget management, responsiveness, reporting quality, problem resolution, turnover, strategic strength, or the accuracy of claims made during the selection process.
Reconcile Major Evaluator Differences
Different evaluators will not always score an agency the same way. Those differences can be useful when they reveal that reviewers interpreted the requirement differently or noticed different evidence.
A multi-evaluator process should identify material scoring differences and discuss the reasoning behind them rather than simply averaging every score automatically. The purpose is to improve judgment, not hide disagreement inside a decimal.
For a formal process, teams can make a separate copy of the Evaluation sheet for each evaluator or duplicate the evaluation blocks before scoring begins.
Avoid Groupthink
Evaluators may benefit from completing their initial ratings independently before the group discusses the agencies.
If everyone scores together from the beginning, early opinions from senior or influential participants can unintentionally shape the ratings of other evaluators.
Independent scoring followed by structured reconciliation can preserve more information about how different reviewers interpreted the proposal.
Avoid Pitch Theater
An entertaining presentation can influence perception far beyond its actual relevance to agency performance.
Strong speakers, polished speculative creative, celebrity leadership, impressive offices, or high production value may create confidence without demonstrating how well the agency will perform the day-to-day work.
The scorecard is designed to bring the discussion back to the assignment, evidence, team, operating model, measurement, value, and delivery capability.
Evaluate the Actual Team
The people presenting the pitch may not be the people assigned to the business after selection.
Organizations should identify which proposed team members will remain involved, their roles, their expected level of participation, and whether substitutions require approval.
This is particularly important when senior leadership plays a large role during business development but routine delivery is delegated after the account is won.
Preserve the Selection Record
The completed scorecard should be stored with the request for proposal or brief, agency submissions, pricing, reference findings, procurement records, approval decisions, and final scope where appropriate.
Preserving the evaluation makes future agency reviews easier. It allows the organization to compare what it expected during selection with what the agency actually delivered over time.
Use the Scorecard for Existing Agency Reviews
The framework can also be adapted for periodic reviews of an incumbent agency.
Criteria can be modified to evaluate strategic contribution, creative quality, media performance, measurement, communication, budget management, delivery, innovation, team continuity, scope control, and other aspects of the ongoing relationship.
For performance reviews, ratings should rely on actual delivery evidence rather than the promises made during the original selection process.
Customize the Scorecard
The criteria and default weights are intended as a starting point. Organizations should adjust them to reflect the specific assignment, its risk, and the capabilities genuinely required for success.
A specialized creative project should not necessarily be evaluated using the same weighting as a full-service agency-of-record search. Likewise, a research engagement, media assignment, website project, or production relationship may require substantially different criteria.
The scoring model becomes most useful when it reflects the decision being made.
Use It With the AAMA Professional & Agency Resources
Use the scorecard alongside the AAMA Client Discovery Questionnaire, AAMA Project Scope Template, AAMA Marketing Budget Template, AAMA Competitive Analysis Worksheet, Campaign Measurement Framework, AAMA Campaign Performance Report Template, and AAMA Post-Campaign Analysis Template.
Together, these resources provide a practical system for understanding the assignment, defining the work, evaluating potential partners, establishing the scope, managing the engagement, measuring results, and documenting what the organization learns from the relationship.

