Vendor Evaluation Scorecard Template
- Contributor
- 3 days ago
- 5 min read
The most dangerous scorecard I ever helped build had Vendor A at 7.7 and Vendor C at 7.3, and the room treated that 0.4 as a decision. It wasn't. Vendor A scored a 9 on features and a 6 on cost; Vendor C was an 8 and a 9. The weighted average had quietly blended away the one fact that actually mattered — that C was materially cheaper and barely weaker — and handed us a number precise enough to stop the argument we should have been having. We picked A. The features we paid the premium for went half-used.
A vendor evaluation scorecard converts qualitative impressions into comparable numbers, and that is exactly why it's dangerous: a number looks like an answer. Its honest job is narrower — to anchor the conversation to specific criteria and specific evidence so the decision doesn't drift to whoever ran the slickest demo or talked loudest in the room. It is structured input to a judgment, never a substitute for one. This guide is a working template plus how to use it without it becoming an illusion of objectivity.
The Working Template
# Vendor Evaluation: [Category]
Date: YYYY-MM-DD
Evaluators: [Names]
## Criteria and Weights
| Criterion | Weight | Why it matters |
|-----------|--------|----------------|
| Must-have features | 30% | Non-negotiable for our use case |
| Integration fit | 20% | Reduces TCO and time-to-value |
| 3-year cost | 20% | Budget reality |
| Support quality | 10% | Affects post-purchase experience |
| Vendor stability | 10% | Risk of vendor going under |
| Implementation timeline | 10% | Aligned with our roadmap |
## Vendor Scores
| Criterion | Vendor A | Vendor B | Vendor C |
|-----------|----------|----------|----------|
| Must-have features (30%) | 9 | 7 | 8 |
| Integration fit (20%) | 8 | 9 | 6 |
| 3-year cost (20%) | 6 | 8 | 9 |
| Support quality (10%) | 8 | 7 | 6 |
| Vendor stability (10%) | 9 | 8 | 5 |
| Implementation (10%) | 7 | 8 | 9 |
| **Weighted total** | **7.7** | **7.7** | **7.3** |
## Notes Per Vendor
### Vendor A
- Strengths: [specifics]
- Weaknesses: [specifics]
- Risks: [specifics]
- References checked: [names, summaries]
[Repeat for each vendor]
## Recommendation
[Specific vendor, with reasoning that doesn't rely solely on the numbers]
## Open Questions
[Anything that would change the recommendation]
The scorecard isn't the decision. It's structured input to the decision conversation.
Picking Criteria
Criteria should be specific to your context, not generic.
Bad criteria (too generic): "Good support," "Easy to use," "Scalable."
Good criteria (specific): "Support response time meets our needs (24h for non-critical)," "Time-to-deploy for our use case under 30 days," "Demonstrated scalability to 10M events/day."
The criteria should be the things that actually matter for your specific decision. Five to ten criteria is workable; more than that and you're including details that should be sub-criteria.
Picking Weights
Weights reflect how much each criterion matters relative to others.
A working approach:
Brainstorm criteria
Have each evaluator independently assign weights (summing to 100%)
Compare; discuss disagreements; align
The discussion is more valuable than the precise numbers. Two evaluators disagreeing about whether "price" is 15% or 25% is a useful conversation about what matters.
Avoid "every criterion gets equal weight." That defeats the purpose. Some things matter more.
Scoring
The 1-10 scale is conventional but rough.
A working calibration:
9-10: exceptional; clear advantage
7-8: good; meets needs well
5-6: acceptable; works
3-4: weak; significant gaps
1-2: unacceptable
Anchor scores to evidence:
"9 because all 7 must-haves verified in demo, vendor showed exact use case"
"6 because 4 must-haves work, 3 require workarounds"
"3 because key feature not available; vendor offered roadmap commitment without date"
Without anchoring, scores reflect mood. With it, they reflect evidence.
Multiple Evaluators
Each evaluator scores independently. Aggregate after.
When evaluators agree: trust the scores.
When they disagree: the conversation about disagreement is where the value is. Often disagreements surface:
Different understanding of what the criterion means
Different exposure to the vendor (one saw a demo the other didn't)
Different priorities (one evaluator weighs ease-of-use higher than scalability)
The disagreement is information. Don't average it away without discussion.
When Totals Are Close
If two vendors have scores within 0.5 of each other, the totals are not telling you which to pick. You need a different decision criterion:
Which has lower implementation risk?
Which has the better reference customers?
Which feels right after exposure?
Which has the team's preference?
The scorecard narrowed the choice; the decision still happens. Don't pretend the scorecard chose for you.
Don't Let the Number Decide
A scorecard total of 7.7 vs. 7.3 is not a meaningful difference — the 1-10 scoring is too coarse to resolve a 0.4 gap, no matter how many decimal places the spreadsheet prints. The Vendor A story that opened this piece is the failure mode: a composite that averages away the specific trade-off you're actually deciding.
So when totals land close, stop reading the total and read the rows. Where does the leader win, where does it lose, and do you care more about the things it wins? A vendor that's a 9 on the criterion you'll live with daily and a 6 on one you'll feel once a year is not the same as its mirror image, even at an identical weighted average. Use the scorecard to focus that discussion, then decide with judgment. The math doesn't replace the team's experience; it organizes it.
References
Reference calls per shortlisted vendor. Standard questions:
What's been working well?
What's frustrated you?
What does implementation look like in reality vs. what they promised?
Where do they fall short?
Would you choose them again?
Capture reference notes in the scorecard. They often shift scoring more than the vendor's own materials.
The Demo Trap
Demos are vendor-controlled. They show the vendor's best case. Treat demos as evidence of what's possible, not what's reliable.
Use scripted scenarios in demos. Give each vendor the same use cases (your real use cases, not their pre-built demos). Watch how they handle yours.
After the demo:
Did they handle your scenarios specifically?
Were they evasive about anything?
Did they answer follow-up questions confidently?
Did the demo match their RFP responses?
Differences between demo and RFP are signals.
Total Cost of Ownership
Cost scoring should reflect TCO, not list price.
Include:
License or subscription
Implementation services
Internal team time
Training
Annual maintenance / support
Per-usage variable costs
Compare on 3-year totals. Vendors who underprice year 1 and have steep year 2-3 ramps look better than they are.
Anti-Patterns
The pre-decided scorecard. Weights tuned to favor a pre-selected vendor. Theater. Skip the scorecard if the decision is already made.
The single evaluator. One person's scores aren't a scorecard. Multiple evaluators are the point.
The score-only decision. Total decides. Misses context.
The scoring without notes. Pure numbers, no anchoring. No one can interpret the result later.
The frozen scorecard. Filled out in the demo phase, no update after references and contracts. Reality keeps shifting.
After the Decision
The scorecard's job doesn't end at selection.
Keep it as record of why you picked who you picked
Reference at the 6-month review: did the chosen vendor live up to the scoring?
Reference at renewal: have things changed enough to re-evaluate?
Pattern over many scorecards: are you systematically wrong about certain criteria?
The scorecard from three years ago is useful when the team has changed and the original reasoning needs to be remembered.
Key Takeaway
A vendor evaluation scorecard converts impressions into structured comparison. Pick specific criteria, assign honest weights, have multiple evaluators score independently with evidence-anchored scores, then discuss disagreements. The numbers focus the conversation; the decision is still made with judgment. Use TCO not list price, scripted demo scenarios, and reference calls. Keep the scorecard as record. Don't let a 0.4-point margin decide a multi-year commitment.


