Vendor Evaluation Scorecard: A 90-Day Post-Award Review Template
TL;DR: A vendor evaluation scorecard should measure what happened after a supplier won the business, not simply repeat the assumptions used during selection. Use a 30/60/90-day review cycle to assess delivery, quality, commercial accuracy, service, compliance, and improvement. Define every metric before the first purchase order, assign evidence owners, score exceptions consistently, and convert the final result into a clear action: expand, maintain, improve, restrict, or exit the relationship. The template and operating model below give procurement teams a practical way to make supplier decisions with evidence rather than anecdotes.
Procurement teams often evaluate suppliers carefully before an award and then become strangely informal once the first purchase order is issued. The sourcing file contains weighted criteria, approvals, quotes, and negotiation notes. The post-award review, meanwhile, becomes a handful of emails: operations says the supplier is difficult, finance reports invoice errors, and the category manager remembers that pricing looked competitive.
That is not vendor performance management. It is organizational memory with a spreadsheet attached.
A vendor evaluation scorecard closes the gap between supplier selection and supplier reality. It gives purchasing teams a repeatable way to determine whether a new vendor delivered the value, control, and service promised during the RFQ. A 90-day post-award format is especially useful because it catches problems early enough to correct them, while producing enough operating evidence to support a fair decision.
This guide explains how to design the scorecard, set defensible weights, run 30-, 60-, and 90-day reviews, and feed the findings into future sourcing decisions. It is designed for procurement managers, purchase managers, sourcing leads, and supply chain directors working with lean teams that need a useful controlnot a ceremonial quarterly report.
What a vendor evaluation scorecard should accomplish
A vendor evaluation scorecard is a structured record of supplier performance against agreed commercial and operational expectations. Its job is not to produce a colorful dashboard. Its job is to help the business decide what to do next.
A useful scorecard should answer five questions:
- Did the vendor deliver the correct goods or services at the agreed time?
- Did delivered quality match the specification and acceptance criteria?
- Did invoices, prices, quantities, taxes, and commercial terms match the award?
- Did the vendor communicate early and resolve issues effectively?
- Should the business expand, maintain, improve, restrict, or end the relationship?
The last question matters most. If every score produces the same responsecontinue buyingthen the process is administrative theater.
Pre-award and post-award evaluation serve different purposes. A sourcing scorecard estimates which bidder is most likely to succeed. A post-award vendor evaluation scorecard tests whether that prediction was correct. The first uses proposals, demonstrations, references, samples, and commitments. The second uses delivery records, inspection results, service tickets, invoices, corrective actions, and stakeholder evidence.
Do not blend the two without labeling them. A supplier can submit the best bid and perform poorly. Another can have a modest proposal but execute reliably. Procurement needs visibility into both the selection decision and actual performance.
The 90-day window is not a universal contract period. It is a control point. For high-frequency materials, 90 days may contain dozens of deliveries. For capital equipment or professional services, it may contain milestones rather than shipments. The scorecard structure stays consistent, but the evidence and review frequency should reflect the category.
Use the scorecard for decisions such as:
- Releasing more volume to a new supplier
- Moving a vendor from conditional to approved status
- Launching a corrective action plan
- Adjusting safety stock or lead-time assumptions
- Reopening a category to competition
- Revising RFQ requirements before the next event
- Creating an evidence-based supplier development plan
That decision orientation keeps the process lean. Every metric must earn its place by influencing risk, cost, continuity, compliance, or future sourcing.
Build the scorecard around six weighted dimensions
The following model works as a practical starting point for many direct-material, indirect-material, and service suppliers. Adjust the weights for category risk, but keep the total at 100 points and document why each change was made.
| Dimension | Default weight | What it tests | Typical evidence |
|---|---|---|---|
| Delivery reliability | 25 | Whether commitments were met in full and on time | PO dates, receipts, milestone records, advance delay notices |
| Quality and specification conformance | 25 | Whether output met defined requirements | Inspection results, defects, returns, acceptance records, rework |
| Commercial accuracy | 15 | Whether the awarded economics survived execution | Quote, PO, invoice, freight, taxes, credits, price variances |
| Service and responsiveness | 15 | Whether the supplier communicates and resolves issues | Response timestamps, escalation logs, resolution time, stakeholder feedback |
| Compliance and documentation | 10 | Whether required controls and records remain valid | Certificates, insurance, licenses, data, audit findings, acknowledgments |
| Improvement and collaboration | 10 | Whether the vendor prevents recurrence and creates value | Corrective actions, root-cause reports, improvement proposals, savings evidence |
Delivery reliability: 25 points
Delivery should measure the promise that matters to the business. For goods, that is usually on-time, in-full delivery against the confirmed datenot the date the supplier later wished had been agreed. For services, it may be milestone completion, staffing availability, response time, or deliverable acceptance.
Define the tolerance explicitly. If a delivery is considered on time within one business day, state that. Decide how partial shipments, buyer-requested changes, force majeure, and early deliveries are treated. Early is not automatically good if it creates storage or working-capital costs.
A basic metric is:
On-time-in-full percentage = orders delivered on time and in full ÷ total orders due × 100
Quality and specification conformance: 25 points
Quality must connect to the RFQ specification, statement of work, approved sample, service level, or acceptance criteria. Avoid a vague rating such as “good quality.” Use evidence: defect rate, first-pass acceptance, rework hours, rejection quantity, customer complaints, audit results, or milestone acceptance.
Severity matters. One critical safety failure should not disappear inside an average created from twenty perfect deliveries. Add a critical-failure rule that caps the overall score or triggers automatic escalation.
Commercial accuracy: 15 points
This dimension tests whether the commercial agreement survived the journey from quote to purchase order to invoice. Track unauthorized price changes, unexplained surcharges, incorrect quantities, tax errors, duplicate billing, freight variance, missed credits, and payment-term discrepancies.
AuraVMS can preserve the original supplier quotations and comparison context used during an RFQ, giving procurement a cleaner reference when a later invoice or price-change request needs to be tested against the award.
Service and responsiveness: 15 points
Measure response and resolution separately. A supplier can acknowledge a problem in ten minutes and still take ten days to fix it. Useful measures include first-response time, time to contain an issue, time to provide root cause, resolution time, escalation frequency, and completeness of status updates.
Stakeholder ratings can be included, but anchor them with specific questions. “How satisfied are you?” invites mood. “Did the vendor provide an accurate recovery date within four business hours?” produces evidence.
Compliance and documentation: 10 points
Evaluate only requirements relevant to the category and risk. Examples include quality certificates, safety records, insurance, licenses, sanctions screening, information-security evidence, sustainability declarations, data-processing terms, country-of-origin records, and signed policy acknowledgments.
Compliance scoring should not turn mandatory requirements into negotiable points. If a valid license is legally required, an expired license is not a low score; it is a stop condition.
Improvement and collaboration: 10 points
This dimension separates suppliers that merely respond from those that reduce future effort. Look for timely corrective actions, credible root-cause analysis, process improvements, cost avoidance, packaging changes, lead-time reduction, specification suggestions, and transparent capacity planning.
Do not award points for promises. Award them when an action is implemented and its effect is visible.
Define scoring rules before the first review
The easiest way to damage trust in a vendor evaluation scorecard is to invent the scoring logic after performance problems appear. Define the rules when the supplier is awarded, share relevant expectations, and apply them consistently.
A five-point scale is usually enough:
| Rating | Meaning | Evidence standard |
|---|---|---|
| 5 | Exceeds requirement | Measurably outperforms the target without creating a trade-off elsewhere |
| 4 | Meets requirement | Achieves the agreed target with no material exception |
| 3 | Minor variance | Small, contained exception with prompt recovery and no serious impact |
| 2 | Material variance | Repeated or significant miss requiring buyer intervention or corrective action |
| 1 | Critical failure | Serious breach, uncontrolled risk, or failure causing major operational impact |
Calculate a weighted score with this formula:
Weighted points = rating ÷ 5 × dimension weight
If delivery reliability is weighted at 25 and the vendor receives a rating of 4, the supplier earns 20 weighted points for that dimension. Add all weighted points to produce a score out of 100.
Use decision bands, but pair them with override rules:
| Total score | Suggested status | Default action |
|---|---|---|
| 90–100 | Strategic performer | Consider additional volume or longer commitment after risk review |
| 80–89 | Approved performer | Maintain business and target specific improvements |
| 70–79 | Conditional | Use a documented improvement plan with owners and deadlines |
| 60–69 | Restricted | Limit new awards and require senior review before additional spend |
| Below 60 | Exit review | Assess replacement, transition risk, and contractual remedies |
An override rule prevents averages from hiding unacceptable events. Examples include a product safety incident, bribery concern, data breach, forged certificate, repeated unauthorized substitution, or intentional invoice manipulation. A critical event may force restricted status regardless of the numerical total.
Five controls make the scoring defensible:
- Define the data source for each metric. Name the system, report, or record.
- Assign one evidence owner. Procurement can coordinate, but operations, quality, finance, legal, or IT may own particular facts.
- Establish the measurement window. State which orders, milestones, or incidents are included.
- Record exclusions. Buyer-caused schedule changes should not quietly become supplier failures.
- Require a note for every rating of 1, 2, or 5. Outliers deserve an explanation.
For categories with few transactions, do not pretend that one delivery creates statistical certainty. Score the available evidence, mark the sample size, and qualify the decision. A score based on two milestones should not carry the same confidence as one based on fifty receipts.
Run the 30-, 60-, and 90-day review cycle
The review cycle should become more decisive as evidence accumulates. Each checkpoint has a different job.
Day 0: establish the baseline
Before the first order or kickoff, store the final quote, negotiated changes, approved specifications, service levels, delivery assumptions, contacts, escalation paths, and required documents. Confirm what will be measured and who supplies the evidence.
This is where an RFQ record matters. If the awarded promise is scattered across inboxes, attachments, and spreadsheet versions, post-award evaluation begins with an argument about the baseline. AuraVMS helps procurement teams collect supplier quotes without requiring suppliers to create accounts, compare submissions side by side, and retain a consistent basis for the award.
Day 30: detect setup and control failures
The first review should focus on whether the relationship was implemented correctly. Ask:
- Did the supplier acknowledge orders and confirm dates?
- Were specifications, ship-to details, billing instructions, and contacts understood?
- Were mandatory documents complete before performance began?
- Did any price, quantity, tax, freight, or payment-term mismatch appear?
- Were early issues escalated through the agreed route?
Do not overreact to a small first-month sample. The objective is early correction. Record baseline scores, identify gaps, and assign actions. If the first shipment exposes a specification ambiguity, fix the requirement as well as the supplier response.
Day 60: test consistency and corrective action
The second review asks whether performance is stable and whether the supplier learned from early exceptions. Compare the first and second periods. A supplier that missed once and permanently fixed the process may be healthier than one that performs acceptably only because buyers chase every transaction.
Review open corrective actions, recurrence, response time, forecast alignment, invoice accuracy, and stakeholder effort. Measure how much manual intervention procurement or operations needed. Hidden buyer effort is part of supplier cost even when the unit price is unchanged.
Use AuraVMS RFQ records to distinguish a true performance failure from a weak sourcing baseline. If delivery terms were never requested consistently from bidders, the remedy belongs partly in the next RFQ template.
Day 90: make a portfolio decision
At 90 days, calculate the full weighted result and choose an explicit status. Do not close the meeting with “monitor closely.” Translate the score into action:
- Expand: consider more volume, additional locations, or a wider scope.
- Maintain: continue current business and preserve the operating cadence.
- Improve: issue a supplier development plan with measurable deadlines.
- Restrict: pause new awards or cap exposure while the supplier recovers.
- Exit: plan replacement, contractual closure, inventory coverage, and data handoff.
Document dissent. If quality recommends restriction but the business owner wants expansion, record both positions, the decision owner, and the accepted risk. Governance is not the absence of disagreement; it is a traceable decision despite disagreement.
The 90-day review should also produce internal actions. Procurement may need a clearer specification, finance may need a revised invoice channel, or operations may need better receipt discipline. A fair scorecard distinguishes supplier causes from buyer causes.
Turn scorecard evidence into better RFQs and awards
The scorecard creates value when its findings change future sourcing behavior. Otherwise, it becomes another isolated supplier-management file.
Start by comparing pre-award promises with post-award outcomes. Which evaluation criteria predicted performance? Which sounded important but produced no useful signal? Which missing question created expensive ambiguity?
Examples:
- If repeated delays came from unrealistic quoted lead times, require capacity evidence and a lead-time breakdown in the next RFQ.
- If invoice variance came from unclear freight treatment, add a mandatory landed-cost structure.
- If quality failures centered on substitutions, add an explicit no-substitution rule and approval workflow.
- If service resolution was weak, require named escalation contacts and response commitments.
- If documentation expired quickly, make renewal dates and notification duties part of the commercial requirement.
This feedback loop strengthens both templates and supplier selection. AuraVMS supports the front end of that loop by standardizing quote requests, centralizing responses, enabling anonymous bidding where appropriate, and making commercial comparisons easier to review. Procurement still owns performance governance; the platform gives the team a cleaner sourcing record from which to start.
Historical scorecard evidence can also shape bidder shortlists. A supplier with an 84 score and strong recovery may deserve another event. A supplier with an attractive new quote but unresolved critical actions may not. Price should remain visible, but not detached from the cost of defects, disruption, expediting, rework, and buyer intervention.
When running a new competitive event, avoid sharing one bidder’s confidential performance or commercial data with others. Convert lessons into neutral requirements. For example, ask every bidder for a confirmed escalation process rather than revealing that the incumbent repeatedly failed to respond.
For categories with recurring purchases, connect award strategy to performance bands. A team might permit approved performers to retain normal allocation, require conditional suppliers to complete actions before new volume, and review restricted suppliers at leadership level. The exact policy should fit supply risk; the important point is consistency.
AuraVMS starts at $5/month. For lean procurement teams, that creates a practical path away from email-and-spreadsheet RFQs without forcing suppliers through a signup process. The goal is not to automate judgment. It is to keep the quotation evidence organized so judgment is faster, fairer, and easier to defend.
Avoid the scorecard mistakes that destroy credibility
Poorly designed scorecards create activity without control. Watch for these common failures.
Too many metrics
Thirty metrics do not create precision if nobody trusts the source data. Begin with the six dimensions above and add a metric only when it can change a decision. If the team cannot explain the action tied to a measure, remove it.
Vague definitions
“Responsiveness” means nothing until the clock, channel, priority, and expected response are defined. “On time” is equally weak unless the agreed date and tolerance are clear.
Procurement scoring alone
Procurement should own the process, not every fact. Quality should validate defects, finance should validate invoice accuracy, operations should validate delivery impact, and information security should validate relevant controls. Cross-functional evidence reduces political scoring.
Letting averages hide critical risk
A strong price score cannot offset a safety violation. Use mandatory gates and critical-event overrides alongside the weighted total.
Changing weights to justify a preferred outcome
Set the model before reviewing the data. If the category genuinely changes, version the scorecard and explain the change. Never quietly edit weights because a favored supplier scored badly.
Ignoring buyer-caused failures
Late approvals, unclear specifications, forecast changes, blocked site access, and incorrect purchase orders can cause supplier misses. Track cause codes and exclude verified buyer-caused events where appropriate. Fairness improves data quality because suppliers are more likely to participate honestly.
Scoring without action
Every red or amber result needs an owner, due date, evidence requirement, and consequence. “Supplier to improve delivery” is not an action. “Supplier to submit a root-cause report by Friday and maintain at least 95% on-time-in-full performance for the next eight due orders” is actionable.
Confusing low price with strong performance
A low unit price can be erased by expediting, downtime, inspection, rework, duplicate invoices, excess stock, or constant follow-up. Keep commercial competitiveness in the analysis, but measure execution separately.
Automating a broken process
Software cannot rescue undefined measures or absent ownership. Establish the decision model first. Then use tools such as AuraVMS to organize RFQ evidence and shorten manual quote handling. A process that is explicit can be improved; a process that exists only in people’s heads cannot.
A practical vendor evaluation scorecard template
Copy the structure below into your preferred working document. Replace every target and evidence source with category-specific definitions before use.
| Dimension | Weight | Target | Actual result | Rating 1–5 | Weighted points | Evidence and comments | Action owner | Due date |
|---|---|---|---|---|---|---|---|---|
| Delivery reliability | 25 | At least 95% on time and in full | ||||||
| Quality conformance | 25 | At least 98% first-pass acceptance; zero critical defects | ||||||
| Commercial accuracy | 15 | At least 99% invoice and price accuracy | ||||||
| Service and responsiveness | 15 | Acknowledge priority issues within four business hours | ||||||
| Compliance and documentation | 10 | 100% mandatory documents valid | ||||||
| Improvement and collaboration | 10 | Actions closed by agreed dates; verified value evidence | ||||||
| Total | 100 |
Add a short decision record beneath the table:
- Review period and transaction sample size
- Overall status: expand, maintain, improve, restrict, or exit
- Critical-event override: yes or no, with reason
- Top three supplier actions
- Top three buyer actions
- Decision owner and approval date
- Next review date
Keep the review meeting disciplined. Circulate evidence before the meeting, resolve factual disputes first, then discuss ratings and actions. Do not spend senior meeting time hunting for attachments.
For a new supplier, the initial setup can be completed in a week:
- Select category-appropriate metrics and weights.
- Define targets, tolerances, exclusions, and critical events.
- Assign evidence owners and the review owner.
- Capture the award baseline and final commercial terms.
- Schedule day-30, day-60, and day-90 checkpoints.
- Share performance expectations with the supplier.
- Run the first review and correct data gaps immediately.
If your sourcing baseline is still buried in email, use AuraVMS to centralize the next RFQ, invite suppliers without forcing account creation, compare quotations side by side, and preserve the award context. Teams moving from manual RFQ cycles that take three to four days can use a structured workflow to reduce the cycle to about two hours, depending on supplier response time and approval complexity.
Request an AuraVMS demo and bring one live sourcing event. Test how your specification, supplier responses, price comparison, and award evidence would flow into a stronger post-award review.
Frequently asked questions
What is a vendor evaluation scorecard?
A vendor evaluation scorecard is a weighted framework for measuring supplier performance against defined delivery, quality, commercial, service, compliance, and improvement expectations. It turns evidence from orders, receipts, inspections, invoices, and issue records into a decision about the supplier relationship.
How often should procurement evaluate vendors?
New or high-risk suppliers should be reviewed more frequently, such as at 30, 60, and 90 days after award. Established strategic suppliers may move to monthly or quarterly reviews. Review frequency should reflect transaction volume, operational impact, compliance exposure, switching difficulty, and the speed at which risk can materialize.
What is a good vendor score?
There is no universal score because weights and tolerances vary by category. In the model in this guide, 80 or above indicates acceptable performance, 70–79 requires structured improvement, and below 70 triggers restriction or exit review. Critical failures should override the total where necessary.
Should price be included in a post-award vendor scorecard?
Commercial performance should be included, but it should test execution: price accuracy, invoice accuracy, agreed freight, taxes, credits, and contract terms. Market competitiveness can be reviewed separately during sourcing or benchmarking. Do not let a low quoted price compensate for major delivery, quality, or compliance failures.
Who should own the vendor evaluation process?
Procurement should normally own the process and decision record. Evidence should come from the functions closest to performance, including operations, quality, finance, legal, information security, and business stakeholders. One process owner and multiple evidence owners create clearer accountability.
How do you score a supplier with very few transactions?
Score the available evidence but disclose the sample size and confidence level. Use milestone evidence, acceptance results, issue handling, and document compliance where shipment volume is low. Avoid treating a two-transaction result as statistically equal to a high-volume supplier’s record.
What should happen when a supplier disputes its score?
Return to the agreed definition and source evidence. Correct factual errors, document any exclusions, and distinguish disagreement about facts from disagreement about consequences. Allow the supplier to submit evidence by a deadline, but keep the buyer’s governance decision and accepted risk clearly recorded.
How does an RFQ system improve vendor evaluation?
An RFQ system preserves the pre-award baseline: requirements, supplier responses, commercial terms, clarifications, and comparison evidence. That makes post-award reviews less dependent on inbox searches and memory. It also helps procurement convert performance lessons into clearer requirements for the next sourcing event.