Procurement teams evaluate SaaS vendors every quarter, yet AI suppliers introduce risks spreadsheets were never designed to capture: opaque model training data, silent model upgrades, output ownership disputes, and subprocessors that change without notice. When legal asks whether a generative writing platform meets enterprise standards, procurement needs a defensible answer backed by evidence, not a gut feeling from a product demo.
An AI vendor due diligence scorecard is a weighted evaluation framework that scores suppliers across security, privacy, model transparency, intellectual property, and business continuity before contract signature. This guide helps procurement, security, and legal teams assess vendors powering AI writing workflows and AI marketing campaigns. The scorecard aligns with NIST AI RMF Govern 6 third-party risk practices and EU AI Act supply chain expectations without replacing counsel on specific deployments.
Scoring Dimensions and Weights
Score AI vendors across five weighted dimensions: security (25%), privacy and data governance (25%), model transparency and safety (20%), intellectual property and output rights (15%), and business continuity (15%), adjusting weights when the use case processes regulated data or supports high-impact decisions. Fixed weights prevent procurement from overweighting price or feature demos while under-scoring data handling. Document weight rationale in the scorecard template so auditors understand why a marketing copy assistant differs from a customer-facing chatbot.
| Dimension | Default weight | Sample criteria | Minimum passing score |
|---|---|---|---|
| Security | 25% | SOC 2 Type II, encryption, SSO, pen test recency | 70% of dimension points |
| Privacy and data governance | 25% | DPA terms, retention, training opt-out, residency | 75% of dimension points |
| Model transparency and safety | 20% | Model card, change notification, safety testing | 65% of dimension points |
| IP and output rights | 15% | Output ownership, indemnity, acceptable use | 70% of dimension points |
| Business continuity | 15% | SLA, exit data portability, subprocessors | 60% of dimension points |
Use Case Weight Adjustments
Increase privacy weight to 35% when vendors process customer PII; increase model transparency to 30% when outputs influence hiring, credit, or medical triage; decrease business continuity weight only for non-production pilots under thirty days. Record the adjusted matrix in the procurement ticket. Teams using AI writing tools for public blog content may accept lower transparency scores than teams generating personalized sales outreach from CRM data.
Scoring Scale
Use a 0 to 4 maturity scale per criterion: 0 absent, 1 planned, 2 partial, 3 implemented, 4 independently verified. Multiply criterion scores by sub-weights within each dimension. Avoid binary yes/no checklists; partial controls deserve partial credit with documented gaps. Overall vendor scores below 70% trigger escalation unless compensating controls apply.
Evidence Tiers: Self-Attest, Report, Audit
Classify vendor evidence into three tiers: Tier 1 self-attestation (vendor questionnaire responses), Tier 2 third-party reports (SOC 2, ISO 27001, penetration test summaries), and Tier 3 independent audit or on-site assessment. Higher-risk use cases require higher tiers. A Tier 1-only vendor may suffice for internal brainstorming on public topics; customer data processing demands Tier 2 minimum with Tier 3 for strategic platforms.
| Tier | Evidence type | Maximum score credit | Typical validity |
|---|---|---|---|
| Tier 1 | Signed questionnaire, policy PDFs, architecture diagrams | Up to maturity level 2 | 12 months |
| Tier 2 | SOC 2 Type II, ISO cert, SIG Lite, CAIQ | Up to maturity level 3 | 12 months from report date |
| Tier 3 | Customer audit rights exercised, red team, code review | Up to maturity level 4 | 24 months or after major change |
Standard Evidence Request List
Request these artifacts with every AI vendor assessment: latest SOC 2 Type II or equivalent, data processing agreement draft, subprocessor list, model card or system documentation, incident notification SLA, training data opt-out confirmation, and output ownership clause from master agreement. Store artifacts in your GRC platform linked to the vendor record. Reject stale reports dated more than fifteen months unless the vendor provides a bridge letter covering the gap period.
- Completed AI-specific security questionnaire (SIG Lite AI supplement or custom).
- SOC 2 Type II report with relevant trust service criteria highlighted.
- Penetration test executive summary dated within eighteen months.
- Data processing agreement with training prohibition and deletion terms.
- Subprocessor notification process and current subprocessor register.
- Model change notification commitment in contract or SLA.
- Business continuity and disaster recovery summary.
Risk Acceptance and Compensating Controls
When a vendor scores below the passing threshold, document a formal risk acceptance with named approver, compensating controls, expiration date, and conditions that void acceptance. Compensating controls might include network isolation, prompt filtering, human review of all outputs, or prohibition on uploading confidential data. Risk acceptance without compensating controls fails audit scrutiny.
Approval Authority Matrix
Procurement managers approve vendors scoring 70% to 79% with low data sensitivity; CISO or delegate approves 60% to 69% or any processing of confidential data; legal and executive committee approve below 60% or high-risk EU AI Act classifications. Marketing teams adopting AI marketing platforms that generate customer-facing copy should route through this matrix even when contracts fall under marketing budget authority.
| Score band | Data class | Required approver |
|---|---|---|
| 80% and above | Any | Procurement with security review |
| 70% to 79% | Internal or public only | Business owner plus security |
| 60% to 69% | Confidential | CISO and legal |
| Below 60% | Any | Executive risk committee |
Compensating Control Examples
- Disable API training flags and verify via contract plus technical configuration audit.
- Route all outputs through human editor before external publication.
- Deploy vendor only on VDI with clipboard restrictions and DLP monitoring.
- Use enterprise tenant with private endpoint instead of consumer tier.
- Limit integration to anonymized or synthetic data sets during pilot.
Annual Re-Score Triggers
Re-score AI vendors at least annually and within thirty days of material changes: model version upgrades, new subprocessors, security incidents, acquisition, pricing tier changes affecting data handling, or expansion into regulated use cases. NIST AI RMF Manage 3.1 expects ongoing third-party monitoring, not one-time onboarding. Calendar annual re-scores on contract anniversary dates.
Continuous Monitoring Signals
Subscribe to vendor status pages, CVE feeds mentioning vendor products, and subprocessor change notifications; flag any signal for ad hoc re-score before the annual cycle. Procurement should receive automated alerts when vendors publish new terms of service affecting data training or output rights. A model swap announced in release notes without customer notice is a re-score trigger even if the contract is mid-term.
- Annual scheduled re-score on contract renewal date.
- Ad hoc re-score within thirty days of reported security breach.
- Re-score when deployment scope expands to new data classes or regions.
- Re-score after vendor merger, acquisition, or insolvency filing.
- Re-score when internal risk tier for the use case increases.
Scorecard Version Control
Maintain versioned scorecard templates; when criteria change, re-score active vendors against the new template within ninety days and archive prior scores for audit trail. Auditors compare historical decisions against the criteria in force at approval time. Never retroactively alter scored results without documenting a formal re-assessment event.
Mapping Scorecard to Inventory Register
Link every approved vendor scorecard to a corresponding entry in your AI tool inventory register so auditors trace procurement decisions to operational deployments. The register captures system name, owners, and risk tier; the scorecard captures vendor evidence and approval authority. When a vendor re-score drops below threshold, flag the register entry for suspension review within five business days.
Procurement System Integration
Configure your procurement platform to block PO creation for AI vendors without a completed scorecard record ID attached to the requisition. Coupa, SAP Ariba, and similar systems support custom fields for compliance gates. Security teams receive automated notification when procurement opens a new AI vendor request, triggering parallel scoring before contract negotiation advances.
Vendor Tiering After Score
Assign approved vendors to tiers (strategic, standard, pilot) based on score and use case criticality; strategic tier vendors require annual on-site or virtual review in addition to document re-score. Strategic tier includes vendors processing customer PII in production or supporting revenue-critical workflows. Pilot tier covers time-limited evaluations under ninety days with capped data scope and mandatory re-score before production expansion.
Operationalizing the Scorecard
Embed the scorecard in procurement workflow as a mandatory gate before PO issuance: intake form captures use case, data class, and risk tier; security completes scoring; legal reviews contract gaps; approvers sign risk acceptance when needed. Average cycle time from intake to approval should target fifteen business days for standard vendors and thirty for high-risk assessments.
Cross-Functional RACI
Procurement owns the process; security scores technical dimensions; privacy scores data governance; legal scores IP and contract terms; business owner attests use case accuracy. Disputes on dimension scores escalate to the AI governance committee or equivalent body. Single-function scoring breeds blind spots that surface during customer security reviews.
Frequently Asked Questions
How do we score early-stage AI startups without SOC 2 reports?
Accept Tier 1 evidence with enhanced compensating controls: shorter contract terms, limited data scope, mandatory human review, and contractual audit rights upon Series B or first enterprise customer milestone. Require startups to commit to SOC 2 Type II within twelve months as a contract condition. Score security dimension conservatively and cap overall approval at 65% until Tier 2 evidence arrives.
Can we use open-source models without a traditional vendor?
Yes, but assign your organization as the accountable party: score infrastructure provider, model maintainer community health, license compatibility, and your internal deployment controls as vendor substitutes. Open-source models lack SOC 2; compensate with container scanning, model provenance verification, and internal penetration testing. Document the model hash pinned in production and re-score when upgrading versions.
How do we score SaaS products with embedded AI we did not procure separately?
Treat embedded AI as a sub-assessment within the parent vendor scorecard: extract AI-specific criteria into a supplemental section and score against the same dimensions at reduced weight. CRM, design, and productivity suites increasingly bundle generative features. Parent vendor SOC 2 may not cover AI subprocessors; request AI-specific subprocessor disclosures.
Procurement complains the scorecard slows deals. How do we balance speed and rigor?
Pre-score approved vendor tiers for common use cases, maintain a fast-track path for renewals without scope change, and parallelize legal review with security scoring. A pre-vetted catalog of AI writing tools with existing scores cuts new deal time from weeks to days. Speed without scoring returns shadow AI adoption through personal accounts.
Industry-Specific Adjustments
Healthcare organizations add HIPAA BAA compliance and clinical safety criteria as mandatory dimensions; financial services add model explainability and fair lending assessment; government contractors add FedRAMP or IL-level authorization requirements. Adjust default weights rather than adding unlimited criteria. A scorecard with more than seven dimensions becomes unscoreable within reasonable procurement timelines.
Marketing Vendor Considerations
Marketing AI vendors require heightened IP and output rights scoring because generated copy, images, and campaigns may infringe third-party trademarks or reproduce copyrighted training data patterns. Teams evaluating AI marketing platforms should require explicit commercial use licenses for outputs and indemnification clauses covering copyright claims. Score IP dimension at maximum weight for any vendor producing customer-facing creative assets.
Writing Platform Evaluation
Enterprise AI writing platforms should demonstrate enterprise SSO, admin audit logs, data retention controls, and contractual training opt-out before processing internal strategy documents. Consumer-tier subscriptions fail privacy and security dimensions regardless of writing quality. Procurement should reject personal plan reimbursements for business use when enterprise tier with scorecard approval exists.
Scorecard Maturity Roadmap
Phase one (months one to three): publish scorecard template and train procurement and security; phase two (months four to six): score all active AI vendors retroactively; phase three (months seven to twelve): integrate with procurement gates and automate annual re-score reminders. Mature programs benchmark vendor scores against industry peers and share anonymized aggregate data with the AI governance committee for trend analysis.
Red Flags That Trigger Automatic Deferral
Automatically defer vendor approval when any of these red flags appear: no enterprise DPA available, refusal to disclose subprocessors, no SOC 2 or equivalent in progress, training on customer data without opt-out, or history of unresolved security incidents in public databases. Red flags override composite scores; no amount of feature advantage compensates for missing baseline security evidence. Document red flag deferrals separately from low-score risk acceptances.
Scorecards Enable Defensible Procurement
An AI vendor due diligence scorecard succeeds when weighted dimensions reflect use case risk, evidence tiers match sensitivity, risk acceptance documents compensating controls, and annual re-scores catch vendor changes before they become incidents. Procurement, security, and legal share ownership of the framework as living infrastructure, not a one-time RFP attachment.