AI PM TEMPLATES

GenAI Use Case Scoring Template: How to Pick Which AI Use Cases to Build First

By Institute of AI PM·15 min read·Aug 27, 2026

TL;DR

Every AI PM is drowning in use case ideas and short on capacity to build them. This template gives you a repeatable scoring rubric across five dimensions: ROI potential, technical feasibility, data readiness, risk and compliance, and strategic fit. Run every candidate through the same scorecard, surface the top five, and build your sequencing argument on numbers rather than opinions. Includes the full scoring rubric, a worked example, and the organizational questions you need answered before scoring starts.

The AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.

Why Use Case Scoring Exists

When a company commits to building with GenAI, the immediate result is a backlog explosion. Sales wants an AI-powered pitch deck generator. Customer success wants a churn predictor. Engineering wants a code review agent. Legal wants a contract summarizer. Finance wants automated invoice reconciliation. Every one of these is plausible. None of them come with proof.

The standard response is to prioritize by gut feel, by whoever argued loudest in the last meeting, or by which use case a competitor just announced. All three approaches produce the same outcome: a portfolio of AI features with no compounding logic, most of which fail to drive measurable impact.

Use case scoring solves this by applying a consistent framework before any resources are committed. The goal is not to eliminate judgment but to make judgment explicit, comparable, and defensible. When you can show a room of stakeholders that use case A scored 78 and use case B scored 44, the sequencing debate shifts from "my instinct says B" to "here is what would have to be true about B to outrank A."

When to use this template

Use it when you have four or more GenAI use case candidates and limited capacity to build all of them in the next two to three quarters. Use it before running feasibility spikes, not after. The scoring process takes about 90 minutes per use case if done properly. Block that time before the sprint planning calendar fills in.

The Five Scoring Dimensions

Each use case is scored across five dimensions. Each dimension is worth 20 points, for a maximum total score of 100. Adjust the weights if your organization has a strong bias toward one dimension (e.g., a company under regulatory pressure might weight risk and compliance at 30 points and reduce ROI to 10).

1. ROI Potential (0-20)

How much value does this use case generate if it works? Score based on combination of revenue impact, cost reduction, and user metric improvement.

16-20 pts

Clear, measurable, large impact

Reduces customer support handle time by 30% across 5,000 monthly tickets

11-15 pts

Measurable, moderate impact

Reduces time per contract review from 45 min to 15 min for 200 reviews/mo

6-10 pts

Hard to measure or small impact

Improves internal meeting notes quality; no clear revenue or cost link

0-5 pts

No clear ROI case

Cool demo, no identified user or business metric it moves

2. Technical Feasibility (0-20)

How confident are you that current AI capabilities can solve this problem at the quality threshold users need? Score based on task structure, output verifiability, and model availability.

16-20 pts

Proven at similar companies, structured task

Text summarization, classification, extraction from well-formatted documents

11-15 pts

Plausible with current models, some uncertainty

Complex reasoning over mixed-format data, multi-step agentic tasks with defined scope

6-10 pts

Requires significant prompt engineering or fine-tuning

Domain-specific expertise, judgment calls, nuanced tone matching

0-5 pts

Current AI cannot reliably solve this

Real-time physical world interaction, guaranteed accuracy in high-stakes legal or medical decisions

3. Data Readiness (0-20)

Do you have the data the AI feature needs to work? Score based on data availability, quality, and access.

16-20 pts

Data exists, is clean, is accessible

CRM data in a queryable warehouse, internal documents in a shared drive with API access

11-15 pts

Data exists but requires significant preparation

PDFs in file storage needing OCR, legacy databases requiring ETL before use

6-10 pts

Partial data or access constraints

Data is PII-heavy and needs anonymization, third-party data requires licensing

0-5 pts

Data does not exist or cannot be accessed

Would need 6+ months to accumulate ground truth labels, or data is legally off-limits

4. Risk and Compliance (0-20)

How much regulatory, legal, or reputational risk does this use case carry? A higher score means lower risk. Reverse-scored: 20 = minimal risk, 0 = severe risk.

16-20 pts

Minimal risk, internal or low-stakes output

Internal documentation tool, meeting summarizer for internal use

11-15 pts

Moderate risk, human in the loop required

Customer-facing chatbot where agent answers are reviewed before sending

6-10 pts

Significant risk, legal or compliance review needed

Financial advice tool, medical symptom checker, employment screening AI

0-5 pts

High risk, may be prohibited or requires extensive guardrails

Automated credit decisions, autonomous legal filing, biometric identification

5. Strategic Fit (0-20)

How well does this use case align with the company's stated AI strategy, competitive positioning, and roadmap? Score based on alignment to declared priorities.

16-20 pts

Directly supports a stated company-level AI priority

CEO has named this workflow as a 2026 transformation target

11-15 pts

Supports a team or product-level priority

Aligns with this year's product goal but not named at company level

6-10 pts

Neutral or tangential to declared strategy

Good idea but not connected to a specific OKR or strategic bet

0-5 pts

Works against current strategy or creates distraction

Enters a market the company has explicitly decided to exit or avoid

Worked Example: Scoring Four Candidates

The following example applies the scorecard to four real-world GenAI use case candidates at a B2B SaaS company with 250 enterprise customers. The product team has capacity to prioritize one use case for Q4 build.

Use CaseROIFeasibilityDataRiskStrategyTotal

AI contract summarizer (internal legal)

Good feasibility, low risk, moderate ROI since legal is internal-facing

121816181074

Customer churn predictor

High ROI and strategy fit, but CRM data quality issues drag data score

181410141874

AI-assisted support ticket triage

Balanced across all dimensions, most buildable with current stack

161717161682

Automated quarterly business review deck generator

Low risk, but limited ROI and weak strategy fit

10121418862

The AI-assisted support ticket triage scores highest at 82. It wins not because any single dimension is exceptional but because it is the most balanced: high feasibility, solid data readiness, meaningful ROI, manageable risk, and alignment with the company's stated goal of reducing support costs. The churn predictor has higher potential ROI but the data readiness gap (10) would require a three-quarter data quality program before the AI feature could be built. The scoring makes that tradeoff visible.

Learn to Build the AI Roadmap Your Team Will Follow

The AI PM Masterclass covers use case scoring, roadmap sequencing, and the full set of frameworks AI PMs need to make defensible prioritization decisions. Taught live by a Salesforce Sr. Director PM.

Common Scoring Mistakes

The scorecard is only as useful as the discipline applied to filling it out. These are the mistakes that consistently corrupt the output:

!

Scoring the outcome you want, not the outcome you have evidence for

Fix: Require each score to be accompanied by a one-sentence rationale and at least one source of evidence. An ROI score of 18 with no cited data points is an opinion disguised as a number.

!

Letting the sponsor score their own use case

Fix: Score as a cross-functional group: PM, engineering lead, data lead, and one member from the business side who does not own the use case. Divergent scores reveal disagreement that needs to surface before the build decision.

!

Treating the score as the decision

Fix: The score is an input. A use case that scores 85 but faces an absolute regulatory blocker should not be built. A use case that scores 65 but is a dependency for three other high-value initiatives may still deserve prioritization. Document and explain any decision that diverges from the ranked order.

!

Skipping the feasibility spike

Fix: No matter how high a use case scores, run a one-week feasibility spike before committing to a full quarter of build work. The spike should answer: can the AI actually do this at the quality threshold users need? If not, the score was based on faulty feasibility assumptions.

!

Treating the scorecard as a one-time exercise

Fix: Re-score your top candidates every quarter. Model capabilities change, data readiness improves, strategy pivots. A use case that scored 55 in Q1 because feasibility was low may score 72 in Q3 after a new model release changes what is technically achievable.

Questions to Answer Before You Score

The scoring exercise exposes gaps in organizational readiness. Before you run the scorecard, get alignment on the following:

What does success look like for a GenAI initiative at our company?

Without a shared definition of success, the ROI dimension is unscoreable. Is success a 10% efficiency gain? A new revenue line? A net promoter score improvement? Get this in writing from your executive sponsor before scoring.

What is our acceptable risk threshold for AI-driven decisions?

The risk dimension score depends on what your legal and compliance teams consider acceptable. A healthcare company and a consumer app company have completely different ceilings. Know yours before you assign a risk score.

Who owns the data, and do they need to approve access?

Data readiness scoring is often optimistic because PMs assume data access that has not yet been confirmed. Get written confirmation from the data team before assigning a high data readiness score.

What AI infrastructure do we already have?

Technical feasibility is different if your company already has an LLM deployment pipeline, model evaluation framework, and vector database vs. if you are starting from scratch. Know your baseline before scoring.

The Scoring Session: How to Run It

A scoring session for five use cases should take no more than three hours. Run it as a structured working session, not a presentation.

Before the session

  • Circulate the use case list and the scoring rubric at least 48 hours in advance
  • Ask each participant to individually score all five candidates before the session
  • Collect scores anonymously before the group convenes

During the session (3 hours max)

  • Reveal the anonymous individual scores for each dimension of each use case
  • Discuss dimensions with the highest variance first: these signal real disagreement, not noise
  • Converge on a group score for each dimension using the rubric as anchor, not opinion
  • Document the rationale for each dimension score, especially where the group overrode outlier scores

After the session

  • Publish the final scores and rationales in your product wiki or decision log
  • Assign ownership of the top-ranked use case to a named PM
  • Set a re-scoring date for the remaining candidates (next quarter planning cycle)
  • Run a one-week feasibility spike on the top-ranked use case before committing Q capacity

Build an AI Roadmap That Gets Funded

The AI PM Masterclass covers prioritization frameworks, business case writing, and how to get executive alignment on your AI roadmap. Led live by a Salesforce Sr. Director PM.

Before you go: get the AI PM Minute

One tactic to make you a sharper AI PM, twice a week. 60 seconds to read. Free.

No fluff. Unsubscribe anytime.