Why most AI backlogs are wrong
The typical AI backlog inside a 30–200 person business is a mix of vendor pitches, someone's LinkedIn read, and the one thing the CEO saw a competitor announce. It isn't a backlog — it's a suggestion pile. The Scorecard exists to convert that pile into a defensible rank order in about ninety minutes.
The four scoring axes
Impact (1–5). How much revenue, cost, or leadership time does this actually free up in the next two quarters? Not theoretical impact — booked impact.
Effort (1–5, inverted). How many people-weeks between decision and value? A five-week build scores worse than a one-week one, all else equal.
Dependency (1–5, inverted). How many other systems, integrations, or approvals must line up before this ships? High dependency is where AI projects go to die quietly.
Reversibility (1–5). If we ship this and it's wrong, how easily do we back it out? High reversibility earns the right to move faster.
How to run the scoring session
- List every AI candidate — pitches, ideas, vendor demos — on one page.
- Score each on the four axes with two people, independently, then reconcile the deltas.
- Sum the four scores. Rank descending.
- Draw a cut line at the top three. Everything below is not "no," it's "not now."
What to do with the result
The top three become your 90-day AI roadmap. The bottom of the list becomes a parking lot you revisit quarterly. This is exactly the artifact we build inside a Diagnostic — the value of doing it yourself first is knowing whether your own ranking survives contact with an outside read.