Tool

AI Design Evaluation Rubric

A quick scoring checklist for judging a single AI feature or design. Each of the six groups has three checks. Use it in a design critique or before a launch review, then follow the checks that fail back to the laws.

AAgency — surface, don’t decide

  1. A1
    The system presents options, evidence, or analysis rather than silently executing irreversible actions on the user’s behalf.
  2. A2
    A human can see what the AI is about to do before it happens, not only after.
  3. A3
    High-stakes actions (financial, legal, personal data, compliance, employment) always route through explicit human confirmation.

TTransparency — reasoning is legible

  1. T1
    A confidence level or certainty signal is shown, not just a final answer.
  2. T2
    The reasoning behind a suggestion is visible in plain language, not buried in a tooltip or absent entirely.
  3. T3
    A user can trace an output back to the source data or logic that produced it.

HHonesty — limits are disclosed

  1. H1
    The system has an explicit “can’t determine this” state, rather than guessing with false confidence.
  2. H2
    Edge cases and low-confidence outputs are flagged distinctly from high-confidence ones, not blended together in the same visual treatment.
  3. H3
    Product or marketing copy around this feature doesn’t oversell what the AI can reliably do.

EEquity — works across different users

  1. E1
    The design was tested with assistive technology (screen reader, keyboard-only, voice control), not just visually reviewed.
  2. E2
    The system accounts for different user contexts (expert vs. novice, high-stakes vs. routine) rather than a single path for everyone.
  3. E3
    Output has been checked for systematic bias against any user group, language, region, or edge-case identity.

RReal reduction — not hidden complexity

  1. R1
    The AI measurably reduces steps, tools, or time, not just visual clutter while the underlying complexity remains.
  2. R2
    Automation is reversible, a user can undo, edit, or override without starting the whole task over.
  3. R3
    The feature has been validated with real task-timing or usage data, not only a demo or stakeholder walkthrough.

FFailure design — graceful degradation

  1. F1
    There’s a clear, designed path for what happens when the AI is wrong, not just a generic error state.
  2. F2
    Unfixable or ambiguous cases are flagged explicitly and don’t silently block or corrupt unrelated work.
  3. F3
    A human reviewer can always intervene without needing engineering support to do it.

This page lists the checks. Interactive scoring is planned; see the Changelog. Check numbers (A1, T2, and so on) are labels used on this site.