Audit Ready AI Risk Assessment in 6 Steps for Teams

September 14, 2026•11 min read

Isometric six-step AI risk assessment title card

An AI risk assessment is the documented process of identifying, evaluating, and controlling the harms an AI system could cause before and after deployment. The recommended approach follows a fixed lifecycle: inventory the system, map its context, assess likelihood and impact, decide on acceptance or mitigation, then monitor continuously. Anchor the work to the NIST AI RMF and the MIT AI Risk Repository. Some organizations apply this sequence in their advisory engagements.


TL;DR:

  • Most AI risks are missed during assessments because teams focus solely on model accuracy metrics and neglect safety, security, privacy, and fairness concerns.

  • An effective risk assessment must be continuous, involving proper ownership, documented rationale, and regular reassessment triggered by system changes or drift.

  • Tiering risks based on impact and likelihood guides mitigation efforts, with high risks requiring approval from executives and ethics review before deployment.

  • Using established frameworks like NIST AI RMF and MIT AI Risk Repository helps standardize risk categories and improves consistency across teams.

  • Post-deployment monitoring should track data drift, fairness, and incident reports to catch emerging risks and enable timely reassessment.


PROJECT-JTH

projectjth.com

Bring Clarity to AI Risk

PROJECT-JTH provides technical advisory services and risk assessments to help organizations address operational challenges under pressure.

Explore PROJECT-JTH

Table of Contents

Which Frameworks Should Anchor an AI Risk Assessment?

Two references to dominate this field, and each solves a different problem. The NIST AI Risk Management Framework organizes the work into four core functions: Govern (establish policy and accountability), Map (understand context and identify risks), Measure (analyze and track risks with metrics), and Manage (allocate resources to respond). It also defines the trustworthiness characteristics assessors should score against.

The MIT AI Risk Initiative takes a different angle. Rather than prescribing a process, it maintains a living repository of more than 1,700 documented AI risks pulled from 65 existing frameworks, giving teams a taxonomy to check their findings against real incidents.

Both converge on the same trustworthiness criteria:

  • Safety — the system doesn’t produce physically or operationally dangerous outputs

  • Security — resistance to adversarial manipulation, data poisoning, or extraction attacks

  • Privacy — data handling meets consent and minimization expectations

  • Fairness — outputs don’t systematically disadvantage protected groups

  • Explainability — decisions can be traced and justified to a human reviewer

How Do You Actually Run an AI Risk Assessment?

Treat the assessment as a sequence tied to the system’s lifecycle, not a single form filled out once.

  1. Inventory the system. Log the owner, business purpose, training and inference data sources, and every upstream or downstream dependency.

  2. Map the context. Identify who is affected, how data flows through the system, and which impact domains (financial, physical, reputational) are in play.

  3. Identify the risks. Pull from technical failure modes, socio-technical harms, misuse scenarios, and supply chain exposure, then tag each against a trustworthiness characteristic.

  4. Assess likelihood and impact. Score each risk, assign a tier, and write down the reasoning, not just the number.

  5. Select controls and record residual risk. Decide whether the remaining risk after mitigation is acceptable, and document who approved that call.

  6. Produce the deliverables. A completed assessment yields a risk register, a model factsheet, a control list, and a defined escalation path.

Pro Tip: Write the rationale for every risk tier decision in the register itself, not in a separate meeting note. Auditors and new team members need to see why a risk was called “medium” six months later, not just that it was.

What Are the Main AI Risk Categories to Check For?

Most AI risk assessment failures happen because teams evaluate the model but miss the categories that don’t show up in accuracy metrics.

  • Safety — a diagnostic model missing critical cases; watch false-negative rates on held-out data.

  • Security — vulnerability to prompt injection or data extraction; watch for anomalous query patterns.

  • Privacy — re-identification risk from training data; watch for memorized personal information in outputs.

  • Harmful bias — disparate error rates across demographic groups; watch subgroup accuracy deltas.

  • Information integrity — fabricated citations or hallucinated facts in generative systems; watch factual error rates against a verified benchmark.

  • Misuse — the system repurposed for fraud or harassment; watch abuse-report volume.

  • Resilience — degraded performance under distribution shift; watch drift metrics over time.

Generative AI carries its own weight here. NIST’s Generative AI Profile calls out fabrication and hallucination as risks distinct from traditional model error, because a confident, well-formatted wrong answer is harder for a reviewer to catch than an obvious failure.

How Do You Prioritize and Tier AI Risks?

A basic impact times likelihood matrix does most of the work. Score each risk 1 to 5 on both axes, multiply, and sort into tiers.

  • High tier (15 to 25): requires mitigation before deployment and executive sign-off

  • Medium tier (6 to 14): requires a documented control and scheduled review

  • Low tier (1 to 5): requires monitoring but no immediate action

Certain factors should push a risk up a tier regardless of the raw score: regulatory exposure, the scale of the affected population, and whether the harm is reversible. NIST’s own framework material recommends tiering against context and potential harm rather than chasing zero risk across the board. From there, the decision is binary at each tier: accept, mitigate, or avoid deployment entirely. Write the decision and its owner into the register every time.

What Templates and Repositories Should Teams Use?

Building an assessment framework from scratch wastes time that existing resources already cover.

  • MIT AI Risk Repository offers a searchable database of documented risks and their originating frameworks, useful for checking whether a risk you found has already been studied elsewhere.

  • NIST’s playbook and companion profiles translate the AI RMF’s four functions into specific suggested actions, particularly useful for generative AI deployments.

  • A starter checklist covering inventory, classification, assessment, control selection, documentation, and monitoring schedule gets low-maturity organizations moving in a week rather than a quarter.

  • A model factsheet template standardizes what gets documented per system, so registers stay comparable across teams.

Adapt every external template rather than adopting it wholesale: narrow the scope to your actual system inventory, and document why you dropped or added any field.

Who Owns an AI Risk Assessment Inside an Organization?

Credible assessments require named owners, not a shared inbox.

  1. Executive sponsor — accountable for the program and resourcing.

  2. AI risk owner — runs the assessment process and maintains the register.

  3. Model owner — the technical lead responsible for the specific system’s behavior.

  4. MLOps or engineering — implements controls and monitoring in production.

  5. Legal and compliance — checks regulatory exposure and documentation adequacy.

  6. Ethics board or review committee — signs off on high-tier risks before deployment.

High-risk systems need a hard gate: no production release without documented sign-off from the risk owner and, for the highest tiers, the ethics board. Build this checkpoint into procurement and change control, not as a separate step teams can skip under deadline pressure.

What Should You Monitor After Deployment?

An assessment that stops at deployment is already out of date. Track performance drift against the original baseline, shifts in input data distribution, fairness metrics across subgroups, and raw incident counts from user reports.

  • Set a monitoring cadence tied to system risk tier: weekly dashboards for high-tier systems, monthly for medium, quarterly for low.

  • Trigger a full reassessment on any material change: retraining, new data sources, expanded use case, or a logged incident.

  • Log incidents with enough context (input, output, user impact, resolution) that an auditor can reconstruct what happened without pulling in the original engineer.

Pro Tip: Treat drift alerts as a reassessment trigger, not just an engineering ticket. A model that’s drifted enough to hurt accuracy has often drifted enough to change its risk profile too.

What Do Practitioners Get Wrong About AI Risk Assessment?

The most common failure is treating an assessment as a one-time approval instead of a standing responsibility. A system that passed review in January can fail on fairness metrics by June if the underlying population shifts, and nobody owns catching that unless reassessment is scheduled, not just permitted.

Probability severity matrix for AI oversight

A probability times severity matrix helps set the right level of human oversight: high severity or high probability cases need a human in the loop before any action executes, while medium cases can run on spot checks and automated alerts. Cross-functional ownership beats a single sign-off every time. Some firms build this responsibility into their advisory engagements because a risk register nobody updates is worse than no register at all.

What the Frameworks Get Right, and What They Leave Out

NIST and MIT both give teams something the field badly needed: a shared vocabulary and a repeatable structure. That alone fixes the biggest problem in AI risk work before either framework existed, which was every team inventing its own categories and comparing notes was nearly impossible. The MIT repository’s value is specifically in that shared taxonomy, letting a team map an internal finding to a documented external incident instead of arguing from first principles every time.

AI findings mapped to shared risk taxonomy

Where conventional advice falls short is treating framework adoption as the finish line. Reading the AI RMF and building a register is table stakes, not the achievement. The real test is whether the register gets touched after the initial sign-off, whether a drift alert actually triggers reassessment, and whether the ethics board has real authority to block a release rather than a rubber stamp role.

If there’s one priority for a team starting today, it’s this: build the reassessment trigger before you build the initial scoring rubric. Organizations obsess over getting the first assessment perfect and then let it go stale. A mediocre assessment that gets revisited every quarter beats a perfect one that’s frozen in a document nobody opens again.

— Jesse Hart

How PROJECT-JTH Supports Your AI Risk Assessment Program

Technical advisory routes exist for organizations that need an AI risk assessment operationalized, not just written down and filed away. Where a framework document gives you the structure, PROJECT-JTH gives you the implementation: technical advisory engagements to build your risk register from scratch, assessment workshops to train your team on the inventory-to-monitor lifecycle, and hands-on implementation support for the controls that follow.

PROJECT-JTH

Certain production record and contract management tools can play a direct role here too. Teams use it to keep assessment artifacts, model factsheets, and control documentation tied to actual production records instead of scattered across spreadsheets and email threads. If your organization needs a credible, auditable AI risk assessment program built and maintained rather than assembled once under deadline pressure, start with PROJECT-JTH and get the process running before your next high-risk deployment goes live.

Primary Sources for Frameworks and Templates

Start with the NIST AI RMF 1.0 and its Generative AI Profile, the MIT AI Risk Repository, and IBM’s implementation guide for procurement and CI/CD integration steps. For a strategy-side view of framework adoption, see this thought leadership piece on AI frameworks and risk.

Sources

FAQ

What Are the Four Types of AI Risk?

Most frameworks group AI risk into safety, security, privacy, and fairness or bias, with information integrity and misuse increasingly treated as a fifth and sixth category for generative systems.

How Do You Do an AI Risk Assessment?

Follow the lifecycle: inventory the system and its metadata, map stakeholders and data flows, identify and score risks by likelihood and impact, decide on controls or acceptance, and monitor continuously with scheduled reassessment.

Can AI Write a Risk Assessment?

AI tools can draft sections of a risk register or summarize known risks from a repository like MIT’s, but a human owner still has to validate context, assign tiers, and sign off on acceptance decisions.

What Is the 30% Rule for AI?

There’s no single, established “30% rule” in the major frameworks like NIST’s AI RMF or MIT’s repository. If you’ve seen it cited elsewhere, treat it as informal guidance rather than a standard, and rely on documented likelihood and impact scoring instead.

Who Should Own an AI Risk Assessment?

No single person should own it alone. An AI risk owner runs the process, but the model owner, legal and compliance, and an ethics board or review committee all need defined roles for the assessment to hold up under audit.

Back to Blog

Put It Into Practice

Find Your Next Clear Step.

Need help with a technical challenge or a leadership event? Tell Jesse what you are working through and start a practical conversation.

Start a Conversation

PROJECT-JTH LLC · Hoover, Alabama
Calmer Leadership. Better Systems. Practical Tools.

Home   /   Contact