AI Risk Review Services

An Independent AI Risk Review is a scoped, evidence-based review designed to reduce uncertainty and strengthen the basis for a practical decision about deployment, scaling, increased autonomy, controls or remediation.

Two primary review routes are available. They are not basic and advanced versions of the same service. Each addresses a different question and works with a different evidence basis. The appropriate route depends on the decision, the stage of the system and the evidence currently available.

Discuss Your AI System

See How It Works

Two Review Routes

AI System Design Risk Review

Primary question: Does the proposed system design provide enough support for the next step?

Typical stages: concept, design, early prototype and pre-deployment.

Evidence basis: architecture documents, specifications, policies, interviews, planned controls and, where available, a prototype or demo.

Main focus: intended behavior, data and context boundaries, authority and autonomy design, planned human oversight, monitoring, fallback, recovery, assumptions and evidence gaps.

Likely result: a clearer view of design risks, control gaps, unsupported assumptions and the conditions that should be addressed or verified before implementation, pilot or deployment.

Operational AI Risk Review

Primary question: What is the system actually doing, and can it be kept under control under relevant operating conditions?

Typical stages: pilot, representative staging, production, scaling, material system change and increased autonomy.

Evidence basis: configurations, permissions, test results, sanitized logs or traces, observations, interviews and other representative or operational evidence.

Main focus: actual or representative behavior, permissions and authority, monitoring and traceability, human oversight, failure handling, interruption, recovery and controllability.

Likely result: evidence-based findings, operational consequences, control limitations, remediation priorities and a bounded conclusion relevant to the next decision.

Both routes are designed to clarify what is known, what remains uncertain and what should happen before the next decision. They are not a mandatory sequence.

AI System Design Risk Review

Review architecture, intended workflows and planned controls before implementation or deployment

This review is intended for a system that is still being designed, built or prepared for its next operational step. It helps clarify whether the proposed design, controls and available evidence provide enough support for the decision ahead.

The review looks at how the system is expected to work: what information it uses, what it can access or change, and where human approval is required. It also considers how behavior will be monitored and what should happen when something goes wrong.

It may examine:

  • system architecture and workflow design;
  • data, knowledge and context boundaries;
  • model, retrieval, tool and connector roles;
  • permissions, authority and autonomy assumptions;
  • human oversight and approval points;
  • monitoring, fallback and recovery design;
  • dependencies on models, vendors or external services;
  • unsupported assumptions and missing evidence.

Typical inputs may include architecture diagrams, specifications, workflow descriptions, policies, interviews and, where available, a prototype or demo.

The result is a clearer view of design risks, missing controls, unsupported assumptions and what should be corrected, tested or demonstrated before implementation, pilot, deployment or client delivery.

A Design Risk Review assesses the intended design and the evidence supporting it. It does not verify actual production behavior.

Operational AI Risk Review

Review how the system behaves and remains controllable under representative operating conditions

This review is intended for a system in pilot, staging or production, or for a system preparing for scaling, material change or increased autonomy. It helps clarify how the system behaves in practice, where control may be insufficient, and whether the available evidence supports the next decision.

The review looks at what the system actually does: which information and tools it uses, what actions it can take, and where people can approve or stop those actions. It also considers how unexpected or degraded behavior is detected and how operations can be recovered after failure.

It may examine:

  • actual or representative system and workflow behavior;
  • permissions, tool access and authority to change data, systems or business processes;
  • retrieval, model, tool and connector behavior;
  • monitoring, logging and traceability;
  • human oversight, approvals and response mechanisms;
  • failure detection, interruption and containment;
  • fallback, recovery and incident reconstruction;
  • behavior after changes to models, prompts, retrieval, connectors or permissions;
  • evidence supporting decisions about deployment, scaling or increased autonomy.

Typical inputs may include configurations, permission records, test results, sanitized logs or traces, observations, approval events, incident evidence and interviews.

The result is a clearer view of operational risks, control limitations, potential business impact and what should be corrected, tested or monitored before the next step.

Where the evidence supports it, the review may provide a bounded readiness conclusion or recommend conditions for permitted use relevant to the decision being considered.

The strength of any conclusion depends on the completeness, quality and representativeness of the evidence available.

Tailored Review Scope

Some decisions do not require a review of the entire system. A tailored scope can focus on a specific component, workflow, control question, evidence set or decision.

Where needed, it may also combine design and operational questions. The deliverable is defined around the issue being examined and the evidence available.

When More Direct Evidence Is Available

Documentation can reveal intended design, planned controls, important assumptions and likely areas of risk. More direct evidence helps determine whether those assumptions and controls hold in practice.

Depending on the scope, the review may include:

  • a prototype or demo;
  • a sandbox or representative test environment;
  • configuration and permission review;
  • sanitized logs or traces;
  • controlled tests;
  • controlled experiments in a synthetic environment.

This additional depth can reduce uncertainty, compare intended and observed behavior, and provide stronger support for the next decision.

Evidence from a demo, sandbox or synthetic environment applies to the environment that was tested. It does not by itself establish how the production system behaves.

This is an additional depth option within either primary review route.

Focused Entry Option

AI Agent Readiness Check

For an agentic system, a focused initial review can clarify whether the current design, controls and available evidence provide enough support for the next operational step.

The review may examine:

  • the tools, APIs and systems the agent can access;
  • permissions and authority to take or trigger actions;
  • human approval requirements;
  • monitoring and traceability of agent activity;
  • interruption, fallback and recovery mechanisms;
  • evidence gaps relevant to pilot, deployment or increased autonomy.

The result is a clearer view of where operational control may be insufficient, which questions require further evidence, and what should be addressed before the next step is taken.

The AI Agent Readiness Check is a focused entry option, not a third primary review route, certification or deployment approval. Depending on the decision and evidence available, it may remain a narrow engagement or lead to a Design or Operational AI Risk Review.

What You Receive

The exact output depends on the review route, scope, decision being supported and evidence available.

A whole-system view

A concise reconstruction of the architecture, data flows, workflows, permissions, human oversight and control points relevant to the decision.

Evidence coverage and material limitations

A clear distinction between what is documented, reported, observed or tested; what remains unverified; and which conclusions depend on assumptions or further evidence.

Findings, mechanisms and business consequences

An explanation of how technical conditions may create risk, how those risks may affect operations or business outcomes, and where existing controls are effective or insufficient.

A controllability assessment

An assessment of how incorrect actions can be prevented, unexpected behavior detected, the system interrupted when necessary, and operations restored after failure.

Bounded decision support

A conclusion limited to the reviewed system, workflow, environment, decision context and evidence available. It may clarify whether the next step is sufficiently supported and under what conditions.

Prioritized remediation and validation

A practical set of actions ordered by decision relevance, together with the evidence needed to confirm remediation or support further testing.

These elements are brought together in a decision-support deliverable defined by the agreed scope and evidence basis. Depending on the engagement, it may take the form of an AI Risk Review Report, an Executive Decision Support Note or another clearly defined deliverable.

How the Review Is Scoped

The starting point is the decision that needs support, not a preselected service package.

Initial scoping considers:

  • the decision and its potential consequences;
  • the current stage of the system;
  • the architecture and workflows relevant to that decision;
  • the availability and quality of evidence;
  • the system's permissions, authority and level of autonomy;
  • the main control questions and unresolved uncertainties.

This determines the appropriate review route, evidence depth and whether a focused or tailored scope is more suitable.

The scope is confirmed through professional judgment after an initial conversation. It is not generated by an automated readiness score or a standard form.

From Technical Evidence to Decision Support

A seemingly minor technical issue can create significant business consequences.

The review takes a whole-system view, connecting architecture, data, workflows, permissions, human oversight and controls to the decision being considered.

It turns relevant technical evidence into practical decision support by clarifying how the system is intended to work, what the available evidence shows, what remains uncertain, and which remediation or validation actions should take priority.

Decision Support Within Clear Boundaries

The review supports, but does not replace, the judgment of the responsible system owner.

It provides evidence-based findings, decision conditions and prioritized actions within the agreed scope. Final decisions about deployment, scaling, increased autonomy or permitted use remain with the responsible owner.

Conclusions apply only to the reviewed system, workflow, environment, decision context and evidence available.

The review is not a formal audit, certification, legal opinion, compliance approval, penetration test or red-team exercise.

Discuss Your AI System

Begin with a short overview of:

  • what the system or workflow does;
  • its current stage;
  • the decision, change or next step being considered;
  • the main uncertainty you would like to clarify.

You do not need to choose the final review route before making contact.

Detailed technical materials are not required for the first message.

Please do not submit passwords, credentials, API keys, sensitive personal data, confidential documents, unredacted logs, production traces, source files or archives in your initial email.

Discuss Your AI System