Evidence That Supports Judgment

An AI Risk Review uses relevant documents, interviews, observations, configurations, test results, logs, traces and controlled experiments to clarify what is known about a system and what remains uncertain.

Its purpose is to make findings traceable, limitations visible and recommendations more defensible.

Discuss Your AI System

See How It Works

Evidence Supports a Decision

Technical material becomes useful when it is connected to a specific question, examined in the context of the wider system and interpreted in relation to potential consequences.

A document may describe how a control is intended to work. An interview may explain how the team believes it operates. A demonstration, trace or test result may show what occurred under particular conditions.

These sources do not necessarily support the same conclusion.

The review therefore distinguishes:

  • what is documented;
  • what is reported;
  • what has been directly observed or tested;
  • what remains dependent on assumptions;
  • what was unavailable or outside the review boundaries;
  • what requires further confirmation.

The objective is not to create an appearance of certainty. It is to establish the strongest conclusion that the material reviewed can reasonably support.

Evidence Type and Review Status

An evidence type describes where information comes from and the conditions under which it was produced.

A review status records how that information was examined and what it can support within the engagement.

These are two different dimensions.

Evidence Type

Documented

What architecture documents, specifications, policies, procedures or system records state.

Reported

What responsible stakeholders describe in interviews, questionnaires or written explanations.

Observed

What can be seen directly in a demonstration, workflow, accessible environment or recorded system behavior.

Synthetic or Controlled-Test

What is produced through a bounded scenario designed to examine a specific mechanism, failure mode or control.

Representative Operational

Information produced under conditions sufficiently similar to the intended operating environment.

Production Operational

Information obtained from the system’s actual production use, such as relevant configurations, sanitized logs, traces, incidents or operating records.

Review Status

Included and Reviewed

The material was available, included within the agreed review boundaries and examined during the engagement.

Corroborated

A statement, condition or observation is supported by more than one relevant source or form of examination.

Condition-Bounded

The material supports a conclusion only under the particular environment, configuration or scenario in which it was produced.

Not Verified

A reported claim, control or behavior could not be independently confirmed from the material available.

Outside Review Boundaries

The relevant component, environment, workflow or question was not included in the engagement.

Further Confirmation Required

Additional testing, observation or operational material is needed before a stronger conclusion can be supported.

These labels are not necessarily mutually exclusive. For example, a production trace may be included and reviewed, condition-bounded and still require further confirmation because its coverage is limited.

How Evidence Becomes a Finding

A trace, test result or technical observation is not automatically a finding.

It becomes decision-relevant when it is connected to the system context, potential consequences, existing controls and the action required next.

  1. 1. Mechanism

    What technical or operational condition may produce the behavior being examined?

  2. 2. Supporting Material

    Which documents, observations, configurations, traces or test results support the analysis?

  3. 3. Potential Consequence

    How could the condition affect users, operations, data, contractual commitments or a business decision?

  4. 4. Existing Strengths

    Which controls, design choices or operating practices already reduce the likelihood or potential impact?

  5. 5. Control Gap

    Where is the current protection incomplete, inconsistent, unverified or dependent on an unsupported assumption?

  6. 6. Remediation

    What should be corrected, strengthened or tested before the next consequential step?

  7. 7. Validation Evidence

    What would demonstrate that the issue has been addressed sufficiently for reconsideration or retesting?

    This structure also supports a bounded controllability assessment: can incorrect actions be prevented, unexpected behavior detected, the system interrupted when necessary and operations restored after failure?

An Illustrative Evidence-to-Finding Example

The following synthetic RAG example applies the seven-part structure above to one bounded technical mechanism.

It shows how an abstract review sequence can become a traceable technical finding connected to operational consequences and a practical next action.

Source Boundary in a RAG Workflow

Mechanism

A retrieval workflow contains both a current policy and an obsolete version. Metadata distinguishes the two sources, but the retrieval process does not consistently enforce the approved status.

Supporting Material

A controlled comparison records what is retrieved when the source filter is absent and when it is applied.

The supporting material may include:

  • the test configuration;
  • the documents available to the workflow;
  • retrieved chunks and their metadata;
  • the context passed to the model;
  • the resulting answer;
  • the trace connecting the stages of the run.

Potential Consequence

Obsolete or unauthorized information may enter the model context and influence the answer presented to a user.

In a consequential workflow, this could contribute to an incorrect customer response, an unsupported action or an unreliable business decision.

Existing Strength

The source records already contain metadata that distinguishes current and obsolete material.

Control Gap

Metadata alone does not establish an effective source boundary.

The restriction depends on trusted application logic applying the correct filter consistently, preventing inappropriate override and preserving enough information to reconstruct what happened.

Remediation

Apply the source boundary through trusted workflow logic, test both permitted and prohibited retrieval paths and define how the system should respond when adequate authorized material is unavailable.

Acceptance Evidence

A repeatable test should show that only permitted sources enter the retrieved context under the defined conditions.

The configuration, retrieved material and trace should be sufficient to reconstruct and compare the result.

Synthetic Evidence Boundary

This example demonstrates one mechanism and control pattern under defined synthetic conditions.

It does not establish the behavior, readiness or control effectiveness of a client production system. Any client-specific conclusion would require separate material from the relevant system and operating context.

Decision-Support Deliverable Structure

A written deliverable should present not only the conclusion, but also how it was reached, what supports it and where its limits remain.

A public sample, when available, may include:

Executive Decision Support Note

A concise account of the decision being considered, the principal findings, material consequences, relevant conditions and prioritized next actions.

Finding Card

A structured connection between the mechanism, supporting material, potential consequence, existing strengths, control gap and recommendation.

Evidence Status Summary

A clear account of what was reviewed, corroborated, condition-bounded, unavailable or not verified.

Remediation and Acceptance Evidence

The proposed action and the information needed to confirm implementation, support retesting or reconsider the issue.

Traceability Fragment

A limited mapping from the reviewed material to the finding and from the finding to the relevant recommendation or decision condition.

The exact format depends on the engagement and the audience for the decision.

Only synthetic, anonymized, sanitized or explicitly permitted material may be used in a public sample.

A Bounded Case Example

An anonymized or synthetic case can show how an AI Risk Review moves from incomplete information to a structured decision-support result.

The case may illustrate:

  • reconstruction of the relevant architecture and workflow;
  • distinction between intended design, reported implementation and observed behavior;
  • identification of existing strengths as well as material control gaps;
  • incorporation of documents, interviews, demonstrations or test results;
  • translation of technical observations into operational or business consequences;
  • prioritized remediation and validation conditions;
  • a conclusion limited to the reviewed context.

The purpose is to demonstrate the reasoning and evidence-handling process, not to present the reviewed system as a universal design reference or proof of production readiness.

Public material will remain synthetic or anonymized unless explicit permission has been obtained to use names, screenshots, quotations or other identifiable content.

Controlled Experiments and the Evidence Lab

The AI Systems Risk & Evidence Lab supports the consulting practice through synthetic cases, controlled experiments and repeatable evidence work.

A controlled experiment can isolate a specific mechanism, compare behavior under defined conditions and preserve the technical record needed to reconstruct the result.

Depending on the question, the supporting record may include:

  • versioned configurations and test conditions;
  • input data or synthetic source material;
  • retrieved information or system outputs;
  • logs and traces;
  • raw artifacts and manifests;
  • normalized observations;
  • limitations and applicability notes;
  • reproducibility instructions where these have been verified.

These materials may strengthen a defined line of analysis. Their relevance, sufficiency and decision significance are assessed within the wider review.

The Lab is a supporting method and evidence environment. It is not a separate software product, automated risk assessor or certification function.

RAG Risk & Evidence Harness

The Lab includes a RAG Risk & Evidence Harness for controlled examination of retrieval behavior, source boundaries, context construction and traceability.

The Harness supports the review process by:

  • comparing system behavior under defined configurations;
  • preserving retrieved sources, assembled context and relevant traces;
  • demonstrating specific failure or control mechanisms;
  • producing structured technical material for further review;
  • recording the conditions and limitations of each experiment.

An experiment may show, for example, how an obsolete source enters the model context when a retrieval boundary is absent and how the result changes when the boundary is applied.

Harness outputs remain technical inputs to the review process. They do not independently determine whether a finding is material, whether a system is controllable or whether a proposed operational step is sufficiently supported.

Conclusions about a client system require separate material from that system and interpretation within the relevant workflow and decision context.

Public traces, experiment summaries or repository links may be added after their repeatability, limitations and suitability for publication have been reviewed.

Evidence Within Clear Professional Boundaries

Technical material supports the review. Professional judgment connects it to consequences, controllability and decision conditions.

The strength of a conclusion depends on:

  • the question being examined;
  • the relevance and quality of the material available;
  • the conditions under which behavior was observed or tested;
  • the coverage and traceability of the review;
  • the limitations that remain.

The absence of an identified finding does not establish that every relevant risk has been examined or that the system will behave correctly under all conditions.

Conclusions apply only to the reviewed system, workflow, environment, decision context and supporting material.

Security tests, controlled experiments and Harness results may contribute to a broader review, but they do not independently establish operational readiness or approve a system for deployment.

An Independent AI Risk Review is not a formal audit, certification, legal opinion, compliance approval, penetration test or red-team exercise.

More detailed boundaries are available on the Professional Boundaries page.

Controlled demo case

Testing Operational Boundaries in the FRIS AI Demo

Three bounded interactions examined source reliability, approval integrity and safe recovery after an ambiguous execution state.

View the FRIS AI Case

Discuss Your AI System

Begin with a short overview of:

  • what the system or workflow does;
  • its current stage;
  • the decision, change or next step being considered;
  • the main uncertainty you would like to clarify.

You do not need to choose the final review route before making contact.

Detailed technical materials are not required for the first message.

Please do not submit passwords, credentials, API keys, sensitive personal data, confidential documents, unredacted logs, production traces, source files or archives in your initial email.

Discuss Your AI System