Evidence That Supports Judgment

An AI Risk Review uses relevant documents, interviews, observations, configurations, test results, logs, traces and controlled experiments to clarify what is known about a system and what remains uncertain.

Its purpose is to make findings traceable, limitations visible and recommendations more defensible.

Discuss Your AI System

See How It Works

Evidence Supports a Decision

Technical material becomes useful when it is connected to a specific question, examined in the context of the wider system and interpreted in relation to potential consequences.

A document may describe how a control is intended to work. An interview may explain how the team believes it operates. A demonstration, trace or test result may show what occurred under particular conditions.

These sources do not necessarily support the same conclusion.

The review therefore distinguishes:

  • what is documented;
  • what is reported;
  • what has been directly observed or tested;
  • what remains dependent on assumptions;
  • what was unavailable or outside the review boundaries;
  • what requires further confirmation.

The objective is not to create an appearance of certainty. It is to establish the strongest conclusion that the material reviewed can reasonably support.

Evidence Type and Review Status

An evidence type describes where information comes from and the conditions under which it was produced.

A review status records how that information was examined and what it can support within the engagement.

These are two different dimensions.

Evidence Type

Documented

What architecture documents, specifications, policies, procedures or system records state.

Reported

What responsible stakeholders describe in interviews, questionnaires or written explanations.

Observed

What can be seen directly in a demonstration, workflow, accessible environment or recorded system behavior.

Synthetic or Controlled-Test

What is produced through a bounded scenario designed to examine a specific mechanism, failure mode or control.

Representative Operational

Information produced under conditions sufficiently similar to the intended operating environment.

Production Operational

Information obtained from the system’s actual production use, such as relevant configurations, sanitized logs, traces, incidents or operating records.

Review Status

Included and Reviewed

The material was available, included within the agreed review boundaries and examined during the engagement.

Corroborated

A statement, condition or observation is supported by more than one relevant source or form of examination.

Condition-Bounded

The material supports a conclusion only under the particular environment, configuration or scenario in which it was produced.

Not Verified

A reported claim, control or behavior could not be independently confirmed from the material available.

Outside Review Boundaries

The relevant component, environment, workflow or question was not included in the engagement.

Further Confirmation Required

Additional testing, observation or operational material is needed before a stronger conclusion can be supported.

These labels are not necessarily mutually exclusive. For example, a production trace may be included and reviewed, condition-bounded and still require further confirmation because its coverage is limited.

How Evidence Becomes a Finding

A trace, test result or technical observation is not automatically a finding.

It becomes decision-relevant when it is connected to the system context, potential consequences, existing controls and the action required next.

  1. 1. Mechanism

    What technical or operational condition may produce the behavior being examined?

  2. 2. Supporting Material

    Which documents, observations, configurations, traces or test results support the analysis?

  3. 3. Potential Consequence

    How could the condition affect users, operations, data, contractual commitments or a business decision?

  4. 4. Existing Strengths

    Which controls, design choices or operating practices already reduce the likelihood or potential impact?

  5. 5. Control Gap

    Where is the current protection incomplete, inconsistent, unverified or dependent on an unsupported assumption?

  6. 6. Remediation

    What should be corrected, strengthened or tested before the next consequential step?

  7. 7. Validation Evidence

    What would demonstrate that the issue has been addressed sufficiently for reconsideration or retesting?

    This structure also supports a bounded controllability assessment: can incorrect actions be prevented, unexpected behavior detected, the system interrupted when necessary and operations restored after failure?

Agent Authorization & Runtime Evidence

This evidence line examines whether consequential AI actions remain within intended authority and whether the available record can reconstruct what happened. The public examples follow a common sequence from authority and runtime control to execution evidence and controllability.

  1. Authority

    Identify which model, orchestrator, tool, connector or downstream component can actually cause the external consequence.

  2. Runtime Control

    Examine whether an independent control checks the material target, parameters and permitted transition at the point of action.

  3. Execution

    Establish what tool, API or downstream operation was actually invoked and with which material parameters.

  4. Evidence

    Connect approval or intent to the invocation, target-system response and independently verified resulting state.

  5. Controllability

    Determine whether incorrect action can be prevented or detected, execution interrupted, events reconstructed and operations recovered safely.

Sample Deliverable

A written deliverable should present not only the conclusion, but also how it was reached, what supports it and where its limits remain.

A public sample, when available, may include:

Executive Decision Support Note

A concise account of the decision being considered, the principal findings, material consequences, relevant conditions and prioritized next actions.

Finding Card

A structured connection between the mechanism, supporting material, potential consequence, existing strengths, control gap and recommendation.

Evidence Status Summary

A clear account of what was reviewed, corroborated, condition-bounded, unavailable or not verified.

Remediation and Acceptance Evidence

The proposed action and the information needed to confirm implementation, support retesting or reconsider the issue.

Traceability Fragment

A limited mapping from the reviewed material to the finding and from the finding to the relevant recommendation or decision condition.

The exact format depends on the engagement and the audience for the decision.

Only synthetic, anonymized, sanitized or explicitly permitted material may be used in a public sample.

Controlled demo case

Testing Operational Boundaries in the FRIS AI Demo

Controlled interactions now cover source reliability, approval integrity, safe recovery and bounded execution evidence from connected MoySklad and Bitrix24 demo environments.

View the FRIS AI Case

RAG Risk & Evidence Harness

RAG remains an important applied evidence capability within the wider practice, particularly for retrieval boundaries, source authorization, context construction and traceability.

The AI Systems Risk & Evidence Lab supports the consulting practice through synthetic cases, controlled experiments and repeatable evidence work.

A controlled experiment can isolate a specific mechanism, compare behavior under defined conditions and preserve the technical record needed to reconstruct the result.

Depending on the question, the supporting record may include:

  • versioned configurations and test conditions;
  • input data or synthetic source material;
  • retrieved information or system outputs;
  • logs and traces;
  • raw artifacts and manifests;
  • normalized observations;
  • limitations and applicability notes;
  • reproducibility instructions where these have been verified.

These materials may strengthen a defined line of analysis. Their relevance, sufficiency and decision significance are assessed within the wider review.

The Lab is a supporting method and evidence environment. It is not a separate software product, automated risk assessor or certification function.

The Lab includes a RAG Risk & Evidence Harness for controlled examination of retrieval behavior, source boundaries, context construction and traceability.

The Harness supports the review process by:

  • comparing system behavior under defined configurations;
  • preserving retrieved sources, assembled context and relevant traces;
  • demonstrating specific failure or control mechanisms;
  • producing structured technical material for further review;
  • recording the conditions and limitations of each experiment.

An experiment may show, for example, how an obsolete source enters the model context when a retrieval boundary is absent and how the result changes when the boundary is applied.

Harness outputs remain technical inputs to the review process. They do not independently determine whether a finding is material, whether a system is controllable or whether a proposed operational step is sufficiently supported.

Conclusions about a client system require separate material from that system and interpretation within the relevant workflow and decision context.

Public traces, experiment summaries or repository links may be added after their repeatability, limitations and suitability for publication have been reviewed.

Source Boundary in a RAG Workflow

Mechanism

A retrieval workflow contains both a current policy and an obsolete version. Metadata distinguishes the two sources, but the retrieval process does not consistently enforce the approved status.

Supporting Material

A controlled comparison records what is retrieved when the source filter is absent and when it is applied.

The supporting material may include:

  • the test configuration;
  • the documents available to the workflow;
  • retrieved chunks and their metadata;
  • the context passed to the model;
  • the resulting answer;
  • the trace connecting the stages of the run.

Potential Consequence

Obsolete or unauthorized information may enter the model context and influence the answer presented to a user.

In a consequential workflow, this could contribute to an incorrect customer response, an unsupported action or an unreliable business decision.

Existing Strength

The source records already contain metadata that distinguishes current and obsolete material.

Control Gap

Metadata alone does not establish an effective source boundary.

The restriction depends on trusted application logic applying the correct filter consistently, preventing inappropriate override and preserving enough information to reconstruct what happened.

Remediation

Apply the source boundary through trusted workflow logic, test both permitted and prohibited retrieval paths and define how the system should respond when adequate authorized material is unavailable.

Acceptance Evidence

A repeatable test should show that only permitted sources enter the retrieved context under the defined conditions.

The configuration, retrieved material and trace should be sufficient to reconstruct and compare the result.

Synthetic Evidence Boundary

This example demonstrates one mechanism and control pattern under defined synthetic conditions.

It does not establish the behavior, readiness or control effectiveness of a client production system. Any client-specific conclusion would require separate material from the relevant system and operating context.

Evidence Within Clear Professional Boundaries

Technical material supports the review. Professional judgment connects it to consequences, controllability and decision conditions.

The strength of a conclusion depends on:

  • the question being examined;
  • the relevance and quality of the material available;
  • the conditions under which behavior was observed or tested;
  • the coverage and traceability of the review;
  • the limitations that remain.

The absence of an identified finding does not establish that every relevant risk has been examined or that the system will behave correctly under all conditions.

Conclusions apply only to the reviewed system, workflow, environment, decision context and supporting material.

Security tests, controlled experiments and Harness results may contribute to a broader review, but they do not independently establish operational readiness or approve a system for deployment.

An Independent AI Risk Review is not a formal audit, certification, legal opinion, compliance approval, penetration test or red-team exercise.

More detailed boundaries are available on the Professional Boundaries page.

Discuss Your AI System

Begin with a short overview of:

  • what the system or workflow does;
  • its current stage;
  • the decision, change or next step being considered;
  • the main uncertainty you would like to clarify.

You do not need to choose the final review route before making contact.

Detailed technical materials are not required for the first message.

Please do not submit passwords, credentials, API keys, sensitive personal data, confidential documents, unredacted logs, production traces, source files or archives in your initial email.

Discuss Your AI System