RAG remains an important applied evidence capability within the wider practice, particularly for retrieval boundaries, source authorization, context construction and traceability.
The AI Systems Risk & Evidence Lab supports the consulting practice through synthetic cases, controlled experiments and repeatable evidence work.
A controlled experiment can isolate a specific mechanism, compare behavior under defined conditions and preserve the technical record needed to reconstruct the result.
Depending on the question, the supporting record may include:
- versioned configurations and test conditions;
- input data or synthetic source material;
- retrieved information or system outputs;
- logs and traces;
- raw artifacts and manifests;
- normalized observations;
- limitations and applicability notes;
- reproducibility instructions where these have been verified.
These materials may strengthen a defined line of analysis. Their relevance, sufficiency and decision significance are assessed within the wider review.
The Lab is a supporting method and evidence environment. It is not a separate software product, automated risk assessor or certification function.
The Lab includes a RAG Risk & Evidence Harness for controlled examination of retrieval behavior, source boundaries, context construction and traceability.
The Harness supports the review process by:
- comparing system behavior under defined configurations;
- preserving retrieved sources, assembled context and relevant traces;
- demonstrating specific failure or control mechanisms;
- producing structured technical material for further review;
- recording the conditions and limitations of each experiment.
An experiment may show, for example, how an obsolete source enters the model context when a retrieval boundary is absent and how the result changes when the boundary is applied.
Harness outputs remain technical inputs to the review process. They do not independently determine whether a finding is material, whether a system is controllable or whether a proposed operational step is sufficiently supported.
Conclusions about a client system require separate material from that system and interpretation within the relevant workflow and decision context.
Public traces, experiment summaries or repository links may be added after their repeatability, limitations and suitability for publication have been reviewed.