Founder-led service · Limited engagements

Your GenAI controls exist.
What does the evidence let you claim?

Epistamate examines the evidence behind the control claims organisations rely upon across their GenAI workflows.

An initial engagement begins with one consequential workflow, keeping every conclusion tied to a defined system, use, configuration, population and period.

One workflow Three to six control claims Evidence-linked conclusions Founder-led
The evidence gap

Governance may tell you what should happen.
It may not establish what the controls achieve.

You may have a GenAI policy, named owners, defined controls and a review cycle. Those are necessary foundations. They do not, by themselves, establish how reliably the controls operate, whether the evaluation is valid, or whether evidence gathered before a material change still applies.

Control present

Every answer has a citation

A citation can point to an approved document while the cited passage does not entail the generated answer.

Does the evidence establish source support, not just citation presence?

Control present

Every output is human approved

An approval record shows that a person acted. It does not show which errors the person reliably detects.

What has reviewer performance actually demonstrated?

Control present

The evaluator reports a quality score

A score can be consistent without validly identifying the policy inconsistency or material error management cares about.

What criterion was the evaluator validated against?

Control present

The dashboard shows healthy operation

Monitoring can report what it was designed to see while retrieval, ingestion or escalation failures remain silent.

What is observable, and what could fail without an alert?

What counts as a workflow?

The unit of review is not the model.
It is the path from input to consequence.

A GenAI workflow is a bounded sequence.

A defined input is processed using GenAI, reviewed or acted upon, converted into an output or decision, and subsequently monitored, escalated and changed.

01InputQuestion, case, request or document
02ContextRetrieval, data and instructions
03GenerationAnswer, recommendation or draft
04DecisionReview, action, send or escalation
05OperationMonitoring, incident and change

RAG is a mechanism inside a workflow. A policy describes what should happen. A control is intended to reduce a risk. A control claim states what management relies upon that control to achieve. Evidence determines whether that claim is supportable.

Recognisable examples

Workflows where an AI output
reaches a customer or shapes what a human does.

The technology can differ. The review follows the consequential path, including the controls, people, evidence and changes around it.

Customer operations

Customer support and complaint responses

A message is classified, product or policy material is retrieved, an answer is generated, reviewed or sent, and the outcome is monitored.

Business consequence: cost to serve, response quality, retention and complaint risk.

Service operations

Technical support and field service

Equipment history, manuals and service records are used to recommend troubleshooting steps or produce the customer work report after an intervention.

Business consequence: downtime, first-time fix, repeat visits and customer evidence of work.

Revenue workflow

Client proposals and RFP responses

Approved credentials, product facts, requirements and prior language are retrieved to draft an external response that may create a commitment.

Business consequence: bid capacity, turnaround, consistency and unsupported commitments.

Digital customer journey

Customer-facing product or policy guidance

An assistant interprets a customer's need, retrieves catalogue, product or policy information, recommends an option and may hand off or initiate an action.

Business consequence: conversion, self-service, consistent guidance and unsuitable recommendations.

Examples are illustrative. Suitability depends on the workflow boundary, materiality, evidence access and specialist requirements. Regulated individual decisions, medical or legal advice and safety-critical workflows are not automatically accepted.

The engagement

One workflow. A small set of claims.
The evidence that bears on them.

An initial engagement normally covers one intended-use population, one immediate decision context, one configuration and period, and three to six material control propositions.

01

Define the boundary

Confirm the workflow, intended use, material failure, decision to be informed, configuration and any defined tolerance or decision criterion. Broad requests such as “review our AI governance” are narrowed before evidence is assessed.

02

Reconstruct the claims

Translate policies, control descriptions and management assertions into specific propositions about retrieval, human review, automated evaluation, monitoring or change control.

03

Test the evidence

Map samples, logs, configurations, incidents, validation records and changes to each proposition. Assess whether the evidence is direct, valid, independent, complete and current enough for the claim.

04

State the decision

Separate supported claims, demonstrated weaknesses and unestablished assurance. Identify the smallest additional evidence or test that could change each unresolved conclusion.

What the client receives

A decision record, not
a generic maturity score.

Every conclusion remains tied to the reviewed workflow, evidence set, configuration and period. The report can be useful even when the available evidence cannot support management's original claim.

01

Executive decision brief

The decision the evidence supports, the important limitation, and what management should not overstate.

02

Proposition findings

Three to six exact control claims, their evidence, conclusion, scope and current-applicability conditions.

03

Weakness and assurance ledger

Demonstrated operational weaknesses kept separate from missing evidence, uncertainty and interaction hypotheses.

04

Evidence-development roadmap

Prioritised tests, records or validation work that could resolve the material uncertainty without creating a generic programme.

05

Evidence register

Source, date, period, provenance, proposition mapping and integrity limitations for the records relied upon.

06

Scope and currency statement

The assessed system condition, evidence cut-off, material changes and permitted use of the conclusions.

SUPPORTED

The available evidence establishes the bounded proposition for the reviewed conditions.

CONTRADICTED

The evidence demonstrates a material weakness against the proposition.

NOT ESTABLISHED

The claim may be true, but the available evidence does not justify relying upon it.

Evidence access

Well-structured documents help.
They are not a prerequisite for starting the conversation.

The work follows the operation and evidence, not the volume of governance documentation. The engagement route depends on what the organisation can actually show.

Route A · Review-ready

The workflow and records are identifiable

The client can provide a workflow boundary, control descriptions and relevant operating evidence. The engagement can move directly into proposition and evidence assessment.

Route B · Guided reconstruction

The operation exists, but documentation is fragmented

The workflow is first reconstructed from interviews, configurations and operating records. If only policies and assertions exist, the output may be evidence readiness rather than an effectiveness conclusion.

Evaluation samplesGround truth and adjudicationRetrieved passagesHuman edits and rejectionsEscalationsEvaluator promptsValidation recordsIncidents and complaintsMonitoring alertsSource freshnessConfiguration snapshotsChange history
Research currency

Evidence does not
stand still.

A model, prompt, retrieval corpus, evaluator, reviewer population or external research finding can change what an earlier conclusion means. Epistamate can separately track research relevant to the mechanisms examined and flag developments that may warrant another look.

An alert identifies possible relevance

It tells the client what changed in the research and which reviewed proposition may be affected.

An alert does not amend the assessment

The original conclusion remains a dated record of the workflow and evidence reviewed at that time.

A refreshed conclusion requires current evidence

An applicability review or reassessment is separately scoped when the client can provide the changed configuration, logs or supplemental evidence.

The boundary

What this is.
And what it is not.

This is

  • A bounded review of evidence behind specific control claims
  • Proposition-level reasoning tied to one workflow and period
  • A distinction between demonstrated weakness and missing assurance
  • A practical route to the next evidence or test that matters
  • A limited, founder-led professional engagement

This is not

  • An audit, certification or legal opinion
  • A regulatory-compliance determination
  • Organisation-wide AI assurance
  • A general declaration that a system is safe, fair or reliable
  • Exhaustive red-teaming, cybersecurity testing or model validation

Before you bring
a workflow.

What is a GenAI control evidence review?

It is a bounded review of whether the available evidence supports the specific claims management relies upon about controls within a defined GenAI workflow. It does not stop at confirming that policies, controls or review cycles exist.

What counts as one GenAI workflow?

A workflow is the bounded path from a defined input through GenAI processing, review or action, to a consequential output, together with its monitoring, escalation and change controls. A customer-support drafting process may be one workflow. The organisation's entire use of GenAI is not.

Is this a review of our AI policy?

No. Policies help reconstruct what should happen and identify intended controls. The review examines operational evidence about what those controls actually did within the selected workflow and period.

What if our workflow is poorly documented?

The workflow may first be reconstructed from interviews, configurations and operating records. If only policies and assertions exist, an effectiveness conclusion may not be possible, but an evidence-readiness output may still identify exactly what is missing.

What does the client receive?

The client receives an executive decision brief, proposition-level findings linked to evidence, demonstrated weaknesses, assurance limitations, applicability conditions, an evidence-development roadmap and a technical evidence register.

Can we submit another workflow later?

Yes. The initial engagement begins with one workflow so the conclusions remain bounded and defensible. Additional workflows can be separately scoped if the first review proves useful.

Does a research alert update the original assessment?

No. An alert can identify research that may affect a reviewed mechanism. A refreshed conclusion requires a separately scoped applicability review or reassessment using the current workflow and evidence.

How long does it take and what does it cost?

Timing and fee are proposed after a short scoping conversation confirms the workflow boundary, evidence access and any specialist requirements. This keeps the proposal tied to the review effort the workflow actually requires.

Is this an audit, certification or compliance opinion?

No. It is a founder-led professional evidence review, not an audit, certification, legal opinion, regulatory-compliance determination or general declaration that a system is safe or reliable.

Limited founder-led availability

Bring one workflow and
one decision that needs to hold up.

The first conversation establishes whether the workflow is bounded, consequential and supported by enough access to make the review useful. Do not send confidential evidence through the public form.

Discuss a workflow →