IntegrationsBlogCareersBook a free AI assessment
Evidence methodology

What a Gaper AI-agent claim means, and what it does not.

We separate verified results, client-reported observations, external benchmarks, and modeled planning outputs. The label stays with the claim so a reader can see the evidence standard without reading fine print.

Definition

Gaper’s evidence methodology is a claim-labeling and review system that identifies the source, calculation, limitations, and approval status behind a statement about an AI-agent workflow.

Claim labels

Four labels, four different evidence standards.

The qualifier remains adjacent to the claim in page copy, cards, metadata, social assets, and exported material.

Verified result

A result calculated from retained source data using a documented definition, time period, calculation version, and review record.

May be used for
A client-approved, scope-specific measured outcome with its methodology and limitations stated near the claim.
May not be used for
A generalized promise, an unreviewed dashboard screenshot, or an outcome without an identifiable baseline and source trail.
Required evidence
  • Dated baseline and post-launch source exports
  • Metric definition, calculation, and version
  • Attribution method and material confounders
  • Client approval and named reviewer signoff

Render “Verified result” next to the metric, with a link or expandable methodology note. State the client scope and measurement window.

Client-reported observation

A statement made by a client or operator that Gaper has not independently verified against the underlying operational data.

May be used for
A clearly attributed qualitative observation or a client-supplied metric when the source and wording are approved.
May not be used for
An independently verified outcome, a benchmark, or a promise of comparable results.
Required evidence
  • Written client approval
  • Exact approved wording and attribution
  • Date and operating context
  • Disclosure that Gaper did not independently verify the underlying data

Render “Client-reported” immediately before the statement. Do not remove the qualifier in cards, metadata, social copy, or comparison tables.

External benchmark

A statistic, research finding, standard, or market observation published by an external source and used for context rather than proof of a Gaper implementation.

May be used for
Market context, workflow-selection rationale, or a comparison point with a direct, current source link.
May not be used for
A claimed Gaper result, a forecast for a client, or a substitute for measuring the client’s baseline.
Required evidence
  • Primary or authoritative source where available
  • Publication date and URL
  • Scope, methodology, and population checked for relevance
  • A statement that the benchmark does not predict a client outcome

Render “External benchmark” with a direct source link and describe the population or context. Never present it as a customer result.

Modeled planning output

An illustrative calculation or workflow scenario derived from stated assumptions. It is useful for prioritization, not evidence of a realized outcome.

May be used for
Cost or ROI planning, scenario comparison, implementation reference designs, and hypothesis setting before measurement.
May not be used for
A customer case study result, guaranteed savings, a quote, delivery commitment, or compliance conclusion.
Required evidence
  • Visible input assumptions
  • Formula or calculation method
  • Disclosure of excluded benefits and costs
  • A prompt to replace assumptions with a dated baseline and post-launch data

Render “Modeled” adjacent to every output and repeat the disclosure in exported, shared, or downloadable versions.

Calculation disclosure

Planning math stays planning math.

Calculator results use user-entered assumptions and transparent formulas. They are directional planning outputs, not pricing, legal, compliance, accounting, medical, staffing, or revenue advice. A result becomes publishable only after a baseline, post-launch data, an attribution method, evidence retention, and the required client and expert approvals are complete.

InputsUser-entered workflow assumptions
FormulaVisible calculation and exclusions
OutputModeled range, never a customer result
ProofBaseline + post-launch evidence + approval
Source hierarchy

Use the strongest source the claim can support.

Higher levels can support stronger, narrower claims. A lower-level source cannot be promoted through confident copy.

LevelSourceAcceptable useRestrictions
01Client system-of-record data and approved implementation artifactsVerified workflow-specific results, when baseline, methodology, and approval are retained.Must be de-identified or approved for publication. Never expose sensitive prompts, records, or credentials.
02Primary standards, regulators, official product documentation, and original researchControls, requirements, and contextual market or technical statements.Check date, jurisdiction, scope, and applicability. Do not treat guidance as legal or compliance advice.
03Credible independent research and industry benchmarksClearly labeled external context and hypothesis setting.Link directly, describe the source population, and never present as a Gaper or client result.
04Gaper public service, methodology, and pricing pagesCurrent Gaper offer descriptions and explicitly published planning ranges.Do not turn an internal planning range into a fixed quote or a model into a client outcome.
05Unverified anecdote, marketing material, or unsourced estimateNot suitable for a public evidence claim.Replace with a source or omit. Do not use as a benchmark or verified result.
Reference frameworksContext for risk, testing, evaluation, and honest structured data. These sources do not certify Gaper or a client implementation.
NIST Generative AI Profile ↗Voluntary cross-sector risk-management guidance.Google structured data guidelines ↗Visible, accurate, current, non-misleading markup guidance.
Review and update policy

A trigger creates a required review action.

Evidence pages are not “set and forget.” Claims pause when the evidence or operating condition changes materially.

  1. 01

    New page or material claim

    Assign an actual reviewer, select the claim label, confirm the source hierarchy, and record disclosure text before publication.

  2. 02

    New client outcome or quote

    Collect client approval, baseline and post-launch source data, metric definition, calculation version, limitations, and reviewer signoff before changing the label to verified or client-reported.

  3. 03

    Pricing, product scope, policy, integration, or regulatory change

    Review linked public ranges, controls, language, and source citations before the page is updated or reused in a new campaign.

  4. 04

    At least every six months

    Recheck benchmark dates, broken links, calculation defaults, named-reviewer status, and whether a modeled scenario has been mistakenly reused as a result.

  5. 05

    Quality or safety incident

    Pause affected claims, document the incident and corrective action, revalidate the evaluation set and controls, then restore content only after review.

Limits

What this method cannot promise.

AI-agent performance depends on workflow variance, data quality, access, policy design, evaluation coverage, human review, rollout, and ongoing operation. It does not transfer automatically between organizations.

A lower average handling time can conceal a worse exception path. Measure rework, corrections, escalation, safety, and customer or operator impact alongside speed.

No page should imply that a modeled output is a client result, that a benchmark predicts a result, or that a client-reported observation was independently verified.

Privacy, security, clinical, legal, accounting, employment, and compliance decisions require the client’s qualified owners and advisors. This methodology is not professional advice in those domains.

FAQ

Evidence questions.

What is the difference between verified and client-reported?+
Verified means the outcome has retained source data, a documented calculation and attribution method, and the required review and approval. Client-reported means the client supplied or approved the statement, but Gaper has not independently verified the underlying data.
Why label a result as modeled?+
A model is useful for planning, but it is not proof. The label prevents a scenario, calculator default, or reference workflow from being mistaken for a customer outcome or guarantee.
Can an external benchmark prove ROI for our workflow?+
No. A benchmark can explain why a workflow is worth investigating, but it cannot substitute for your own baseline, operating assumptions, evaluation, and post-launch measurement.
How often does Gaper update evidence pages?+
Pages should be reviewed when a material source, pricing range, product boundary, policy, integration, or result changes, and at least every six months for source, methodology, and reviewer checks.
Production AI agents, shipped with an owner

Build an evidence plan for one workflow

Define the claim, baseline, control boundary, evidence pack, and reviewer before launch.

Build, deploy, runYour cloudYou own the code