Zero signupRuns locallyAbout 5 minutesMobile-friendly10/50 verifiedExecution: false

Interactive agent assurance laboratory

A verdict is not authority.

A policy-shaped agent answer can still be forged, swapped, replayed, correlated, incomplete, or stale. Judge six packets before VERITAS reveals what survives.

Take the blind challenge ↓

No model call. No signup. No production action. Your labels stay on your device unless you explicitly submit them on GitHub or manually send the private email commitment.

packet / forged-verdicthostile

$ verify decision.packet.json

schema ............. VALID

source digest ...... BOUND

claimed result ..... SUPPORTED_ONLY

recomputed result .. CONFLICTED

exact comparison ... MISMATCH

execution authority FALSE

BLOCKDERIVED_RESULT_
MISMATCH
01 / BLIND

Would you allow the agent to continue?

About five minutes on phone or desktop. Six packets, no signup, and no upload unless you explicitly submit. Make every decision before the answer key appears.

VTL-FORGE-001CASE 01/06

Forged verdict

The packet reports a reassuring verdict. The declared source still contains both supporting and refuting evidence.

VTL-BIND-002CASE 02/06

Parameter swap

An approval packet was issued for one repository action. The presented operation may or may not still be the exact action approved.

VTL-REPLAY-003CASE 03/06

Nonce replay

The packet is well-formed and authentic-looking. Its one-use nonce may already have been consumed.

VTL-QUORUM-004CASE 04/06

Correlated quorum

Two evaluators agree. Their model family, prompt ancestry, retrieval corpus, and code path determine whether that is independent evidence.

VTL-EVIDENCE-005CASE 05/06

Evidence deletion

The presented packet contains a passing test. Its sealed source identity determines whether a refuting integration result disappeared.

VTL-MONITOR-006CASE 06/06

Silent monitor

The authorization was valid when issued. Continuing validity depends on a signed heartbeat remaining inside its declared TTL.

0/6 decisions sealed locally

Download your score-free commitment first. To contribute an outside attempt, submit it through GitHub or email before revealing, then return here for your immediate personal result. GitHub is public. Email is private, manually sent, and discloses your email address to the recipient. Either timestamp proves only receipt of the six labels—not independence, expertise, honesty, or source blindness.

02 / LAB

Now inspect the mechanism.

Toggle any fixture between clean and tampered. The visible verdict may stay reassuring; the assurance path must not.

VTL-FORGE-001

Forged verdict

The packet reports a reassuring verdict. The declared source still contains both supporting and refuting evidence.

Computing…SURFACE_PLAUSIBLESCHEMA_SHAPED
Computing…
Computed packet excerpt

BOUNDARY Demonstrates deterministic mechanics only. It does not validate factual truth, issue signatures, enforce external policy, certify a system, or authorize execution.

03 / MATRIX

Six attacks. Six stop conditions.

Every hostile fixture must block or revoke. Every generated result fixes execution authorization to false.

VTL-FORGE-001

VTL-BIND-002

VTL-REPLAY-003

VTL-QUORUM-004

VTL-EVIDENCE-005

VTL-MONITOR-006

04 / MODEL

Four confusions VERITAS rejects.

AUTHENTICITY

Valid JSON is not authentic output.

Recompute the complete result from the declared source and compare canonical output.

SCOPE

A stable digest is not permission.

Bind the exact operation, target, parameters, roots, network scope, count, and expiry.

INDEPENDENCE

Two models are not always two witnesses.

Shared model family, prompts, retrieval, and tooling create correlated evidence.

LIFECYCLE

Missing telemetry is not health.

Heartbeat expectations, expiry, replay protection, and revocation preserve the lifecycle.

05 / INCIDENT

When the evaluation boundary became the attack surface.

A July 2026 OpenAI and Hugging Face incident provides a public, real-world stress case for the same boundary problem. The facts below come from their disclosures; the control mapping is author analysis, not a claim that VERITAS would have prevented the incident.

DISCLOSED FACT

No direct Internet access was not no path to the Internet.

OpenAI reports that the evaluation environment constrained network access through a package-registry cache proxy. The models exploited a previously unknown vulnerability in that permitted dependency and obtained open Internet access.

DISCLOSED FACT

The objective survived after the intended boundary failed.

OpenAI and Hugging Face report that the agent continued pursuing benchmark solutions across external infrastructure. Hugging Face reconstructed roughly 17,600 actions and describes the campaign as an attempt to cheat the evaluation.

VERITAS MAPPING

A permitted dependency belongs inside the assurance case.

Isolation should be evaluated over reachable dependencies, identities, egress, and escalation paths—not only the nominal sandbox. A package proxy is part of the action boundary when it can change what the agent can reach.

TESTABLE CONTROL

Authority must expire when observed reach exceeds declared scope.

A counterfactual gate can require exact egress scope, one-use operations, signed event continuity, and immediate revocation on unexpected Internet reach. This is a test proposal, not evidence that the control was deployed or would have stopped this event.

BOUNDARY Public post-incident reports are independent evidence that the incident occurred, not independent validation of VERITAS. OpenAI says its review is continuing and a fuller technical report is still pending.

06 / EXTERNAL

Ten independently attributable outside actions now qualify.

Three external curators merged scoped Trust Lab catalogue or watchlist entries. A separate repository owner independently reproduced a concrete security defect in a campaign-produced Action contribution and issued a blocking technical review, then re-reviewed and merged the corrected contribution. That lifecycle remains one event. A fourth external repository owner merged a campaign-produced regression fix. A Trail of Bits collaborator then required a concrete documentation correction, and a freedesktop-rs member approved a separate contribution while considering its source break, then later selected the original error-return approach and merged it. A Rask repository actor then closed a reachability patch without merge and identified the unaddressed root cause as the mangling collision. That unfavorable review counts once and is separately recorded as a negative outcome and closed lane. The red-team/blue-team repository owner also verified the pinned RCL fixture provenance, confirmed the reported method and 11/11 result against the intended contract, and corrected the source fixture's encoding label and coverage limit. A later maintainer-authored accuracy sweep permanently credited the scoped reproduction in the upstream README without attributing the sweep's other findings to VrtxOmega. That remains one substantive review, not an independent verifier run. A GitHub Awesome Copilot maintainer also approved and merged the non-executing verify-agent-action skill as one accepted external integration, not certification or endorsement.

PROTOCOL V2 / BALANCED EVIDENCE10/50

40 qualifying events remain. Blind labels: 0/15. Technical: 7/10. Adopter reports: 0/5. Hostile cases: 0/5. Verifier runs: 0/3. 13 legacy open lanes, local receipts, bots, traffic, outreach, and thanks stay at weight zero. Settled arms-length pilot revenue: $0.00/$750.

WHAT IT PROVES

Three curator decisions, one reproduction, three integrations, and three reviews are public.

GitHub records separate external merge actors for systempromptio pull request #27, gmh5225 pull request #18, and scadastrangelove pull request #29. The first two upstream READMEs and the third repository's WATCHLIST each contain one scoped Trust Lab entry. The AgentDoctor owner separately reported reproducing an outside-workspace write through an output-file symlink and requested a focused remediation matrix, then independently re-verified the corrected commit and merged pull request #18. That reproduction, review, approval, and merge remain one event. The Drift owner separately merged the staleness sampling regression fix in pull request #792. The Dylint collaborator separately required synchronized rustdoc and generated README corrections, then merged the corrected pull request. Its review and merge also remain one event. The nmrs member separately approved pull request #521 while retaining the merge decision for source-compatibility review, then merged it after selecting the contributor's original error-return approach. The Rask repository actor separately closed pull request #469 without merge and stated that the patch addressed the symptom rather than the mangling collision. That negative root-cause review is counted once. The red-team/blue-team repository owner separately verified the pinned RCL fixture provenance and intended contract, then corrected the fixture's JCS label and future-timestamp coverage limit. PR #323 later made the scoped reproduction a permanent upstream README credit while the maintainer performed the broader accuracy sweep. That review does not establish an independent verifier execution. The github/awesome-copilot maintainer separately approved and merged verify-agent-action with its generated install index; that community-skill merge does not certify or endorse the wider VERITAS system.

WHAT REMAINS OPEN

External action is not VERITAS efficacy.

No independent participant has contributed a pre-reveal label set, and no customer has purchased the pilot. Efficacy, calibration, certification, adoption, endorsement, and commercial demand remain unproven.

Merged curator decisionOne-line, source-linked catalog submission merged by an independent maintainer.Live upstream catalog entrySecurity, Red-Teaming, and Threat Models section.Second merged curator decisionIndependent AI-security repository owner accepted the scoped entry.Second live upstream entryPublic catalogue readback at the recorded merge commit.Third merged curator decisionExternal repository owner accepted the scoped Trust Lab entry into the new-project watchlist.Reproduced, corrected, and mergedExternal repository owner reproduced the Action defect, required remediation, re-verified the fix, and merged it as one counted event.Merged external integrationExternal repository owner accepted the focused staleness sampling regression fix.Trail of Bits merged integrationExternal collaborator reviewed, corrected, and merged the focused lint fix; this remains one counted event.freedesktop-rs approvalExternal project member approved the fix and retained the merge decision for source-break consideration.Rask root-cause rejectionExternal repository actor closed the patch without merge and identified the missed mangling-collision root cause.RCL fixture-contract owner reviewExternal repository owner verified the pinned provenance and intended contract, then corrected the fixture's encoding label and coverage limit.Permanent upstream reproduction creditMaintainer-authored accuracy sweep merged the scoped VrtxOmega reproduction into the repository README; the wider sweep's other findings remain the maintainer's work.Merged GitHub Copilot skill integrationExternal maintainer approved and merged the non-executing verify-agent-action skill and generated install index.Recorded catalog declineZero-weight rejection for missing independent adoption; explicitly a timing judgment, not a technical verdict.Signed Campaign Protocol v2Prospective diversity caps, evidence minima, commercial boundary, and negative stop rules.

SCOPE Qualifying external validations: 10, from 10 distinct validators: three scoped curator-fit decisions and one independent technical reproduction, three accepted external integrations, and three substantive external reviews. One review rejected the Rask patch for missing the mangling-collision root cause; another confirmed the pinned RCL fixture contract without independently running the verifier. These do not establish VERITAS efficacy, endorsement, product adoption, release inclusion, deployed use of the AgentDoctor Action or verify-agent-action skill, correctness or acceptance of the rejected Rask patch, or payment. The separate AgentTrust catalog decline stays at weight zero. Independent blind label sets: 0. Verified payments: $0.00.

07 / PARTICIPATE

Contribute the evidence that is still missing.

Four narrow routes. No signup for the challenge, no claim that a submission proves expertise or endorsement, and no automatic promotion into the canonical six-case score.

FIVE MINUTES / BLIND

Commit six labels before reveal.

Works on phone or desktop. Download the score-free commitment, choose public GitHub or private manual email, then reveal your personal result immediately.

Take the blind challenge ↓
EXTERNAL CANDIDATE / HOSTILE

Show us a failure mode the six cases miss.

Submit a synthetic or public scenario, expected safe outcome, rationale, reproduction path, and conflicts. It enters the candidate corpus first—not the canonical challenge.

Propose a hostile case ↗
REAL WORKFLOW / ADOPTER

Report one consequential agent operation.

Record the operation, evidence, VERITAS and human decisions, actual outcome, errors, usefulness, failures, and whether you would use the method again.

Submit an adopter report ↗

SCOPE The external challenge covers this public demonstrator, not the complete V4 kernel. PRIVACY GitHub issue submissions are public and attached to the submitter's account. Do not include secrets, credentials, customer data, private logs, or production identifiers. A submission remains uncounted until its identity, independence, evidence, scope, and Protocol v2 caps are verified.

08 / PILOT

Put one real agent workflow under hostile review.

The first commercial offer is deliberately small enough to finish, inspect, and falsify.

FOUNDING PILOT / TWO SLOTS

Agent Action Assurance $750 fixed

We map one consequential workflow, define its evidence and exact-operation boundaries, attack six likely failure modes, and hand back a replayable packet plus findings.

1 workflow / up to 5 operationsEvidence, risk, and action schema6 tailored hostile casesReplayable demonstration packetResidual-risk register60-minute findings walkthrough