EVIDENCE & VALIDATION SCOPE
Clear claims.
Defined boundaries.
A credible evaluation distinguishes what has been demonstrated internally from what still needs to be proven in your environment.
Status: September 2026. Internal engineering validation and pre-pilot discussions are not a certification, a partner endorsement or validation in a customer's environment.
What a partner evaluation should establish
One agreed command path is integrated with Gatekeeper. The partner defines the authorized action, scope and operating conditions. Testing then checks release behavior and the associated evidence.
| Test condition | Required behavior on the tested path |
|---|---|
| Valid authorization for an approved task | ALLOW: release the authorized request after verification. |
| Changed task or unauthorized requester | DENY: do not release the new request. |
| Replayed authorization | Do not permit an additional release using the previous authorization. |
| Verification unavailable or incomplete | HOLD: withhold the new request until the conditions for release are satisfied. |
| Invalid evidence | Withhold the request; record the verification outcome. |
AI agents, humanoids and industrial systems
These application areas describe proposed validation scopes, not completed integrations. An AI-agent evaluation can test authority over tool or API calls; a humanoid evaluation can test task release; an industrial evaluation can test production-job release. Each requires an agreed interface, coverage of the relevant execution routes and its own acceptance evidence.
Execution authorization is separate from the quality of a model’s reasoning. Gatekeeper evaluates whether an action may be released on the integrated path; it does not certify every model output as correct or safe.
Internal engineering evidence
The Gatekeeper API and command-release mechanism have been implemented and tested internally for the NAVIS integration. Internal tests covered authorized single-request release and withholding when authorization or verification conditions were not satisfied.
These tests do not establish behavior in a partner’s software or physical system. A technical review should examine the test setup, accepted and rejected inputs, observed downstream requests and the evidence records. A scoped evidence package can be discussed without exposing source code or proprietary implementation.
Partner discussions
Pre-pilot technical discussions with AgileX Robotics concern a bounded Gatekeeper validation for NAVIS. This is a discussion stage, not a claim of an accepted or running pilot, endorsement or external certification.
Recorded internal measurements
Source: controlled technical validation record v0.3, dated 1 September 2026. The record reports repeated internal measurements on the production host; it is not independent certification.
| Measured scope | Recorded result |
|---|---|
| Authorization gate plus proof | 319–320 ns |
| Complete protected authorization path | 494–497 ns |
| Decision-to-proof gap | 0 ns reported at the governed boundary |
| Complete-path throughput | Approximately 2.01 million decisions/s per CPU core |
Method: repeated batch-averaged direct execution, without a per-call timer wrapper. The ranges are recorded batch-average observations, not percentiles, worst-case latency or guarantees for individual requests. End-to-end application, network and machine response must be measured separately on the target system.
Observed release and evidence behavior
The same record describes a synthetic payment-release test with a simulated downstream receiver. A valid request was released; HOLD and DENY cases were withheld; an altered evidence copy was rejected by a logically separate verifier. No live customer endpoint, production funds or customer credentials were involved.
The reported zero decision-to-proof gap concerns the linkage between an authorization decision and its evidence in that test. It does not mean zero processing time or zero latency throughout a distributed system. The record does not establish performance on a partner’s hardware, humanoid or swarm.
For a partner evaluation, the hardware, software build, workload and timing method must be agreed alongside the results to be reproduced. A sanitized evidence review can be arranged without disclosing source code or internal mechanisms.
What this does not replace
- Functional safety, emergency-stop circuitry or certified motion control.
- A complete assessment of all routes through which a machine can be commanded.
- The customer's task policy, operating procedures or accountable operator.
- Testing of actual interfaces and release behavior on the target system.
Withholding a new command is not the same as stopping a machine already in motion. Claims apply only to the integrated and tested command path.
Meduza: defense, swarm coordination and recovery
Meduza’s capability scope extends from protecting individual runtimes to coordinating distributed agents and autonomous machines. It includes active defensive response, containment, swarm coordination, self-repair and recovery under disrupted operating conditions. Its protective operation does not require an LLM.
A partner evaluation can assess observable outcomes: how the integrated group responds when a member becomes unavailable, whether affected components are contained, and whether protected functions recover within the agreed operating limits. Connectivity conditions, platform coverage and acceptance criteria are agreed for each evaluation.
Gatekeeper remains the independent authorization boundary for consequential actions on integrated execution paths. Group coordination does not replace platform safety or establish that every requested action is authorized.
Validate outcomes. Protect the implementation.
Public material describes capability and application scope. Source code, internal coordination methods, defensive decision rules and recovery mechanisms are not published. Qualified partners can discuss a bounded black-box evaluation under an appropriate confidentiality agreement, using synthetic inputs and agreed observable results.
Swarm behavior, disruption tolerance, recovery and performance claims require evidence for the particular deployment and test conditions. They are not universal guarantees across every platform or failure scenario.
What you receive from a scoped validation
- An agreed description of the action and authorization conditions.
- A synthetic test matrix, expected outcomes and observed results.
- Decision-bound authorization records and a review procedure.
- A concise account of assumptions, limitations and next integration steps.