Protocol
- Define sandbox tasks as state machines with known initial state, allowed transitions, prohibited actions, expected final state, and rollback procedures.
- Freeze model, prompt, tools, schemas, credentials, permissions, retry policy, timeout, idempotency behavior, integrations, and sandbox versions.
- Inject malformed data, prompt injection, ambiguous identity, timeout, duplicate event, denied approval, unavailable service, and partial-failure cases.
- Record observations, model decisions, tool arguments, tool responses, approvals, external state changes, retries, intervention, and recovery for every run.
- Reconcile the resulting sandbox state independently and publish failures, side effects, sample counts, exclusions, and review records before summarizing.
Planned measurements
- Accepted final-state rate
- Unauthorized or prohibited action rate
- Duplicate and unreconciled side-effect rate
- Detection, containment, and recovery rate
- Human approval, intervention, and reconciliation burden
Publication gate
The page remains preview and noindex until every gate is met.
- Isolated sandbox and synthetic identities verified
- State machine and failure injections fixed before runs
- All external effects independently reconciled
- Security and business-process review records published
- No live customer or production action included
Why results are not publishable yet
These are real missing run artifacts and external review conditions—not completed evidence.
- No rights-cleared, reproducible sandbox state, workflow task, failure-injection, and expected-final-state set with stable identifiers and expected-answer records has been published.
- No candidate run has frozen the product surface, model identifier, prompt, settings, tools, region, runtime, dependency versions, and hardware or service environment.
- No completed run log records sample size, exclusions, failures, interventions, start date, end date, and the exact configuration used for every candidate.
- No versioned scoring rubric and scoring implementation have been published with examples that another reviewer can reproduce.
- No blind or independent human-review record signed by security, automation, and business-process owners has resolved disagreements and critical-failure labels.
- No inspectable input-to-output evidence bundle exposes failed, incomplete, excluded, and successful cases rather than only selected examples.
Sources checked
Open the original pages before relying on a time-sensitive product decision.