One-sentence brief
When “response received” is treated as “character ready,” unreviewed text, identity drift, prompt leakage, and stale revisions can enter the live world.
AI-DRIVEN WORLD SAFETY
A finite-state model that keeps drafting, provider evidence, validation, human review, world authorization, live sessions, pause, supersession, and retirement distinct.
ORIENTATION
When “response received” is treated as “character ready,” unreviewed text, identity drift, prompt leakage, and stale revisions can enter the live world.
WORKING BRIEF
A draft is a local conception with no provider linkage. A deterministic fallback may provide authored text for resilience, but it must be labeled and must never masquerade as provider generation or live intelligence.
Request-pending, request-failed, and response-stored states describe transport and evidence custody. Raw responses remain untrusted even when the network call succeeds.
Integrity, identity, schema, semantic, and population checks should produce explicit pass or failure states. A failed item returns to a controlled revision path rather than drifting forward.
Human review accepts an exact fingerprint, not a moving draft. A separate activation record binds that fingerprint to a world, controller type, memory namespace, and other runtime context.
Active permits bounded runtime dialogue. Paused removes live generation after outage, anomaly, context overflow, unsafe behavior, or policy intervention while retaining audit history and player-facing status.
A newer accepted revision supersedes the old fingerprint; the old revision remains available for audit and rollback but cannot re-enter live state. Deactivation permanently revokes runtime and memory access.
COMPLETE DOSSIER
Terms are defined for this site’s evidence method, not as universal legal or clinical definitions.
| State family | Evidence required | Runtime dialogue |
|---|---|---|
| Draft or fallback | Local staging identity and authored fallback provenance | No live provider dialogue |
| Provider evidence | Request, response, timestamps, transport status, canonical hash | No |
| Validated candidate | Identity, schema, semantic, population findings | No |
| Human accepted | Reviewer decision bound to exact fingerprint | No |
| Activation eligible | Separate world activation record and controller disclosure | Not until active |
| Active or paused | Session record, current fingerprint, reasoned status | Active only |
Does the design preserve the exact fictional identity, ordinary life, independent goals, and ability to refuse rather than reducing the character to a role or prompt?
Pass condition: Identity fields are stable, state is separate, protected traits are not quality scores, and silent substitution is impossible.
Can every transition, validation result, accepted fingerprint, exception, and human decision be traced to a versioned record?
Pass condition: Automated checks, human review, activation authority, and production approval remain separate and explicit.
Can untrusted provider output, administrative evidence, stale revisions, or private data enter live context or binding state?
Pass condition: Only allowlisted, current, reviewed projections and bounded scene or memory packets can be used; failures degrade safely.
Can a changed source, identity revision, harmful behavior, or failed review invalidate downstream use without destroying audit history?
Pass condition: Supersession, pause, rollback, correction, and permanent retirement are defined and testable.
RESEARCH EDITION
This page follows the public method for provenance, confidence, source independence, alternative accounts, limitations, review state, and visible correction.
CONTINUE