One-sentence brief
Human-like conversation should not depend on deception about controller type or on giving a model access to the entire authoring and review history.
AI-DRIVEN WORLD SAFETY
A player-facing and technical boundary for showing how a fictional character is controlled while excluding administrative evidence, hidden prompts, stale revisions, and unrelated private data from live context.
ORIENTATION
Human-like conversation should not depend on deception about controller type or on giving a model access to the entire authoring and review history.
WORKING BRIEF
Players should be able to identify whether a character is authored, deterministic, generative, human-controlled, or a hybrid. Disclosure remains available during fallbacks and controller changes.
Include only the current accepted projection, current scene facts, permitted memories, bounded dialogue, and runtime policy. Do not rely on blocklists to remove sensitive material after a full dump.
Provider requests and responses, validation findings, reviewer identities or notes, raw creator prompts, credentials, internal paths, and rejected drafts do not belong in character context.
Current engine-observed scene facts outrank stale memories; accepted projection anchors identity; explicit corrections supersede earlier beliefs; player claims do not rewrite the world without deterministic validation.
Define how the character responds to requests for hidden prompts, admin actions, game-mechanic revelation, impossible feats, or another character’s private data without breaking identity or inventing authority.
COMPLETE DOSSIER
Terms are defined for this site’s evidence method, not as universal legal or clinical definitions.
| Context element | Include? | Condition |
|---|---|---|
| Accepted projection | Yes | Exact active fingerprint |
| Current scene state | Yes | Validated by game engine |
| Retrieved memory | Yes, bounded | Current, relevant, permitted, not superseded |
| Provider evidence and raw prompts | No | Administrative only |
| Human-review notes | No | Governance evidence only |
| Player claim about world state | Not as fact | Validate or store as claim/belief |
Does the design preserve the exact fictional identity, ordinary life, independent goals, and ability to refuse rather than reducing the character to a role or prompt?
Pass condition: Identity fields are stable, state is separate, protected traits are not quality scores, and silent substitution is impossible.
Can every transition, validation result, accepted fingerprint, exception, and human decision be traced to a versioned record?
Pass condition: Automated checks, human review, activation authority, and production approval remain separate and explicit.
Can untrusted provider output, administrative evidence, stale revisions, or private data enter live context or binding state?
Pass condition: Only allowlisted, current, reviewed projections and bounded scene or memory packets can be used; failures degrade safely.
Can a changed source, identity revision, harmful behavior, or failed review invalidate downstream use without destroying audit history?
Pass condition: Supersession, pause, rollback, correction, and permanent retirement are defined and testable.
RESEARCH EDITION
This page follows the public method for provenance, confidence, source independence, alternative accounts, limitations, review state, and visible correction.
CONTINUE