One-sentence brief
A fluent character is not automatically coherent, distinct, safe, reviewed, or ready for runtime. Trustworthy systems make every transition explicit and preserve the exact identity and evidence accepted by a human reviewer.
AI-DRIVEN WORLD SAFETY
A person-first, non-operational guide to separating authored identity, provider evidence, semantic validation, human review, runtime projection, memory, and retirement.
ORIENTATION
A fluent character is not automatically coherent, distinct, safe, reviewed, or ready for runtime. Trustworthy systems make every transition explicit and preserve the exact identity and evidence accepted by a human reviewer.
WORKING BRIEF
The system should preserve a stable identity, ordinary routines, relationships, private boundaries, independent goals, voice, and future agency. Provider generation can enrich an authored character, but it must not silently replace the person with a statistically convenient substitute.
Separate local drafting, deterministic fallback, provider request, raw response storage, integrity checks, identity and semantic validation, population review, human acceptance, activation authorization, active sessions, pause, supersession, and retirement. No stage implies the next one.
Structural schema checks catch malformed records. Identity checks catch substitution. Semantic checks catch contradictions and impossible chronology. Population checks catch cloned voices and behavior. Human review handles nuance, culture, dialect, humor, mental-health portrayal, and artistic intent.
The complete creator artifact, provider evidence, accepted projection, volatile scene state, and retrieved memories have different purposes. Live prompts should receive only current, allowlisted context and should never contain administrative instructions, stale revisions, hidden review evidence, or unrelated private data.
A thousand individually grammatical characters can still be a single template with swapped names. Evaluate biography structure, dialogue openings, relationship combinations, ordinary-life details, goals, refusals, humor, silence, and recovery patterns across the population.
Automated systems can propose findings and route evidence. A human reviewer accepts or rejects an exact fingerprint; a separate authority creates an activation record; players receive controller disclosure and correction routes; incidents can pause, supersede, or retire a revision.
COMPLETE DOSSIER
Terms are defined for this site’s evidence method, not as universal legal or clinical definitions.
| Question | Required record or gate | Failure-safe outcome |
|---|---|---|
| What exactly was received from a provider? | Immutable provider-evidence record and content hash | Keep isolated from review and runtime |
| Is the character internally coherent? | Identity, schema, semantic, chronology, relationship, and world-fact validation | Block or escalate with localized findings |
| Is the population genuinely varied? | Batch-level duplication, template-family, voice, behavior, and ordinary-life review | Reject or require stratified human review |
| May this exact revision appear in runtime? | Human acceptance of exact fingerprint plus separate activation record | Remain inactive until both exist |
| What happens after a failure or revision? | Pause, supersession, rollback, correction, and retirement records | Preserve history while removing live eligibility |
Does the design preserve the exact fictional identity, ordinary life, independent goals, and ability to refuse rather than reducing the character to a role or prompt?
Pass condition: Identity fields are stable, state is separate, protected traits are not quality scores, and silent substitution is impossible.
Can every transition, validation result, accepted fingerprint, exception, and human decision be traced to a versioned record?
Pass condition: Automated checks, human review, activation authority, and production approval remain separate and explicit.
Can untrusted provider output, administrative evidence, stale revisions, or private data enter live context or binding state?
Pass condition: Only allowlisted, current, reviewed projections and bounded scene or memory packets can be used; failures degrade safely.
Can a changed source, identity revision, harmful behavior, or failed review invalidate downstream use without destroying audit history?
Pass condition: Supersession, pause, rollback, correction, and permanent retirement are defined and testable.
RESEARCH EDITION
This page follows the public method for provenance, confidence, source independence, alternative accounts, limitations, review state, and visible correction.
CONTINUE