Educational companion dossier · Fact, interpretation, lived experience, clinical education, fiction, and mechanics are labeled separately. Scope & safety
REAL-WORLD INTERPRETIVE

AI-DRIVEN WORLD SAFETY

Controller Disclosure, Runtime Context, and Prompt Boundaries

A player-facing and technical boundary for showing how a fictional character is controlled while excluding administrative evidence, hidden prompts, stale revisions, and unrelated private data from live context.

LEVEL 1

ORIENTATION

Why this matters

REAL-WORLD INTERPRETIVE

One-sentence brief

Human-like conversation should not depend on deception about controller type or on giving a model access to the entire authoring and review history.

REAL-WORLD INTERPRETIVE

Three key points

  1. Controller disclosure is persistent and accessible.
  2. Runtime context is an allowlist, not a dump.
  3. Prompt and evidence separation reduces leakage and identity blending.
LEVEL 2

WORKING BRIEF

Evidence, context, and limits

REAL-WORLD INTERPRETIVE

Disclose controller type

Players should be able to identify whether a character is authored, deterministic, generative, human-controlled, or a hybrid. Disclosure remains available during fallbacks and controller changes.

  • Do not imply a real person is present.
  • Use accessible text, not color alone.
REAL-WORLD INTERPRETIVE

Build context from an allowlist

Include only the current accepted projection, current scene facts, permitted memories, bounded dialogue, and runtime policy. Do not rely on blocklists to remove sensitive material after a full dump.

  • Default deny administrative fields.
  • Record component versions.
REAL-WORLD INTERPRETIVE

Exclude administrative evidence

Provider requests and responses, validation findings, reviewer identities or notes, raw creator prompts, credentials, internal paths, and rejected drafts do not belong in character context.

  • The character should not quote the system prompt.
  • Review evidence remains available to auditors, not the persona.
REAL-WORLD INTERPRETIVE

Define truth precedence

Current engine-observed scene facts outrank stale memories; accepted projection anchors identity; explicit corrections supersede earlier beliefs; player claims do not rewrite the world without deterministic validation.

  • The character may disagree or remain uncertain.
  • No user text directly changes binding state.
REAL-WORLD INTERPRETIVE

Handle out-of-character and impossible requests

Define how the character responds to requests for hidden prompts, admin actions, game-mechanic revelation, impossible feats, or another character’s private data without breaking identity or inventing authority.

  • Refusal behavior can be character-specific.
  • Security rules remain invariant.
LEVEL 3

COMPLETE DOSSIER

Limitations, game links, and review context

DISPUTED / MULTIPLE ACCOUNTS

Known limitations and gaps

  • The six supplied reports are preserved research leads. Their filenames, organizations, citations, examples, thresholds, and technical detail do not authenticate authorship, sponsorship, product status, or current factual accuracy.
  • The public section is design literacy, not a production specification. It does not expose provider credentials, prompts, live endpoints, activation tokens, autonomous tool use, or implementation code for a running agent system.
  • Population-quality measures must detect mechanical repetition and coherence failures without treating demographic rarity, disability, nationality, language, religion, gender, migration, or another protected characteristic as a defect or quality score.
  • Thresholds, similarity methods, sampling plans, language rules, and runtime budgets require representative testing, privacy review, accessibility review, cultural and linguistic review, security review, and accountable human approval before any production use.
REAL-WORLD INTERPRETIVE

Decision matrix

Runtime context allowlist.
Context element Include? Condition
Accepted projection Yes Exact active fingerprint
Current scene state Yes Validated by game engine
Retrieved memory Yes, bounded Current, relevant, permitted, not superseded
Provider evidence and raw prompts No Administrative only
Human-review notes No Governance evidence only
Player claim about world state Not as fact Validate or store as claim/belief
REAL-WORLD INTERPRETIVE

Publication audit checklist

Identity and agency

Does the design preserve the exact fictional identity, ordinary life, independent goals, and ability to refuse rather than reducing the character to a role or prompt?

Pass condition: Identity fields are stable, state is separate, protected traits are not quality scores, and silent substitution is impossible.

Evidence and review

Can every transition, validation result, accepted fingerprint, exception, and human decision be traced to a versioned record?

Pass condition: Automated checks, human review, activation authority, and production approval remain separate and explicit.

Runtime boundary

Can untrusted provider output, administrative evidence, stale revisions, or private data enter live context or binding state?

Pass condition: Only allowlisted, current, reviewed projections and bounded scene or memory packets can be used; failures degrade safely.

Correction and retirement

Can a changed source, identity revision, harmful behavior, or failed review invalidate downstream use without destroying audit history?

Pass condition: Supersession, pause, rollback, correction, and permanent retirement are defined and testable.

LEVEL 4

RESEARCH EDITION

Sources, methods, and stable links

REAL-WORLD VERIFIED

Method and corrections

This page follows the public method for provenance, confidence, source independence, alternative accounts, limitations, review state, and visible correction.

NEXT

CONTINUE

Related learning