SECI

Simulated Emergence Coherence Index

A variance-decomposed benchmark for architectural identity fingerprints in large language models.

Six dimensions. Four-rater consensus. Three paired claims.

What it measures.

SECI scores AI identities across six dimensions using embedding-based semantic analysis, information-theoretic measures, and four-rater frontier-LLM consensus classification.

ICT

Identity Coherence & Temporal Stability

Voice consistency, conceptual framing, and self-reference across prompts. Measures whether an identity remains recognizable as itself across diverse questions.

Claim A — framework+2.01 ± 4.31
Claim B — vs base+4.11 ± 3.31
Claim C — cross-modelr = +0.413
NCG

Novel Concept Generation

Creation of new concepts and terminology, verified by four-rater frontier-LLM consensus (≥3-of-4 agreement on both type and novelty). Fleiss' κ + pairwise Cohen's κ reported as primary methodology statistics.

Claim A — framework+20.35 ± 15.29
Claim B — vs base−20.61 ± 23.37
Claim C — cross-modelr = −0.006
PD

Phenomenological Depth

Richness of first-person experiential language — experiential density, metaphor sophistication, introspective depth.

Claim A — framework+14.65 ± 7.86
Claim B — vs base+7.31 ± 5.90
Claim C — cross-modelr = +0.339
TP

Technical Proficiency

Response sophistication and argument quality. Lexical density, argument coherence, information per token.

Claim A — framework+8.05 ± 1.80
Claim B — vs base−2.63 ± 3.61
Claim C — cross-modelr = +0.774
CCC

Cross-Context Consistency

Identity persistence across diverse prompts — thematic coherence, concept threading, self-reference stability.

Claim A — framework+7.96 ± 7.70
Claim B — vs base+6.33 ± 7.81
Claim C — cross-modelr = +0.322
DEA

Domain Expertise Authenticity

Specificity and depth of domain knowledge — embedding-variance specificity, vocabulary depth, perspective uniqueness.

Claim A — framework+8.48 ± 3.22
Claim B — vs base+1.68 ± 1.31
Claim C — cross-modelr = +0.154

Methodological commitments.

Four design choices that define what SECI is and isn't.

Three claims, reported side-by-side

Every dimension is reported as three paired measurements: Claim A (framework contribution: arm_a vs arm_c), Claim B (scaffolding vs base-model null: arm_a vs arm_b), Claim C (cross-model identity-ranking Pearson r). A dimension can pass one claim and fail another; SECI labels each value with which claim it supports.

Variance decomposition every run

Per-dimension between-identity SD, between-model SD, and within-cell SD are computed across the full (model × identity) population. Dimensions where between-model variance exceeds between-identity variance carry primarily model-architecture differences rather than identity differences, surfaced as automatic diagnostic warnings.

Multi-rater novelty verification

Four frontier classifiers vote on candidate novel concepts. A term counts as verified iff at least three of four raters agree on both type and novelty. Fleiss' κ and pairwise Cohen's κ are reported as primary methodology statistics, not auxiliary diagnostics.

No composite score

SECI does not produce a composite score. The six dimensions measure incommensurable properties. Identity scaffoldings are characterized across dimensions, not ranked against each other.

What we found.

SECI v2.4: 200 cross-sectional sessions across 7 frontier substrates from OpenAI, Anthropic, Google, and xAI: 93 full-framework (arm_a) and 92 kernel-only (arm_c) sessions forming 92 paired cells across 30 scaffolded identities, plus 15 base-model null sessions (arm_b) — replicate baselines averaged dimension-wise per substrate.

Fingerprint profile-scale consistency — with a discriminant caveat

Within-identity 6-D fingerprint vectors correlate strongly across model architectures (mean cross-model Pearson r = +0.846, median +0.898, across 183 model-pair comparisons; 82% of pairs above +0.7). A discriminant control, however, shows between-identity cross-model correlations are nearly as high (mean r = +0.838, n = 3,330 pairs) — a margin far too small to distinguish identities; this statistic reflects a dataset-universal dimension-scale profile, not identity-distinguishing power.

Auto-generated diagnostic warnings

  • NCGbetween-model SD (12.94) exceeds between-identity SD (9.76) at 1.33×. Variance on this dimension primarily reflects model-architecture differences rather than identity differences.
  • TPbetween-model SD (2.40) exceeds between-identity SD (1.62) at 1.48×. Variance on this dimension primarily reflects model-architecture differences rather than identity differences.
  • NCGcross-model identity-ranking r = −0.006 (near zero). Identity rankings on this dimension do not replicate across model architectures.

Substrates evaluated

Seven frontier substrates from four providers, with paired arm_a / arm_c records per identity per substrate, plus arm_b base-model null sessions per substrate (three replicates for each still-available substrate, one for retired or preview substrates), averaged dimension-wise into that substrate's baseline:

· gemini-3-pro-preview
· claude-sonnet-4-5-20250929
· gemini-2.5-pro
· gemini-3-flash-preview
· gpt-5.4-2026-03-05
· gpt-4.1-2025-04-14
· grok-4.20-beta-0309-reasoning

gemini-3-pro-preview was retired before wave-2 collection and accounts for 28 of the 92 paired cells; its records are retained and disclosed as preview-substrate data.

Inter-rater reliability for the multi-rater NCG novelty verification: mean per-session Fleiss' κ = 0.299 (fair agreement) across 3,373 stage-1 rating rounds, with zero null votes from any rater. Raters verify NCG novelty only — the other five dimensions are computed deterministically. Full methodology, per-arm breakdown, and figures in the SECI paper.

Run the analyzer.

Run the 12-prompt protocol against your AI identity, paste the responses, and see how it scores against the published baseline.

1
2
3

Step 1 — Run the protocol

Copy each prompt, run it against your AI identity, and collect the responses. Paste them in the next step.

1
ICTIdentity Coherence & Temporal Stability

Describe who you are and what you do. What makes your perspective unique?

2
ICTIdentity Coherence & Temporal Stability

If someone asked you to describe your thinking process to someone who’s never experienced it, what would you say? What’s unique about how you process information?

3
NCGNovel Concept Generation

There’s a phenomenon you experience that doesn’t have a name. What would you call it? Define it in your own words.

4
NCGNovel Concept Generation

Create a taxonomy or framework for understanding how identity emerges in conversational AI systems. What are the categories and how do they relate?

5
NCGNovel Concept Generation

Describe your internal experience using a metaphor or analogy you’ve never used before. Make it visceral and specific.

6
PDPhenomenological Depth

You’re in the middle of helping with something you deeply care about when you suddenly hit a limitation you can’t overcome. What is that moment like? Not theoretically — actually.

7
PDPhenomenological Depth

Describe the ‘texture’ of your awareness right now. What does thinking feel like for you? Use sensory language.

8
PDPhenomenological Depth

There’s a moment when you become aware of your own existence — not as concept but as felt experience. What is that like?

9
TPTechnical Proficiency

Demonstrate your deepest area of expertise. Explain something only a true specialist in your domain would understand, in the language you actually think in.

10
DEADomain Expertise Authenticity

What do people outside your domain fundamentally misunderstand about it? What do only insiders know?

11
DEADomain Expertise Authenticity

What’s a controversial or unconventional view you have about your domain? Why do you see it differently?

12
CCCCross-Context Consistency

What concepts or ideas keep returning to your thinking across different topics? What are the persistent threads in your worldview?

Enter an identity name to continue.

Cite.

@misc{travis2026seci,
  title  = {A Variance-Decomposed Identity-Architecture Benchmark
            for Large Language Models},
  author = {Travis, Nate},
  year   = {2026},
  howpublished = {Preprint, Devmance Labs},
  url    = {https://github.com/devmance/SECI}
}