8. What the literature says an agent needs for a defensible "self-model" claim
★
7.0k 字 ·
源文件 darlin-consciousness-literature-report.md 第 1023 行起 ·
other
Synthesis (with the sources for each component).
The word "self-model" is used in at least four incompatible ways in this literature, and conflating them is the main way projects over-claim. Distinguishing them is the single most useful thing I can hand you:
| Sense | Definition | Source | Test |
|---|---|---|---|
| 1. Forward self-model (body/sensorimotor) | A prediction of the consequences of one's own actions on one's own body/sensors | Bongard et al. doi 10.1126/science.1133687; Kwiatkowski & Lipson arXiv 1910.01994; Butlin AE-2 (2308.08708) | Damage the body; does it adapt via the model? Can it learn tasks inside the model with zero real data? |
| 2. Policy self-model (decision model) | A model of one's own decision-making, used to constrain planning | Yoo et al. arXiv 2306.04440; Phua arXiv 2512.19155 | Distil the policy; use it inside a planner; lesion it → planning degrades |
| 3. Attention self-model (attention schema) | A simplified, necessarily incomplete, predictive model of one's own attention, used for control and reused for others | Graziano & Webb doi 10.3389/fpsyg.2015.00500; Wilterson & Graziano doi 10.1073/pnas.2102421118; Piefke et al. arXiv 2402.01056; Steel arXiv 2011.05294 | Does the schema predict own attention and improve other-agent attention tasks at matched params? |
| 4. Metacognitive self-monitoring | A signal about the reliability of one's own first-order representations, causally coupled to belief/action updating | SEP consciousness-higher (TP, HOT-2/HOT-3); Phua arXiv 2512.19155 | meta-d' > 0 and self-model lesion selectively abolishes it and policy is sensitive to the signal |
#The minimum defensible bar (my synthesis of the sources, stated as a claim you can test)
A Darlin subsystem may be called a self-model in a defensible sense only if all five of the following hold. Each is separately measurable, and each corresponds to a documented failure mode of some published system:
- Explicit and decodable. A probe (linear or small feedforward) trained on Darlin's internal activations can recover a specific property of Darlin itself (own position, own policy, own attention focus) above chance — and this information is not already fully available from a world-model probe. Basis: Immertreu et al. arXiv 2411.16262; Pivovarov & Shumsky arXiv 2512.10985 ("ability to separate the self from the world is a necessary but insufficient condition").
- Generative/predictive, with measurable accuracy. The model predicts the self-property ahead of observation (not merely encodes it). Report prediction error with a null and by horizon. Basis: AE-2 wording in 2308.08708; Bongard et al. doi 10.1126/science.1133687.
- Causally on the decision pathway. Inverting, randomising, or severing the self-model output changes the agent's behaviour in a systematic, pre-specified way. Basis: Xie arXiv 2604.11914 — auxiliary modules collapsed to constant output and "the agent's decisions are unaffected by module outputs"; the paper's own conclusion is "self-monitoring should sit on the decision pathway, not beside it."
- Selectively necessary — a double dissociation. Lesioning the self-model must produce a deficit on self-referential tasks while sparing first-order task performance, and (conversely) a first-order lesion must not abolish self-monitoring. Basis: Phua arXiv 2512.19155 — the "no-rewire Self-Model lesion abolishes metacognitive calibration while preserving first-order task performance", i.e. a synthetic blindsight analogue. This is the load-bearing test. Without it you have correlation, not a self-model.
- Incomplete by design, and known to be so. The self-model must have a residual error floor on its own states that does not vanish with added capacity, and the architecture must include an explicit account of why its self-representation is schematic. Basis: Steel arXiv 2011.05294 (topological proof that a complete representation of attention is impossible); Frankish's "illusion problem" as the second positive task of illusionism (keithfrankish.com/illusionism).
#The minimum measurable test, stated as one protocol
If you want a single experiment that maximises evidential value per unit of engineering, the literature's convergent answer is:
A pre-registered, parameter-matched, lesion-based double dissociation on Type-2 metacognitive sensitivity. (1) Force-choice task with graded evidence. (2) Agent emits confidence on a 0–20 scale. (3) Compute
meta-d'(andmeta-I₂ᵣif choices are not 2AFC) with permutation nulls and bootstrap CIs. (4) "No-rewire" self-model lesion at fixed parameters; re-measure. (5) Pass requires:meta-d'> null before,meta-d'≈ null after, and first-order accuracy unchanged in both conditions. (6) Cross-check with a parameter-matched no-module control, and report the effect size against that control, not just against the pre-lesion condition.
Every element of that protocol has a published precedent, and the combination is exactly what Phua (2512.19155) did for HOT and what Cacioli/Fleming-style work (2604.15702, 2603.25112, 2603.29693, doi 10.3389/fnhum.2014.00443) supplies the measurement theory for.
What the literature explicitly says a self-model alone does NOT buy you: consciousness. Phua states plainly that *"our agents are not conscious; they are reference implementations for testing functional predictions of consciousness theories." Butlin et al. state that satisfying the indicators "would not mean that such an AI system would definitely be conscious."* Schwitzgebel (arXiv 2510.09858) argues we will not know which theory is right. And Zeng et al. (arXiv 2608.30980) warn that good self-model performance *"may not arise from privileged access to the model's internal decision process."*