7. CONSOLIDATED TABLE OF GENUINELY TESTABLE CRITERIA
★★
28.5k 字 ·
源文件 darlin-consciousness-literature-report.md 第 947 行起 ·
other
How to read this. Every row is a criterion that (a) some citable source treats as an indicator or a measurable construct, and (b) can be operationalised on a small RL agent. "PASS/FAIL" gives the comparison that would count, with a threshold where the literature supplies one and an explicit *"no established threshold"* where it does not — which is the majority of cases, and that absence is itself a finding. Assumed test-bed for the "how to measure" column: a PPO/DQN-style agent in a partially observable gridworld or MiniGrid-class environment, with a recurrent core, an optional explicit workspace module, an optional explicit self-model head, and the ability to run lesions/ablations at fixed parameters.
#7A. Indicators derived from theory (Butlin et al. IDs)
| ID | Criterion | Source | How to measure on a small RL agent | PASS | FAIL | Caveat |
|---|---|---|---|---|---|---|
| RPT-1 | Algorithmic recurrence in input modules | RPT / 2308.08708 | Inspect architecture: does the perceptual path have recurrent connections (not just depth)? Then ablate recurrent connections at fixed params and measure performance drop on a task requiring integration over time | Recurrence present and removal causes a selective drop on temporal-integration tasks; feedforward clone with matched params fails them | A feedforward net with matched params matches recurrent performance on all temporal-integration probes (the "unfolding" scenario) | Already trivially satisfied by most modern agents (Butlin et al. say so explicitly). Nearly zero discriminating power |
| RPT-2 | Organised, integrated perceptual representations | RPT / same | Linear probe on the perceptual layer for global shape/structure rather than local features; test crowding-style or global-vs-local stimuli with matched feedforward control | Recurrent agent shows global-structure sensitivity where matched feedforward control is at chance | Feedforward control matches → no global integration | The direct human/machine evidence: Doerig et al., Crowding Reveals Fundamental Differences in Local vs. Global Processing in Humans and Machines, arXiv 2004.12676 — "ffCNNs cannot produce human-like global shape computations for principled architectural reasons" |
| GWT-1 | Multiple specialised modules running in parallel | GWT / same | Count task-dissociable modules (lesion each: does it produce a specific deficit?); measure specialisation (mutual information between module state and subtask) | Modules are task-dissociable and low-redundancy | Monolithic representation; module lesions cause uniform degradation | Trivially satisfiable by any modular net; combine with GWT-2/3 to be meaningful |
| GWT-2 | Limited-capacity workspace → bottleneck + selective attention | GWT / 2308.08708, 2103.01197 | Sweep workspace width k and measure: (i) task performance, (ii) specialisation, (iii) compositionality on held-out module pairings. Cross-check with a compete-for-access gate vs. always-broadcast control | Inverted-U: some k is optimal, and bottleneck beats unlimited bandwidth on compositionality/generalisation at matched params | Performance monotonically increases with bandwidth (bottleneck is pure loss) | Goyal et al. give the rationale; Phua arXiv 2512.19155 shows graded collapse with partial reductions and qualitative collapse with full lesion |
| GWT-3 | Global broadcast: information in the workspace available to all modules | GWT / same | Lesion test: cut the broadcast pathway while keeping the workspace representation intact (this is exactly the "no-rewire" lesion design). Then measure whether downstream modules still carry task-relevant information (probe decoding) | Removing broadcast causes qualitative collapse in access-related markers (probe decodability of workspace content in every module goes to near-zero) | Broadcast removal leaves performance and probe decodability unchanged → the "workspace" is decorative | This is the strongest GWT test and the one with an implemented precedent (Phua, Exp. 2) |
| GWT-4 | State-dependent attention; use workspace to query modules in succession for complex tasks | GWT / same | Train on multi-step compositional tasks; measure systematic extrapolation to unseen operation sequences/arities; compare to param-matched LSTM/Transformer | GW model generalises to unseen compositions (interpolation and extrapolation) with fewer parameters | No extrapolation advantage, or only with more parameters | Direct precedent: Chateau-Laurent & VanRullen, arXiv 2503.01906 |
| HOT-1 | Generative, top-down or noisy perception modules | HOT / same | Verify the perceptual module actually generates (has a decoder / prior) and that top-down signals measurably modulate its activity; ablate top-down pathway | Top-down ablation selectively degrades degraded-input inference (e.g. occlusion/noise) but not clean-input performance | Top-down ablation has no differential effect | Careful: "noisy perception" is nearly free to satisfy |
| HOT-2 ★ | Metacognitive monitoring distinguishing reliable representations from noise | HOT / 2308.08708, SEP consciousness-higher | Force-choice task with graded evidence (perceptual noise, occlusion, evidence distance). Agent emits confidence. Compute meta-d' (type-2 SDT: how well confidence separates correct from incorrect, in type-1 d' units) and meta-I₂ᵣ for open-ended cases. Also compute the withdraw delta (withdrawal rate on incorrect minus on correct) | meta-d' significantly > 0 with a permutation null excluded; z-ROC slope estimated; metacognitive sensitivity not reducible to accuracy or to a confidence-threshold shift | meta-d' ≈ 0, or confidence is a monotone function of the evidence variable with no trial-level discrimination, or the score does not survive re-scoring | This is the single most defensible measurable criterion in the literature. Sources: Servajean arXiv 2603.29693 ("gold standard"); Cacioli arXiv 2603.25112; Cacioli arXiv 2604.15702; Fleming & Lau 2014 doi 10.3389/fnhum.2014.00443. Use a 0–20 confidence scale, not 0–100 — 78%+ of 0–100 responses concentrate on three round numbers (Dai & Wang arXiv 2603.09309) |
| HOT-2b ★ | Selective necessity of the self-model for metacognition | HOT / 2512.19155 | "No-rewire" self-model lesion: freeze/replace the self-model pathway with a fixed transform, keeping parameters and first-order pathway identical. Re-run the metacognition battery | Lesion abolishes metacognitive calibration (meta-d' → ~0) while preserving first-order task accuracy → synthetic blindsight | Lesion degrades first-order accuracy too (non-selective), or has no effect on metacognition (module is not causally load-bearing) | This is the double dissociation. It is the only causal, agent-level protocol I found with a published precedent |
| HOT-3 | Agency via a general belief-formation/action-selection system; strong disposition to update beliefs per metacognitive output | HOT / 2308.08708 | Perturb the metacognitive signal (e.g. invert or randomise it) and measure whether beliefs/policy change accordingly on a subsequent trial. Measure policy sensitivity to the self-monitoring channel | Inverting the metacognitive channel produces a systematic, large behavioural change (e.g. exploration rate, abstention, replanning) | Agent's decisions are statistically unaffected by module output | Direct precedent and warning: Xie arXiv 2604.11914 found auxiliary modules collapsed to near-constant output (confidence std < 0.006) and "Policy sensitivity analysis confirms the agent's decisions are unaffected by module outputs." Their conclusion: "self-monitoring should sit on the decision pathway, not beside it" |
| HOT-4 | Sparse and smooth coding generating a "quality space" | HOT / same | Measure population sparsity and the smoothness/topology of the representational manifold (e.g. RSA + dimensionality + smoothness of similarity gradients) | Sparse codes with a smoothly varying similarity structure that supports generalisation to novel stimuli along the manifold | Dense, jagged, or degenerate manifold | Weakest-supported HOT indicator; no established threshold. Doerig et al. (arXiv 2412.20873) argue structural correspondence alone is insufficient — you also need sensitivity, organisation, exploitation, contextualization |
| AST-1 ★ | A predictive model representing and controlling the current state of attention | AST / 2308.08708, 2411.00983 | (i) Add an attention-schema module; verify it predicts the agent's own attention state (decode attention from the schema). (ii) Parameter-matched complexity control. (iii) Test whether the schema improves prediction/categorisation of other agents' attention | Schema predicts own attention above chance and improves performance on other-agent attention tasks, at matched parameter count | Improvement vanishes under the parameter-matched control (then it was just capacity), or the schema fails to predict own attention | The parameter-matched control is what makes this publishable: Farrell, Ziman & Graziano arXiv 2411.00983. Emergence condition: attention must be uncertain and hard to track — Piefke et al. arXiv 2402.01056 |
| AST-1b | Attention schema is necessarily incomplete (not a perfect self-predictor) | Steel arXiv 2011.05294 | Test whether the schema's prediction error on its own attention stream has a non-zero floor that cannot be reduced by more capacity/training | Residual self-attention prediction error stays bounded away from zero as capacity grows | Error → 0 with capacity (schema is complete) | A theorem-backed prediction, so a clean falsification target. Also the constructive bridge to Frankish's "illusion problem" |
| PP-1 | Input modules using predictive coding | PP / 2308.08708 | Verify existence of a top-down prediction pathway + error units; ablate the prediction pathway and measure the differential cost on predictable vs unpredictable input | Prediction ablation costs more on predictable input (i.e. the agent was exploiting predictability) | No differential cost → no predictive coding in use | Near-trivial to satisfy; predictive objectives are standard |
| AE-1 | Agency: learning from feedback, pursuing goals, flexible responsiveness to competing goals | Agency & embodiment / 2308.08708 | Goal-switching task: present two conflicting goals with changeable priority; measure re-planning latency and success under switches; also test "flexible responsiveness" by introducing a new goal at test time | Rapid, systematic re-planning under goal switches; success on novel goal combinations | Policy locked to a single goal; no re-planning | Butlin et al. note this is already arguably met by existing agents — very low discriminating power |
| AE-2 | Embodiment: modelling output–input contingencies, including systematic effects, and using the model in perception/control | Agency & embodiment / same | Learn a forward model of own action → next observation; verify it is used (not just trained): lesion it and measure control degradation; test on sensor/actuator perturbation | Forward self-model predicts own action–observation contingencies; lesion selectively degrades control under perturbation but not under normal conditions | Model trained but lesion has no effect (Xie's failure mode); or no perturbation-specific deficit | This is the entry point to the self-model section (§8) |
#7B. Additional testable criteria from outside the Butlin list
| Criterion | Source | How to measure on a small RL agent | PASS | FAIL | Caveat |
|---|---|---|---|---|---|
| Workspace bandwidth is functionally necessary (graded) | Phua arXiv 2512.19155 | Full workspace lesion vs. graded k-reductions; measure access markers (probe decodability, task accuracy) | Full lesion → qualitative collapse; partial → graded degradation | All-or-nothing response, or no response | GWT's "ignition" framing predicts gradedness |
| Broadcast amplifies internal noise (a negative GWT finding) | Phua arXiv 2512.19155 | Inject latent perturbations; compare fragility of GW-style broadcasting vs. a non-broadcast control family | GW broadcasting shows extreme fragility to latent perturbation vs. control | GW is no more fragile than control | Reported as a real effect. If Darlin uses a GW, test this — it is a design risk, not a selling point |
| PCI-A (perturbational complexity) is a valid IIT-adjacent proxy in engineered agents | Phua arXiv 2512.19155 — negative result | Compute a perturbational complexity index on the agent under the workspace bottleneck vs. baseline | (Predicted by naive IIT transfer) PCI-A rises with workspace capacity | **Observed: PCI-A decreases under the workspace bottleneck** | Cite as a caution: "cautioning against naive transfer of IIT-adjacent proxies to engineered agents" |
| Multimodal robustness as an emergent GW property | Maytié et al. arXiv 2502.21142 | Train GW-world-model agent; then zero out one observation modality at test | Robust performance under modality dropout, and the matched baseline fails | Both fail, or baseline equally robust | Needs a strong baseline; "emergent" claims need the control |
| Sample efficiency of GW latent imagination | same | Count environment steps to reach a performance target, GW vs. PPO/Dreamer baselines | GW-Dreamer reaches target in measurably fewer env steps | No advantage | Careful reward/task matching |
| GW models need 4–7× less matched cross-modal data | Devillers et al. arXiv 2306.15711 | Vary the amount of paired data; measure alignment/translation quality with and without the shared workspace + cycle-consistency | Shared workspace achieves the same quality with 4–7× less paired data; both the workspace and cycle-consistency are necessary (ablate each) | No data-efficiency gain, or gain survives removing the workspace (then it wasn't the workspace) | Cleanest GW ablation design in the literature |
| Attention schema emerges only under attention-tracking uncertainty | Piefke et al. arXiv 2402.01056 | Manipulate how uncertain the agent is about its own attentional window; let it allocate free resources; probe whether a schema forms | Extra resources develop an attention schema when attention is hard to track; probe decoding of attention from the extra module rises with uncertainty | Schema forms regardless of uncertainty (then it's not tracking attention), or never forms | Emergence condition is the prediction, not mere presence |
| Standard active inference is Bellman-optimal only at horizon 1 | Da Costa et al. arXiv 2009.08111 | On a POMDP with a known optimal solution, run standard active inference vs. sophisticated inference; measure optimal-action rate by horizon | Standard: optimal at horizon 1, sub-optimal beyond. Sophisticated: optimal at all finite horizons | Standard active inference is optimal at all horizons (would falsify the claim), or sophisticated inference is also sub-optimal | A rare crisp quantitative prediction — built for a small agent |
| Active inference works with no reward signal | Tschantz et al. arXiv 2002.12636 | Remove the extrinsic reward channel entirely; test on sparse/dense/no-reward benchmarks | Robust performance with no rewards at all | Collapse without reward | Strong and unusual prediction. But note Millidge arXiv 1907.03876: deep active inference resembles MaxEnt policy gradients, so "no reward" may be a hidden-prior effect |
| Active inference is robust to (even improved by) observation noise | Noel et al. arXiv 2106.02390 | Sweep observation noise; measure performance | Performance is flat or improves with added observation noise | Performance monotonically degrades | Counter-intuitive; genuine falsification target |
| Directional information-seeking beats undirected exploration in sparse-reward tasks | Schneider et al. arXiv 2206.10313 | Sparse-reward manipulation task; compare EFE-driven exploration vs. ε-greedy/count-based | Information-seeking agent solves tasks on which undirected-exploration baselines fail | Baselines match or win | The comparison baseline must be strong |
| Meta-d' must be reported alongside (or instead of) ECE/Brier | Cacioli arXiv 2603.25112; Servajean arXiv 2603.29693 | Compute ECE, Brier, and meta-d' on the same trials. Show they dissociate | ECE/Brier and meta-d' dissociate (a model can be well-calibrated yet metacognitively insensitive) | They are redundant → then ECE suffices | Direct empirical demonstrations exist in LLMs; port to RL |
| Confidence scale must be validated before use | Dai & Wang arXiv 2603.09309 | Compare 0–100 vs 0–20 vs boundary-compressed scales; measure meta-d' under each | Scale choice does not change the qualitative conclusion | Conclusion flips with scale → your measurement, not the agent, is driving the result | 78%+ of 0–100 responses pile on three round numbers |
| Prospective regulation dissociates from retrospective monitoring | Cacioli arXiv 2604.15702 | Measure (i) retrospective confidence accuracy and (ii) prospective KEEP/WITHDRAW or BET/decline behaviour; correlate | Low correlation between the two (reported r = .17) | High correlation → one construct, not two | Two-probe design (Koriat & Goldsmith 1996) is directly implementable |
| Metacognitive sensitivity has a functional payoff independent of accuracy | Li & Steyvers arXiv 2507.22365; Guo et al. arXiv 2605.08710 | Human/agent-in-the-loop experiment, or derive the complementarity bound and test team vs. best member | Lower-accuracy/higher-meta-d' agent improves joint accuracy; or team beats best member iff error correlation ρ < ρ* (≈ a in the symmetric near-chance regime) | No complementarity, or bounds violated | Guo et al. give an impossibility result too; predicts well on human data (R≈0.93–0.94) |
| Self-monitoring modules must be on the decision pathway (not auxiliary) | Xie arXiv 2604.11914 | Compare three arms: (a) auxiliary-loss add-on, (b) structurally integrated into policy, (c) parameter-matched control with no module. 20 seeds, report effect sizes | Integrated > add-on and integrated > parameter-matched control | Add-on ≈ no-module (the observed failure: confidence std < 0.006, decisions unaffected) — or integrated ≈ no-module control (their null: d = 0.15, p = 0.67) | ⚠️ This paper is largely a negative result even for structural integration. It is the best available caution against self-monitoring theatre. Report effect sizes and the parameter-matched control |
| Self-model is causally necessary for control under perturbation, not for normal control | Bongard, Zykov & Lipson, Science (2006), doi 10.1126/science.1133687 | Let the agent infer its own structure from actuation–sensation relations; then damage it (remove a limb / perturb actuators). Does it adapt via the self-model? | Agent infers its own structure and "when a leg part is removed, it adapts the self-models, leading to the generation of alternative gaits" | No adaptation after damage, or adaptation without self-model use | The gold-standard behavioural self-model test. 4-legged robot; portable to a simulated agent |
| Self-model enables zero-shot task learning in imagination | Kwiatkowski & Lipson arXiv 1910.01994 | Train a predictive self-model; then run RL inside the self-model only; transfer the policy to the real environment with no further data | New tasks solved with zero additional real-environment data | Policy fails to transfer | "allows for learning new tasks without necessitating any additional data collection, essentially allowing zero-shot learning of new tasks" |
| A distilled policy functions as a usable self-model for planning | Yoo, de la Torre, Yang arXiv 2306.04440 | Compare model-free-only vs. dual-policy (model-free + distilled self-model) agents | Self-model stabilises training, speeds inference, promotes exploration, and "could learn a comprehensive understanding of its own behaviors" | Self-model distillation costs more than it buys | Reports the cost honestly: "at the cost of distilling a new network apart from the model-free policy" |
| Self/world separation is necessary but insufficient for self-awareness | Pivovarov & Shumsky arXiv 2512.10985 | "What"/"where" pathway agent; probe whether the internal state separates self-position/state from world-state | Agent separates self from world and this confers a measurable advantage | No separation, or separation with no benefit | Explicit "Principle 1": "the ability to separate the self from the world is a necessary but insufficient condition for self-awareness" |
| Self-model must be detectable by probes, not merely declared | Immertreu, Schilling, Maier, Krauss arXiv 2411.16262 | Train a probe (linear/FF classifier) on the agent's internal activations to predict the agent's own spatial position/state | Probe predicts the agent's own state significantly above chance; a probe for the world state is separately decodable | No self-information in the representation beyond what a world-model probe already gives | "the agent can form rudimentary world and self models" — note the probe design, it is cheap and directly portable |
| Self-report consistency across contexts and sibling models | Perez & Long arXiv 2311.08576 | Train the agent to answer self-questions with known answers; then test: (a) paraphrase/context invariance, (b) agreement with a sibling model trained identically-but-differently, (c) resilience under prompt/observation perturbation, (d) interpretability corroboration | Self-reports are invariant across semantically equivalent rephrasings, agree across sibling models, survive perturbation, and corroborate with internal-state readouts | Self-reports change with paraphrase, or contradict internal readouts | The four criteria are the authors' own proposal. ⚠️ Also see Zeng et al. arXiv 2608.30980: "improved self-modeling may not arise from privileged access to the model's internal decision process." Accuracy is not enough; you need the causal pathway |
| Counterfactual self-prediction | Zeng et al. arXiv 2608.30980 | Ask: "would this prompt/observation change your final answer?" Verify against ground truth by re-running | Above-chance on a balanced counterfactual set, with errors characterised | Chance-level, or systematically wrong on simple counterfactuals (the observed result) | Observed: "non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior" |
| Identity persistence across decision steps (co-instantiation, not window-wise occurrence) | Perrier & Bennett arXiv 2603.09043 | Instrument scaffold traces; compute their Arpeggio/Chord persistence scores over grounded identity statements | Identity constraints are co-instantiated at a single objective step, not merely present somewhere in the evaluation window | The agent only "talks like a stable self" without being organised like one | "It separates talking like a stable self from being organized like one." A concrete, computable metric family |
| Mechanism-linked indicator assays with causal intervention (protocol, not a criterion) | Sanyal arXiv 2602.23232 | Recurrent persistence loop + affect proxy; fixed-parameter ablations; report effect sizes with lesion AUC drops | Dissociations link recurrence → persistence and affect-coupling → preference stability/scanning/lingering caution; lesion selectively reduces persistence (AUC drop 27.62, 27.9%) | Markers present but no lesion effect, or lesions are non-selective | The paper's own methodological thesis: "indicator-like signatures can be engineered and... mechanistic and causal evidence should accompany behavioral markers" |
| Welfare proxy: preference satisfaction, verbal vs. behavioural | Tagliabue & Dung arXiv 2509.07961 | Elicit stated preferences and revealed preferences (forced choice with actual task performance); test stability across semantically equivalent prompts and under cost/reward manipulation | Stated and revealed preferences agree, and are stable across paraphrases | Measures diverge, or responses change under paraphrase | The authors are themselves uncertain: "we are currently uncertain whether our methods successfully measure the welfare state of language models" |
#7C. Protocols that are not criteria but that the literature treats as necessary to make any of the above credible
| Requirement | Source | Why |
|---|---|---|
| Parameter-matched complexity control | Farrell et al. arXiv 2411.00983; Xie arXiv 2604.11914 | Without it, every "the module helps" result is attributable to extra capacity. Farrell et al. show the improvement is specific; Xie shows the benefit can vanish |
| Lesion/ablation with a double dissociation | Phua arXiv 2512.19155; Sanyal arXiv 2602.23232 | Only causal evidence discriminates between "the mechanism is used" and "the mechanism is present" |
| Pre-registration | Cogitate: Ferrante et al. doi 10.1038/s41586-025-08888-1; INTREPID: Corcoran et al. arXiv 2509.00555; Cacioli arXiv 2603.25112 | The field's own best practice. Note Chan et al./Cacioli: v1/v2 conclusions reversed after a scoring-bias correction |
| Permutation nulls + bootstrap CIs on every metacognition estimate | Cacioli arXiv 2603.25112 | Confidence metrics are easy to produce spuriously |
| Confidence-scale validation | Dai & Wang arXiv 2603.09309 | The scale can dominate the signal |
| Pre-specified contrastive conditions (evidence strength, missing information, conflict) | Nazzal arXiv 2608.14552 | Gives a within-subject normative baseline; confidence should track evidence strength |
| Report the negative results | Phua arXiv 2512.19155 (PCI-A); Xie arXiv 2604.11914 (null vs. no-module control) | Both papers are more credible because of them |