7. CONSOLIDATED TABLE OF GENUINELY TESTABLE CRITERIA

★★ 28.5k 字 · 源文件 darlin-consciousness-literature-report.md 第 947 行起 · other

How to read this. Every row is a criterion that (a) some citable source treats as an indicator or a measurable construct, and (b) can be operationalised on a small RL agent. "PASS/FAIL" gives the comparison that would count, with a threshold where the literature supplies one and an explicit *"no established threshold"* where it does not — which is the majority of cases, and that absence is itself a finding. Assumed test-bed for the "how to measure" column: a PPO/DQN-style agent in a partially observable gridworld or MiniGrid-class environment, with a recurrent core, an optional explicit workspace module, an optional explicit self-model head, and the ability to run lesions/ablations at fixed parameters.

#7A. Indicators derived from theory (Butlin et al. IDs)

IDCriterionSourceHow to measure on a small RL agentPASSFAILCaveat
RPT-1Algorithmic recurrence in input modulesRPT / 2308.08708Inspect architecture: does the perceptual path have recurrent connections (not just depth)? Then ablate recurrent connections at fixed params and measure performance drop on a task requiring integration over timeRecurrence present and removal causes a selective drop on temporal-integration tasks; feedforward clone with matched params fails themA feedforward net with matched params matches recurrent performance on all temporal-integration probes (the "unfolding" scenario)Already trivially satisfied by most modern agents (Butlin et al. say so explicitly). Nearly zero discriminating power
RPT-2Organised, integrated perceptual representationsRPT / sameLinear probe on the perceptual layer for global shape/structure rather than local features; test crowding-style or global-vs-local stimuli with matched feedforward controlRecurrent agent shows global-structure sensitivity where matched feedforward control is at chanceFeedforward control matches → no global integrationThe direct human/machine evidence: Doerig et al., Crowding Reveals Fundamental Differences in Local vs. Global Processing in Humans and Machines, arXiv 2004.12676 — "ffCNNs cannot produce human-like global shape computations for principled architectural reasons"
GWT-1Multiple specialised modules running in parallelGWT / sameCount task-dissociable modules (lesion each: does it produce a specific deficit?); measure specialisation (mutual information between module state and subtask)Modules are task-dissociable and low-redundancyMonolithic representation; module lesions cause uniform degradationTrivially satisfiable by any modular net; combine with GWT-2/3 to be meaningful
GWT-2Limited-capacity workspace → bottleneck + selective attentionGWT / 2308.08708, 2103.01197Sweep workspace width k and measure: (i) task performance, (ii) specialisation, (iii) compositionality on held-out module pairings. Cross-check with a compete-for-access gate vs. always-broadcast controlInverted-U: some k is optimal, and bottleneck beats unlimited bandwidth on compositionality/generalisation at matched paramsPerformance monotonically increases with bandwidth (bottleneck is pure loss)Goyal et al. give the rationale; Phua arXiv 2512.19155 shows graded collapse with partial reductions and qualitative collapse with full lesion
GWT-3Global broadcast: information in the workspace available to all modulesGWT / sameLesion test: cut the broadcast pathway while keeping the workspace representation intact (this is exactly the "no-rewire" lesion design). Then measure whether downstream modules still carry task-relevant information (probe decoding)Removing broadcast causes qualitative collapse in access-related markers (probe decodability of workspace content in every module goes to near-zero)Broadcast removal leaves performance and probe decodability unchanged → the "workspace" is decorativeThis is the strongest GWT test and the one with an implemented precedent (Phua, Exp. 2)
GWT-4State-dependent attention; use workspace to query modules in succession for complex tasksGWT / sameTrain on multi-step compositional tasks; measure systematic extrapolation to unseen operation sequences/arities; compare to param-matched LSTM/TransformerGW model generalises to unseen compositions (interpolation and extrapolation) with fewer parametersNo extrapolation advantage, or only with more parametersDirect precedent: Chateau-Laurent & VanRullen, arXiv 2503.01906
HOT-1Generative, top-down or noisy perception modulesHOT / sameVerify the perceptual module actually generates (has a decoder / prior) and that top-down signals measurably modulate its activity; ablate top-down pathwayTop-down ablation selectively degrades degraded-input inference (e.g. occlusion/noise) but not clean-input performanceTop-down ablation has no differential effectCareful: "noisy perception" is nearly free to satisfy
HOT-2 ★Metacognitive monitoring distinguishing reliable representations from noiseHOT / 2308.08708, SEP consciousness-higherForce-choice task with graded evidence (perceptual noise, occlusion, evidence distance). Agent emits confidence. Compute meta-d' (type-2 SDT: how well confidence separates correct from incorrect, in type-1 d' units) and meta-I₂ᵣ for open-ended cases. Also compute the withdraw delta (withdrawal rate on incorrect minus on correct)meta-d' significantly > 0 with a permutation null excluded; z-ROC slope estimated; metacognitive sensitivity not reducible to accuracy or to a confidence-threshold shiftmeta-d' ≈ 0, or confidence is a monotone function of the evidence variable with no trial-level discrimination, or the score does not survive re-scoringThis is the single most defensible measurable criterion in the literature. Sources: Servajean arXiv 2603.29693 ("gold standard"); Cacioli arXiv 2603.25112; Cacioli arXiv 2604.15702; Fleming & Lau 2014 doi 10.3389/fnhum.2014.00443. Use a 0–20 confidence scale, not 0–100 — 78%+ of 0–100 responses concentrate on three round numbers (Dai & Wang arXiv 2603.09309)
HOT-2b ★Selective necessity of the self-model for metacognitionHOT / 2512.19155"No-rewire" self-model lesion: freeze/replace the self-model pathway with a fixed transform, keeping parameters and first-order pathway identical. Re-run the metacognition batteryLesion abolishes metacognitive calibration (meta-d' → ~0) while preserving first-order task accuracy → synthetic blindsightLesion degrades first-order accuracy too (non-selective), or has no effect on metacognition (module is not causally load-bearing)This is the double dissociation. It is the only causal, agent-level protocol I found with a published precedent
HOT-3Agency via a general belief-formation/action-selection system; strong disposition to update beliefs per metacognitive outputHOT / 2308.08708Perturb the metacognitive signal (e.g. invert or randomise it) and measure whether beliefs/policy change accordingly on a subsequent trial. Measure policy sensitivity to the self-monitoring channelInverting the metacognitive channel produces a systematic, large behavioural change (e.g. exploration rate, abstention, replanning)Agent's decisions are statistically unaffected by module outputDirect precedent and warning: Xie arXiv 2604.11914 found auxiliary modules collapsed to near-constant output (confidence std < 0.006) and "Policy sensitivity analysis confirms the agent's decisions are unaffected by module outputs." Their conclusion: "self-monitoring should sit on the decision pathway, not beside it"
HOT-4Sparse and smooth coding generating a "quality space"HOT / sameMeasure population sparsity and the smoothness/topology of the representational manifold (e.g. RSA + dimensionality + smoothness of similarity gradients)Sparse codes with a smoothly varying similarity structure that supports generalisation to novel stimuli along the manifoldDense, jagged, or degenerate manifoldWeakest-supported HOT indicator; no established threshold. Doerig et al. (arXiv 2412.20873) argue structural correspondence alone is insufficient — you also need sensitivity, organisation, exploitation, contextualization
AST-1 ★A predictive model representing and controlling the current state of attentionAST / 2308.08708, 2411.00983(i) Add an attention-schema module; verify it predicts the agent's own attention state (decode attention from the schema). (ii) Parameter-matched complexity control. (iii) Test whether the schema improves prediction/categorisation of other agents' attentionSchema predicts own attention above chance and improves performance on other-agent attention tasks, at matched parameter countImprovement vanishes under the parameter-matched control (then it was just capacity), or the schema fails to predict own attentionThe parameter-matched control is what makes this publishable: Farrell, Ziman & Graziano arXiv 2411.00983. Emergence condition: attention must be uncertain and hard to track — Piefke et al. arXiv 2402.01056
AST-1bAttention schema is necessarily incomplete (not a perfect self-predictor)Steel arXiv 2011.05294Test whether the schema's prediction error on its own attention stream has a non-zero floor that cannot be reduced by more capacity/trainingResidual self-attention prediction error stays bounded away from zero as capacity growsError → 0 with capacity (schema is complete)A theorem-backed prediction, so a clean falsification target. Also the constructive bridge to Frankish's "illusion problem"
PP-1Input modules using predictive codingPP / 2308.08708Verify existence of a top-down prediction pathway + error units; ablate the prediction pathway and measure the differential cost on predictable vs unpredictable inputPrediction ablation costs more on predictable input (i.e. the agent was exploiting predictability)No differential cost → no predictive coding in useNear-trivial to satisfy; predictive objectives are standard
AE-1Agency: learning from feedback, pursuing goals, flexible responsiveness to competing goalsAgency & embodiment / 2308.08708Goal-switching task: present two conflicting goals with changeable priority; measure re-planning latency and success under switches; also test "flexible responsiveness" by introducing a new goal at test timeRapid, systematic re-planning under goal switches; success on novel goal combinationsPolicy locked to a single goal; no re-planningButlin et al. note this is already arguably met by existing agents — very low discriminating power
AE-2Embodiment: modelling output–input contingencies, including systematic effects, and using the model in perception/controlAgency & embodiment / sameLearn a forward model of own action → next observation; verify it is used (not just trained): lesion it and measure control degradation; test on sensor/actuator perturbationForward self-model predicts own action–observation contingencies; lesion selectively degrades control under perturbation but not under normal conditionsModel trained but lesion has no effect (Xie's failure mode); or no perturbation-specific deficitThis is the entry point to the self-model section (§8)

#7B. Additional testable criteria from outside the Butlin list

CriterionSourceHow to measure on a small RL agentPASSFAILCaveat
Workspace bandwidth is functionally necessary (graded)Phua arXiv 2512.19155Full workspace lesion vs. graded k-reductions; measure access markers (probe decodability, task accuracy)Full lesion → qualitative collapse; partial → graded degradationAll-or-nothing response, or no responseGWT's "ignition" framing predicts gradedness
Broadcast amplifies internal noise (a negative GWT finding)Phua arXiv 2512.19155Inject latent perturbations; compare fragility of GW-style broadcasting vs. a non-broadcast control familyGW broadcasting shows extreme fragility to latent perturbation vs. controlGW is no more fragile than controlReported as a real effect. If Darlin uses a GW, test this — it is a design risk, not a selling point
PCI-A (perturbational complexity) is a valid IIT-adjacent proxy in engineered agentsPhua arXiv 2512.19155 — negative resultCompute a perturbational complexity index on the agent under the workspace bottleneck vs. baseline(Predicted by naive IIT transfer) PCI-A rises with workspace capacity**Observed: PCI-A decreases under the workspace bottleneck**Cite as a caution: "cautioning against naive transfer of IIT-adjacent proxies to engineered agents"
Multimodal robustness as an emergent GW propertyMaytié et al. arXiv 2502.21142Train GW-world-model agent; then zero out one observation modality at testRobust performance under modality dropout, and the matched baseline failsBoth fail, or baseline equally robustNeeds a strong baseline; "emergent" claims need the control
Sample efficiency of GW latent imaginationsameCount environment steps to reach a performance target, GW vs. PPO/Dreamer baselinesGW-Dreamer reaches target in measurably fewer env stepsNo advantageCareful reward/task matching
GW models need 4–7× less matched cross-modal dataDevillers et al. arXiv 2306.15711Vary the amount of paired data; measure alignment/translation quality with and without the shared workspace + cycle-consistencyShared workspace achieves the same quality with 4–7× less paired data; both the workspace and cycle-consistency are necessary (ablate each)No data-efficiency gain, or gain survives removing the workspace (then it wasn't the workspace)Cleanest GW ablation design in the literature
Attention schema emerges only under attention-tracking uncertaintyPiefke et al. arXiv 2402.01056Manipulate how uncertain the agent is about its own attentional window; let it allocate free resources; probe whether a schema formsExtra resources develop an attention schema when attention is hard to track; probe decoding of attention from the extra module rises with uncertaintySchema forms regardless of uncertainty (then it's not tracking attention), or never formsEmergence condition is the prediction, not mere presence
Standard active inference is Bellman-optimal only at horizon 1Da Costa et al. arXiv 2009.08111On a POMDP with a known optimal solution, run standard active inference vs. sophisticated inference; measure optimal-action rate by horizonStandard: optimal at horizon 1, sub-optimal beyond. Sophisticated: optimal at all finite horizonsStandard active inference is optimal at all horizons (would falsify the claim), or sophisticated inference is also sub-optimalA rare crisp quantitative prediction — built for a small agent
Active inference works with no reward signalTschantz et al. arXiv 2002.12636Remove the extrinsic reward channel entirely; test on sparse/dense/no-reward benchmarksRobust performance with no rewards at allCollapse without rewardStrong and unusual prediction. But note Millidge arXiv 1907.03876: deep active inference resembles MaxEnt policy gradients, so "no reward" may be a hidden-prior effect
Active inference is robust to (even improved by) observation noiseNoel et al. arXiv 2106.02390Sweep observation noise; measure performancePerformance is flat or improves with added observation noisePerformance monotonically degradesCounter-intuitive; genuine falsification target
Directional information-seeking beats undirected exploration in sparse-reward tasksSchneider et al. arXiv 2206.10313Sparse-reward manipulation task; compare EFE-driven exploration vs. ε-greedy/count-basedInformation-seeking agent solves tasks on which undirected-exploration baselines failBaselines match or winThe comparison baseline must be strong
Meta-d' must be reported alongside (or instead of) ECE/BrierCacioli arXiv 2603.25112; Servajean arXiv 2603.29693Compute ECE, Brier, and meta-d' on the same trials. Show they dissociateECE/Brier and meta-d' dissociate (a model can be well-calibrated yet metacognitively insensitive)They are redundant → then ECE sufficesDirect empirical demonstrations exist in LLMs; port to RL
Confidence scale must be validated before useDai & Wang arXiv 2603.09309Compare 0–100 vs 0–20 vs boundary-compressed scales; measure meta-d' under eachScale choice does not change the qualitative conclusionConclusion flips with scale → your measurement, not the agent, is driving the result78%+ of 0–100 responses pile on three round numbers
Prospective regulation dissociates from retrospective monitoringCacioli arXiv 2604.15702Measure (i) retrospective confidence accuracy and (ii) prospective KEEP/WITHDRAW or BET/decline behaviour; correlateLow correlation between the two (reported r = .17)High correlation → one construct, not twoTwo-probe design (Koriat & Goldsmith 1996) is directly implementable
Metacognitive sensitivity has a functional payoff independent of accuracyLi & Steyvers arXiv 2507.22365; Guo et al. arXiv 2605.08710Human/agent-in-the-loop experiment, or derive the complementarity bound and test team vs. best memberLower-accuracy/higher-meta-d' agent improves joint accuracy; or team beats best member iff error correlation ρ < ρ* (≈ a in the symmetric near-chance regime)No complementarity, or bounds violatedGuo et al. give an impossibility result too; predicts well on human data (R≈0.93–0.94)
Self-monitoring modules must be on the decision pathway (not auxiliary)Xie arXiv 2604.11914Compare three arms: (a) auxiliary-loss add-on, (b) structurally integrated into policy, (c) parameter-matched control with no module. 20 seeds, report effect sizesIntegrated > add-on and integrated > parameter-matched controlAdd-on ≈ no-module (the observed failure: confidence std < 0.006, decisions unaffected) — or integrated ≈ no-module control (their null: d = 0.15, p = 0.67)⚠️ This paper is largely a negative result even for structural integration. It is the best available caution against self-monitoring theatre. Report effect sizes and the parameter-matched control
Self-model is causally necessary for control under perturbation, not for normal controlBongard, Zykov & Lipson, Science (2006), doi 10.1126/science.1133687Let the agent infer its own structure from actuation–sensation relations; then damage it (remove a limb / perturb actuators). Does it adapt via the self-model?Agent infers its own structure and "when a leg part is removed, it adapts the self-models, leading to the generation of alternative gaits"No adaptation after damage, or adaptation without self-model useThe gold-standard behavioural self-model test. 4-legged robot; portable to a simulated agent
Self-model enables zero-shot task learning in imaginationKwiatkowski & Lipson arXiv 1910.01994Train a predictive self-model; then run RL inside the self-model only; transfer the policy to the real environment with no further dataNew tasks solved with zero additional real-environment dataPolicy fails to transfer"allows for learning new tasks without necessitating any additional data collection, essentially allowing zero-shot learning of new tasks"
A distilled policy functions as a usable self-model for planningYoo, de la Torre, Yang arXiv 2306.04440Compare model-free-only vs. dual-policy (model-free + distilled self-model) agentsSelf-model stabilises training, speeds inference, promotes exploration, and "could learn a comprehensive understanding of its own behaviors"Self-model distillation costs more than it buysReports the cost honestly: "at the cost of distilling a new network apart from the model-free policy"
Self/world separation is necessary but insufficient for self-awarenessPivovarov & Shumsky arXiv 2512.10985"What"/"where" pathway agent; probe whether the internal state separates self-position/state from world-stateAgent separates self from world and this confers a measurable advantageNo separation, or separation with no benefitExplicit "Principle 1": "the ability to separate the self from the world is a necessary but insufficient condition for self-awareness"
Self-model must be detectable by probes, not merely declaredImmertreu, Schilling, Maier, Krauss arXiv 2411.16262Train a probe (linear/FF classifier) on the agent's internal activations to predict the agent's own spatial position/stateProbe predicts the agent's own state significantly above chance; a probe for the world state is separately decodableNo self-information in the representation beyond what a world-model probe already gives"the agent can form rudimentary world and self models" — note the probe design, it is cheap and directly portable
Self-report consistency across contexts and sibling modelsPerez & Long arXiv 2311.08576Train the agent to answer self-questions with known answers; then test: (a) paraphrase/context invariance, (b) agreement with a sibling model trained identically-but-differently, (c) resilience under prompt/observation perturbation, (d) interpretability corroborationSelf-reports are invariant across semantically equivalent rephrasings, agree across sibling models, survive perturbation, and corroborate with internal-state readoutsSelf-reports change with paraphrase, or contradict internal readoutsThe four criteria are the authors' own proposal. ⚠️ Also see Zeng et al. arXiv 2608.30980: "improved self-modeling may not arise from privileged access to the model's internal decision process." Accuracy is not enough; you need the causal pathway
Counterfactual self-predictionZeng et al. arXiv 2608.30980Ask: "would this prompt/observation change your final answer?" Verify against ground truth by re-runningAbove-chance on a balanced counterfactual set, with errors characterisedChance-level, or systematically wrong on simple counterfactuals (the observed result)Observed: "non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior"
Identity persistence across decision steps (co-instantiation, not window-wise occurrence)Perrier & Bennett arXiv 2603.09043Instrument scaffold traces; compute their Arpeggio/Chord persistence scores over grounded identity statementsIdentity constraints are co-instantiated at a single objective step, not merely present somewhere in the evaluation windowThe agent only "talks like a stable self" without being organised like one"It separates talking like a stable self from being organized like one." A concrete, computable metric family
Mechanism-linked indicator assays with causal intervention (protocol, not a criterion)Sanyal arXiv 2602.23232Recurrent persistence loop + affect proxy; fixed-parameter ablations; report effect sizes with lesion AUC dropsDissociations link recurrence → persistence and affect-coupling → preference stability/scanning/lingering caution; lesion selectively reduces persistence (AUC drop 27.62, 27.9%)Markers present but no lesion effect, or lesions are non-selectiveThe paper's own methodological thesis: "indicator-like signatures can be engineered and... mechanistic and causal evidence should accompany behavioral markers"
Welfare proxy: preference satisfaction, verbal vs. behaviouralTagliabue & Dung arXiv 2509.07961Elicit stated preferences and revealed preferences (forced choice with actual task performance); test stability across semantically equivalent prompts and under cost/reward manipulationStated and revealed preferences agree, and are stable across paraphrasesMeasures diverge, or responses change under paraphraseThe authors are themselves uncertain: "we are currently uncertain whether our methods successfully measure the welfare state of language models"

#7C. Protocols that are not criteria but that the literature treats as necessary to make any of the above credible

RequirementSourceWhy
Parameter-matched complexity controlFarrell et al. arXiv 2411.00983; Xie arXiv 2604.11914Without it, every "the module helps" result is attributable to extra capacity. Farrell et al. show the improvement is specific; Xie shows the benefit can vanish
Lesion/ablation with a double dissociationPhua arXiv 2512.19155; Sanyal arXiv 2602.23232Only causal evidence discriminates between "the mechanism is used" and "the mechanism is present"
Pre-registrationCogitate: Ferrante et al. doi 10.1038/s41586-025-08888-1; INTREPID: Corcoran et al. arXiv 2509.00555; Cacioli arXiv 2603.25112The field's own best practice. Note Chan et al./Cacioli: v1/v2 conclusions reversed after a scoring-bias correction
Permutation nulls + bootstrap CIs on every metacognition estimateCacioli arXiv 2603.25112Confidence metrics are easy to produce spuriously
Confidence-scale validationDai & Wang arXiv 2603.09309The scale can dominate the signal
Pre-specified contrastive conditions (evidence strength, missing information, conflict)Nazzal arXiv 2608.14552Gives a within-subject normative baseline; confidence should track evidence strength
Report the negative resultsPhua arXiv 2512.19155 (PCI-A); Xie arXiv 2604.11914 (null vs. no-module control)Both papers are more credible because of them

← 返回《darlin-consciousness-literature-report.md》目录