5. The testable-criteria literature
★
22.7k 字 ·
源文件 darlin-consciousness-literature-report.md 第 585 行起 ·
other
#5.1 Butlin, Long et al. — the indicator-property rubric
Butlin, Long, Elmoznino, Bengio, Birch, Constant, Deane, Fleming, Frith, Ji, Kanai, Klein, Lindsay, Michel, Mudrik, Peters, Schwitzgebel, Simon, VanRullen, *Consciousness in Artificial Intelligence: Insights from the Science of Consciousness*, arXiv 2308.08708 (v3, Aug 2023). Full front matter, executive summary, method section and indicator table retrieved via ar5iv HTML.
The computational functionalism assumption, verbatim:
"First, we adopt computational functionalism, the thesis that performing computations of the right kind is necessary and sufficient for consciousness, as a working hypothesis. This thesis is a mainstream—although disputed—position in philosophy of mind. We adopt this hypothesis for pragmatic reasons: unlike rival views, it entails that consciousness in AI is possible in principle and that studying the workings of AI systems is relevant to determining whether they are likely to be conscious."
They note the limits honestly: *"computational functionalism does not entail that any substrate can be used to construct a conscious system (Block 1996). As Michel and Lau (2021) put it, 'Swiss cheese cannot implement the relevant computations.'"*
Their method is explicitly a credence-update, not a test:
"To a first approximation, our confidence that a given system is conscious can be determined by (a) the similarity of its computational processes to those posited by a given scientific theory of consciousness, (b) our confidence in this theory, (c) and our confidence in computational functionalism."
And the crucial disclaimer: *"Our claim, however, is that AI systems which possess more of the indicator properties are more likely to be conscious." The v3 footnote to the abstract records the walk-back: "A previous version of this sentence read '…but also shows that there are no obvious barriers to building conscious AI systems.' We have amended it to better reflect the messaging of the report: that satisfying these indicators may be feasible. But satisfying the indicators would not mean that such an AI system would definitely be conscious."*
They deliberately exclude IIT: *"We do not consider integrated information theory, because it is not compatible with computational functionalism."*
THE FULL INDICATOR TABLE (Table 1), verbatim:
| Theory | ID | Indicator property |
|---|---|---|
| Recurrent processing theory | RPT-1 | Input modules using algorithmic recurrence |
| RPT-2 | Input modules generating organised, integrated perceptual representations | |
| Global workspace theory | GWT-1 | Multiple specialised systems capable of operating in parallel (modules) |
| GWT-2 | Limited capacity workspace, entailing a bottleneck in information flow and a selective attention mechanism | |
| GWT-3 | Global broadcast: availability of information in the workspace to all modules | |
| GWT-4 | State-dependent attention, giving rise to the capacity to use the workspace to query modules in succession to perform complex tasks | |
| Computational higher-order theories | HOT-1 | Generative, top-down or noisy perception modules |
| HOT-2 | Metacognitive monitoring distinguishing reliable perceptual representations from noise | |
| HOT-3 | Agency guided by a general belief-formation and action selection system, and a strong disposition to update beliefs in accordance with the outputs of metacognitive monitoring | |
| HOT-4 | Sparse and smooth coding generating a "quality space" | |
| Attention schema theory | AST-1 | A predictive model representing and enabling control over the current state of attention |
| Predictive processing | PP-1 | Input modules using predictive coding |
| Agency and embodiment | AE-1 | Agency: Learning from feedback and selecting outputs so as to pursue goals, especially where this involves flexible responsiveness to competing goals |
| AE-2 | Embodiment: Modeling output-input contingencies, including some systematic effects, and using this model in perception or control |
That is 14 indicators (2 RPT + 4 GWT + 4 HOT + 1 AST + 1 PP + 2 AE). Note they also discuss, in §2.4, midbrain theory and unlimited associative learning without elevating them to indicators.
Their own assessment of current systems: *"This work does not suggest that any existing AI system is a strong candidate for consciousness." And they flag which indicators are already trivially met: "There are some properties in the list which are already clearly met by existing AI systems (such as RPT-1, algorithmic recurrence), and others where this is arguably the case (such as the first part of AE-1, agency)."*
They explicitly reject behavioural tests: *"we are sceptical about whether behavioural approaches to consciousness in AI can avoid the problem that AI systems may be trained to mimic human behaviour while working in very different ways, thus 'gaming' behavioural tests (Andrews & Birch 2023)."*
#5.2 ★ The published critique of Butlin et al.
Koch, Calibration and transfer in indicator-based assessments of artificial consciousness, arXiv 2603.27597, published in Neuroscience of Consciousness 2026(1), niag061, doi 10.1093/nc/niag061. This is a direct, published critique. Verbatim:
"Research on artificial consciousness increasingly shifts evaluation from behaviour to internal architecture. Theory-based indicators are used to update probability assignments. This improves on behavioural tests but raises two distinct problems. First, these assignments cannot currently be calibrated against independently established artificial consciousness outcomes. Second, their evidential relevance is transferred from biological cases without independent support that indicator-consciousness relations remain stable across substrates. This commentary distinguishes calibration from transfer and adapts the iterative natural-kind strategy by proposing a preliminary, theory-relative comparative space for cross-substrate assessment."
★★ This is the single most important methodological criticism for Darlin to internalise. The indicator approach cannot be calibrated (you have no labelled positive/negative AI cases to check it against) and its transfer from biology to silicon is unvalidated. So the honest output of a Butlin-style assessment is a relative ranking of architectures on theory-relative indicators, not a probability of consciousness.
Also relevant, an earlier and more general checklist paper: Doerig, Schurger, Hess, Herzog, *Hard criteria for empirical theories of consciousness, Cognitive Neuroscience* (2020), doi 10.1080/17588928.2020.1772214. Verbatim: *"we suggest that one reason for this abundance of extremely different theories may be the lack of stringent criteria specifying how empirical data constrains ToCs. First, we argue that consciousness is a well-defined topic from an empirical point of view and motivate a purely empirical stance on the quest for consciousness. Second, we present a checklist that, we propose, empirical ToCs need to cope with. Third, we review 13 of the most influential ToCs and subject them to the criteria."* (Abstract only — the 13 criteria themselves are in the PDF, which I could not read.)
#5.3 Higher-Order Theories: do they give a testable criterion?
- **SEP, Higher-Order Theories of Consciousness, first published 2001, substantive revision 9 Jan 2026, plato.stanford.edu/entries/consciousness-higher. Full text retrieved. The core criterion, the Transitivity Principle (TP)**, verbatim: *"A conscious mental state is a state whose subject is, in some way, aware of being in it." General form: "A phenomenally conscious mental state is a mental state that is (or is disposed to be) the object of a higher-order representation of a certain sort."*
- Lau & Rosenthal, Empirical support for higher-order theories of conscious awareness, *Trends in Cognitive Sciences* 15(8) (2011), doi 10.1016/j.tics.2011.05.009. Citation and authorship verified via OpenAlex; abstract not available there, so I do not quote it.
- The SEP's §8 ("HOT Theory and the Prefrontal Cortex") is where testability is cashed out: *"A 'prefrontal HOT theory' says that the prefrontal cortex (PFC), or at least sub-areas of the PFC, is the likely site of HOTs in the brain and PFC activity is essential to having conscious mental states. Some evidence comes from neuroimaging studies which have systematically found increased activity in the prefrontal and parietal cortex when comparing conscious versus unconscious conditions, often even when performance capacity is controlled for (Lau & Passingham 2006; Lau & Rosenthal 2011; Odegaard et al. 2017; Boly et al. 2017). Rounis et al. (2010) find that transcranial magnetic stimulation (TMS) directed at the PFC, which disrupts its activity, has a significant impact…"* → the testable content of HOT is (i) a metacognitive monitoring signal and (ii) a causal lesion/intervention on the module carrying it.
★ The most important HOT result for the project, and it is an empirical one on an artificial agent: Phua, Can We Test Consciousness Theories on AI? Ablations, Markers, and Robustness, arXiv 2512.19155. Verbatim:
"In Experiment 1, a no-rewire Self-Model lesion abolishes metacognitive calibration while preserving first-order task performance, yielding a synthetic blindsight analogue consistent with HOT predictions. In Experiment 2, workspace capacity proves causally necessary for information access: a complete workspace lesion produces qualitative collapse in access-related markers, while partial reductions show graded degradation, consistent with GWT's ignition framework. In Experiment 3, we uncover a broadcast-amplification effect: GWT-style broadcasting amplifies internal noise, creating extreme fragility… We also report an explicit negative result: raw perturbational complexity (PCI-A) decreases under the workspace bottleneck, cautioning against naive transfer of IIT-adjacent proxies to engineered agents. These results suggest a hierarchical design principle: GWT provides broadcast capacity, while HOT provides quality control. We emphasize that our agents are not conscious; they are reference implementations for testing functional predictions of consciousness theories."
★ This is a template Darlin should copy: synthetic blindsight via self-model lesion is exactly the "metacognitive sensitivity without first-order impairment" dissociation, and it is the single cleanest HOT-flavoured test that has been run on an engineered agent. Note also their explicit disclaimer and the negative PCI-A result.
#5.4 Machine metacognition: meta-d' and confidence calibration as measurable proxies
The methodology paper. Fleming & Lau, How to measure metacognition, Frontiers in Human Neuroscience
8:443 (2014), doi 10.3389/fnhum.2014.00443. This is the canonical
statement of meta-d' as the type-2 SDT measure of metacognitive sensitivity (how well confidence
discriminates correct from incorrect responses, in units of type-1 d'). I retrieved the Frontiers page but the
fetcher returned only the page shell; I did not read the body. The definitional content below is
corroborated by the applied papers, not by direct quotation of Fleming & Lau.
The 2025–2026 machine-metacognition literature (all abstracts retrieved):
- Servajean & Servajean, Measuring the metacognition of AI, arXiv 2603.29693. Verbatim: *"This paper is primarily a methodological contribution arguing for the adoption of the meta-d' framework as the gold standard for assessing the metacognitive sensitivity of AIs—the ability to generate confidence ratings that distinguish correct from incorrect responses. Moreover, we propose to leverage signal detection theory (SDT) to measure the ability of AIs to spontaneously regulate their decisions based on uncertainty and risk."* Tested on GPT-5, DeepSeek-V3.2-Exp, Mistral-Medium-2508.
- Cacioli, Do LLMs Know What They Know? Measuring Metacognitive Efficiency with Signal Detection Theory, arXiv 2603.25112. 224,000 factual QA trials, 4 LLMs. Verbatim on the key conceptual point: *"Standard evaluation of LLM confidence relies on calibration metrics (ECE, Brier score) that conflate two capacities: how much a model knows (Type-1 accuracy) and how well its confidence signal tracks that knowledge (Type-2 metacognitive sensitivity)." Findings: "metacognitive information varies by a factor of 1.98 across models and is not predicted by accuracy"; "the confidence signal has model-specific unequal-variance structure (z-ROC slopes 0.78 to 1.18) invisible to calibration metrics"; "efficiency is domain-specific, weakest in Science & Technology for every model"; "temperature dissociates accuracy from metacognitive information, which stays near-flat for three of four models while accuracy falls"*; and *"metacognitive information tracks the accuracy gain from confidence-based abstention exactly (rho = +1.00) while accuracy does not."* ⚠️ Note the version history: v1/v2 reported an inverse accuracy–efficiency coupling which "does not survive relabelling" after correcting a scorer length bias. This is a live methodological hazard: automated correctness scoring can manufacture or destroy metacognition results.
- Cacioli, The Metacognitive Monitoring Battery, arXiv 2604.15702. 20 frontier LLMs, 10,480 evaluations, grounded in Nelson & Narens (1990) and using KEEP/WITHDRAW + BET probes from Koriat & Goldsmith (1996). The key metric, verbatim: *"the withdraw delta: the difference in withdrawal rate between incorrect and correct items." Result: three profiles — "blanket confidence, blanket withdrawal, and selective sensitivity. Accuracy rank and metacognitive sensitivity rank are largely inverted." Also: "Retrospective monitoring and prospective regulation appear dissociable (r = .17...)."* → a directly implementable battery for a small agent: force a choice, then ask the same module to KEEP/WITHDRAW and to BET/decline, and measure the differential withdrawal rate.
- Trinh, Pham, Pham, Nguyen, Metacognitive Sensitivity for Test-Time Dynamic Model Selection, arXiv
2512.10451: uses
meta-d'computed on the fly as a bandit signal to choose which expert model to trust. → This is the strongest demonstration thatmeta-d'is a causally load-bearing internal variable, not a post-hoc statistic. - Dai & Wang, Rescaling Confidence: What Scale Design Reveals About LLM Metacognition, arXiv 2603.09309. Critically important measurement-artefact finding, verbatim: *"verbalized confidence is heavily discretized, with more than 78% of responses concentrating on just three round-number values... a 0–20 scale consistently improves metacognitive efficiency over the standard 0–100 format, while boundary compression degrades performance and round-number preferences persist even under irregular ranges."* → if you measure confidence on a 0–100 scale you are largely measuring the scale.
- Li & Steyvers, Beyond Accuracy: How AI Metacognitive Sensitivity improves AI-assisted Decision Making, arXiv 2507.22365: *"an AI with lower predictive accuracy but higher metacognitive sensitivity can enhance the overall accuracy of human decision making"*, confirmed behaviourally. → a functional payoff for metacognitive sensitivity, independent of accuracy.
- Nguyen, Colombatto, Fleming, Posner, Hawes, Bhattacharyya, arXiv 2508.03293, N=100 user study on robot teleoperation: *"pairing poorly-calibrated AI-DSS with humans hurts performance instead of helping the team."* (Note Fleming is a co-author — the human-metacognition group is now running machine metacognition experiments.)
- Nazzal, Large Language Models Show Metacognitive Sensitivity in Medical Reasoning, arXiv 2608.14552: 135 trials, accuracy 93.5%, mean confidence 78.4%, AUROC2 = 0.876; confidence increased with evidence distance from the diagnostic boundary and decreased with missing information. *"These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence."* → a concrete example of a well-powered, small-N, fully-specified metacognition assay (evidence-strength manipulation + missing-information manipulation).
- Guo, Wu, Yiu, arXiv 2605.08710: derives *"a complementarity theorem (teams outperform individuals iff error correlation ρ_HM < ρ...)" and *"minimax bounds showing gains scale as Θ(√Δd) with metacognitive sensitivity difference"*, plus an impossibility result. Predictions fit observed team accuracy (R = 0.94 on ImageNet-16H, R = 0.91 on CIFAR-10H). → a quantitative theory linking metacognitive sensitivity to a downstream benefit, with an impossibility theorem. This is the most "physics-like" result in the machine-metacognition literature.
- Liu, Caciularu, Yona, Szpektor, Cohan, arXiv 2606.32032: RL with metacognitive feedback (RLMF) improves "faithful calibration" — aligning expressed with intrinsic uncertainty — "surpasses standard RL by up to 63%."
Anti-hype counterweight (important): Zeng, Assis, Wang, Evaluating and Improving LLM Self-Modeling, arXiv 2608.30980 (EMNLP '26 Main). Verbatim: *"Current models show non-trivial but limited self-modeling skill, and make systematic mistakes on simple counterfactual questions about their own behavior... These gains, however, do not seem to constitute introspection consistently: improved self-modeling may not arise from privileged access to the model's internal decision process."* → **good self-model performance ≠ privileged access. You must test for a specific causal pathway, not for accuracy on self-questions.**
#5.5 Illusionism (Frankish) and its consequences
Keith Frankish, Illusionism (author's own overview page), keithfrankish.com/illusionism. Full text retrieved. Key claims, verbatim:
"Illusionism rejects phenomenal realism—the view that conscious experiences possess special 'phenomenal' properties (or 'qualia') that are irreducibly subjective... Illusionism holds that this entire conception is mistaken. While conscious experiences are real, their supposed phenomenal properties are an introspective illusion."
"Phenomenal properties, illusionists propose, are merely intentional objects—things that can be represented but not instantiated, like impossible objects in mathematics."
"What distinguishes conscious experiences from nonconscious ones is not the presence of mysterious phenomenal properties but detectable functional differences. I use the term 'quasi-phenomenal properties' for the functional properties distinctive of experiences that are conscious."
The two positive tasks he sets: (1) *"developing a functionalist theory of consciousness, which tells us what quasi-phenomenal properties are", (2) "the illusion problem": explaining how and why introspection generates the compelling but misleading impression that our experiences possess phenomenal properties."* His own current proposal: Reactivity Theory (consciousness = "complex patterns of psychological reactivity") plus Reactivity Schema Theory (*"self-monitoring systems track these global reactivity patterns and create simplified, schematic models of them"). His analogy: "As optical effects caused by the reflection and refraction of sunlight by airborne water droplets, rainbows are real; as spatially located multicolored arcs, they're illusory. Similarly... consciousness is real; but as an irreducibly subjective world of phenomenal properties, it is illusory." And he explicitly rejects the "dogmatic materialism" charge: "I want to explain consciousness, and phenomenal realism is an explanatory dead end—it posits properties forever beyond scientific investigation and calls that an explanation."*
Primary sources he cites for himself: Illusionism as a theory of consciousness, *Journal of Consciousness Studies 23(11–12): 11–39 (2016); What is illusionism?, Klēsis 55 (2023); The ethical implications of illusionism, Neuroethics* 17(2): 28 (2024), doi 10.1007/s12152-024-09562-5. Also relevant: his response to the IIT pseudoscience affair (above), where he argues IIT is "a metaphysical theory mispresented as science."
★ Consequence for Darlin. Illusionism is the position that makes the project tractable, and it is not a concession — it is a research programme with specific deliverables:
- Replace "is it conscious?" with "what functional differences mark quasi-phenomenal states?"
- Then solve the illusion problem — i.e. build an explicit account of *why the system's self-monitoring would misrepresent its own states*. This is a concrete architectural requirement: an agent with a self-model that is known to be lossy and schematic by design. That is precisely AST's "necessarily incomplete" attention schema (Steel, arXiv 2011.05294) and Frankish's reactivity schema.
- It is compatible with Butlin et al.'s computational functionalism, and it is the position that makes falsifiable claims possible at all — because it identifies consciousness-talk with a functional target.
Note the SEP entry on Chinese Room lists illusionism among the four major metaphysical positions and says of all four: *"there is work for the science of consciousness to do on all four of these metaphysical positions... According to illusionism, neuroscience can explain why the illusion arises, and in particular why it arises in connection with some brain states but not others."* link