9. Honest assessment: which theories are unfalsifiable and should NOT be built on
★
8.3k 字 ·
源文件 darlin-consciousness-literature-report.md 第 1099 行起 ·
other
#Tier 1 — Do not build on these
IIT. The indictment is triple-sourced and converging:
- Aaronson (blog): max-Φ in Vandermonde/polynomial-evaluation systems that "don't even come close to doing anything that we'd want to call intelligent, let alone conscious"; Φ in such a system can exceed any plausible bound for a human brain; and the definition of Φ is non-robust (making Φ well-defined required reducing information integration by 2×).
- Doerig, Schurger, Hess, Herzog (doi 10.1016/j.concog.2019.04.002): the link between causal structure and consciousness is "either false or outside the realm of science."
- Hanson & Walker (arXiv 2006.07390): IIT is *"simultaneously falsified at the FSA level and unfalsifiable at the CSA level."*
- The 2023 pseudoscience preprint (doi 10.31234/osf.io/zsr78) with Fleming, Frith, Goodale, Lau, LeDoux, Michel, Owen, Peters, Slagter among the signatories; and Frankish's independent conclusion that "At best, IIT is a metaphysical theory mispresented as science." (doi 10.31234/osf.io/uscwt)
- Practically: Φ is not computable for realistic systems (Aaronson's NP-hardness conjecture) — so even if it were the right theory, you could not evaluate Darlin with it.
- Empirically: the Cogitate adversarial collaboration found IIT's claim that *"network connectivity specifies consciousness" contradicted by "a lack of sustained responses within the posterior cortex."* (doi 10.1038/s41586-025-08888-1)
- Butlin et al. excluded IIT on compatibility grounds: *"We do not consider integrated information theory, because it is not compatible with computational functionalism."* (2308.08708)
- And the specific port of an IIT proxy to an engineered agent was tested and failed: *"raw perturbational complexity (PCI-A) decreases under the workspace bottleneck, cautioning against naive transfer of IIT-adjacent proxies to engineered agents."* (Phua, arXiv 2512.19155)
Residual caveat: an IIT proponent would object that Aaronson's construction assumes IIT 1.0's normalization and that IIT 4.0 replaced the measure. IIT 4.0 (arXiv 2212.14787) claims testable predictions. But IIT 4.0's abstract does not address the unfolding argument or falsifiability, and I could not read the body — so I cannot rule out that it answers these. My Tier-1 placement is based on the indictment of the framework's causal-structure premise (which is not normalization-specific) plus the empirical failure of the one IIT-adjacent proxy tested on an engineered agent.
The Free Energy Principle as a universal principle. Colombo & Wright (doi 10.1007/s11229-018-01932-w) record that it *"has been called a postulate, an unfalsifiable natural law or an imperative."* Biehl, Pollock & Kanai (arXiv 2001.06408) *"prove by counterexample that the original free energy lemma, when taken at face value, is wrong" and that the Bayesian-inference interpretation "is not sufficiently justified."* Bruineberg et al. (doi 10.1017/s0140525x21002351) show the metaphysical Markov-blanket move *"requires additional premises that cannot be justified by appeal to the success of the mathematical framework alone."* And the defenders' own replies retreat to "it holds regardless" for the general case (arXiv 2205.07793) while conceding the formulation is inappropriate for "certain simple systems" — the classic signature of an unfalsifiable framework. → Do not build Darlin's claims on the FEP. Use active inference as an algorithm with the specific, falsifiable predictions in §7B.
Physical/substrate requirements that don't cash out computationally (Type-A biological naturalism). Klatzmann & Doerig (arXiv 2606.02121): *"this dissociates consciousness from behaviour, making Type-A-BN untestable."* If a silicon/biological distinction makes no difference to information processing, no experiment can address it. → **Do not build on it; if you want a substrate argument, you must specify a Type-B (information-processing) difference, which then is testable.**
#Tier 2 — Usable, but only for functional claims, never for consciousness
| Theory | What it can support | What it cannot |
|---|---|---|
| GWT | Capacities: multimodal integration with less data, systematic composition generalisation, sample efficiency, modality robustness. Bandwidth as a design trade-off. | Any claim that a workspace makes Darlin conscious. Even GNWT failed its own adversarial test on ignition-at-offset and prefrontal representation (Nature 2025) |
| HOT | The best-specified measurable construct in the field: metacognitive sensitivity (meta-d'), selectively lesioned. Also the synthetic-blindsight dissociation. | Any claim that metacognition = consciousness. Metacognitive sensitivity is fully explainable as a quality-control function |
| AST | Attention modelling with a specific, controlled benefit (modelling others' attention at matched params); and a theorem-backed prediction of necessary incompleteness. | Consciousness. AST is by Graziano's own framing a control-engineering theory; the "awareness" claim rides on the schema being misrepresented as non-physical — which is an illusionist move, not a consciousness-generating one |
| Predictive processing / active inference (as algorithm) | Predictive coding payoff on predictable input; epistemic foraging; sparse-reward performance; reward-free operation; the horizon-1 Bellman prediction. | The FEP as metaphysics, and any "free-energy minimisation ⇒ sentience" inference |
| Agency & embodiment | Forward self-models, damage adaptation, zero-shot planning in imagination, self/world separation, identity persistence metrics | Both are explicitly hedged indicators in Butlin et al., and both are already arguably met by existing systems (so they discriminate almost nothing) |
| Illusionism | The most productive frame for Darlin. It converts an unfalsifiable question into two deliverable engineering tasks: (i) a functional theory of quasi-phenomenal states, (ii) the "illusion problem" — an explicit account of why the self-monitoring system misrepresents its own states. Both are testable. | It does not tell you whether anything is "really" conscious — but it argues that question is ill-formed |
#The meta-point
★ There is no single validated test for machine consciousness, and the literature says so in its own voice. Butlin et al. (the paper with the most explicit rubric) say satisfying the indicators *"would not mean that such an AI system would definitely be conscious"*; the published critique of their approach (Koch, arXiv 2603.27597) says the assignments *"cannot currently be calibrated against independently established artificial consciousness outcomes"* and that indicator–consciousness relations are transferred from biological cases without independent support; Schwitzgebel (a co-author of that very rubric) says "We will not be in a position to know which theories are correct" (arXiv 2510.09858); and Anwar & Badea (arXiv 2404.15369) note *"a marked lack of consensus around what constitutes consciousness and... an absence of a universal set of criteria."*
So the defensible research strategy is not "test whether Darlin is conscious". It is: build mechanisms whose functional consequences are pre-registered and falsifiable, test them with parameter-matched controls and lesions, and report the results as claims about those mechanisms. Everything above Tier 2 is that strategy. Any Darlin claim that goes beyond it is not supported by this literature.