A Regress Argument Concerning Autonomous Self-Correction in AI Agents 

Julian Lee-Sursin & Hong Joo Ryoo

École Normale Supérieure & John’s Hopkins University

This paper advances a regress argument against treating “self-healing” as a coherent property of artificial intelligence agents. By self-healing we mean an agent’s capacity to detect and correct its own failures without external intervention, a notion that has acquired prominence in work on large language model based agentic systems. We contend that this concept of self-healing, which is increasingly discussed in the emerging literature on AI, generates an infinite regress. In broad brushstrokes, our argument goes as follows. An agent that self-heals with respect to a task must assess its performance against some benchmark; yet assessment is itself a task admitting of failure; if self-healing extends to assessment failures the requirement iterates, yielding a hierarchy of meta-tasks that no finite system can traverse. We consider several candidate regress-stoppers, among them externalist proposals that ground assessment in environmental feedback, foundationalist appeals to primitive capacities, pragmatic views that tolerate imperfect self-healing, and deflationary accounts that dissolve the regress by revising the concept. Each strategy exacts a price. The upshot is that self-healing, though it may initially seem to be a coherent property, raises questions about the nature of reflexive capacities in artificial systems and imposes constraints on the design of agent architectures.

Chair: tba

Time:

Location:


Posted

in

by