Does Mechanistic Interpretability Research Provide Mechanistic Explanations?  

Joachim Stein

Heidelberger Akademie der Wissenschaften / University of Tübingen

Machine learning, in particular neural networks (NNs), is a powerful automation of induction and learning in general. Given a vast amount of data, machine learning algorithms can learn a prediction function. However, the learned prediction function as well as the learning mechanism themselves remain relatively incomprehensible to us ¬ they are termed opaque.

The need to overcome opacity has led to the emergence of the field of explainable artificial intelligence (XAI). Researchers in this area develop post-hoc explanation techniques for established machine learning methods. In recent years, there has been a proliferation of efforts to understand NNs through their internal mechanisms, giving rise to the field of mechanistic interpretability (MI). Both machine learning researchers and philosophers have argued that MI methods facilitate understanding by providing mechanistic explanations. Specifically, NN behavior is explained by decomposing the system into its constituent parts, identifying the functions of these parts, and showing how their interactions causally produce the phenomenon under investigation.

Several authors have argued that complexly structured systems cannot be understood through mechanistic explanations because they resist decomposition. At the same time, the structural organization of NNs appears to instantiate precisely such complexity. This creates a tension between the apparent complexity of NNs and the claim that MI provides understanding by offering mechanistic explanations. One might therefore conclude either that NNs are not as complexly structured as assumed and can, in fact, be decomposed, or that MI does not provide genuine mechanistic explanations. I argue that there is empirical evidence against the first conclusion. I then present case studies showing that central examples in MI do not, in fact, yield mechanistic explanations, supporting the second conclusion. Together, these arguments challenge the optimistic view that MI methods facilitate understanding via mechanistic explanations, which is prevalent in both the machine learning literature and its philosophical discussions.

Chair: tba

Time:

Location:


Posted

in

by