Build on our work
Our tools for susceptibilities, local learning coefficients, and SGMCMC sampling are open source in the devinterp library.
Work with us
Timaeus has merged into Resolution. Open roles are now posted on the Resolution careers page.
Timaeus is merging into Resolution. Read the announcement
These notes introduce the theory of susceptibilities for interpreting neural networks. The susceptibility of an observable to a data perturbation is defined as a derivative of a posterior expectation, which by the fluctuation-dissipation theorem equals a posterior covariance. Different choices of observable yield different objects: per-sample losses give the influence matrix (the Bayesian influence function), while component-localized observables give the structural susceptibility matrix that pairs model components with data patterns. The susceptibility matrix is (up to a factor) the Jacobian of the map from data distributions to structural coordinates; its pseudo-inverse provides a linearized solution to the patterning problem: finding data perturbations that produce a desired structural change. We motivate the theory from its statistical-mechanical foundations, then give a detailed exposition of susceptibilities, their empirical estimators, and their connection to the geometry of the loss landscape.
Our tools for susceptibilities, local learning coefficients, and SGMCMC sampling are open source in the devinterp library.
Timaeus has merged into Resolution. Open roles are now posted on the Resolution careers page.