Epistemic Status: Short note of what are basically shower thoughts.
It occurred to me the other day that training model capabilities on math seems plausibly differentially beneficial for alignment over capabilities, amongst other effects. This is because largely capabilities progress does not seem hugely bottlenecked on mathematical knowledge and skill. Rather, capabilities seem to be mostly about increasing compute, scaling up existing recipes, and empirically tweaking existing recipes, and finding ways to get better data and extract better signal from the data that we have. Obviously these all involve some degree of math, but that is largely not the core bottleneck here, and often the math it requires is fairly basic college-level linear algebra, calculus, numerical methods etc. Conversely, a lot of alignment agendas, especially in agent foundations, seem extremely bottlenecked on mathematical skill and capability, since we need to be able to come up with novel methods to prove things about unknown agents in highly general settings.
Conversely, capabilities such as coding, ML research, etc seem extremely strongly differentially beneficial for capabilities. Then there are third categories of capabilities such as generic office work, protein folding, playing video games etc, which seem to neither be differentially helpful for capabilities or alignment, and thus are neutral at least to first order. To second order, insofar as they increase AI lab revenue or increase prestige/media attention due to their innate interestingness and hence ultimately lead to greater lab funding, then they obviously differentially favour capabilities.
Another way to think about this, at a higher level, is that math capabilities primarily help increase understanding rather than execution. Understanding is obviously helpful for capabilities too. If we had a fantastic mathematical theory of deep learning, for instance, that would presumably help capabilities. But at the same time, empirically ML seems to progress heavily through experiments and scaling things in a mostly a-theoretical way. So for capabilities, understanding is nice to have but is not strictly necessary. Conversely, for alignment, having a deep understanding of what is actually going on seems extremely important. It is possible that we muddle through by just using prosaic alignment techniques and hoping that the magic of deep learning generalization saves us, but this seems rather unlikely and a very precarious position to be in. So it seems that to have a deep and predictive theory of alignment, we need to have a serious mathematical theory of things like agency and RL in the idealized limit, how neural networks empirically learn and generalize, how they learn and internalize various goals through different types of training etc. Obviously some of this is deeply intertwined with capabilities too – if we understand how neural networks learn and generalize, then we also probably understand how to make them better (unless our a-theoretical current approach is somehow optimal). Nevertheless, this seems like an unavoidable externality since it seems to me that we need a deep understanding of how neural networks work to ever hope to be serious about aligning them.
Perhaps the cleanest way to think about this then is that math capabilities primarily boost theoretical understanding. Understanding is, I claim, effectively necessary for alignment and helpful but unnecessary for most capabilities research, and hence differentially accelerates alignment, at least insofar as there are actually people working on using the math capabilities of these models to actually accelerate alignment.
A potential complication here is that AI safety math is potentially very underformalized compared to a lot of math, and it is possible that contemporary models are lagging behind here. I.e. when we see models prove famous conjectures, they are typically proving things well within a known and extensively developed mathematical system. There haven’t yet been any results where a model has struck off and invented a whole new field of math by itself1. It is possible that the kinds of progress we need in AI safety is more similar to this than e.g. proving conjectures or developing the consequences of mathematical models within existing fields of math. I am somewhat skeptical of this. It feels like our existing primitives of Bayesian inference, statistics, random matrix theory, information geometry, POMDPs, model theory, logic etc, seem amply sufficient to express both the problems and potential solutions to the problems that we are likely to face. But then again this could just be an unknown unknown and we don’t know what we might face. A similar argument also applies to hypothetical capabilities improvements from a theory of deep learning. Here it seems likely that we can acquire a suitable theory of deep learning from existing mathematical fields. The hard part is knowing what are good toy models which nevertheless adequately represent the dynamics, and then understanding which features of these models have the most importance at realistic scales. If we have a good theory of deep learning, including generalization and RL, then it feels like this should sit close in latent space to nontrivial insights into alignment. But nevertheless, perhaps not all of the insights. And so this gap could remain problematic.
-
Although if it did, would we even be able to recognize it as a contribution? I think a lot of this is the labs are primarily optimizing for extremely legible and verifiable math metrics like ‘solve famous conjecture X’ while if the AI invents some new subfield that is much harder to verify and much harder for everybody, including them, to understand since it requires actually reading and understanding the methods the AIs used. ↩