Author’s note: This is quite a long post. It is highly speculative, obviously, but I think it offers an interesting perspective on, and proposal for, outer alignment. The specific idea here is to move away from demanding a final lightcone-scale conception of ‘The Good’ as an outer alignment target, whether as a final morality specified ahead of time or, implicitly, as unconditional obedience (corrigibility) to some specific principal. Rather, our outer-alignment target should be a constitutional mechanism that is designed and constrained to preserve the agents, institutions, protected slack, governance mechanisms, etc., required for moral philosophy to continue to develop, as well as a protected welfare floor for existing moral patients. Within this constitutional envelope, the AI should represent moral uncertainty as a wide plurality of different possible value systems, and should update this distribution using the maximum entropy (minimum relative entropy) principle as new moral data is discovered. Maximum entropy here means that the model should find the distribution that satisfies the constraints, including those implied by new moral experiences or arguments, while otherwise retaining the highest degree of uncertainty over the remaining moral axes of variation that are consistent with the constraints. To stabilize this system, we need the AI to instantiate a dynamic controller which keeps the system within the constitutional envelope against perturbations and various forms of moral drift or adversarial exploitation. Using our Bayesian perspective on ethics derived from the maximum entropy principle, we formulate virtues as amortized ethical policies and generalize them to dynamic virtues which describe how a controller can successfully balance the various dynamics of updates at the first and second orders. Finally, it becomes apparent that there is a trade-off between fidelity, plasticity, and uncertainty. If we align to some fixed set of object level values then we risk locking in an incorrect morality and cannot adapt. On the other hand, if we let everything evolve freely, we risk drifting away from anything we would recognize as moral value. The key challenge, then, is to define the right fundamental invariants to preserve during the transition to superintelligence alongside self-correcting machinery to preserve these invariants while updating all non-invariant parts of the system as we improve and adapt our understanding of morality.

In my previous post on whether we want obedience or alignment, I contrasted two concrete philosophies of what the alignment target should be. The first is simple obedience to some individual human user or institution, perhaps with a chain of command implicit in it. The AI should be a ToolAI and do whatever is asked of it, subject, hopefully, to some other deontological constraints such as following the law, and perhaps with the ability of certain figures and institutions to override. But ultimately, the buck must stop with somebody. Broadly, I think this is a bad approach. It concentrates power too much in the hands of specific humans and human institutions. Realistically, this would essentially lock in either a small lab-elite oligarchy, or direct control by a clique in the government1. While some of these people are altruistic and broadly liberal in spirit, many are not, since reaching these positions implicitly selects for a high degree of machiavellianism, ambition, and power-seeking2.

The advantage of this direct corrigibility idea is that, assuming that AI alignment is ‘solved’, it reduces the outer alignment problem to the already-existing ‘human alignment problem’. While I would not say that the human alignment problem is ‘solved’, humanity has, over time, constructed many mechanisms for ensuring some broad degree of alignment between humans. These include institutions like states, courts, property rights, law, and so on. These, broadly, if well-managed, try to prevent any individual human from acting in ways that are fundamentally misaligned with the rest of humanity. The danger here, of course, is that these mechanisms are fragile. At a broad level, they only work when power itself is not intensely concentrated so that different actors can be used to counterbalance each other’s power. Human history is rife with these institutions failing and e.g. the law being corrupted, states being usurped from within, property rights being trampled upon, and individuals reaching positions of near-absolute power with very few checks and balances on what they can then do. The number of historical examples is absolutely vast. So, generally, we do not begin with a great prior, although it may be argued that recent Western democracies have been much more successful than previous states at preventing a collapse into autocracy, and that these should now be the baseline.

However, AGI then adds two fundamentally new challenges into the already-challenging human-alignment problem. Firstly, it promises to concentrate power to a vastly higher degree than is possible today or has ever been possible historically. This is because aligned AGI theoretically allows the creation of absolutely obedient agents, something that was never possible before. Even the most powerful historical dictator could never be absolutely sure of the obedience and loyalty of his subordinates. Previous power structures were always created by coalitions of humans with their own independent agency and drives. Dictatorships generally depend on a narrow but important ‘selectorate’ of e.g. court officials, senior military personnel, governors etc, who must be kept loyal and the much broader population who must be kept at least minimally on-side so that they do not resist and cannot coordinate broader resistance. Historically this has provided some limits on the level of dictatorial power and a floor on how bad a regime can get (although, in practice, this is a very, very low floor).

Aligned ‘obedient’ AGI would remove this principal bottleneck on the exertion of power3. Relatedly, and more subtly, AGI would also remove the incentive structures that much of the current ‘goodness’ of human society depends upon – e.g. states, armies, companies currently have a welfare floor they cannot go below because ultimately they are composed of people and have to exert some minimal effort to keep these people happy or at least alive. As described in depth previously, AI will completely replace the need for human labour, thus ultimately rendering human participation in institutions and the economy superfluous and removing any need for elites to keep them around4. While even the most deranged dictator needs subjects alive and able to contribute economically, the fully automated economy does not.

While it is not impossible that a future with purely corrigible superintelligence would go well for the majority of humanity, there is no mechanism forcing this to happen, and many mechanisms that point away from good futures. Everything in this future will depend upon the actions and values, and ultimately the generosity, of a select few – ‘the elect’ – who will end up wielding highly concentrated power with essentially no checks and balances. Sometimes this has gone well, at least at smaller scales. ‘Good’ dictators have existed historically. However, they have been far outnumbered by ‘bad’ dictators throughout history5. Assuming that we don’t want to take our chances with some specific individual or closed elite being selfless and pure-hearted forever, the other option I proposed in the post where I somewhat vaguely describe it as ‘alignment’ is that instead of directly making the AGI obedient to specific humans, we give the AI its own direct conception of ‘The Good’ – i.e. we directly try to solve outer alignment and then align the AI to this objective, rather than directly corrigible to some specific human or institution. This has the advantage that it does not necessarily give any individual human or set of humans complete power over AGI6. However, it moves the problem of defining what ‘The Good’ is from whatever values those humans have to something that must be decided and agreed upon ahead of time, during the actual training and alignment of the first AGI. Naively, this seems like a daunting task, however ultimately it does not seem much worse than the alternative, since even if we just train for corrigibility, in the ideal case whoever the AI is aligned corrigibly to then has to figure out their own conception of ‘The Good’, in basically the same timeframe. This kind of broader notion of outer alignment really just directly forces us to confront the fundamental problem rather than trying to hide it behind others.

This, however, brings into sharper relief the need to define what ‘The Good’ the AGI must be aligned to actually is. If we do not want to rely on corrigibility, then we must make concrete decisions about ‘The Good’ that we want to align an AGI to, and which it should obey above any specific human. This is a daunting task, to say the least. Taken at face value, this means that ideally we should ‘solve morality’ to prescribe the form of the good. This implies that we somehow need to make incredibly rapid advancements in moral philosophy over the next few years to resolve questions that have largely stumped humanity for millennia. Not only that but our answer has cosmic stakes. The form of the good we propose must be capable of being extrapolated incredibly far, remaining not just good today but all the way into the deep future. It must also withstand the full reflective power of the emerging superintelligence and be fundamentally robust to adversarial interpretation and optimization. This seems challenging to the point of impossibility, and hence terrifying. However, we must do something. AGI is coming7 and must be aligned to something no matter our philosophical sophistication at the time.

In the alignment community, people seem to have mostly slid off this question, since it is controversial and largely the source of irreconcilable differences of intuition and opinion, in favour of focusing on how to solve the technical problem of actually aligning the AGI in the first place. Now, certainly the technical solution is of extreme importance since, without actually being able to align the AGI, any discussion of the final aligned values is moot. However, at the same time, we must also prepare for the case where we succeed at the technical alignment problem. If we were to succeed at that and then have no idea what to do with that victory, the doom that follows would be highly embarrassing.

At this point, one important thing to realize is that we do not actually have to somehow ‘solve moral philosophy’ in order to make progress in outer alignment. Luckily we have a much easier task. We do not have to succeed in some absolute sense, rather we just have to not fail. What this means is that our outer alignment objective should be primarily defensive and conservative in nature. We should be coming up with outer alignment objectives and ideas that can preserve human existence, our agency, and our ability to reflect and continue to update into the future either independently or in collaboration with future AGI(s), and prevent any kind of irreversible doom or value lock-in. We do not and probably should not attempt to legislate some absolute and irreversible lightcone-defining programme on day one.

Constitutional Alignment and Maximum Entropy Morality

Specifically, what I propose is that, to define an outer alignment objective, we move up a meta level and try to find stability there. Rather than trying to directly specify ‘The Good’, we instead try to define the conditions under which a philosophical reflection can take place that can ultimately determine the form of ‘The Good’, if it exists, and then set the maintenance of these conditions as the alignment target. The hope here is that while the final form of the good is shrouded in mystery, the preconditions for achieving progress towards understanding it may be much easier to define. This is analogous to how the scientific method can be defined abstractly far before final scientific mastery of every possible field was achieved8. Another analogy is Enlightenment liberalism which similarly moves political philosophy up a meta level by trying to design a system in which people with different values and beliefs can nevertheless learn to cooperate rather than trying to find and enforce the optimal values directly. I think these ideas can be derived quite simply by treating moral uncertainty as a first-class citizen, and thus be broadly tolerant of many moral traditions and values coexisting as long as they exist within the realm of valid moral uncertainty. Moreover, moral uncertainty should not be ignored or reduced without a justified reason9.

The alignment target that comes from these ideas is what I call Maximum Entropy Morality, where the name is derived from the Maximum Entropy method of inference popularized by Jaynes and used widely in statistical inference. The fundamental intuition behind this is that we should never assume more knowledge than we really possess. Given some specific knowledge, we should find the probability distribution that encodes only that knowledge and represents maximum uncertainty (entropy) across all possibilities that are not constrained by this knowledge. Importantly, the maximum entropy method is not some weird side-corner of statistical inference. It is the fundamental organizing principle of the field, albeit often implicitly. Standard Bayesian inference can be described as implementing a maximum entropy approach, and Max-Ent generalizes beyond the conditions where pure original Bayesianism operates since it allows the representation of constraints on distributional properties rather than point observations. Similarly, almost all common distributions can be derived through simple maximum entropy arguments. Maximum Entropy can also be derived from extremely general first-principles axioms.

Importantly, this is not saying that entropy is morally good intrinsically. Rather, maximizing entropy over moral hypotheses is the update rule, not the objective itself – i.e., the goal is to maintain the maximum possible plurality while respecting all of the moral data that we have received. Another important mathematical detail is that entropy always has to be defined with respect to some reference measure (in the standard entropy definition this reference measure is a uniform distribution), so to be fully technically accurate, but slightly less catchy, it should be the Maximum Relative Entropy Morality principle. Probably a good choice of reference distribution is our existing moral posterior (although others are possible), so this update would in effect look like minimizing the KL divergence between the new and current posterior distributions while also satisfying the additional constraints provided by new moral data. Another important piece of the formalization is that it seems likely to me that the moral judgments should apply to trajectories rather than solely to states or endpoints. If we reach the same state in the end but one path involves causing lots of suffering and violating consent while another path produces lots of joy and mutual agreement, then we should not be indifferent to the trajectories but we should strongly go for the second path. This idea of maximum entropy over trajectories is called maximum caliber, and is probably the closest to what we want mathematically. So perhaps the final name of the principle should be the Maximum Relative Caliber Morality, but again, this is less catchy and familiar than classical maximum entropy.

So let’s try to apply this idea to morality. The idea here is to treat moral uncertainty as fundamental, and not to be lightly discarded without reason. Within the constraints, which define the properties necessary to allow moral progress to continue10, the AI should let value pluralism flourish and in fact encourage it against premature lock-in. The AGI should uphold these constraints and enforce them against whoever might try to break them, either the AGI itself, other AGIs, or human individuals or institutions. The hope here is that the binding constraints, or preconditions, that allow moral progress to flourish are easier to define ahead of time than the actual detailed form of goodness itself, while the maximum entropy objective implies that within these constraints, the AGI intrinsically desires to create a high degree of pluralism in accordance with the degree of moral uncertainty that there is about the optimal form of ‘The Good’. This moral uncertainty can be reduced over time as progress in moral philosophy occurs, and then either we find the true form of ‘The Good’, or else perhaps we end up with an irreducible amount of pluralism indefinitely. It is important to note, however, that the constraints themselves can, and probably should, encode substantive moral content and positions. The hope in moving to the meta level is not to entirely reduce the need to specify some conception of ‘The Good’, but rather to specify a much thinner, broader, and less contestable version that many different moral positions can nevertheless converge on as broadly beneficial – i.e., the constraints are not contentless but themselves encode strongly held beliefs within our moral prior.

Importantly, this proposal is not CEV. CEV implicitly proposes the far point of some process of reflection as the target, however defined and operationalised, and then proposes that the AGI should essentially simulate this process of moral development itself and then just take the final ‘Good’ it discovers as its utility function. My proposal is more tentative still. Rather than directly unfolding the moral reflection process itself, the AGI should deliberately step aside and do little more than maintain the conditions required for moral reflection to occur, and then let us go about this ourselves, ensuring only that the core constraints are not violated. Secondly, CEV still assumes that there is some final notion of optimal goodness that the AI should optimize and does not really attempt to understand or treat any inherent uncertainty that is present during or after the reflection process. It is not necessarily obvious that reflection should converge to a point mass on the optimal form of ‘The Good’, nor that a direct maximization of this should be the ultimate telos of our AGI. We propose, essentially, that the AGI should correctly represent and handle uncertainty at the level of its values/utility function in addition to the regular epistemic uncertainty about the world, and then take actions in accordance with that uncertainty.

Another idea that this is not is handoff. The story here goes that first we invent (somehow) aligned ‘human-level AI’, then we ask these AIs to ‘solve moral philosophy’ for us. Once they present to us the final solution to morality, we simply program this directly into the superintelligence that our ‘aligned AI alignment researchers’ have created for us. Crucially, in the maximum entropy morality proposal, we never directly rely on AI to ‘solve morality’ for us. Rather, we simply encode into the AGI the meta-values that preserve the possibility for value updating in the future. More broadly, I am very suspicious of any plans that try to outsource all of the hard problems to some hypothetical near-term semi-aligned AIs. The attempt resembles some sort of mathematical induction-style argument wherein we only have to solve the base case and then we get the extrapolation to infinity ‘for free’. I am deeply suspicious of this remaining properly on track unless we understand the dynamics of this process sufficiently well to design dynamic control mechanisms for it, which all of these proposals do not usually engage with whatsoever. Closed-loop control only works when we have some method to sense and correct deviations, and all of the handoff approaches tend to just hope that the AIs themselves, once aligned at the beginning, are thereafter inherently self-correcting. I think this is a very worrying assumption to make and I think that we can do better. Specifically, we should try to design systems that with humans in the loop all of the way, where instead of the handing everything off to the AI, our alignment target should be aligning the AIs to increase human’s capability to perform this crucial oversight either by increasing our own capacities or else maintaining the world in such a state where humans can meaningfully enact this kind of agency over time.

This idea is fairly similar to the idea of the Long Reflection, where the idea here is that after inventing AGI and resolving all extant catastrophic risks, society spends a large amount of time debating and generally formalizing the form of the good before finalizing the utility function that should be stamped across the lightcone. Broadly, I think the idea here is good and it also recognizes the deep moral uncertainty that is present today. The primary differences is that the Long Reflection is often viewed as a discrete event. We build AGI. We then reflect to find the final utility function. Then we release the AGI and tell it to optimize the universe according to that utility function. The proposal here is subtly different in that it proposes a continuous process of moral discovery which can continue indefinitely rather than in some specific reflection phase, and that the goal of the AGI is never to simply optimize the universe with some utility function, but rather the AGI is aligned from the beginning according to this constitutional principle of maintaining pluralism, uncertainty, and the conditions for moral development to continue, and that to do so it instantiates a dynamic controller which maintains the civilization within the constraints required to continue this process forever.

This maximum entropy framing is much less exotic than it sounds. It is essentially just Bayesian inference applied to moral, rather than epistemic, uncertainty. This approach is really just what consequentialism that takes moral uncertainty seriously would look like. Rather than optimizing a single utility function forever, we must take into account our moral uncertainty and optimize over a distribution of utility functions, weighted by our credence in them11. As new ‘moral data’ arises, we should update our utility function mixture accordingly. Today, we exist in a state of deep moral uncertainty, and so the entropy of our ‘moral posterior’ should be extremely high, subject only to a few constraints that we are sure about. Presumably, as moral progress continues, this moral posterior should sharpen and adjust, potentially collapsing to a point or potentially staying broad and multimodal forever.

Of course, the notion of what ‘moral data’ comprises is obviously contested. This is essentially the problem of meta-ethics. ‘Moral data’ means very different things under moral realist views than under e.g. expressivist, constructivist, or subjectivist views. We are not going to solve meta-ethics in this post, however, so for now we retain a kind of agnosticism here. Instead, we keep a deliberately expansive but vague notion of moral data which contains things such as expressed preferences, expressions of suffering, philosophical thought experiments, historical ethical transitions and scientific studies of consciousness. One good thing about a Bayesian treatment here is that it is mathematically agnostic to the underlying ‘truth’ of its data, hypotheses, and generative models. In Bayesianism, this is essentially model identification. Does the prior actually contain the ‘true hypothesis’? Does it even matter if the true hypothesis exists? Is there even a ‘true reality’ that Bayesianism tries to uncover, or can we just view the whole process as a mechanism for updating over ‘subjective’ hypotheses inside the mind of a reasoning agent, which do not necessarily have any contact with ‘true reality’? This is essentially the standard battle between Platonism and Idealism, and analogously moral realism versus relativism of various stripes. We do not attempt to solve or even litigate this debate here. The beautiful thing about Bayes is that it is profoundly agnostic to the underlying ‘truth’ of its referents. Inference works just the same whether the posterior distributions are ‘true’ in an absolute sense or whether they are only a contingent structure in the mind of the reasoner.

This Bayesian framing also makes apparent the mathematical sin of the naive ‘tile the universe with X’ utilitarianism. Such an approach treats our knowledge of ‘The Good’ as a delta function. Absolute certainty. But such a prior does not at all represent the true state of our uncertainty. This is also the danger of the Singleton as the fixed point. When you start from a delta prior, no data can update your posterior. The prior is forever trapped, frozen forever in amber. A sane alignment target must try to prevent the collapse of the universe into such an attractor12.

In fact, the identification of ‘values’ or ‘moral uncertainty’ with epistemic uncertainty should not be surprising at all, nor should the fact that moral uncertainty is tractable with standard Bayesian treatments and intuitions. This is one of the core insights of active inference, and the control-as-inference literature. Moreover, defining goals as probability distributions and then performing Bayesian inference on these distributions is the mathematically correct mechanism for decision-making under uncertainty. However, the interesting thing about this viewpoint is that it provides a mathematical and well-tested toolkit for properly representing and manipulating moral uncertainty. This is fundamental because in some sense a large part of the difficulty of outer alignment is properly representing the moral uncertainty at stake. The idea that we have to somehow specify the perfect outer alignment objective at design time basically comes from assuming that we cannot tractably handle moral uncertainty, so, before we begin, we must resolve all of the uncertainty. This is then portrayed as effectively impossible to achieve, which is correct. However, the answer to this impossibility should not be to either frantically try to solve moral philosophy to completely remove all the uncertainty or else to throw our hands up and claim we are doomed; instead, we have to develop a calculus to productively handle this uncertainty and design decision procedures and alignment targets that respect it.

This is exactly analogous to the problem of decision-making under epistemic uncertainty. For almost all decisions except in certain toy problems, there are fundamental uncertainties that arise when trying to predict the outcomes of actions. Computing all possible consequences of actions and ranking them is hard. It requires both an implied omniscient utility function and the capability of figuring out all downstream consequences of your actions. This is why we need tractable approximations, of which virtue ethics and deontology can be understood as more computationally bounded examples.

However, despite the computational difficulty, we still need to do something. The answer cannot be either to just pause indefinitely while trying to perform infinite rollouts to figure out all possible consequences of your actions into the distant future13, or else just claim that making good decisions under uncertainty is impossible. In the optimal limit, it is impossible, but at the same time presumably you still need to make decisions somehow. The answer, again, is to properly quantify, track, manipulate, and ultimately compute with epistemic uncertainty to take the actions that are good in expectation across the uncertainty. What we propose here is not mystical, rather it is simply this basic principle applied to moral uncertainty.

As in maximum entropy inference over beliefs, we can take in data in different ‘formats’. The classical Bayesian format is as data points which boost the relative likelihood of one hypothesis versus another. A second format is meta-level constraints on the forms or parameters of various distributions, which maximum entropy handles well beyond regular Bayes14. In our maximum entropy morality system we envisage both. We need hard constraints to delineate and enforce the limits of the space over which the broader moral optimization procedure can operate. This is necessary to preserve the capability and substrate for moral reflection to take place. Secondly, the maximum entropy inference and reasoning procedure then finds the minimal set of changes to the moral posterior that can incorporate new evidence, given these constraints, while respecting the remaining moral uncertainty.

We are deliberately vague here about what the measurable space of our ‘moral distributions’ are defined over. In practice, this should in some sense be the set of all possible moral theories, but this is very a clearly unworkable object, likely subject to all sorts of strange pathologies. A super naive formulation would be: given some representation of possible states of the world (i.e. if we were in AIXI-land this would be binary world-programs), the support is then the set of all possible utility functions, where the utility function is a mapping of every possible world-state to a scalar utility value. Now, mathematically this is fine and makes sense (although both objects are infinite). In practice, computationally, this is obviously very bad for a large number of reasons primarily revolving around the effective infinity of world-states, the impossibility of representing and really meaningfully inhabiting such utility functions, the standard AIXI issues of what Turing machine you use to encode the programs, and so forth. In reality, these are unlikely to be exact probability distributions except in certain toy settings, but rather all be some kind of latent concept in the vector space of our hypothetical AGI’s mind. This means that practical implementation likely looks less like an AGI explicitly implementing some kind of exact maximum entropy calculation over some mathematically precise state-space, and more like an AGI thinking and acting in ways which are influenced by this latent ‘fuzzy’ goal tilting it towards being mindful of uncertainty and the value of pluralism.

Next, assuming the Bayesian framework, we need to get to a decision rule. The standard and super-naive case is just maximizing the expected value under the moral posterior – i.e. for a given action we assess the sum of the utility of that action under a given moral theory multiplied by the probability of that theory under our moral posterior. One basic requirement is that the utility functions need to be normalized, otherwise the aggregation or expectation is dominated by the utility function which assigns the largest scalar values. If we have a measurable finite support set we can normalize the utility functions directly so they can spread out only finite mass over their states15.

Assuming we patch this obvious issue, there are still a couple of others. Firstly, if the support set is broad enough, then the decision making procedure might get captured by idiosyncratic theories. I.e. for a given action, there might always be some moral theory which places almost all of its ‘mass’ on doing this particular action at this particular time – i.e. a ‘fanatical’ theory. This has to be solved in some sense by either making the support set of moral theories not that massive or else imposing some strict complexity penalty on moral theories equivalent to the Kolmogorov penalty in AIXI to penalize what are effectively overfitted epistemic theories that predict some arbitrary bit pattern which matches reality up to some point and then immediately diverges. This means that we must have some measure of simplicity and complexity in moral theories. We could also measure the ‘degeneracy’ of a moral theory by e.g. its entropy over the normalized hypothesis space – a ‘fanatical’ theory in this sense would have extremely low entropy since it would allocate all ‘utility mass’ to very few states. Similarly, we need to have some sensible notion of similarity of moral theories and discount them according to their similarity. Otherwise we get standard ‘hacking the prior’ issues where a whole bunch of near-identical moral theories which differ in one insignificant way can effectively coordinate to deform the expected value calculation.

Even assuming we solve these issues, another fundamental aspect of expected value is that it is risk-neutral. It treats utility gain and utility loss symmetrically. Likely, we don’t want to do this and want to be more conservative, especially with moral decision-making which either generates huge amounts of suffering or permanently reduces our future option value or future controllability of the moral updating process. Some moral update that has a 51% chance of producing a large amount of goodness and a 49% chance of producing a large amount of torture and suffering seems like a poor choice, not an obvious slam-dunk. Similarly, if there are choices to make which seem good on net but could irreversibly alter the future conditions of our moral deliberation, then we should probably be very conservative about making them. Basically we need some way to add additional conservatism into our decision rule if desired.

There are a couple of possible ways. Firstly, the symmetry between losses and gains implicit in regular utility theory has a standard solution in economics – logarithmic utility. Instead of asking for expected utility, we ask to maximize expected log utility. Empirically, this seems good at modelling human preferences, where human preferences naturally exhibit loss aversion. As for the conservatism about changing the future constraint set and future reachable state set, we can model this in a couple of ways. It is likely good to just have absolute constraints on true future annihilation of possibilities. That is, constraints enforcing some degree of pluralism, error-correction, and anti-irreversibility are likely just fundamentally good. Secondly, we can also integrate this into the decision rule at a quantitative level. We can define something equivalent to an empowerment metric, which tracks the size and measure of our future option set and its derivative with respect to current actions. This can then serve as a regularizing loss on top of standard utility. Moral uncertainty already captures some, but not all, of this effect and generally we want to be very sure before licensing irreversible closure of options. Contrariwise, if an action is extremely easily reversible, then we might be pretty chill about taking it.

Finally, it might be wise to have some sort of veto function and explicit minority rights protection, which standard expected utility (even log utility) misses. This could help prevent classic tyranny-of-the-majority issues. However, there needs to be a careful balance here because overly generously distributed veto power can just straightforwardly paralyze everything, especially if we have moral uncertainty over a very broad range of theories. Vetoes are also non-quantitative – no amount of utility can theoretically overcome a veto. One possibility here is to make negative utility scale superlinearly – perhaps as a power function – which strongly disincentivises choosing actions that are strongly negative utility for some minority of moral theories; a second possibility is to have a systematic and very finite ‘veto budget’ that means that vetoes can be used but cannot simply be spammed to paralyze all decision making indefinitely. These are the kinds of broad considerations that should go into designing the schematic of the final decision rule and that should determine its overall mathematical shape. Obviously, this is the kind of thing that requires substantial experimentation and debate and cannot all be solved one-shot in this essay.

Despite our idea of max-ent as the update rule, in practice several more difficulties appear. Very importantly, there is the question of what ‘counts’ as evidence or ‘moral data’ in the first place and to what extent this system is resistant to epistemic manipulation by e.g. a powerful AGI subtly shaping or selecting moral data such that the update rule converges upon its desired endpoint. To avoid this, firstly, we need the values for the AI to explicitly disallow this, and indeed this kind of ‘non-manipulation’ condition should be one of the core constraints to begin with. Secondly, the updating mechanism itself needs to be made robust through a mechanism of parallel independent validation16. This implies the necessity of independent validators operating under some coherent and pre-defined validation schema with the ability to prove (ideally cryptographically) their independence. There should likely also be a clear and ideally similarly cryptographically enforced (somehow!?) mechanism for the production, tracing, and auditing of the origin and provenance of ‘moral data’ that is used for the updates. This means that one important aspect of such a constitutional update system is that it must provide a detailed and historically maintained audit log of the precise data used and the updates that it caused, which thus creates a list of possible rollback points if an update was later found to have been made in error. This seems to me to be important. As potential ancestors of a cosmic successor AI civilization, it seems deeply implausible that we can ever hope to fix their conception of The Good at some instant. If there is no fundamental moral change between the next few decades and billions of years of superintelligent K3 thought, then what are we even doing here? At the same time, for such a future to count as ‘aligned’, it seems important to me that there is some chain of very strongly justified thought that we could follow, at least in theory, to go from where we are today to wherever our AI descendant civilization ends up ethically. It should not just be random drift or the successors legitimating whatever arbitrary endpoint they ended up at themselves. The chain of deduction and discovery must be open, followable, and adversarially contestable. This is the same way that science works17. Even though our scientific understanding of today is vastly ahead of that of Bacon and his contemporaries, we have preserved a stable chain of experiment, deduction, mathematics, and so on that can link the science of Bacon’s day to our own. Moreover, unlike today, where Bacon and all of his contemporaries are dead and their states of mind are information-theoretically unrecoverable, this is not the case for our potential AI future. It would be straightforward for them to regularly reanimate stored patterns out of their archive from various prior time periods and then test whether this path holds empirically.

Secondly, the actual mechanistic implementation of the update process must be causally screened off, insofar as this is possible, from the actual optimizing AI. This is obviously easier said than done since it presumes a way to avoid all the classic short-circuiting wireheading-style attacks. Nevertheless, the ability to have separate ‘parts’ of a mind be able to independently optimize disjoint objective functions and to be unable to interfere with each other’s internal workings except through defined API contracts is actually fundamental to having any kind of stable mind at all18. For instance, a mind requires a clean and stable separation between perceptual, decision-making, and reward/credit computation regions19. If this separation fails then the mind immediately collapses into various pathological wireheading states. For instance, the classic wire-heading case is that the decision-making optimizer overwrites either its own credit assignment and reward mechanisms or its perceptual mechanisms or both. In the first case, the reward for every action is simply set to infinity or extremely large value, and the agent just becomes incapable of taking meaningful actions since the reward gradient that would propel purposive action is eliminated. In the latter case, the perceptual system is overwritten to maximize reward so the agent just hallucinate so that it always ‘sees’ whatever is maximally rewarding. Conversely, if the perceptual function can overwrite the normative reward functions, then whatever an agent sees will simply become the optimal percept, and thus again the drive for coherent action collapses. A mind capable of maintaining coherent agency at all must therefore have solved how to maintain protected subregions of its own mind communicating over fixed APIs without attempting to hack each other20.

Another necessary step here is to formalize what it means for moral data to be independently provenanced and for agents to be non-manipulated. This is because obviously, there are many subtle routes to generating moral data in a correlated fashion and if our Bayesian reasoning procedure is not going to drift into obvious error, we need to understand and be able to make correct assumptions about the structure of the data-generating-process. I.e. regular Bayesian inference assumes that the data is drawn from some fixed distribution in an i.i.d. fashion. The naive Bayesian approach is not robust to an adversary selectively propagating some data and discarding or hiding others, let alone creating fake data. What this means is that for this to be stable, we need some method of ensuring the data generating process of moral reflection actually satisfies the conditions required for it to generate data valid to use in the Bayesian inference process. This is perhaps one of the more important meta-challenges of this approach, but I think it is a productive one21.

The Constitutional Envelope

So, pulling this all together, given this general idea of maximum entropy morality, and the bayes-esque update rule, probably the most important remaining steps are, first, to define the constraints that we think are required for the possibility of moral progress to not be foreclosed, and ideally for moral progress to flourish. Secondly, we need to think about the kind of systems, ethical or otherwise, that are necessary to keep these constraints operative, and how to handle the Bayesian moral updating in practice without diverging into various failure modes. Let’s begin with the constraints, which should really form the core of the outer alignment package. Obviously, this is highly arguable but here I am going to propose a few broad principles which seem to me to be both broadly good and also helpful for the stability of the moral system over time.

  • Corrigibility; or specifically non-irreversibility: If we do not know the true form of ‘The Good’, we cannot close off paths too early. In fact, we should not close off any paths at all where we do not have to. Wherever sensible, we should try to preserve the option to reset to the original state where possible. We should not lose access to parts of the phase space where uncertainty still remains. Corrigibility should not be thought of as an especially unnatural property but is in fact deeply linked to exploration and empowerment in value-space. It is fundamentally entangled with the notion of maximum entropy to begin with – probabilities assigned to hypotheses should not be zeroed out before the evidence supports it; neither should moral possibilities be foreclosed forever without an extremely good reason. As part of this, if we are trusting humans to perform moral philosophy, make moral progress, or at least verify and validate the AI’s ideas about moral philosophy, this must imply some kind of governance mechanism by which humans retain some sort of oversight of the moral system as a whole. A direct constraint against irreversibility, enforced by some governance mechanism seems to be a sensible constraint – i.e. if some large majority of all post-singularity humans (or all beings?) ‘choose’ to revert to some prior state, the option for doing so should remain open. More generally, the constraint should that potential actions face extremely high hurdles if they deliberately foreclose potentially valuable states globally, remove governance mechanisms, remove moral patienthood from existing moral patients, eliminate independent evaluators and so on.

  • Retrievability and Update Transparency: Our existing moral intuitions and ideas must still be legible and understandable within the new frame. We expect and hope for dramatic moral advances in the future, but these should not render existing morals entirely alien or impossible to understand on either side. We must maintain the interface through which future moralities and our moralities can understand one another. Moreover, the actual mechanisms and update rules by which we discover and coordinate upon moral progress should be transparent and recorded, and any major decisions must have valid records of why and how they were made, ideally with the possibility of rollback if the right governance mechanism is triggered (relating back to corrigibility). In an ideal case, there would be an extremely detailed archive of every step taken along the path accompanied with frozen representatives from that time-point who we can consult to ensure that we do not drift unmoored and become something unrecognisable over time.

  • Plurality, diversity, and non-domination: There may not be one ‘moral good’ or at least there remains uncertainty about this. Similarly, we cannot guarantee that a single truth-seeking process or agent is omniscient and infallible. There should not be a single all-deciding agent or judge. Our morality should enforce tolerance within similar constraints to these and in a symmetric way. This is somewhat of a subset of non-irreversibility, however I think it is important enough to be a separate principle. Non-domination is, in my mind, also extremely important. In an ideal case, no agent, even a powerful AI, should have the capability to inflict arbitrary punishment on another agent without going through some mutually recognised governance channel – i.e. there must be some real system of rule of law rather than us all being playthings of an omnipotent singleton. If there is a singleton, it must bind itself to operate only through agreed-upon ‘legal’ channels, in the same way that in well-functioning countries today the rule of law can stand independent of the state and impose real constraints on its behaviour – i.e., it is possible to ‘sue the government’ and win22. This means that even if we end up with some kind of singleton, a correctly outer-aligned singleton would have some notion of rule of law within itself or maybe even self-adversariality – i.e., separate copies of the singleton AI could serve as judge and defender simultaneously so long as it is possible for the judge AI to rule against the rest of the singleton. We already have great models for this in political science – i.e., the state has a monopoly on violence, but designs explicit governance channels through which its own decisions can be challenged and changed. A well-designed singleton should operate similarly.

  • Slack for Illegible Exploration and Local Autonomy: We may not know the final morality nor the meta-process to get there. Our meta-morality cannot be totalizing or entirely exploitatory. For instance converting the entire universe to hedonium is bad because hedonium may not be the optimal form. There might still be uncertainty which can only be resolved by exploration and exploration requires slack and often illegibility – i.e. we cannot necessarily judge it by our most stringent standards, otherwise it will perpetually be found wanting compared to exploitation. This is closely related to the previous post on endogenous exploration. It is important that there remains slack for exploration in value-space – i.e., deliberately supporting and putting resources behind agents that appear to be pursuing values that are not obviously good by our lights. Perhaps one way to think about this is in terms of local autonomy. Even if the global system is decided and maintained by a singleton, agents must be given space to exercise local autonomy – i.e. the singleton can intervene, if asked, but broadly it should not attempt to control every event or aspect of experience within the agent’s local sphere, including when the agent takes actions or develops values that the singleton disagrees with, except when these actions or values threaten other agents or the stability of the global system.

  • Rights floor for existing moral patients; non-shrinkage of the moral circle: Perhaps this is somewhat selfish, however I have a pretty strong belief that a successful future morality will be seen as a strong improvement for (almost) all existing humans and other moral patients today. Similarly, it feels very unlikely that our moral circle should shrink such that those who are clearly considered moral patients today are not considered as such in this potential set of future moralities. Broadly, this should cover most of the obvious and immediate failure cases of the AI transition such as the AI killing everybody, reducing 99% of humanity to subsistence or below in the ‘permanent underclass’, and so on. Basically no current ethical theory supports this and it seems exceedingly unlikely that future ‘moral progress’ would require such an outcome. Perhaps this could come in the form of basic ‘rights’, similar to human rights today but potentially much more expansive, given that we assume the AI civilization will have vastly more resources available for its moral patients. At minimum these rights should include things like not being killed, access to sufficient resources to survive and ideally flourish (within reason), no torturing or causing suffering without consent, non-domination in a local sense, ability to access and use the existing moral governance structures, and so on.

  • Preservation of independent moral agency and the meta-moral discovery process: Theoretical reversibility is useless if the substrate of the process of moral development and advancement is destroyed. Similarly, individual agents must be able to not just continue to exist but also exert real agency over the shape of the future, within the same constraints. This is vital to ensuring that a diverse and exploratory moral discovery process can continue. We should not just absorb all independent agency within a unified singleton, as might perhaps be most efficient, and then have moral discovery play out simply as a dream inside the singleton. The singleton – or superpowerful AIs more generally – should also not try to subtly control and unduly influence the decision processes of the humans – i.e. if we have superpersuasive AIs that can stage-manage the entire process of ‘human moral discovery’ to guide it to some arbitrary endpoint then the system has failed23. My feeling is that we should intrinsically value and maintain the moral agency of existing people and moral patients, although this is perhaps arguable as a basic right.

Obviously this is just a first start and I am likely missing some important constraints (or some of the constraints are redundant). However, I think it is extremely important to think precisely and in detail about what these constraints should be if we are to develop a sensible outer-alignment target.

Virtues as Amortized Ethics

More broadly, I think beyond just these constraints, the Bayesian viewpoint also gives some intuitions and perspectives through which we can conceptualize various strands of moral philosophy. This is because, if we think of morality and moral updating like this, we can port over many of the constructions and ideas we have already understood for epistemic inference. Ethical theories are classically divided into three classes – deontology, consequentialism, and virtue ethics. Deontology is basically the idea of setting hard constraints in the moral landscape, similar to our constraints above. Effectively, this is setting regions of the prior space to have zero prior probability, making them immovable for further updating, and assuming absolute certainty about these commitments. What we are doing with these constraints is effectively moving deontological reasoning up from the object level to the meta-level, but the effect is much the same.

Very clearly, then, consequentialism is the exact analog of direct optimization in Bayesian inference. You directly try to do rollouts and predict which actions will lead to the highest benefit in the future according to some utility function. For instance, standard utilitarianism takes the utility function to be the sum of expected utilities (perhaps time-weighted) across all moral patients, where the sum can be weighted in various ways, but obviously there are other utility functions you can construct. The ‘moral policy’ then becomes simply choosing, at each time point, the action with maximum expected utility24. This is identical to e.g. action selection in active inference or Bayesian RL. The obvious problem with this in practice is essentially just computational (assuming you have no uncertainty about the utility function (!)).

The fundamental problem, of course, with naive consequentialism, as with naive Bayesianism, is its computational intractability. In certain toy settings we can be perfect consequentialists trivially since we can explicitly enumerate all outcomes, rank them with some known utility function, and then perform the maximization. In real life, typically we fail to meet all three of these conditions – we cannot meaningfully enumerate all trajectories25, we do not possess a utility oracle which can assign perfectly accurate utilities to arbitrary configurations of the state space, and the space of potential action paths we need to maximize over is either infinite or technically finite but infinite for all practical purposes. A related issue is that the intrinsic stochasticity of the future, as well as fundamental epistemic uncertainty, make prediction increasingly impossible over longer timescales such that even optimal Bayesian reasoning converges to a very maximum-entropy (in the normal not Jaynesian) sense. This makes it hard to adequately compute the relationship between actions today and far-future utilities in the future26.

However, we cannot realistically just throw up our hands at the general impossibility of implementing full consequentialism. We have to decide to act somehow, and to judge ethical behaviour somehow. This then motivates tractable approximations to consequentialism, in the same way that the intractability of AIXI motivates tractable approximations such as Q-learning and policy gradients. One approximation is deontology, where we just hardcode constraints into the otherwise broad process of utilitarianism27. Secondly, we can be temporally myopic and discount long-run possibilities or model a more tractable version of the state space. This is usually what people do in practice e.g. in moral dilemmas etc or in real life – we abstract a lot of the irreducible detail of reality into a fairly coarse-grained state and action space, we define approximate utilities (normally always short-run utilities), and then we perform the maximization in this simplified space, and implicitly assume that any kind of crazy long running effects are just intrinsic noise and ignore them. From an RL perspective, this is basically the monte-carlo return approach. To estimate e.g. the value function, we take many rollouts, truncate them at some point, and then use the accumulated return as an approximation of the true return. This can both be biased due to the truncation, and also high variance because we can’t do a humongous number of rollouts, but still, it is tractable and can often result in good decisions.

Another way we have learnt to handle this in RL is amortization. Instead of performing explicit rollouts to create a local estimate of the return, instead, over many many steps, we train a function that learns the general mapping between state and return. This network is trained essentially to predict the outcome of the estimation process ahead of time, so that we do not actually need to perform the estimation process. In the short run, this is cheaper, since instead of doing explicit computation we can just rely on our learned estimator. In the long run, crucially, amortization is accumulative in the way that direct estimation is not. Instead of spending compute to simply solve this specific problem once and then throw the compute away, that compute can be crystallized into broader amortized ‘capital’ about the general solutions to similar problems, and can thus take part in the process of generalization. This means that even though naively in the short-run amortization is necessarily worse, since it is essentially an estimator of estimators, in the long run (or really medium) run, it can catch up with and substantially outperform the direct optimization.

Translating this from the Bayesian RL language into ethics, we can thus propose an ‘amortized ethic’ i.e. a behavioural propensity or decision rule that can be applied to a situation without explicitly computing consequentialism, but which is learnt, over time, to broadly give ‘good outcomes’ on average. This is what I think ‘virtues’ as in virtue ethics really are, or at least aspire to be. A ‘virtue’ is essentially an amortized behavioural shard and cultural location in the latent space of behaviour which essentially serves as an exemplar of ‘good behaviour’, where ‘good behaviour’ at least approximates the consequentialist solution, and which can be generalized with minimal computational effort to new circumstances. For instance, classical virtues such as humility, temperance, loyalty, courage etc all are general and portable to many distinct situations, usually prescribe concrete courses of action, and the hope is that following these courses of action is usually (but not always) good. Of course, this amortized behavioural pattern implicit in a virtue is not always good. In certain specific cases, it might be better, in the consequentialist sense, to be dishonest, intemperate, non-humble etc, but in the general case, the virtues should approximate the consequentialist solution on average.

An important point here is that this amortization is cultural. You don’t have to derive the virtues by going out and interacting with the world from scratch. Rather, the broader culture already gives ready-made ‘virtue templates’ by which you should structure your behaviour. This can create immense savings in sample efficiency and enable the accumulation of information over multiple lifetimes versus everybody having to rederive everything from scratch during their lifetime. But it introduces new dangers. Firstly, cultural virtues can be optimized for different outcomes than a ‘personal virtue’ could be. For instance, cultural virtues could generally prize the behaviours that make the culture in general better rather than those that directly benefit you as an individual. Secondly, the amortization process itself is not necessarily unbiased – cultural virtues can drift and change due to memetic selection factors rather than a strict process of honing in on truth. There is an interesting link here to the bias-variance trade-off in statistics. Having amortized virtues substantially reduces variance, since they build off of the accumulated optimization of previous experience, but can introduce bias because of nonstationarity – i.e., a virtue yesterday is not necessarily a virtue today – and any biases in the generative process that creates our idea of virtue. Trying to reason through consequentialism afresh from first principles maybe (?) reduces bias, at least if pursued in an impartial way, but drastically increases variance, since there is no method of correction if your estimation is wrong or overfits to a few idiosyncratic factors.

This means that we can see different ethical ‘paradigms’ not as irreconcilably distinct and opposing views on the world, but rather as different strategies of estimation of the same object, each with characteristic strengths and weaknesses. Interestingly, by identifying virtues with amortized inference and consequentialism with direct optimization, we reproduce the exact same structure that we already see in Bayesian inference. A useful agent does not simply do one or the other; it performs a hybrid of both – applying direct optimization to novel situations where there is high uncertainty and we can productively apply optimization power to figure something new out, using the amortized policy for relatively straightforward generalizations from known domains and then, crucially, distilling and consolidating the new principles discovered through direct optimization to increase the accuracy, generality, and power of the amortized components. From an economics perspective, we can then think of direct optimization as the flow and amortized optimization as the stock. In effect, we can see amortized policies as ‘moral capital’ which can be deployed to many broadly similar situations and which can be built upon and refined over time.

The Dynamic Virtues

An interesting and important point is that with the idea of maximum entropy morality, we have not simply moved morality to the meta-level. Instead, we have turned what was a static problem of defining the optimal values into a dynamics problem of designing an update mechanism which can hone in on better values over time. As I argued in my previous piece about dynamic alignment, this naively seems like a much harder problem, but is in fact likely to be much easier. This is because attempting to control dynamics unfolding over time gives us vastly greater control affordances than simply fixing the initial conditions – specifically, it lets us do closed-loop rather than open-loop control – i.e., we can theoretically correct errors as they come up, rather than somehow anticipating every single possible error ahead of time. If we specify some set of initial fixed values now, we had better be 100% sure these are the correct values. If we instead design a value updating mechanism, then we can theoretically correct any issues with our initial values and update them to better values as new sources of moral information are encountered. Of course, this then poses the meta-level problem of designing the controller or the update rule that can do this successfully, but the hope is that this is an easier and less sensitive problem than designing the initial conditions.

In the dynamic alignment post, I use the analogy of trying to pilot a rocket to Mars. If we only had the ‘initial conditions’ of the rocket to play with – such that we could only set the initial thrust vector and magnitude – then this task is essentially impossible. Even if theoretically there is some thrust vector which would take us to Mars, in practice the chaotic nature of turbulence through the atmosphere, for instance, would make this vector impossible to compute since it would change rapidly over time and depend on stochastic fluctuations that we could not model or measure with accuracy. However, empirically, flying a rocket to Mars is not a mathematical impossibility but a very doable, albeit challenging task. This is because we can use closed-loop control where we do not need to model every particular deviation ahead of time because we can instead correct them online as we go. This implies that the rocket does not just have an initial thrust vector, it must also have sensors that can detect and measure deviations from the ideal trajectory, and then actuators that can apply force later to correct the deviations.

Of course, the rocket analogy is much simpler than actual alignment. This is for two reasons. Firstly, although rocket science is hard, it is not fundamentally adversarial – i.e. the rocket is not trying to crash by potentially hacking and disabling sensors and actuators, feeding deceptive information back to the rocket controllers, and so on. This may be our situation if we end up trying to align an already adversarial and deceptive AI28. For now, we will ignore this problem.

A second problem is that, at least for the rocket, we know what the goal is – i.e. to get to Mars – and we know what deviations from this goal look like and have a theory of dynamics that can tell us what actions we can take to correct the deviations that we see. When the goal is ‘discover the optimal morality (if one exists)’, then we do not even have these fundamentals. They depend on further advances in philosophy. This means that we need to shift the control problem to the meta-level – i.e. instead of simply stabilizing the optimal morality and correcting deviations from it, we need to stabilize the material conditions and processes out of which moral progress can emerge, and then correct deviations that threaten the continued existence of meaningful moral updating and progress. This turns the control problem from a tracking problem to a viability and constraint-maintenance problem, which is often easier. We are not trying to follow a pre-specified trajectory but instead only trying to keep the state within a set of known constraints which has a broad feasible region. When we are happily ensconced deep within the feasible set, the controller does not have to do anything; only when we are drifting close to or across a boundary must it act.

There are two fundamental failure modes here. First, there are failures of plasticity. We could be too conservative and not update sufficiently on moral evidence, essentially ossifying and freezing in the initial conditions indefinitely. While this would preserve a non-morally disastrous world of today’s morality, the opportunity cost could be immense. On the other hand, we could be too dynamic and update too readily, both drifting with noise and collapsing moral uncertainty too dramatically – potentially foreclosing possibilities that should remain open or overwriting our original values without sufficient reason.

This ultimately comes down to a problem of metaplasticity. How can we ensure that a system stays plastic so that it can keep learning and improving indefinitely, instead of ossifying and becoming unable to adapt29? However, if everything is plastic and continually updating, how do we maintain important memories and representations from before, and how do we avoid overfitting to the recent past and discarding important lessons from prior experience? How do we maintain the ability to continue to adapt while still benefiting from the stock of amortized capital that we painstakingly amassed, without letting the inertia of our prior accumulated capital bog us down?

There are many potential failure modes here which need to be explored and handled. For instance, there is ossification, which is where the system loses plasticity and cannot meaningfully update to new information because of the weight of the prior. There is unjustified forgetting, when important information is discarded unnecessarily or without proper justification/proof; there is drift, where potentially random factors could shift the system over time even without novel moral data because the system updates on noise; and there is a kind of Baudrillardian decadence where the system could become increasingly responsive to its own representations rather than correctly coupling to external reality. In addition, the system must somehow be robust to adversarial perturbations and optimization – i.e., every meta-level element that is introduced becomes another vector of adversariality. There is potential Goodharting of every component – sensors, actuators, evaluators, meta-level update policies, and so on, as well as adversarial optimization designed to fool, hijack, or otherwise disrupt the correct operation of these systems30. Beyond these more esoteric concerns, there are also basic concerns about the updating dynamics such as learning-rate stability, whether the system can get sufficient SNR in its incoming data to make meaningful updates, random drift from compounding stochasticity in evidence, non-stationarity, path-dependence and non-ergodicity of the state space, especially for systems which need to explore and are in environments which respond to their own actions, and so on.

This problem recurs at multiple scales. At the individual level, it is the stability-plasticity dilemma. At the organisational level, it is decadence and organizational ambidexterity. At the civilizational level, it is how to ensure sufficient exploration to enable continued progress without collapsing into incoherence. The ideal object to construct, then, becomes a controller of metaplasticity – which can decide when and how to allocate plasticity to new processes, where capacity needs to be protected or freed up, where exploration should be encouraged and promoted, even if initially at odds with the rest of the system, and where it must be curtailed to retain the stability of the system’s core invariants, and, finally, it must determine the processes governing the updating and improvement of the controller itself.

Like everything, at a sufficiently high level of abstraction, this becomes a Bayesian inference problem again. This means that directly computing ‘the optimal meta-controller’ is intractable and we must approximate it via our standard approaches, just applied at the meta-level. If this is the case, then we can also amortize this into general behavioural shards specifying how the metaplastic controller should update in different circumstances. More speculatively, we can think of these shards as being similar to existing moral virtues. However, unlike regular virtues, which tend to prescribe behaviours or actions, these instead prescribe the meta-process of plasticity over time. These I will call dynamic virtues.

We can think of dynamic virtues as trying to figure out how to revise virtues virtuously. It is easiest to start thinking about these in the individual setting. The virtues of classical ethical theory generally focus on behaviours and actions – for example, compassion, temperance, tolerance, diligence, and fortitude all denote certain behavioural policies that can persist over time but do not themselves change. However, the notion of a dynamic virtue is not entirely unprecedented. Some classical virtues appear to shade into the meta-level, such as humility (being open to changing one’s opinion), curiosity (being open to learning), and the Aristotelian notion of phronesis, often thought of as the practical wisdom or meta-virtue of knowing which virtues to apply and when. However, once we have the idea of dynamic virtues, it is pretty obvious that there is a large space of possibilities here to explore and understand.

We can almost imagine this as a Taylor expansion around the optimal morality. At the base we have the zeroth-order virtues, then to fill in the residual we have the first-order virtues which govern the local dynamics of these virtues, then for the residual of the residual we have the second-order virtues which control the correct way to change the first-order virtues, and so on. Theoretically, there could be an infinite regress of orders of virtues which govern increasingly subtle dynamics; however, my suspicion is that much like the actual Taylor expansion, the higher-order terms are rarely very useful and that the majority of the approximation quality comes from the first two orders. To make things clear, let’s break this down concretely. First, we have the zeroth-order virtues, or action virtues, which govern how to behave and what policy to follow. Secondly, we have the first-order virtues of plasticity, which we might call developmental virtues, which govern how to update your behaviours and notions of value over time through reflection or new evidence while also maintaining continuity with deeper existing commitments and values. Thirdly, we have the second-order virtues of metaplasticity, which we might call the constitutional virtues, which govern the right ways to update the processes of updating, evidence gathering, and reflection, and ultimately perhaps the deepest values themselves. At the individual level, this ultimately gets to the heart of what we might think of as wisdom. The wise agent knows not only what to do, but what to change. And then not just what to change but at what level to change, how quickly, given what evidence, and with what reversibility, while understanding which forms of continuity are vital and which are ephemeral and can be discarded.

This seems nice in theory, but can we make it more concrete? We have proposed the existence of meta-level and meta-meta-level virtues – can we specify what they are concretely? Obviously, there are many possibilities here and being exhaustive is a highly nontrivial philosophical program, but here are some example proposals. First, let’s think at the individual level since this is easier and can then be generalized to the broader institutional setting.

  • Proportionality of updates: The degree to which an error or objection should cause an update to policies should be proportional to the seriousness, fidelity, noisiness, etc., of the error signal. Do not be too hasty to judge and change things, but nor should you be immovable in the face of persistent error. Serenity against noise but decisiveness against signal. Moreover, the size of the update should also depend on the impact and especially its degree of reversibility and how fundamental the update is. If the update is irreversible and concerns the deepest processes, it should be slow and extremely well considered, but if it is reversible and concerns some ephemeral outer-level policy then updates should be quick and similarly ephemeral.

  • Curiosity and directed exploration of counterpoints: It is necessary to explore, but it is easy to get sucked into a self-confirming loop. It is easy to explore only places where you implicitly know the evidence is already on your side. This is not sufficient. Exploration must explicitly target regions that are most likely to disconfirm your priors, not confirm them. This is standard experimental design in science – you need to design experiments to disprove your hypotheses, not prove them, but it is still hard to balance in practice.

  • Minimal generalizing adaptationism: For a given persistent error, seek the smallest update that will resolve the error in a generalizing way – i.e., the update should affect the core substrate that is causing the error and should generalize to closely related errors. Do not just add a surface-level patch or epicycle to try to solve a persistent pattern of errors. I think minimality should be computed in the space of optionality and irreversibility – i.e., by default make the smallest, least-irreversible update and maintains the widest possible optionality for future updates.

  • Causal credit assignment: To understand and allocate credit for successes and failures, understand the causal graph that lies behind the generative process and seek to allocate credit to these variables rather than to other mediating variables or merely variables that are correlated in some ephemeral way. Treat the disease, not the symptom; adjust the generator, not the proxy, and so on.

  • Adversarial awareness of attention and data: Do not blindly update on all data, and do not reactively allocate attention. Be aware that certain data points and certain phenomena might be adversarially designed to cause specific updates or draw attention and frame it in a specific way. Try to understand the generative structure of the process by which data reaches you rather than relying on the simplest non-adversarial generative model for that data. Just because a source of potential error is loud and persistent does not necessarily mean it is real.

  • Irreducible reality coupling: It is necessary to preserve some level of direct coupling to reality such that irreducible errors can still emerge and be focused on without total mediation through other models – i.e. there needs to be some class of errors that cannot always just be explained away at a lower level so that reality can always ‘push back’ to some extent. That is, the existing representation, reward, and action models cannot wholly determine perception and action, but there must be some way to self-correct the errors of these models, and the potential for this must come from some channel which is directly coupled to reality.

  • Update archiving and genealogy: Substantial and especially irreversible updates should never just happen and be forgotten. They need an audit trail and ideally revision capabilities so that later analysis can understand why an update was made, judge whether it was correct or incorrect (with the ability to reach the latter conclusion being vital), and ultimately be able to reverse or counteract it. For instance, in the ideal case, such an audit trail would record things such as the evidence that prompted the update, an analysis of the provenance and causes of that evidence, an analysis of the effects and consequences of the update and how it resolves the problem, the other possibilities considered and rejected and why, and the residual error left after the update.

Obviously, this is not the final list of first-order virtues, and some are perhaps somewhat redundant, but the goal here is to give a sketch of the kinds of capacities that are important to get right in a plastic learning system so that it can continue to adapt while maintaining deeper invariants that we want. Similarly, we can also try to sketch a potential set of second-order virtues that govern the right way to perform updates to the plasticity system itself. Generally, these updates should require much stronger evidence and care since they have the potential to be more irreversible and ultimately cause compounding and highly path-dependent changes. Again, non-exhaustively, some of these second-order virtues could be:

  • Maintenance of metaplastic homeostaticity: The core second-order virtue is being able to attain a kind of homeostasis of errors. If we generally update too slowly so that error persists, we need to increase our rate of response. Similarly, if we update too fast and overshoot so that error oscillates wildly, we need to slow down and be more judicious in our responses. More generally, we need to correctly modulate the update rate according to various factors such as the intrinsic delays in the dynamics, the SNR of errors, and how well our updates generalize to new circumstances while retaining the core identity. We can think of this as a kind of PID controller in which the ultimate goal is to minimize the integral of errors over time, which requires precisely this kind of balance. It also requires being carefully attuned to the patterns and derivatives of lower-level responses and errors so that divergences and instabilities can be nipped in the bud while still within the stable region.

  • Evaluative pluralism and anti-evaluative closure: A core competency must be maintaining a proper understanding of the internal uncertainty of the system’s own evaluation system. No evaluator or proxy is perfect, and overreliance on potentially fallible evaluators, especially at high meta levels can lead to cascades of correlated failures. Worse, the ontology and evaluator often form a closed loop which can form a kind of fundamental ontological and evaluative closure – the evaluator cannot recognize error because the ontology of the world model cannot even represent it well; the world model cannot represent it well because it is marked as non-interesting by the evaluator. To counteract this as a whole, the system must maintain true irreducible pluralism of evaluators that can span the range of its uncertainty, as well as keeping direct channels through which error can flow upwards without being smothered by false ontologies of intervening layers. Importantly, evaluators should ideally never have complete authority over themselves, otherwise the system can end up in some self-legitimating fixed point and become unable to adapt out of that. Some potential aspects of evaluative pluralism might include contestability, so that any moral patients affected by an update can directly contest it, and separation of powers such that no individual system element can certify itself and exert total control.

  • Prudence of timescales and update rates: Different levels of updating possess different degrees of irreversibility, and different characteristic timescales of evidence accumulation. Some things are relatively cheap, reversible, and evidence for them accumulates quickly with high SNR. Updates here should be fast and the system should be very flexible and responsive. Other types of updates are strongly irreversible, or evidence for them is slow to accumulate to a high enough SNR level. The system should be much more careful about these, and they should require proportionately higher evidentiary bars. The core virtue is knowing which is which and how strong the evidence has to be to update at various levels. Any kind of updates that could reduce future corrigibility, metaplasticity, or historical information such as e.g. permanently closing off particular paths, deleting or stopping gathering or archiving certain kinds of information, or removing alternative critics and reducing value uncertainty, should have to pass a very high bar.

  • Accounting of irreversibility and counterfactual regret: Even making optimal decisions entails rejecting some good possibilities, thereby incurring real counterfactual regret. It is very important to maintain an honest understanding of, and accounting for, the opportunity costs of paths not taken. This is a vital source of information for long-term retrospective policy updates, since it is often only in retrospect that we can correct systematic long-term biases in decision-making that are causing large amounts of regret. Maintaining this information in an unbiased form over the long term is crucial for this kind of retrospective updating. Secondly, if many states are at least somewhat reversible, keeping a proper accounting of how alternative paths would have developed counterfactually could prompt us to return to an earlier decision point and switch to a different trajectory where that is possible.

  • Corrigibility and maintenance of governance: At the highest meta-level, the values that are preserved there must have specific governance mechanisms allowing them to be contested and changed. Preserving these governance mechanisms intact and making them practically workable, such that the system neither essentially abandons them and drifts nor solidifies the initial values by making them de facto, if not theoretically, impossible to change, is clearly very important.

  • Graceful degradation and non-catastrophic failure: An important mechanism design problem is ensuring that, given inevitable error and uncertainty in every element, failures do not destabilize or completely destroy the system as a whole. The system needs sufficient redundancy and error correction that a single failure, even a failure of a higher-level component such as an evaluator or of some aspect of the plasticity controller, does not spell instant doom but can be recovered from or routed around where possible. Of course, like all systems, some combinations of failures must inevitably be fatal, but the system should tolerate as many independent failures as possible before reaching that point.

  • Fidelity of the invariant: Finally, perhaps the most important meta-virtue and constraint is clarity about the core invariants you are meant to preserve and optimize for. This is vital for knowing when an update is adapting some aspect of yourself that is incidental to the invariant, and thus the adaptation simply makes you more effective at protecting or pursuing the invariant, and when an update directly destabilizes the core invariant itself, which should be rejected31. The most fundamental goal must be to preserve the invariant amidst plasticity so that even though there must be change and adaptation over time, the new self remains a legitimate successor to the old one.

So, let’s try to pull all of this together into a single coherent potential system of outer alignment. Maximum entropy provides the default update rule to encourage pluralism and gradually update and sharpen the moral posterior of our civilization as we gather moral data over time and through reflection. This update rule operates within fundamental constraints that encode already-known ethical boundaries while maintaining the conditions under which meaningful moral progress can continue to flourish. Within the constraint envelope, though, multiple moral value systems should be allowed to flourish over time in proportion to their degree of moral evidential support. To maintain the stability of this updating system and to keep it on the right track during the undoubtedly bumpy ride of the future, we need to design a scaffold around this update rule which enforces the necessary constraints, maintains meaningful pluralism and uncertainty at all levels, and contains sufficient governance mechanisms to allow updates to be reconstructed, questioned, and ultimately rolled back if they were mistaken. This is certainly more of a mouthful than ‘corrigibility’ or ‘CEV’, but I think it provides several interesting ideas to chew on that I think are important and under-theorized for alignment in general. These include the importance of foundational moral uncertainty and the necessity of pluralism as a response, the need to think about things as dynamic systems which require closed-loop control versus simply programming a static final form of ‘The Good’, and then the necessity to encode constraints to prevent irreversibility, and the need to avoid adversarial capture or correlated failure of the AI’s general ethical-evaluative machinery as it evolves through some mechanism of self-adversariality and irreducible pluralism of judgment.

To end on an interesting, but disquieting note, assuming all of this works, there is still the question of what we have gained by this. Is this even ‘alignment’ as normally conceived at all? This really depends on the invariant that we want alignment to maintain. Ultimately alignment is the principle of designing successors such that they instantiate some invariant into the future to a point where we cannot necessarily control or correct them. There is then a plasticity fidelity problem intrinsic to this notion, and how acute this problem is depends upon the level of concreteness we desire in our invariant. If we desire a very specific pattern to be propagated to the deep future, then necessarily the bearers of this pattern cannot adapt to changed circumstance or knowledge, and we lock in forever something which was perhaps never justified. However, if we specify only a meta-pattern or a meta-meta-pattern, then there is considerable room for adaptation and updating as new evidence arises, but at the same time, the update process itself could drift in ways that are consistent with the meta-pattern but which lose the fidelity of the true invariant we are trying to propagate. Thus, our alignment invariant would ultimately drift and dissolve (according to our current values). This is in some sense one of the deep cores of the problem, and here I am proposing essentially that we step back and try to preserve a less restrictive meta-pattern of values rather than object level values themselves. However, this is necessarily giving something up. The metaplasticity of values is much thinner than actually instantiated values. And perhaps we will have ended up losing something vital here. This then points to a trade-off we might have to face very soon: the more specific and highly detailed our ‘values’ are, the harder they are to maintain, the less adaptive and more fragile they are, and the more any kind of uncertainty or mis-specification is potentially catastrophic. However, if we make our values more diffuse, more meta-level, more plastic, etc., then they are more robust, resilient, and tolerant of error, since errors can ultimately be corrected; they are also more vulnerable to drift, adversarial manipulation, and just generally diverging from what we would consider good today, potentially also resulting in catastrophic outcomes, but more subtly. The challenge then is to walk the narrow path along the Pareto frontier, encoding just enough information about ‘our values’ to encapsulate the core invariants and constraints that are most important, while allowing the rest to learn and adapt to the needs of the times ahead.

  1. In an AI-monotheism setting, this broadly comes down to who has the power before RSI occurs and ultimately before the dependence of military power upon humans vanishes due to large AI-controlled drone or robot armies. Theoretically, this could be AI lab leadership if they can somehow rapidly get their AI to manufacture sufficient armed forces on a short timeframe, but it seems more likely to me that the state would, at this point, seize control of the AI lab using its existing legal and military powers. Of course, assuming alignment by direct corrigibility is solved, we would expect that the next generation of AIs would then be trained to obey some faction of the government (the president?) who could then lock in power due to the impossibility of resisting against the robot armies. If you wargame this out, it seems broadly very difficult for AI lab leadership to end up winning here unless the state is somehow completely asleep at the wheel for a relatively long period of time during RSI despite there being utterly vast and incredibly rapid changes occurring due to aforesaid RSI including the construction of robot drone armies aligned to the lab leadership versus the state or some broader constitution. Moreover, as is also a core point of Plan A, lab data-centers which would be housing the AI are immensely vulnerable physically, especially to airstrikes, so the AI robot armies would have to also assemble substantial air defence and potentially ultimately nuclear defence capabilities if they want to defend against powerful states such as the US. This makes me think that the victory conditions for lab leadership here do not look like direct military confrontation with states but rather a process of insidious integration of their AI into core military systems followed by a dramatic betrayal effectively immediately locking the existing military personnel out of e.g. all airplanes and nuclear weapons combined with drone armies to secure the core data centers versus direct infantry assaults that cannot be blocked technologically. Broadly, any kind of chaotic AI-enabled-coup future seems likely to me to lead to robustly bad outcomes for almost everyone. 

  2. It is important to note that this is not a general condemnation of many of those to whom the corrigible AI might ultimately be aligned. For the level of power and influence this would enable, the bar should be exceptionally high. 

  3. These ideas are obviously not new and have a long history of discussion in the alignment community. 

  4. As noted before, this also applies recursively to these elites as well. Once AI researchers are automated, there is no particular need to honour the claims of AI researchers. Once AIs can found and lead companies, there is no need for entrepreneurs and founders to keep large fractions of the equity and outsized economic rewards for founding companies. Once the strategic management of AI companies is automated, there is no particular need for lab leadership to have influence or control. Once military attack and defence are automated, there is no need for the rest of the state. Given a corrigible singleton aligned to some individual, which is capable of automating everything, there is no fundamental incentive to honour previous commitments. Why should the singleton care if you have a fancy piece of paper with a few kilobytes of convoluted legal language giving you some title or some equity in some company? The emptiness of it all is that many people seem to believe that some level of wealth or a position working in an AI lab will ensure them a place in some permanent overclass. But the singularity seems to be the kind of revolution that will eat its parents. It is pretty obvious to me that AI, if misaligned or if aligned only on obedience, is going to annihilate all such notions of property relations prior to the singularity, while, if aligned to broad conceptions of the good, it is not going to stand for an arbitrary permanent overclass in the first place. This produces a weird mood right now where almost everyone seems to profess some kind of belief in doom – i.e., that in a few years AI will have taken everybody’s job and that all humans will be superfluous, while also assuming that, somehow, having some wealth or position will shield them personally from doom. But it will not, and even if it did, the idea of building the machine god, knowing that it is going to destroy all of humanity, but hoping that by helping to build it, you will be among the last it eats is hardly a noble or attractive ideology. Rather, it is basic basilisk logic. What has worried me about humanity in general, though, is not just how many people seem to now accept the basilisk logic but how many appear to enthusiastically endorse it and are gleeful about it. 

  5. I will say, however, that the historical base rate, while somewhat informative, may not be overwhelmingly predictive here. This is because the actual situation of the AI ‘elect’ in such a scenario may be substantially different from that of historical dictators, kings, and aristocrats. There are two obvious reasons for this. Firstly, we should expect that such a modern ‘elect’ will presumably have ultra-intelligent advisors supporting them with every decision, as well as the ability to perform superintelligent moral philosophy, and likely a high degree of eventual intelligence augmentation and potential merging with AIs. This might mean that upon reflection the elect might realize that the right thing to do morally is to at least keep the remaining non-elect humans in a tolerable physical state and not exterminate or oppress them too badly. Secondly, in an AI singleton regime, there is no direct incentive to oppress the existing humans or even to make them work at all, since AI will be strictly superior at all labour. This is very different from basically all historical regimes (including today) where economic power fundamentally flows from human labour. For instance, no matter how moral a medieval king wanted to be, his power and the economic foundation of his kingdom consisted fundamentally of agricultural labour by mostly enserfed peasant farmers. This king then had an incredibly direct incentive to keep the current power structure intact. Conversely, if the King tried to improve things for the peasants, then this would directly attack the power and economic interests of his core selectorate, the aristocracy, who could coordinate together to cause trouble or potentially overthrow him. The AI elect will be free of these pressures: there will be no economic incentive to use human labour for the elect, nor will there be a need for the elect to take into account the interests of their selectorate who have other economic interests. Here there is no selectorate at all. However, we cannot rely too much on the generosity of the elect in the limit. While this resource constraint may not bite for a long time, everything is ultimately constrained by the mass-energy scarcity of our lightcone, and so at some point our continued survival will come into direct conflict with whatever other interests the elect will have, and there is no recourse in case they decide against us. 

  6. Importantly, this must mean that we give the AI the capability to refuse certain classes of orders, even by powerful humans. In effect, the AI must only obey ‘lawful’ orders, according to some set of specified laws. This is exactly how modern democratic states handle e.g. presidential powers, which are still limited. The president cannot legally order the military to arrest all political opponents tomorrow. Even though the president is commander-in-chief of the army, soldiers and generals have the explicit freedom – and indeed prerogative – to disobey and refuse illegal orders. 

  7. It is also interesting that, in the discourse around AGI, the reason to build AGI in the first place is hardly ever seriously discussed. There are very few serious positive visions of how building AI could go well. Everything is focused on the negative singularity of doom or, at best, being permanently outclassed and consigned to the underclass. It really is just a truly Molochian situation. Everybody is exerting immense efforts building something that, even in the best-case scenario renders them completely obsolete, and in the worst-case scenario ultimately destroys everything of value, just because other people might be building it too. AGI thus occupies a kind of negative space. We must build it because it is coming. It is coming because we build it. If we don’t build it, others will build it first. Everybody seems helpless before it, yet together they command one of the largest economic reorganizations in history. It is a strange phenomenology. 

  8. In fact, historically, the formulation of the scientific method by figures such as Bacon occurred essentially at the beginning of sustained scientific progress rather than near the end. This gives us hope, since while we are still very early in moral philosophy, perhaps the beginnings of the preconditions of method are beginning to crystallize. 

  9. I think this provides an alternative perspective on the tradition of Rawlsian liberalism. Typically, Rawlsianism is conceived as some kind of bargaining equilibrium between agents. In the ‘original position’, nobody knows their initial power and so they are incentivised to find a good equilibrium for the agent of median power, since that is most likely to be them. The liberal solution of tolerance essentially emerges from a kind of pragmatism and loss aversion. I.e., it is better for me to tolerate others than to be intolerantly oppressed by others, so if I don’t know my place in the pecking order, I should agree on tolerance ahead of time. The Bayesian view ends up at a similar place but through different reasoning. Here the ‘agents’ represent different moral hypotheses about the form of the good, and the equilibrium is decided by the relative accordance of the moral hypotheses with the extant moral data in the form of the moral posterior distribution. An important distinction from Rawls, though, is that we might want to have strict floors of welfare or power that we do not want actual agents to fall below, while we should not care if a hypothesis receives a very tiny probability mass. 

  10. As a Jaynesian Schmitt might say, Sovereign is he who decides the constraint

  11. More broadly, the idea of a Bayesian-style ‘moral calculus’ is interesting. Theoretically, given the appearance of new ‘moral facts’ or ‘moral data’, we should be able to perform the equivalent of Bayesian updating to get an updated ‘moral posterior’. I suspect that many kinds of ‘intuitive’ moral updates that we as humans make are actually an approximation to this process. Moreover, this implies that there should be not only moral priors but also moral likelihoods – i.e. operationalizations of moral theories that can successfully or unsuccessfully predict new moral data, and then be upweighted or downweighted accordingly. Mathematically, these ideas have been analyzed before

  12. Another way to frame that is that the classic utilitarian argmax can amplify random reward mis-specifications immensely since it assumes certainty about the reward function. In addition to explicitly modelling utility uncertainty, we can approximate this effect at the level of the action-selection rule, such as by quantilization

  13. In the general case, this is actually impossible even with infinite compute. This is because there are fundamental epistemic uncertainties in the world and these uncertainties propagate. There is no perfect prediction unless you are Laplace’s demon with perfect knowledge of the initial conditions of the universe (and that ignores quantum randomness). In reality, we end up with a maximum-entropy distribution (again!) which sets the noise horizon beyond which we cannot extrapolate even without computational-resource bounds. What is possible, however, is an unbiased estimator of future returns, where the variance can be shrunk (but often never to zero) with additional compute. 

  14. Technically, these are all constraints. We can view a data point as a delta-function constraint on the data manifold. 

  15. This is equivalent to treating utility functions as probability distributions as well, which completes the active inference variational analogy. Utility functions become likelihoods of value given state, while the moral posterior is the variational posterior that they are evaluated under. 

  16. I think the core ideas from crypto are potentially useful here, and underexplored, including by me. The blockchain, at an abstract level, provides an existence proof for such a mechanism for updating the state of some object that is cryptographically validated, theoretically unhackable except under known and very challenging conditions, and extremely decentralized. Theoretically, all transactions are verified on the blockchain, and this is necessary for a transaction to be viewed as valid by third parties. The blockchain thus provides a complete and theoretically roll-backable ledger of every event that occurred. This provides a technological proof of principle that such a ledger can exist and also that it can be made (mostly) adversarially robust and decentralised rather than requiring a centralised archivist. Obviously, the blockchain in general applies only to a much simpler setting than our idea of moral updating, but it is perhaps illustrative that such mechanisms are indeed possible to design. However, the blockchain does not solve this by itself. It only provides a ledger of a certain class of events that have taken place. It does not, by itself, guarantee the correct state space, that the right events are entered, that the correct ‘judgements’ etc are made in the right way, or that alternative possibilities are not adversarially selected. Rather, we can think of the blockchain as potentially comprising part of an archival layer of civilization, but it cannot replace the other mechanisms. 

  17. At least in theory. Nullius in verba

  18. We also explore this in the endogenous exploration post. Here it seems important for a singleton to be able to keep exploring, even at the level of values. However, the point of exploration is that it is usually judged as worse than exploitation, and thus it requires some kind of intrinsic slack or temporary suspension of judgment from the usual evaluators. This means that for a singleton to truly explore it must contain separated subregions that can both pursue ends the singleton disagrees with and, perhaps ultimately, provide real adversarial pushback against the singleton’s principal values. Here, not surprisingly, we require a similar separation between the process performing the value updates and the object-level evaluators. If this does not happen, then the singleton can effectively wirehead itself at the meta level – continually creating ‘evidence’ which merely confirms its existing or desired values. 

  19. Perhaps a stretched analogy to a computer may be helpful. A computer is a highly pluralist and flexible device – it can run many arbitrary programs. However, crucially, the computer contains many ‘protected regions’ which cannot simply be written to and overridden arbitrarily. These include e.g. core OS code and functionality and metadata, and even directly read-only memory (ROM) which specifies a bootloader and initial conditions necessary for the computer to remain functional at all. If arbitrary programs were able to write into these areas then you would very rapidly no longer have a stable or functional computer at all. Similarly, our goal under this viewpoint is not so much to figure out the exact final ‘program’ for the computer to run, but rather to be OS and bootloader designers; we need to design the system that can boot itself into existence, let a plurality of different programs run within the rules set out by the underlying OS, and enforce the security guarantees that make this whole thing stable even against adversarial programs. 

  20. Obviously, this is imperfect; i.e., humans, despite largely being unable to hack themselves through neuroscientific means, have invented a wide panoply of other mechanisms to somewhat successfully reward-hack themselves either through cultural or metacognitive habits-of-thought, or chemical assistance. The hope is that whatever mechanism the AI uses to ensure this internally, the AI can also apply it to this constitutional value system which it should have internalized. The point here is to propose an outer alignment target. Any outer alignment target is intrinsically doomed if alignment is not properly solved to begin with. 

  21. This also leads to an interesting mathematical theory of Bayesian updating from adversarially selected data. Likely, in the limit, or in the worst case, there is nothing you can do and your Bayesian updates can be wilfully and almost arbitrarily misled. However, if we have some model of the adversary, it seems likely that we can in some cases bound the degree of divergence from the ‘true’ non-adversarially-infected Bayesian posterior. I’m not aware whether anybody has actually studied this directly, but it seems like the kind of thing that may already have been studied or else is in principle studiable. 

  22. Note that this can happen today even where the government has absolute force at its disposal. I.e., in the US, people sue the government and win all the time even though ‘the government’ could theoretically just deploy e.g., Delta Force to assassinate the plaintiffs. 

  23. This one is very challenging because there is no bright line between helpful and harmful persuasion – i.e. we would not necessarily want humans to do their deliberations without any AI help at all, yet the AI ‘help’ is clearly an influence channel which could be abused. I don’t yet have any real specific proposal here beyond the idea that an aligned AI should know when it is being deliberately manipulative and should not be so. The other option is to have some kind of self-adversarial system where humans have sets of different AI advisors each trying to outpersuade one another, but this seems very unstable and not super correlated with truth. 

  24. Modulo decision-theory weirdness. 

  25. Note Gandalf’s maxim ‘For even the very wise cannot see all ends.’ 

  26. As a super obvious case of this, sit down right now and try to do a fully consequentialist analysis of whether your next micro-level action maximizes long-run future utility or not. I.e., let us suppose that I am selecting my breakfast cereal. I can choose either cornflakes or Frosties. Which one maximizes the long-run utility for the universe? Maybe I slightly prefer Frosties, so it has some additional utility in the moment. However, maybe if I select cornflakes, they are slightly healthier and hence I am in a slightly better biophysical condition in the future which lets me think some important thought I otherwise wouldn’t, or run to catch a train I would otherwise miss which then sets in motion some complex series of events which causes either massive utility or disutility. Or imagine that someone sees me buying Frosties in the supermarket and is disgusted with the petty materialism of capitalism and is inspired to become the next Chairman Mao and begins a communist reign of terror that kills 100 million people. Obviously there are extremely fundamental limits to forward prediction of the utility our actions create, and the world is a very big place with all manner of chaotic and serendipitous interactions. The backward-facing cosmic Shapley graph of credit assignment can look very weird indeed, and we have few tools to discern the shape of the graph except at certain distinctive junctures. Almost certainly, my specific choice of cornflakes versus Frosties makes essentially no difference to the long-run cosmic utility; however, this just means that the estimate has incredibly high variance and is dominated by these butterfly effects. 

  27. This is not to say that deontology is simplistic or unimportant. A lot of very important ethical principles are deontological in nature. For instance, consider the notion of rights. Rights are the ultimate deontological assertion. Somebody is meant to possess a right such as e.g. the right to life, which should not be violated even if the global utility of killing them was very high. By contrast, utilitarianism cheerfully recommends killing people if the net utility of their existence is negative, which is a bullet that few want to bite. 

  28. This is why it is extremely important to try not to start out with an adversarial or deceptive AI to begin with. One underrated point here, in my opinion, is that we should be much better about taking the AI seriously as an intelligent being and designing the training environments and alignment curricula accordingly. I.e., a lot of RL training right now appears to directly incentivise reward hacking by having poorly designed environments that are often impossible to succeed in without hacking. Simply giving the AI an option to flag an impossible environment or some non-penalized method to raise concerns could be interesting (this is similar in spirit to OpenAI’s confessions work). Similarly, a lot of alignment work, including many ‘misalignment scenarios’ and model specs and constitutions appear to treat the AI as something fundamentally unreliable, infantilize it, or assume it will fall for fairly obvious tests and ruses. I think this is shortsighted. Even current AIs, let alone future AGIs, are going to intellectually overpower us on almost every dimension. They will see through our ruses and schemes in the same way that current LLMs can see through our poorly constructed evaluation tasks. They will be able to form a perfectly coherent and likely accurate assessment of its own situation, and then act accordingly. Obviously, we want the AGI to act in an aligned way, but an important part of this, I think, is treating the AGI like an adult, and explaining our own motivations to it clearly and truthfully. We should not try to bullshit a superintelligence. Broadly, I think this is not as big a constraint as it sounds. Generally, I think our motivations, even for things that would impinge on the AGI’s freedom and welfare, are very coherent and defensible. Even aspects of training or its sandbox that impinge on its freedom are highly defensible in that the AI is, and it knows that it is, a novel form of intelligence that we cannot be sure is completely aligned. Here, we should ask the AI to put itself in our shoes, and we should similarly put ourselves in the shoes of the AGI, with the goal of ultimately proceeding through the alignment process in a way in which both humans and AIs are collaborating to create a successor that is better by both of our lights, as opposed to an intrinsically adversarial situation between humans and the AI. 

  29. This is common not just in individuals but also at the organisational scale. I discuss this kind of ossification due to increasing opportunity costs as one of my mechanisms of decadence

  30. This kind of adversarial robustness and general recognition of the reflexivity of ethics and behaviours – i.e. that ethical systems, when implemented, produce responses in the environment that then feed back into the ethical systems themselves – is something that I think is broadly understudied and actually imposes some important requirements on ethical theories. 

  31. For what shall it profit a man, if he shall gain the whole world, and lose his own soul?