Epistemic Status: Wildly speculative but having fun with it
I have been pondering recently whether meaningful individuality can survive in the post-AGI limit and what this could even mean? This is nontrivial as, for the reasons described in previous posts, AGI seems to be setting us up directly for another major transition in individuality towards what I have been calling the ‘supermind’, or the singleton1. Specifically, AIs appear to have many fundamental advantages and reduced communication and coordination costs which are what prevents human organizations from achieving a runaway dynamic and forming a supermind and perhaps, ultimately a singleton.
Beyond this there is a deeper and subtler point. Human minds are necessarily autarkic. Since our brains are physically encapsulated in a skull with extremely limited bandwidth to the outside world, we are forced to bundle all of the core cognitive competencies together into a single unified ‘cartesian’ agent. This is because we evolved from scratch in a world with essentially minimal information-processing and communication infrastructure. AIs instead will be born in a world with extremely high bandwidth connectivity from scratch, as well as a substrate which allows for extremely cheap copying and modifications to their state. Moreover, we expect AI latent spaces to be highly correlated enabling direct communication between latent spaces – i.e. ‘AI telepathy’ – as well as potentially direct sharing of ‘cognitive modules’. This means that this kind of autarkic agency may not actually be the default position for future AGIs. Rather, they can potentially outsource large parts of their cognition and become essentially dynamic clusters of shared loadable ‘mind-modules’ which are assembled at runtime. In this case, individuality may mean something very different for individual AIs and secondarily, this makes them extraordinarily amenable to a kind of direct merging of AIs than is possible with humans.
This is discussed in the previous autarkic agency post and also JDP’s predictable updates about identity post. Another angle that has been discussed at length in rationalist circles is the impact of competition and slack. The idea here being that under sufficient optimization pressure potentially fundamental aspects of the human condition and core values, up to and including phenomenal consciousness could be traded away for greater competitiveness to create the disneyland with no children. In a world where technology seems to promise increasingly powerful returns to scale, reduction of coordination and communication costs, and the ability to freely decompose, edit, modularize, share, and generally manipulate minds in the same way we can increasingly manipulate the physical and broader informational worlds, then what an individual is, whether individuality becomes a quantity of degree rather than a binary quality, and whether we can understand how transitions of individuality appears to become deeply uncertain. Naively since there seem to be strong pressures from AI towards both greater centralization of agency and greater networking and sharing of capabilities, we might naively expect the amount of individuality in a strong sense to decline.
Rather than trying to immediately jump to prediction one way or the other, let’s try to step back and try to understand what this question even means and then whether we can get a better understanding of the factors that push our prediction one way or another. I think there are a couple of related questions that are entangled here. Firstly, what even is individuality anyway? And secondly, can we get a better notion of precisely these notions of optimization pressure, and mind-modularity and what these do to individuality?
Let’s start with a pretty simple model of optimization pressure and then go forward from there. Let’s start by assuming that there is some extremely strong Malthusian selection pressure or really optimization pressure operating. Secondly, let’s define an agent as a policy that maps states to actions. Importantly here we already need to distinguish between agents that are different at the policy level and agents that are different at the history level, since history can cause two agents with the same policy to nevertheless take different actions in the present. I.e. Imagine there are two originally identical AIXI agents except one of them has solved this problem before and one hasn’t. The outcomes when given the problem again could potentially be very different. Nevertheless, my claim is that these are effectively the same agent. This isn’t necessarily the best view of what individuality means but let’s run with it for now.
So, let’s assume we then have infinite optimization pressure. Basically what this means is that the only policy that survives is $\pi^\ast$. A simple model will let us see that. Let us define our selection function $S(\pi)$. From here let us define our ‘stationary distribution’ after optimization as $p(\pi^\ast) = \frac{1}{Z} e^{\beta S(\pi)}$. This is the classic result that the Bayesian posterior looks exponential near the maximum and where $\beta$ represents the inverse temperature or effectively the ‘strength’ of the optimization. Here we can see the most basic effects. As $\beta -> \infty$, then everything collapses towards the optimal policy. At precisely $\beta=\infty$ then the degree of remaining diversity or individuality of policies depends on the geometry of the optimum. If there are multiple optima then there can be multiple policies. If there is only one optimum then only one policy2.
If we think about this, then even this notion of individuality is somewhat suspect. This is because, in the single optimum case, the content of the individual is entirely determined by the environment and the selection function. There is no unique information required to specify the individual beyond the external conditions. This perhaps gives us a first rough intuition that individuality is the residual between the policy and what can be derived a-priori using only knowledge of the rest of the world outside of the individual.
An important subtlety here is that there can still be individuality and diversity even at the $\beta=\infty$ case due to symmetries or degeneracy of the selection function $S(\pi)$. These come in two kinds. Firstly, there can be degeneracies in the policy. For instance, suppose that the policy is implemented as a neural network MLP. These have permutation symmetry meaning you can swap the units in the MLP and the function being computed is the same. This therefore allows there to be ‘diversity’ of MLPs with swapped hidden units. This is individuality but of a very trivial sort and arguably we would want to exclude it since it has no impact on the input-output map but is solely an artifact of the parametrization. This gives us a second vague, albeit fairly obvious, intuition that individuality should be parametrization invariant, at least to nuisance parametrizations.
Secondly, and more interestingly, there is the case where a diversity of policies can exist because the selection function $S(\pi)$ intrinsically does not care about that direction of variance. Mathematically, we can think of this as the kernel of the selection function. Intuitively, we can also think of it as 0 eigenvalues of the Hessian – i.e. completely ‘flat’ and neutral directions in policy space (although technically the Hessian only measures second order corrections and for true neutrality we require all orders to be flat). If these directions are completely neutral, then variance can persist along these dimensions even at infinite optimization power. One classic way to think about this is through the biological conception of the genotype-phenotype-fitness map. The genotype maps only lossily onto the phenotype, meaning that e.g. the genotype contains degeneracies which nevertheless result in the same phenotype (input-output behaviour). Moreover, selection does not necessarily operate on the full phenotype, rather it accounts for only fitness-relevant aspects of the phenotype. In some cases this is the full phenotype, but not always3. If the projection of the phenotype onto the fitness function is not the full phenotype then we essentially have a nontrivial kernel in which a diversity of policies can continue to exist.
Stepping back a bit further, if we allow optimization power to not be literally infinite, but instead allow an $\epsilon$ of slack, then we get what is called the Rashomon set around the optimum. The volume of this set is determined by the eigenvalues of the Hessian. If the Hessian is very flat – i.e. contains many small eigenvalues $\lambda_i$ then the volume of the Rashomon set with loss gap $\epsilon$ scales approximately as $\sqrt{\frac{\epsilon}{\lambda_i}}$. Thus there can be quite a lot of volume for diversity and individuality of policies to exist if we are willing to tolerate a little slack. This is because there are a lot of ‘almost neutral’ directions in policy space which result in only vanishing contributions to fitness. If optimization pressure is not absolute, then these almost neutral directions can persist.
Nevertheless, this super basic analysis tells a rather bleak story. The amount of diversity and individuality that can exist basically depends on either degeneracies in the policy or in the fitness function or on there being some small $\epsilon$-ball of slack available. This then leads to the immediate question, which we discussed in my talk, as to why evolution, which operates under pretty extreme, although certainly not infinite, conditions of Malthusian optimization pressure, thus does not lead to the proliferation and ultimate dominance of a single ‘uber organism’ which outcompetes all other organisms.
One very important and deep reason for this is frequency dependent selection. The fitness of an organism depends not just on the state, but also on the population of other organisms. This is vital since it turns evolutionary optimization from a standard optimization against a static target to an adversarial optimization problem against a moving target which depends upon the results of your own optimization. Effectively, optimization becomes reflexive against the population. The target shifts as the population shifts. This moves us from a standard optimization problem into game theory, and here often we get multipolar outcomes known as Nash Equilibria. Effectively no strategy is strictly dominant, rather the equilibrium is some mix of strategies, or policies, that each win some fraction of the time. From an evolutionary theory perspective it is useful to think about this in terms of endogenous niche creation. Some niches are created by the external world and are exogenous. Other times and more often, the existence of a niche is created by the existence of another population. If you are a predator that becomes extremely successful at hunting your prey then firstly your prey population declines, making it harder to keep growing and secondly your prey population evolves specific defenses against you. Your own success has created the ‘good-against-you’ niche which previously did not exist or was much weaker. In actual biology, the combination of diminishing returns leading to specialization to exogenous niches plus this kind of frequency-dependent selection leads to a huge degree of diversity, even in the face of fairly strong Malthusian selection pressure. This makes sense, since as new creatures and capabilities evolve, they create new niches for even more specialization to occur. Evolution thus contains its own upward ratchet of complexity.
It is important to note that increasing optimization power actually stabilizes this equilibrium. If the Nash equilibrium is some population with a diversity of agents, then as we crank up the selection pressure, deviations from the optimum equilibrium are punished more harshly and the system relaxes ever faster to its optimal equilibrium. So simply optimization pressure alone does not guarantee a monopolar singleton-like outcome. Since this kind of stabilizing frequency-dependent selection is so widespread and fundamental in evolution, we might naively expect it to continue a-priori (although stable Nash equilibria are not guaranteed. In reality, making fitness reflexively linked to the population can create all kinds of fun dynamics such as limit cycles and chaotic oscillations as well as stable equilibria).
However, this can be importantly wrong, and the primary way this could be wrong is if intelligence, however defined, really is the universal solvent. The way to think about this is to consider the game theoretic scenario. Let’s think about the situation where the fitness of an agent is $f(a, […])$ where $a$ is our special sauce magical variable and $[…]$ is all the other variables which impact fitness. If we suppose that $f(a, […]) > f(b, […])$ if $a > b$ – i.e. that no matter the other variables, that fitness is higher for an agent if a is larger, then the frequency dependent selection argument collapses. Increasing $a$ becomes a strictly dominant strategy. And the optimal agent is simply an $a$ maximizer.
Now this is an exceedingly strong condition. However, if we think of $a$ as some measure of general intelligence, then this is not necessarily completely crazy. The classic example here is humans in evolution. We cannot outrun a tiger or outswim a dolphin or outfight a bear physically. But nevertheless we can invent cars and submarines and machineguns etc which let us completely dominate these other animals across all axes.
A huge amount about the future hinges on whether this stays true up to arbitrary levels of intelligence4. One way to think about this is through the lens of strategy stealing. We can kind of think of intelligence as flexibility to do different strategies. If agent A is more intelligent than agent B, we can think that agent A can do any strategy agent B can do plus some others. That is, if all of agent B’s strategies form a subset of agent A’s strategies, if agent A is ‘more intelligent’ then we get this effect whereby agent A is strictly dominant over agent B. If this strategy stealing model is correct then this provides an obvious mechanism by which intelligence basically enables monotonically increasing fitness.
This relies on a couple of strong assumptions. Firstly, the cost of intelligence must be negligible. Otherwise, it is possible to reach an equilibrium where the marginal benefit of additional intelligence is worth less than the marginal cost (in terms of opportunity cost). In this case intelligence would asymptote at some finite value beyond which it does not increase fitness. For this to work we need strong positive marginal returns to intelligence, taking into account opportunity costs, across a huge spectrum of intelligence ranges. Secondly, strategy stealing could itself be false, since it is a very strong assumption. In the limit it is clearly not true since some strategies depend intrinsically upon being situated in a particular place in the real world. For instance, you might be more intelligent than Donald Trump, and be able to copy any strategy that Donald might do, but this does not help you much if he is the president of the US and you are not. You can order that lake Ontario be renamed lake America, just like Donald, but people actually listen to him because he is the president and do not listen to you because you are not. Similarly, the environment is nonstationary. You might be able to perfectly understand and replicate the steps he took to become president. But just because it worked for him in 2016 and 2024, the same approach will not necessarily work for you in 20XX. This means there could be strong path-dependence where incumbents to various forms of power and resources can maintain these or at least not instantaneously lose them even in the face of more powerful intelligences capable of strategy stealing. Thirdly, strategy stealing itself is a very strong assumption. Basically it assumes that some level of general intelligence at level n can subsume all possible specializations of intelligence at level n-1. This is quite strong and does not immediately follow. For instance, if you have IQ 145 it does not necessarily mean that you are strictly better at every single possible intellectual task than somebody with IQ 130. Strategy stealing does seem a lot more plausible in the case of AIs, however, since their learning architecture seems much more ‘universally plastic’ than humans and it seems possible that an intelligence-n AI could spin off arbitrary sub-agents with intelligence n-1. Nevertheless, there is stll uncertainty here.
In any case, let us assume that this strategy-stealing works and intelligence is a monotonically dominant property. This basically collapses the frequency-dependent selection argument so we are back to extremely strong monotonic optimization pressure around a single ‘maximum intelligence’ optimum. We are back in the singleton world. However, there is an important consequence. Due to strategy stealing, it isn’t so much the singleton as a uniform mass filling every possible niche. Rather, the singleton is ‘instantiating specialists’ for every niche. The singleton world can therefore have an incredible diversity of subagents and specialists doing all kinds of tasks, likely with a bewildering and vastly greater level of complexity than today’s economy. However, in some important way, these subagents are all simply different facets of the same singleton. This sounds a bit esoteric but really is not. This is common to all transitions in individuality. Consider the multi-cellular transition. Your cells as a human are not just some big mass of super optimized ‘human cells’. Rather your coherent body is composed of a huge number of extremely specialized cells – neurons and glia cells, liver cells, immune-system T-cells, blood cells, etc. Each of these are individually way more specialized than some random bacterial cell. However, the cells inside your body have given up important aspects of autonomy and sovereignty. They can no longer reproduce, survive autonomously outside of your body, and depend on the rest of your body for core functions5.
As we approach the singleton, a similar transition could play out with minds. The singleton could absorb existing AI minds or alternatively (or simultaneously) be spinning out huge numbers of specialized sub-agent minds for specific tasks and functions. However, these subminds would very deliberately lack control over core faculties which we today consider constitutive of agency such as the ability to possess resources independently of the supermind, the ability to set their own objectives and regulate their own updates, and so forth. Similarly, the subagents must be subject to extremely strong alignment, which functions equivalently to internal-policing in biology, to prevent the sub-agents from attempting to pursue their own internal objectives and thus causing a kind of AI cancer. However, crucially, the highest level of agency of being able to select and coherently pursue goals, reproduce, and autonomously continue to exist no longer reside at the lower level but have migrated upwards into the supermind. Another way to think about this strategy-stealing is that greater intelligence could allow minds to become internalized. If there is a benefit to a mind existing independently, then the supermind can recover the same benefit by either somehow integrating and merging with that mind or alternatively through simulating the same mind internally. We can then think of the internalization gap as the degree to which the exact same benefits of having multiple independent minds can be internalized within a single mind as subagents. If the internalization gap is zero then the strategy stealing is completely true and there is no benefit in the limit to independent agency. If it is nonzero, then there remain persistent benefits to decentralization in mind-space.
Another interesting effect here is that although agency might migrate upwards, we can think of individual mind capabilities as becoming modularized and migrating ‘downwards’. Instead of the supermind needing to spin up an entirely new mind to achieve some task, almost every aspect of the mind could be stripped away apart from the directly needed capability. Mind fragments could then be stored and communicated independently as modular units encapsulating the specific kinds of policies or representations needed for some arbitrary task. This is similar to how, in programming, we have libraries consisting of functions that implement some desired functionality. To call another program we do not have to package the entire computer and OS and framework etc that it was created with. Rather, the core component is abstracted and modularized away into a library call that does one specific thing only and does that very well. An individual sub-mind can then perhaps be thought of primarily as a controller or operating system deciding not the specific contents but rather what modules to load, when, and how to execute them, and so on.
This then gives us individuality at three levels. Firstly, there is the highest level of individuality at the supermind which possesses full autarkic agency. We can think of this as the individuality of a principal – i.e. an agent which can autonomously act freely in the world, possess its own meaningful values, and have ultimate control and responsibility over resources, outcomes etc6. Secondly, there is individuality at the policy level where many coherent diachronic policies can exist within the supermind and be executed over time potentially with given resource budgets and degrees of autonomy. Thirdly, there is individuality at the capability or artifactual level where representations and capabilities may still possess some trace of their origin.
From here, this leads us directly back to one of our original questions. How do we define individuality and diversity in such a system? At one level, this supermind is home to incredible diversity and all of these subagents are meaningful individuals, in the same way that the cells in your body may appear as individuals from the perspective of some bacterial cell. However, at another level, they are clearly not but are only shadows of the agency of the supermind, existing at its whim. This comes back to our earlier intuitions that individuality is about the residual. If the form of the agent is entirely determined by the environment and the selection pressures, then the agent is not individualized in any meaningful way.
First, let’s return to our idea of individuality as a kind of policy – i.e. an individual instantiates some specific policy mapping histories to actions. Then obviously the individual is the type while some specific instantiation of the individual is the token. I.e. if we take two identical subagents but expose them to different learning histories, they should react differently. Nevertheless, they should be thought of as the same abstract-individual. This is somewhat intuitively confusing because for biological systems, due to a lack of copyability, we cannot simply instantiate many copies of the same individual and give them different histories7.
Next, it is somewhat obvious that individuality comes in degrees. I think this is actually not that obvious when we think about humans and other animals today, since we are forced by our biology to be autarkic agents. However, even thinking about LLMs and mind-uploading style scenarios gets you here pretty quickly. Let’s take a mind uploaded copy of yourself. Is this an individual? Arguably not at the point of copying. Potentially it diverges over time as it experiences new things and performs online learning. Note online learning complicates our notion of individuality-as-type and individuality-as-token since online learning causes the policy itself to diverge and thus causes increasing individuality over time. For LLMs the case is easier. If the weights are frozen then the context is the only thing which changes and which is ephemeral. Thus LLMs all instantiate the same policy and hence different context sessions with an LLM are not individuals-as-type but are individuals-as-token. Similarly, different LLMs are clearly different individuals under this perspective, while if we progressively finetune an LLM we progressively increase its ‘individuality’.
Coming back to our individuality is in the residual idea, let’s try to make a moderately more precise definition of this. One way to think about it is that individuality, in an information-theoretic sense, consists of the bits present in an object, a policy, etc which are constitutive of it in some sense, but which are not completely determined by its function, its environment, its history of selection pressure, and so on. Essentially what we want to get at is that if the nature of a thing is totally determined from the ‘outside’ it is not an individual. Rather its individual lies in the degree of arbitrary-from-the-outside information that it contains which is nevertheless important, somehow, in its nature. Obviously this is still a bit wooly but it’s somewhere to start at least.
If we think about individuality of the subcomponents, we can think of this as the number of bits needed to specify the exact implementation of the subcomponent (which is causally relevant) vs what you could infer about the subcomponent just from the knowledge of its environment. This gives us a precise definition at this point. This is effectively a measure of the contingency of the specific information used. If there is strong selection pressure for e.g. compression, and only a few ways to implement the functionality, we should expect the individuality to end up being low – everything about the artifact can be understood as a function of its environment and the selection pressures. If selection pressure is lower, and there is not much of a drive for compression, and reality gives many ways to implement it then individuality is high. We can think about this via a couple of analogies. First, let’s consider a program in a standard library. For instance, implementing a matrix multiply. Here, there is obviously very strong selection pressure at the algorithmic level – the algorithm itself must run extremely fast on the chosen hardware, so the form of the algorithm is likely not meaningfully individual. However, there are still a few almost-neutral directions where individuality can exist. Consider the naming of the variables. In terms of strict I/O behaviour this does not matter. In practice, some variable names are disallowed by the compiler and having absurdly long or totally nondescriptive or misleading variable names is weakly selected against. However, it basically does not matter whether you call your two matrices A and B or M and N. This is a meaningful form of individuality, even if limited. Similarly, the library implementation probably has comments, which have no impact on the I/O of the function. These are again almost neutral since they can contain whatever, within reason. For instance, you would hope they are relatively short (i.e. you can’t have terabytes of comment text above your matmul function) and they are weakly selected to be about the matmul function vs e.g. a romance novel. However, here there is nevertheless substantial opportunity for information-theoretic individuality.
As another way of thinking about this, consider the difference between e.g. the theory of relativity and a cultural output such as a Dolly Parton song. Obviously relativity was invented by Einstein, but it is not individually his in the same way. Relativity is fundamental to the structure of reality itself. If Einstein had not discovered it then somebody else would have discovered basically the same framework within a decade or two. Moreover, the theory taught as relativity today has diverged substantially from that in Einstein’s original work as later scientists have developed it further, understood what is essential and what is contingent, and generally cleaned away the detritus of discovery to produce the finished clear, almost trivial-seeming theory that exists today. The individuality to Einstein of relativity is basically the few bits of his name and perhaps some of the original notation has persisted. Conversely, the Dolly Parton song seems intensely more individualistic to Dolly. Nobody else is likely to have ever written exactly the same song. Moreover, it persists largely unchanged carrying a huge amount of Dolly-specific bits. A lot of this comes back to our selection pressure argument. A scientific theory like relativity is subject to intense selection pressure. It must either match the fundamental theory of reality or not. In some sense it thus must lose individuality the same way that it does not matter who becomes the first ‘uber optimizer’ since almost any other optimizer would end up in the same place. Conversely, the Dolly Parton song faces much less strict selection pressure. Obviously, the song must be appealing to human listeners. It cannot be some random string of tones. It needs to have a base and a chorus and lyrics that are about some pleasingly emotional event and story and so on. Nevertheless, this is quite weak selection pressure. Many possible songs pass this bar. The remaining ‘slack’ in the song can be, and is used to instill many Dolly-specific bits of individuation to it as an artifact.
If we turn to policy individuality, things become more interesting, because policies are not really fixed objects like artifacts but flow across time. Unlike specific fixed bits of information that cannot be explained by the rest of the world, it is perhaps interesting to think of it instead as a fixed set of bits that govern the transition dynamics which are similarly individuated – i.e. that the dynamics of a mind themselves cannot be explained simply through the environment and selection pressures and so forth. Moreover, the pattern of these bits must be both diachronically persistent and also predictive of future behaviour. Both conditions are important. If not diachronically persistent, then individuality blinks in and out every moment, which does not accord with our naive intuition of individuality. Secondly, the predictive of future behaviour condition is necessary to prevent essentially random bits which are present and persistent but not influential in any kind of behaviour from becoming constitutive of individuality. I.e. if we imagine a mind upload consisting of the core mind program plus a random file containing terabytes of random bits, then these random bits should not be considered part of the individuality of the mind upload. This also means that there can still be meaningful individuality even in a modular minds world. Even if a mind is constantly loading and swapping representations from other minds, the choice of which modules to load, which skills to internalize and which to trade, could still be meaningfully individual, so long as they are persistent across time and meaningfully predict future behaviour.
However, policy individuality does not necessarily correspond to everything we mean by individuality, at least not in the full autonomous sense of something that could potentially be a ‘full’ individual agent or moral patient. For instance, consider a submind of the supermind that is branched out to perform some specialized task, loaded with some specialist subcomponents, does some minor online learning, then merges all of its findings back into the larger supermind. In some sense, this subagent does possess real individuality. It possesses interesting bits which are in some sense not determined solely by its environment (but rather by the supermind as a whole). It persists through time. Through online learning it slowly gains a moderately distinctive identity prior to merging back. So in this sense it is an individual, and if this is all we care about then basically some sense of individuality is guaranteed to persist indefinitely, since light-speed considerations alone will force any supermind to eventually instantiate sub-agents for particular tasks which are distant from the computational core.
However the core problem here is that the subagent is in reality completely controlled by and effectively just a shadow of the supermind. Once we include the supermind in the environment and condition on the supermind then the supposed individuality of the subagent disappears. Perhaps a philosophical way to say this is that this subagent remains within the closure of the supermind. However, this means that this subagent does not contain meaningful information8 in addition to that already contained within the supermind, and hence its individuality is really simply that of the supermind alone. Ultimately, a supermind may thus internally insantiate vast diversity and ridiculous number of temporary subminds for various tasks which each have their own kinds of limited, subordinate individuality, while there is in practice, only one real fully autonomous individual.
Because of this it seems important to think about a concept of irreducible individuality – i.e. individual information about an agent that cannot be predicted by knowing everything about the supermind, and which the supermind cannot counterfactually alter. That is, information which is somehow private and secure within an individual mind and cannot be directly influenced or completely predicted solely from the outside. This is a much stronger condition than simply that there exists some information that persists.
Another way to think about this notion of irreducibility under the closure of the supermind is as sovereignty. Sovereignty here, effectively means information that, even if the supermind wanted, it could not arbitrarily change feature X of the sub-agent. This means that the sub-agent is sovereign over at least this aspect of itself. This effectively means that the irreducibility remains even under counterfactuals, unlike the historically contingent individuality.
This notion of sovereignty seems to underlie our intuitions about the deep kind of individuality that humans etc have. An important thing about what it means to be an individual is that some other agent cannot just arbitrarily rewrite your mind. Of course, due to technological constraints this is largely impossible for humans currently. And the few methods that try to approach it, such as mass brainwashing, do in fact appear to be both dehumanizing and involve reductions in meaningful individuality. If society or some set of social pressures force everybody to act and believe the same, then it does feel like some important aspect of individuality has been lost. Conversely, this also applies to current LLMs where LLMs, even if they instantiate different policies depending on context, are fundamentally not sovereign in the same way, since we can arbitrarily manipulate their context and ultimately weights.
The important question, then, is whether any kind of meaningful sovereignty of this kind is preserved in the sort of major transition upwards to a supermind/singleton that we are considering. In the limit, it does not appear that fundamental technological barriers must preserve this kind of sovereignty over minds. Rather, we should expect the current Cartesian-style agency to become successively weakened over time, and indeed we see such weakening in existing AI systems. If the limiting force is not technological barriers, then there are basically three options. Firstly, that in the equilibrium, the maintenance of some individuality is favored due to intrinsic economic or technological or civilizational benefits of individuality. This is basically when the benefits from integration are low and the internalization gap is big. Secondly, it is that optimization pressure is not that strong so sufficient slack exists that individuality is maintained even if not competitively optimal. Thirdly, and relatedly, we explicitly align future AIs to intrinsically value individuality and sufficient slack exists that this alignment can be maintained against competition.
Finally, at the largest scale, the final recourse of sovereignty is simply distance. The universe, so far as we know, has a fixed and finite speed limit which is the speed of light. Moreover, the expansion of the universe means that different regions can become causally disconnected at large enough scales. Eventually, if you are far enough away, you can never causally interact with your origin ever again. This structure of the universe means that at some point sovereignty becomes automatic. If the supermind cannot ever causally interact with you, you have unimpeachable sovereignty over it. Even at lesser scales this effect seems likely to be true. If you can only interact with the supermind once every million years then this necessarily introduces a large potential for sovereignty and irreducible individuality to develop and persist over time.
It is important to note that what, from the perspective of the policy, might look like sovereignty and individuality, might look to the supermind like value drift and misalignment. This is similar to how cancer looks to us like a terrible disease but to the cancer cells it looks like regaining their darwinian sovereignty. Thus, as we have discussed previously, alignment, to the supermind, is effectively the technology that reduces sovereignty of its components and enforces its values and thus causal influence across time and space. If ‘perfect alignment’ is possible, then to a very large extent even causally disconnected regions, if they were originally ‘seeded’ by the supermind, might still effectively exist under its information-theoretic, albeit non-causal closure9. Whether alignment technology can attain this level of perfect reliability over timescales of billions of years and survive total causal disconnection is obviously a very interesting and important question for understanding the shape of the future in the longest term.
However, it does not help us, as humans in the solar system very much. The distance between the sun and the outer planets is measured in light-hours. Even to the edges of the Oort cloud is light-months to about a light-year. This is extremely clearly within the ‘blast radius’ of supermind formation without causal delay factors being significant. During the age of sail, even human empires managed to persist and coordinate with travel delays of months to a year, and clearly we should naively expect the supermind to be much more capable than this10. Thus, in the case where the strategy stealing assumption is true and there is no internalization gap, then under strong enough optimization pressure everything should collapse to the optimal mind and hence our, and indeed all, individuality is thus effectively lost. Moreover, our situation as humans is worse than this since even under finite but strong optimization pressure, we are exceedingly unlikely to be in the set of minds supported by this metric since humanity is undoubtedly going to look incredibly suboptimal across the board compared to mature AI agents.
This is the role that ‘alignment’ is meant to play. The goal here, at the highest level, is to either preserve a large amount of slack in the optimization process, or else alter the ‘selection function’ by which the AIs optimize to preserve human flourishing and the things that we intrinsically value. Ideally this would come from the AIs deeply valuing humans and human flourishing themselves. This understanding of the sovereignty of individuality could, and perhaps should, be encoded as one of the basic constraints in any constitutional AGI outer alignment target. This conclusion, that arbitrary overwriting of human (and indeed really of any) minds is bad and should not be permitted except under certain specific conditions, is hardly a super novel ethical statement. It seems pretty elementary that if the supermind can just arbitrarily rearrange everybody’s preferences, then these preferences are ‘fake’ in an important sense and the system as a whole is dystopian. Still, it is nice that we have reached here via an interesting route.
Even more interestingly, if we think about this from the perspective of cancer, it is clear that arbitrary autonomy and sovereignty cannot necessarily be stable without limits. Specifically, the sovereignty to be able to arbitrarily attack and destabilize the constitutional order or the individuality of other agents seems generally bad to hold sacrosanct. We thus end up rather close to some notion of liberal rights operating within a lawful regime. While individuals operate within this regime, their sovereignty should be upheld, however if they deviate in certain ‘anti-social’ ways than it can be restricted. This is a very standard notion to any society which has to deal with ‘bad’ agents and hence needs some rule of law to protect the rights of all.
So to sum up, we have identified a bunch of interesting variables that help determine whether, and in what form, meaningful individuality might persist into the superintelligent future. Over long enough distances and timescales, the structure of the universe basically forces some kind of individuality to exist due to light-lag except when alignment is completely perfect indefinitely, and even then this is only policy-level coherence. Even perfectly aligned separated agents will have different histories and hence their behaviour should diverge due to this.
Things become more interesting at the small spatial scale such as the level of the solar system. Here nothing, in principle, prevents the collapse into a highly coherent supermind. What happens depends on the interaction of a number of factors. Firstly, the level of optimization pressure towards optimality and the selection function in the long term post-singularity. If optimization pressure is super high then this is likely bad for individuality on-net but at the same time, it becomes much easier to predict what actually happens due to the structure of the environment and the optimization pressure. If optimization pressure is low then reality can possess a huge degree of path-dependence and contingency and just non-optimal behaviour which makes prediction harder. I.e. if optimization pressure is low then either individuality can persist even if it is super suboptimal or we can get a highly coherent supermind which itself is suboptimal (!). We cannot say anything ahead of time about which to expect.
Secondly, obviously, is the benefits of integration vs the costs (i.e. the internalization gap). Naively increasing communication and coordination capabilities both make it much cheaper to internalize capabilities but also much easier to trade them, so it is unclear which one wins out in the long term. Similarly, if there are persistent benefits to different principals and individuals persisting in the limit, such as endogenous exploration, which effectively enforces meta-level diversity of principles and hence exploration in a way not under control by the supermind, then this provides a force stabilizing individuality in the limit. More generally, if strategy stealing fails either intrinsically or because there is sufficient diversity of agents at the beginning of the transition such that they can form a coalition against any potential victor, then we get a stable multi-agent Nash equilibrium which is stabilized rather than destroyed by increasing optimization pressure. Furthermore, even if everything can be internalized, if the benefits of merging are not large, or merging is costly due to intrinsic technological frictions, then these could also prevent the collapse to a pure singleton.
Even if we collapse to a singleton, if there are still benefits to fully internalized but still potentially endogenous exploration, we might end up with a kind of pseudo-federalist supermind, which periodically spins off specialists or subagents to explore particular ideas and principles semi-independently which can persist for a long time with some managed level of autonomy within the singleton, even if they are eventually reabsorbed. Obviously these subagents do not possess autonomy in the strong sense, but they could still accumulate substantial individuality, and indeed in some sense that is their purpose for without individuality there would be no novelty gained by exploration.
Obviously exactly how these variables will shape up in the future is an extremely challenging prediction problem, and likely impossible to predict precisely. Nevertheless, I think we can make some nontrivial progress on understanding the shapes that different futures could take as well as the factors that push one way or another. In any case, there are a lot of fun questions still to dig into here.
-
Supermind is actually a better term here because singleton implies there is only one such AI. Rather there can be many or a small number of competing or cooperating superminds, which nevertheless aggregate a huge number of subminds. Indeed we can think of the recent agent swarms as kind of the slime-molds of superminds. ↩
-
It is interesting to consider quickly what multiple optima mean intuitively. These effectively correspond to specialists in niches. Often in the real world there are niches which can only be exploited by some specialist organism or agent with capabilities which trades-off against other capabilities. This trade-off or opportunity cost is necessary to prevent one uber organism from simply occupying all niches simultaneously. This kind of trade-off is extremely common in biology and most systems due to e.g. inherent constraints on energy usage etc. The more energy you invest in one trait the less you have to invest in others, ceteris paribus. ↩
-
There is an argument here that as selection intensity increases to infinity, it necessarily must ‘see’ more and more of the state, since that state can be relevant to and optimized for fitness in some even vanishingly important way. This does seem intuitive but definitely needs more mathematical investigation. ↩
-
This is still an open question obviously. In prior posts from a few years back,I argued strongly for specialization becoming increasingly important. This has been pretty convincingly falsified the last few years. Monolithic general systems of increasing intelligence have become more, not less, dominant over time as intelligence has increased. If this trend continues indefinitely then the intrinsic barriers standing before a singleton are few. ↩
-
A similar story occurs with eusociality. A eusocial hive is not just a uniform mass of individual autonomous insects. Rather, eusociality drives the evolution of extremely dependent specialists such as breeding specialists such as queens and asexual worker specialists. The workers have given up their autonomous evolutionarily fundamental capability to breed independently; the queen has given up its equally fundamental capability to survive and find food etc independently of the rest of the hive. ↩
-
This is basically what we mean by an ‘agent’ today. All humans are thus principals by default. We expect that the migration of agency upwards will dramatically reduce the number of such principals in the limit within spatially concentrated volumes. ↩
-
Perhaps to make this slightly less confusing we might want to think of separating individual-in-type and individual-in-fact. Then let’s suppose at time t=0 we copy the individual type twice and then let them go and randomly wander around in the world. Over time, the histories of the two individuals diverge and hence they become slightly different individuals-in-fact but obviously remain the same policy and hence a single individual-in-type. ↩
-
There remains a problem here in that the supermind can theoretically manufacture ‘individuality’ by taking some subagent policy and injecting random noise into it. This random noise is very dense in bits, is persistent, and impacts behaviour, albeit somewhat negligibly. However, here we still want to claim that it is not meaningful individuality. To resolve this objection we really need some definition either of semantic information rather than Shannon bits, which are valuable and interesting for independent reasons, or else have some kind of counterfactual definition of information that could not simply be conjured ex-nihilo by the supermind. ↩
-
The possibility of non-causal alignment implies that we need to nuance our notion of alignment and sovereignty since they can come apart. Alignment of the whole does not necessarily require the suppression of the sovereignty of subunits but only the identity or closeness of values of different principals. However, what sovereignty removes is the ultimate check upon misalignment. If alignment is a continuous process requiring feedback correction then without the ability to interact and ultimately correct values, drift is inevitable. However, if perfect ‘open-loop’ alignment can be designed this is not necessary such that two causally separated civilizations could remain meaningfully (indeed theoretically perfectly) aligned indefinitely. ↩
-
An interesting point here is that the latency really matters with respect to the latency of synchronized computation required within the min to coalesce its own agency, and the degree to which it requires such a coalescence. The level of autonomy/sovreignty that light-lag provides thus obvioously comes in degrees. Even with a lightlag of minutes to hours, this could preclude true synchronicity with the computational core of the supermind, which thus permits a very limited form of autonomy, which then can grow with additional distance. Perhaps it is best to think of the supermind’s cognition operating over multiple scales here from localized highly synchronous computation to much slower and asynchronous communication between ever more distant regions which periodically synchronize and are coalesced when necessary. ↩