The other day I was rereading my Thermopylae poem and got to thinking seriously about the information dynamics of history and how these might evolve in the limit. It turns out we can go quite far in this thinking from some very basic premises and it relates somewhat deeply with our other questions about individuality and changing levels of agency. So let’s dive in and have fun.

First, the interesting question is what, exactly, is history? Let’s take the most basic setting. Let’s imagine the universe is some gigantic MDP evolving according to some transition function. From this, let’s imagine trying to reconstruct a ‘perfect history’ – i.e., to maintain all of the information about the past that existed. There are basically two constraints here. Firstly, the actual dynamics of our universe appear to be irreversibly lossy1 in practice for bounded agents. This means at some level of fidelity, there is always an information-theoretic barrier. No matter how much compute we have, without impossible perfect knowledge of the entire state of the universe, we cannot necessarily perfectly reconstruct the past. This immediately shows that the original Fyodorovist dreams of universal resurrection cannot be realized. Of course, with enough compute, a supermind could produce incredibly detailed reconstructions of minds existing at some particular time, but it can only simulate the maximum entropy distribution given the then-current state of the universe, not the true minds that existed at the precise moment being simulated. By chance, it might simulate exactly the right mind amidst uncountable billions of others, but it can never know which one is the true historical mind.

At this point, it is important to realize that the amount of information preserved about the past is not fixed. We can alter the dynamics of the universe to preserve specific information or not. History fundamentally depends upon durable traces which maintain information over time. Some of these are intentional, such as monuments, pyramids, tablets, books etc. Others are unintentional and just byproducts of civilization, such as pottery shards, chemical levels in the atmosphere, and general detritus. The entire field of archaeology shows us that we can retrieve a surprising amount of information about the past from such miscellaneous historical debris. However, this is not an exogenous process, and we as a civilization can decide what and how much information to preserve. This can include things such as storing information on durable media, making many redundant copies of interesting information, storing objects in places where they decay very slowly such as in cold regions, or in outer space, explicit storage centers such as seed and gene banks, cryonic preservation of individuals and so on.

However, there is a limit to this. Historical information must be physically instantiated somewhere. And if these bits are being used for preserving historical information through time, then they are not changing in time which means they are not being used directly for present purposes. In this sense, there is a direct information-theoretic tradeoff between preserving the past and utilizing that state for present or future computation. If the past states are frozen perfectly, then they form a direct conduit into the future, however their entropy cannot be utilized for other purposes during this time. This is largely irrelevant at the current time, and likely for a long time into the future, since we are just incredibly far away from the pareto frontier of using all microstates the universe gives us for value, however once we reach our K4 civilization at the end of time, these considerations may become important. At that point, there will be the capacity to use a substantial fraction of the information capacity of the universe for computation, if desired, and there will be an explicit trade-off between using bits for computation or storage2.

There is a second important and obvious constraint here: the storage space for this historical information. Let us consider here the classic case of Borges’ 1:1-scale map. In some sense this map is the most perfect map possible – it preserves all of the information in the original exactly. Conversely, the map is impossible because the map is the same size as the territory. The storage size of this 1-1 map is identical to the entire ‘resource capacity’ of the world. Attempting to preserve all of history is essentially trying to make a 1:1-scale map, but worse. This is because, unlike the map, history evolves over time. Indeed, that is its core essence. This means that, as time passes, there is more history to record. History increases at the rate of 1 year per year3. This means that if we have any fixed storage capacity, which eventually we must since the universe itself has fixed microstate capacity, then our historical record must become increasingly lossy over time. Specifically, even ignoring the intrinsically lossy dynamics, there is the storage constraint that we must lose information about history proportionately to the amount of history we want to record. That is, historical information must scale, at best, as $\frac{F}{T}$ where $F$ is the fraction of the universe dedicated to historical preservation and $T$ is the total amount of elapsed history that we care about4.

There is thus the interesting question of whether, for future historians, studying our contemporary civilization is likely to be storage-limited or information-limited. I think it is extremely likely that the current historical era, and indeed almost all of human history, will be information-limited, meaning that whatever survives from this era will have substantial cognitive and storage effort dedicated to it and that future civilization will ‘wish they had more’ early 21st century (and historical) material.

This follows immediately from the fact that all of human history up to the present is simply not that big. The entirety of the internet – i.e. common crawl – is only in the hundreds of terabytes once cleaned, deduplicated etc. Probably all of the text information written down by humans to date on the internet, in books etc, is on the order of tens of petabytes. If we include e.g. all the video data on youtube then that is still only in the tens of exabytes and that can almost certainly be compressed substantially without meaningful loss. Moreover, almost all of this information is very close to the present day. If we think about information coming to us from before the 21st century, it is only a tiny fraction of the information the 21st century internet has created. Moreover, these numbers, terabytes, exabytes are minuscule compared to the computational resources of any nontrivial future civilization. This means, we are likely extremely far from the regime where future civilizations want to or must discard information about our time. Rather, any efforts that increase information bandwidth to the future are likely to be seen as very helpful by them. To think productively from here, we need to consider two natural but distinct objects. First, there is the historical state $H$ which we can think of as all of the historical information about all previous times that exists at time $t$. Second, there is the historical transition operator $T$ which takes the current world state and the historical state and updates the historical state. By basic data-processing inequalities, the historical state can never gain information about previous historical moments, only maintain or lose it (although a historian or archive can gain information over time by uncovering new sources of bits about the full historical state). However, this historical state also grows over time by new information being added to it about the current time.

Bounded agents can never realistically absorb the full historical state. Instead they must learn approximations to it, and from our understanding of deep learning, we know what these approximations must look like. First, perhaps the most important thing is understanding the distributional and spectral structure of this historical state. Obviously, there is a lot of uncertainty here but I think we can say some things with some reasonable degree of confidence. Firstly, the historical state is a huge naturalistic ‘real-world’ dataset (perhaps the ultimate ‘real-world’ dataset). From what we know about the general structure of these in practice, which we discussed in depth in this post, we should thus expect it to have certain high-level properties. Specifically, we expect the historical state to have a fractal-like connectivity/importance structure with a power-law decay of importance of different historical modes. This just comes from the structure of the world in which there is a meaningful hierarchy of causes and states, that some causes are much more important and highly central than others, and there is essentially infinite fractal detail – at every lower level you look at you never ‘run out of detail’ which is not trivially predictable from the higher level. Crucially, the power-law importance structure can be thought of as the necessary conditions for indefinite learnability. If we think of this fractal hierarchical structure as a tree, what the power-law structure is saying is that the level of predictive information is the same at every scale, or equivalently every level of the tree. This is vital since the number of potential leaves of the tree, or potential detail at each scale, increases exponentially at each level. If predictive information was e.g. constant with the number of features, then all information would concentrate exponentially at the lowest level of scale, thus leading everything to appear to simply be irreducible noise. If predictive information declined superexponentially, then all predictive information is essentially absorbed by the relatively few modes at the highest scale and thus everything about the system is trivially learnable with some level of scale. This property of natural datasets is what gives the standard power-law scaling we see in ML and generally in naturalistic datasets.

Crucially, by all accounts, history itself appears to have this structure (and it would be extremely surprising if it didn’t). There is clearly large-scale and scale-free structure in history, as well as irreducible fractal-like detail. Some historical events matter hugely. Others are super unimportant. There are many more unimportant than important events. There is never a level at which you ‘run out of history’, and it is not like this additional history is all trivial and pointless detail. Small things can matter and compound into big things later on. At the same time, clearly, there is a hierarchical structure of compressibility. You can say a surprising amount about the broad arc of history at the highest level in relatively few words. You can cover the rough history of the world in a few paragraphs. At the same time, in any particular event, or at any level, there is a massive amount of detail that is intrinsic to that level and which has real predictive information propagating downwards5.

This means that a very interesting way to consider the historical state is through the spectrum of its features. From the perspective of the historian, there is thus going to be a power-law ‘scaling law’ between computational effort (either compute or storage) in historical information-gathering, processing, and general analysis and the amount of ‘historical understanding’ a civilization maintains6. This then brings us to the question of the historical transition operator. In some sense this is just the regular laws of physics acting on the universe state. However, if we take a historian’s perspective on this we can think through a couple of important properties and implications. First, let’s start by thinking about the historical transition operator as an information channel. We have a ‘message’: the historical state at the present time. Through repeated transitions through time this ‘message’ is encoded and sent through the channel. Then it is ‘decoded’ at the other end by the future. Crucially, this channel is extremely lossy – only a tiny fraction of total potential ‘historical information’ about the present time can be meaningfully transferred. This happens for two reasons. Firstly, the laws of physics increase entropy and hence information about the past is irrevocably lost as it falls beneath the noise floor. Secondly, even if the information has persisted and is somehow present, future historians still need to desire to read it, process it, and retain it.

In the present and historically, we are almost certainly primarily bottlenecked simply by information preservation. Having enough storage to store all of history right now is trivial, and huge efforts are spent puzzling over fragmentary and incomplete evidence. This may change in the future as highly advanced digitized civilizations naturally generate enormous quantities of data and the preservation of the data, if desired, can be straightforward. However, today, explicit preservation of information may be the easiest method to send information down the channel, trusting essentially that future historians with much greater computational power will find basically any information we send interesting. This is essentially equivalent to increasing the rate of the channel itself. If we increase the information transmission rate or reduce the noise, then we can send more information down the channel without distortion.

We can think of this, pseudo-mathematically, as affecting the eigenmodes of the historical transition operator. We can think of these as essentially giving decay rates, or half-lives, of historical transmission modes. For instance, information on ephemeral media decays extremely quickly. If your computer is powered off you lose what is in RAM. Information on digital media in general may decay quickly – almost no commercial hard-drives today are designed to last for thousands of years and even if they are the supporting infrastructure of file formats, reading devices etc is required to be able to understand what the bits represent. Other information is much longer lasting. For instance, if stored in suitable climates, books can last for hundreds of years, cuneiform tablets can last for millennia with the information still (mostly) readable, and many artifacts, even unintentionally dumped ones, can persist for thousands of years or longer. Specially designed historical transmission mechanisms such as the Rosetta Disk and the Billion year archive project can produce artifacts designed to transmit information across millions or billions of years. If texts are repeatedly read, processed, and copied they can last essentially indefinitely with error-correction. Similarly, oral history and myth has a half-life over generations since the error-correction properties are much weaker than in text, since each generation must learn and repeat them anew, but nevertheless, they are capable of preserving a few canonical stories very well for hundreds of years. Perhaps the most underrated form of preservation is direct physical preservation of miscellaneous evidence. Archaeology manages to derive shocking amounts of information from very fragmentary physical evidence which, crucially, is not selected by literary/cultural importance unlike almost all of our textual information (it is selected by other, somewhat decorrelated forces such as the differential preservation of materials by different climates, geographies, vagaries of war, and so on7).

More formally, let’s suppose our ‘universal MDP’ has a transition operator T. Then, we can view the process of history as taking the initial state and progressively applying T to it. The transition operator has eigenmodes8 – that is it accentuates elements of the state and diminishes others. These eigenmodes have half-lives. Some historical information is almost certainly lost. Other information persists for a very long time. Other information is lost explicitly but persists as path-dependence of the state – i.e. it locks the state into some persistent region of the state-space and out of others. By choosing to deliberately preserve information we are effectively altering these eigenmodes towards our desired directions, preserving the information that we want, erasing the information that we do not want. What ends up at historical time in the future t then becomes the integral over this evolution.

However, the historical transition operator merely tells us what information is successfully transmitted across the information channel of time. What we want to understand is both the transmission and the reception. That is there is not only a channel but also a decoder. Will future civilizations actually process, understand, study, and ultimately transmit the historical information even deeper into time? Will future civilizations actually care about us or our history?

For instance, Scott Alexander imagines a world where vast superintelligences in distant galaxies exhaustively analyze every aspect of pre-singularity earth. I happen to agree that there is quite a strong chance of something like this occurring, but this outcome is by no means necessary. Presumably this AI civilization of Jupiter brains would in fact be generating absurd amounts of historical information themselves9, which their historians could study instead of ancient Greece or 21st century San Francisco. Moreover, in the limit, the entire universe itself is generating vast amounts of historical information all the time. Why study 21st century singularity history when you could be studying the previous thoughts of your own copies from a thousand years ago, or the historical evolution of dust grains inside an asteroid somewhere between Sol and Alpha Centauri?

If we thus think about the question of which features of the historical state are maintained and processed, or ‘kept in mind’, by future civilizations, then this is basically a product of two factors – the objective ‘feature rank’ and then the relevance of this historical feature to the interests of future civilizations. We can think of the ‘objective rank’ as a kind of structural centrality and salience. How much causal information is some historical bit upstream of? How many descriptions of events or models of later events must include this event? Another way to think about this is how broad is the distribution of future decoders which assign this feature a high rank. Certainly not all possible decoders will, but if a historical event is truly salient, a much larger proportion of decoders must assign it a high rank. The second factor – decoder relevance and value – is basically a question of the distribution of interests of future decoders. The eyes of the future are drawn in their own way. The interesting question is to what extent this distribution of future decoding interest is predictable ahead of time. Certainly, there are certain structural factors we might be able to predict a-priori. For instance, we might predict that causal contribution towards the development and existence of the future civilization itself will be interesting to them. Secondly, we might expect irreducible informational novelty to be interesting. Much of history is frankly extremely boring in that e.g. the motion of the planets in the past is extremely predictable from the laws of physics and the current positions. This can be reconstructed. Contingent decisions in non-ergodic parts of the state-space are interesting, since they create path-dependence in the historical trajectory. Finally, essentially tautologically but almost forgotten, the present information must be preserved into the future so that future civilizations can potentially read and understand it, and use it to reconstruct the past according to their own purposes.

In general, I suspect that human history in general and especially the period around the singularity are likely to be interesting to a large number of potential future civilizations, although not all of them. This is because we are almost certainly going to be the origin point and the era that ‘birthed’ the future civilization in a nontrivial way. Moreover, it seems like the future here is becoming increasingly non-ergodic. There may be substantial path-dependencies arising around this point, which did not exist in prior eras. Both of these imply potentially substantial interest. Of course this is not guaranteed. For instance, if we all get eaten by a paperclipper, the paperclipper has zero interest in history beyond what helps more paperclips, for which almost all of our history seems likely to be marginal. However, a broad distribution of potential values has some interest in history, and then from there our specific history seems likely to be unusually interesting (and well-preserved) a-priori. This is even true if we somehow go extinct due to AI or other existential risk and an alien civilization subsumes the solar system – they might well be interested in the data about what happened on earth during this time. We certainly would if we came across a dead alien planet with clear evidence of prior technological civilization.

Importantly, future historians and decoders being selective is, in some sense, inevitable due to the storage arguments. In the limit, once the information-rate of the historical channel is sufficiently high, not all the resultant historical output can be stored and processed. Future civilizations must choose which history to preserve. They must care more about certain aspects of history than others. They don’t want to create a 1-1 map of the entire universe since most of the universe is ‘boring’. Scott assumes that 21st century history around the singularity will thus be much more interesting to the hypothetical jupiter-brain civilization than the dust-grains in the asteroid. This is not at all unreasonable, but this is implicitly assuming knowledge about the preferences of future civilizations. To do this reliably, we need to have some model of and distribution over what these preferences are likely to look like (albeit inevitably extremely uncertain) and then the question becomes: are there ‘objective’ facts about what might make some historical information interesting to a future civilization or not?

Information theory already has a well-developed apparatus for dealing with these kinds of questions – rate-distortion theory. The idea here is that if the information-rate of a process is higher than the information-rate of your channel, you cannot perfectly transmit everything. Instead, you have to choose how to approximate the signal, and how you do this is a decoder-specific question, which is encapsulated in the ‘distortion function’. This is essentially for each bit of information, how much do you care about losing this bit? This obviously depends on the problem and your objective. In many cases, you can lose a huge amount of bits and still approximate the process very well. If you only care about certain aspects of the approximation, then you can lose even more. Future historians thus necessarily run rate-distortion encoding on the historical information they receive. Moreover, there is not one set of future historians but rather future history is a decentralized process evolving over time with different selection mechanisms occurring at different times. To persist into the deep future, historical information must survive under the integral of many different rate-distortion curves.

This makes understanding the distribution of interests/selection functions of future historians critical. Future historians could care about very different things than us. This should not be surprising. Today’s historians care a lot about things which previous historians did not. For instance, classical historians focused heavily on military and political history – who was king when, which wars were fought, etc., – and cared very little about things like social and technological history and thus did not record things like the daily mode of life for most people, or about technological advancements outside the highly prestigious canons of philosophy, literature, and mathematics. Instead, almost all the information we have about these aspects of classical history comes from either random asides in the recorded textual histories, or else from archaeological findings which preserve random evidence in a way that is decorrelated from the judgements of ancient historians. For instance, almost no ancient historians studied things such as the nutrition, average lifespan, or rates of injury of ‘normal people’ of their era. However, due to this information being implicitly preserved in bones from burial sites we know a surprising amount about this, while almost all of this information would have been lost forever if we only had written records of the time. More generally, historians and people in general have a strong importance heuristic biasing them towards highly salient events while ignoring completely normal ‘day-to-day’ texture because they are completely normalized, while future historians might think differently.

For instance, future historians might care intensely about things like the precise metallic composition and geographical distribution of paperclips and other office stationery during the 21st century, even though only a vanishing fraction of our total ‘historical’ output is about this. While the paperclip example is somewhat silly, one area which I think a broad distribution of future historians might care a lot about is the phenomenology of people alive today. We expect future historians to have vastly more compute and storage than we possess today and phenomenology is a massive source of bits. Moreover, phenomenology is irreducible in the sense that it contains large amounts of new and exogenous information for a potential future supermind. It is quite easy for the supermind to know explicit historical events and predict things like technological and economic trends. It might be much harder for it to predict or simulate the precise phenomenology of somebody living through these changes. This also puts cryonics into an interesting light from a historical perspective. Even if current preservation techniques are insufficient for reanimation or continuous consciousness of those cryonically preserved, it seems very plausible that a nontrivial amount of information is nevertheless preserved with some degree of fidelity, and that future historians might be very interested in this information. I.e. even if they cannot perfectly revive the full person, they might be able to read certain memories or discover certain aspects of personality or learnt motor skills that could be extremely historically informative for future civilizations, in the same way that we know all kinds of interesting things about the ancients from their bones and other preserved remains, while nobody at the time thought these remains would contain and propagate useful historical information.

This perspective provides interesting implicit advice for archivists and others trying to make life easier for future historians. Specifically, the best archive is relevant to the query distribution of future historians. If we know that all future historians only care about one thing, we should make sure to archive all of that and only that. However, in practice, we are deeply uncertain about the query distribution of the future. Because of this, it makes sense to be very broad about the kinds of information that are preserved. The future can find uses for things we might never anticipate. We should try to preserve both highly structured information that we think is useful and that future historians will think is useful and, crucially, a wide spectrum of raw information that we might think is initially useless. While obviously we cannot all the raw historical information, since we would run out of storage capacity, we can try to provide a strong uniform sampling over the raw information we have available to us. This ultimately becomes a question of query robustness – we must assume some broad (maximum entropy?) distribution over the future ‘queries’ of future historians and design an archiving scheme that is maximally robust under this query distribution. Secondly, if we assume that future historians will have vastly greater computational resources than the present, it implies that archiving should focus on irreducible information and not on storing the result of past computations, since the future will be able to trivially rerun these computations itself if it desired. Rather, the inputs to computation and other information underivable by pure compute/memory alone should be preserved. Crucially, this should also include artifacts that contain information that we cannot yet access. A fun example of this is the herculaneum papyri – carbonized books that were engulfed in the pyroclastic flow after mount Vesuvius’ eruption – which we could only begin to read with modern machine learning techniques a few years ago, hundreds of years after their discovery. More generally, if computation becomes more abundant in the future, this means progressively more information can be unscrambled if it is present in any object at all. What matters is to preserve the information even in scrambled form so that the future can see it.

It is also worth noting at this point how this viewpoint illuminates several different philosophies of historiography. If we zoom out, we get the following picture: First, there is the present state and the flow of time. The present state ultimately generates traces that propagate forwards in time. These traces can be material artifacts, or texts, or cultural information. Over time, increasing amounts of these traces are irrevocably lost due to the entropy-increasing nature of the dynamics10. Present day and future actions can impact these traces and indeed a lot of these traces, especially things such as texts, archives, monuments, museums etc are explicit creations of our own agency. Then, future historians look back upon the traces that have been preserved up till then, and use them to construct their own world models and historical understanding. This builds upon the historical evidence but is inevitably shaped by the processes and priorities of future times which are, from our perspective in the present, unknown. Moreover, a crucial function of a historian and historiographer is to both preserve the traces that have reached us in the present from the past and also perform intellectual work on these traces such as by contextualizing them, coming up with theories, hypotheses, and ideas, and by synthesizing a broad sweep of historical information into narratives, models, and compressed representations. Finally, these representations themselves, if written down or stored somehow, or else if they influence the beliefs and actions of present ages, generate further traces which continue to propagate forwards as well. This makes history and historiography fundamentally reflexive11.

From here, we can see that different theories or ideas of historiography simply illuminate different aspects of this process. First, there is the idea that power deeply influences history. This is an ancient idea and comes back to the aphorism that the victor writes the histories. Our perspective lets us precisely understand this. We can think of power as essentially causal intervention capability in the present. The primary way this can influence later history is firstly, and crucially, by determining which traces can survive to the future. For instance, archives, texts, peoples are all physically instantiated entities that can be either destroyed or preserved. Secondly, to some extent causal interventions today can influence the future distribution of historians which then alters the distribution of future traces these historians construct about the present. However, there are limits to this power – interpretation and material propagation of historical information has always remained highly decentralized and has never, thus far, passed entirely through the causal bottleneck of a single individual agency12.

Secondly, and more interestingly, there is the notion that the historians control history13. When stated like that, it is almost tautological but is surprisingly subtle in practice. The incentives of ‘historians’ as a class often systematically deviate from just pure ‘unbiased representations of history’. This is affected by e.g. political power but is not downstream of it in a trivial way. Intellectual subcultures and the dynamics of intellectual history create their own forces through which historians move. Moreover, we have repeatedly stressed that history is never neutral in an objective sense. The interests and thoughts of future historians are both fundamental and endogenous to historical processes themselves that are impossible to fully predict or control. This is essentially reception theory. Finally, the intrinsic information-structure of both the compression channel and the decoder matter deeply. The information-channel itself primarily preserves information that is preservable. When stated like that it seems trivial, but this is essentially a natural-selection-style argument. Information that is tied to long-lived physical substrate, that is widely copied and elaborated upon, that becomes central to other historical information nodes – i.e. higher ‘feature rank’ in our spectral decomposition, and which is highly compressible in its essentials will get propagated. Information that is primarily instantiated on non-durable media, that does not enter into literary culture of civilizations which produce a lot of traces, that requires very precise incompressible information to be transmitted, and that is not of interest to successive generations of historians, will generally not be propagated successfully.

Finally, and perhaps most interestingly, future historians do not necessarily value arbitrary things. Rather, there are specific points in history that are objectively interesting, central, and easily compressible according to a very broad conception of ‘objective’. Historical events are not only recorded in raw form, but rather the purpose of the historian in some sense is to synthesize these raw data into models and narratives. This allows for both immense compression and generalization beyond the raw historical evidence. Rather than needing to know some specific detail without tying it to anything, instead we can step back and ask meta-historical questions such as what were the causes of X? What general patterns occur across empires in decay? What happens to a feudal aristocracy when it starts to industrialize? and so on. In fact, this is a core reason we care about history at all – to extract generalizable lessons from the past that we can use to shape and improve the present and future. History is not merely an archive kept around for curiosity. History is the mechanism by which past actions and information changes present and future action. A core mode of this kind of model building has been narrative. Rather than a medley of details, it is often much better to compress historical events into a narrative that is both memorable, compressible, and easy to generalize from, even at the necessary expense of detailed correctness in all particulars.

This is essentially Hayden-White’s point about emplotment. History is in some sense inseparable from narrative since narrative is a form of model, and we cannot attempt to understand everything in history without some kind of model14. Raw data alone is never enough; you always need a model. However, importantly, the narratives we choose and the models we build are not and cannot be arbitrary. They are constrained both by the data themselves but also, more subtly, by the structure of the latent space in which they operate. For instance, there are only a few core classes of ‘narrative primitives’ – heroes and villains, betrayals and friendships, the power structure vs the people, and so on. Obviously such narratives must inevitably compress a huge amount of detailed historical information and contingency, but at the same time they are generative. Moreover, crucially, many of these narratives are objective in that they reoccur many times and can be used as first-order models of many different situations – that is, these narrative primitives become archetypal. The reason for these specific primitives is not arbitrary. Rather, similar structures and situations emerge again and again because they are broad basins of attraction in our dynamical landscape of many interacting and learning agents which generate natural incentive structures and information-propagation dynamics which produce them. This is essentially a theory of archetypes – archetypes are highly compressed characters and mechanisms which emerge as natural fixed points of the kind of multi-agent game-theoretic interactions which comprise almost all of the history of our civilizations.

At this point, an interesting question becomes to what extent the future distribution of decoders is endogenous. That is, as historically situated agents, we are not completely powerless before whatever future historian-decoders decide. Rather, through our actions today, we have the potential to shape the future distribution of historians. That is, even history itself is reflexive. Future historians are themselves historically situated observers whose existence depends, in a very real way, upon our actions. The question, then, is how many bits of information can we (a): send through our one-way channel of the historical transition operator to impact the future historian distribution and (b): how many deliberative value-related control bits can we send. That is, can we shape the distribution of future historians knowingly. The second is much, much harder than the first. In any highly non-ergodic chaotic world the initial state can have strong impacts on subsequent states and ultimately the distribution at distant future times. However, being able to send specific bits to realize certain preferences over the future historian distribution is much harder. In fact, naively, these two bitrates even seem in tension with each other. If the dynamics of the historical transition operator are truly chaotic and exquisitely sensitive to initial conditions, then we have a very strong causal channel to project bits into the future, but we cannot realistically predict the consequences of those bits. Thus we cannot meaningfully optimize to achieve high ‘value’ far into the future by our actions today. However, if the universe is very orderly but hard to budge, so that almost all variation decays away over time, then we might be able to send very few meaningful causal bits down the channel, but we would be able to predict what these bits entail in terms of consequences and thus have nontrivial control over our future value distribution. This is not specific to future historians; rather, it is a general principle of bounded intelligent agency. The amount of bits of ‘optimization pressure’ an agent can apply to the future, on average, is bounded by the amount of bits it has access to which can be used to predict and discriminate the value of future consequences.

If we assume, sensibly, that the consequences of our actions become harder to predict as we extend the time horizon, then even for an agent which theoretically cares equally about the distant future, a special kind of myopia is still relevant. If all utility-relevant consequences of all actions collapse to a maximum entropy distribution by some time t then the utility beyond this becomes equal for all actions, leading to effective myopia even if we theoretically value different states equally in time. This leads to another interesting consequence – that if actions differ nontrivially in their maximum entropy horizon then predictable actions become more valuable, often quite strongly, with actions that predictably exploit the nonergodicity of the historical transition operator by pushing the state-space permanently one way or the other become supremely valuable.

Super naively, we can model this as two quantities per action, the causality-decay half-life and the predictability-decay half-life. Generally, we expect the causality-decay half-life to be necessarily longer than the predictability-decay since if actions have no causal impact then it is meaningless to predict them beyond that point. We can then think of the causality-predictability decay gap. If this gap is large and the causal decay is low then myopia becomes rational and generally it is hard to optimize for the long term meaningfully even when actions causally propagate. If both causal and predictability decay fast – i.e. the universe almost immediately relaxes to whatever predetermined state it has – then long-horizon optimization is generally pointless. If causality-decay is slow and predictability-decay is also slow – i.e. the gap is small – then we are in a very interesting situation where long-horizon agency is not only possible but overridingly important for non-myopic agents. Crucially, these causality and predictability decays are not uniform across the state-space. In the vast majority of the state-space they are very rapid. For almost all actions in almost all times and spaces, the meaning of any action will dissipate incredibly rapidly. However, rarely and largely (but not entirely) unpredictably, the world arranges itself such that certain actions can resonate far into the future. The transmission channel is open and receiving, and the eigenmodes stretch far indeed. We should expect this to happen primarily around points of non-ergodic branching in the state-space. If a decision is irreversible and causes large-scale change then this pushes history one way or another and this change then becomes reinforced by the rest of the dynamics. One falling rock produces an avalanche, even if the kinetic energy of the avalanche is vastly disproportionate to the energy of the initial rock. Another way this can happen is through evidence preservation. Almost all historical acts leave no preservable trace. Even mundane actions that leave historical traces can redound for a long period, albeit in highly unpredictable ways. The question then moves up a meta-level – how can an agent realize when these are the situations with long vs short historical half-lives and then what exactly does each bit it is transmitting actually do to the future? These are intrinsically very challenging and interesting questions, and of course they are essentially unsolvable in detail. The future remains intrinsically unpredictable. Historiography, especially predictive historiography, is hard.

Nevertheless, it is also interesting to think about things again in terms of coding theory. Suppose an agent can inject certain bits through a historical channel – how can it encode these bits? First, again, one strategy is increasing the temporary bit-rate of the channel – simply transmitting more information, producing more traces, preserving them for longer, etc. Secondly, the agent can attempt to transmit through amplitude. Certain events and actions are more important or striking than others in a near-objective sense. ‘Larger’ actions generate more traces, may cause more path-dependence, and so on. Here, coming back to the idea of the fractal spectrum of history, the goal is to raise the ‘feature rank’ such that it becomes a more important feature in the spectral space. Another approach is error correction and redundancy. Don’t send many bits, rather send only a few bits which are highly compressible, redundant, and using an error-correcting encoding so that their meaning cannot easily be lost. For instance, the agent might encode information that is extremely surprising, or has a high likelihood ratio such that it is very strongly distinguishable from the baseline, and which can be repeated in many redundant ways.

This then all loops back around to individuality15. Suppose an agent desires not some specific causal intervention to be propagated but rather for some number of the specific individuality bits it has to be propagated into the future, what does this imply? The same coding strategies apply directly. Increase the information bandwidth to the future, preserve as many traces as possible including, ideally, the full phenomenology, increase amplitude and hence salience to future historians under some necessarily imperfect reflexive model of the query distribution of future historians, make the crucial bits redundant and error-correcting when possible, of course.

However, zooming out a bit from here something very interesting happens at the macroscale with regard to individuality. We discussed the fractal hierarchical structure of the historical state and its feature spectrum. Crucially, here, if we think about the most ‘important’ predictive features they are almost always extremely coarse-grained. ‘the industrial revolution’; ‘abiogenesis’; ‘The invention of the wheel’ and so on. These have lost almost all individuality; they are determined almost entirely by the structures of the universe. Similarly, even at the narrative level the coarsest-grained ‘features’ in the literary latent space are archetypes which, again, are almost forced by the intrinsic structure of the latent. Archetypality systematically strips individuality16. More broadly, we observe that individuality primarily lives towards the leaves of our tree, since the leaves preserve the particularities which have low predictive information. The coarse graining aggregates the predictive information of many particulars but must strip them of precisely these particulars – this is intrinsic to the essential nature of the abstraction and generalization that this model of feature learning carries out. Importantly, then, if we think of the historical transition operator as primarily pruning from the lower levels of the tree, as it must do, since it operates essentially as a low-pass filter, then historical information loss is disproportionately deleting individuality from the historical record. Individuality is thus intrinsically fragile information. You do not need to store much historical information to know that ancient Egypt existed. You need to store a lot to understand the detailed individual phenomenology of some particular ancient Egyptian. This implies that the integral of the duration of any specific individuality is inevitably very low, even if individuality continues to persist into the limit. In effect, history must compress and thus turn individuality into abstraction. Perhaps the most interesting possibility for the future here is that an uploaded K2+ civilization will enable vastly higher information bandwidth for information to persist into the future, thus raising the ‘channel capacity’ of the historical transition operator sufficiently to enable meaningful individuality to persist for long periods whereas at present it almost uniformly dwindles instantly. This, I think, is an underrated thing to think about when we are considering the deep future. Namely, that will be a long period of vastly increasing historical capacity where our current information transmission bottleneck is removed, but before the actual long-term limits to storage and retrieval start to bite. During this period, vast amounts of history will be produced and persist for far longer than has ever been possible before.

  1. At the lowest level, physics still primarily has time-reversibility – i.e. unitarity of quantum state transitions. However, in practice, even if the information is theoretically preserved by physics, the scrambling nature of physical dynamics mean that recovering that information requires increasing amounts of precision in the knowledge of the state, which must thus eventually be lost. 

  2. Aestivation might complicate this a fair bit since there the constraint is cooling and hence using minimal energy. In this case, passive storage media is totally abundant since you only ever want to use a vanishingly small fraction of your total matter endowment for computation such that it can asymptotically cool to the extraordinarily low temperatures required for maximum efficiency. 

  3. Technically, the important quantity here is actually the entropy rate. If the initial state is known and all dynamics are deterministic, then technically the entire historical trajectory can be predicted perfectly ahead of time so there is no increase in ‘historical knowledge’ every year. In practice, this kind of prediction is not possible, and hence the true historical state grows over time, but the amount of information by which it grows is technically dependent upon the state of the universe. I.e. suppose large amounts of the universe were locked in some kind of stasis, then by definition there would be less historical information accumulating over time. Nevertheless, the one year per year is a decent approximation. 

  4. The biggest problem for our hypothetical K4 civilization at the end of the universe might not even be storage but retrieval. If you want some specific very detailed historical event, the raw information is probably stored somewhere. But actually finding the right information, amongst the uncountable trillions of other stored historical evidence would become nontrivial. Also, the speed of retrieval could be slow due to simply information storage limits meaning that information must be distributed across space meaning eventually nontrivial travel time of orders of thousands or millions of light-years to find the correct archive. Almost all vital information would be redundantly copied locally but if you really want to push out to one of the furthest leaves of historical information, the retrieval would become highly nontrivial. 

  5. Interestingly, this provides a strange but compelling (to me) critique of the paperclipper and ‘tiling-the-lightcone-with-hedonium’ type worlds. These worlds have become extremely simple from an information-theoretic and individuality-centric sense. They produce no meaningful history. As more and more of the universe is converted into paperclips or hedonium, the descriptive complexity of the world does not increase. We simply change a fixed scalar in our ‘fraction of the universe converted to X’ meter. This seems deeply unsatisfying and I think it relates to the fractal information structure of the world we discussed in whence the fractals. These ‘tile-the-universe-with-X’ proposals all, importantly, move us off the fractal information curve to a world where all informativeness concentrates at the head nodes rather than being spread approximately equally throughout the tree. In such a world, there are only a few latent variables that matter, principally the nature of the thing being tiled and the extent of the tiling, and that is it. The learnable information about this universe has collapsed. There is no meaningful intermediate latent structure between the highest level of what is being tiled, and the lowest level about the arrangements of the hedonium atoms. This seems intrinsically bad, at least to me, which itself is interesting. It tells us that our intuitions about goodness are not macro-structure independent in an important way and do not cleanly aggregate in the utilitarian way. Even if some hedonium is good, it does not mean we get linear returns on more hedonium indefinitely until the entire lightcone is tiled with hedonium. 

  6. This assumes, of course, that the civilization has some pure compression/information-focused objective which can allocate utility indefinitely to new but rarer historical features since if a civilization does not care about history or only cares about some very specific finite subset then the rest is irrelevant for this. This is identical to how NTP/reconstruction/compression objectives are incentivized to prefer all features but e.g. classification objectives can be bounded by the difficulty of the classification task. Even if the data has structure, it is the objective-relevant information that determines the scaling laws. 

  7. Often, surprisingly, for archaeological preservation wars are often a good thing since when cities are burnt down a huge amount of fragmentary physical evidence is preserved in the ash. Similarly, almost all of our surviving cuneiform tablets are explicitly fired which hardens the clay so that it survives down to the present. 

  8. Technically, since the transition operator is nonlinear, it does not directly have eigenmodes. However, using Koopman theory, we can instead map the finite dimensional transition operator to an infinite dimensional linear kernel over ‘observables’, rather than the state directly, which does have eigenmodes, and we can think about the decay rates precisely in these terms. 

  9. While today almost all historical information is rapidly lost, future civilizations could deliberately engineer their civilization to preserve vast quantities of historical information over time – i.e., increase their own information bandwidth to the future – such that they will naively run up against the storage limits. As an example of this, just consider the vast increase of digital information available since the creation of the internet with all of recorded history prior to that point. Likely, the information produced this year exceeds all information produced before the 21st century. This does not mean, of course, that contemporary information is more important than all prior centuries. Rather the opposite. However, if this trend continues, and as future civilizations will exist almost entirely within vast deliberately designed computational edifices, we should expect this rate of historical information storage to increase dramatically. This likely includes things like maintaining large archives of historical information, making redundant copies of important information on durable media, alongside various self-contained potential decryption keys and format data for the information, and generally setting things up such that information is preserved sensibly through time. Of course, coming back to our 1-1 map issue, a very large fraction of information must still be lost, but a much much larger fraction of ‘important’, however defined, information will likely be deliberately preserved by such a civilization. One especially interesting avenue where future civilizations can improve over our own is the ability to directly store copies of minds with full phenomenology and episodic memories archived at a particular time. 

  10. In fact this is essentially guaranteed by the data-processing inequality. Information about the initial state variables can only be lost and never increased through transformations. 

  11. This property of reflexivity seems important in differentiating what Dilthey called the ‘human sciences’ from the material sciences. Human sciences are intrinsically reflexive. They study systems of agents which can themselves both understand the system they are in and also understand and be affected by the theory of their own domain. People read history and economic theory and then change their behaviours as historical or economic objects. This is very different to pure sciences where we don’t have e.g. rocks reading geological theory and then changing their behaviour to either conform to the theory or deliberately refute it. This makes these subjects more challenging because of intrinsic nonstationarity. The object of study changes as you study it. 

  12. This is perhaps another very interesting aspect of the singleton. The singleton is assumed to have immense power over the present, perhaps extending to comprehensive causal power over historical information transmission. If this is the case, then even if the singleton itself, somehow, does not persist, all information and interpretation future historians post-singleton could work with might necessarily pass through the singleton. That is, the singleton might be able to completely determine its future distribution of historians in a way that is infeasible for any bounded agent today or in the past. In this way, the singleton could implicitly control the future even without persisting into that future. 

  13. This thesis is proposed in a more memorable, but more opinionated, way by Scholar’s Stage here

  14. It is important to note here that the historical narrative selection operator does not actually necessarily reflect the true causality super well. Many things are humongously causally important but are not at all recorded historically. Similarly, historical preservation can spend disproportionate ‘bandwidth’ coding information that is not that causally relevant. A lot of our explicit knowledge goes through fairly strict explicit dramatization filters where historians describe information that is narratively interesting to them and which identifies causality with a few compressed players rather than the actual much messier more decentralized actual causality. 

  15. This also relates to a point from JDP about identity becoming easily predictable. There remains a very interesting, and somewhat disturbing question, about how much irreducible information is actually present in a given person’s personality. The answer here could simply be not that much. General personalities cluster into predictable types, as shown by the Big Five and other personality methods. Potentially individuality itself has the same fractal information structure we have been discussing before. The largest contributions to individuality such as Big Five personality types are not actually individual at all – they are samples from a well-known manifold of possible types. Again true ‘individuality’ is pushed further down the slope of feature rank, until we are resolving much finer-grained features than extravert vs introvert, conscientious or neurotic, and so on. But then this implies that individuality scales essentially inversely with predictive information – there are many bits about you that distinguish you from others, but these bits can only detail increasingly fine-grained information which by itself is not that predictive. The broad predictive factors that are easy to infer are actually not individual at all – but highly general variance across large populations. In some sense this might be inevitable – if minds span a roughly shared subspace, the largest degrees of variation of this subspace are the easiest to observe and model, and are also thereby the most predictable. In fact, this effect is almost tautological: the information that makes you individual is that which cannot be predicted from the outside. However, much behaviour is actually very predictable, on average, from the outside. Thus, the core predictive features of your behaviour cannot be particularly individual at all. This may be a flaw in our definition of individuality or it may gesture at a deeper truth that we should not expect to find meaningful individuality at the coarse-grained level at all. 

  16. This rapidly leads to an interesting direction of thought on literature specifically because good literature itself embodies the same fractal information spectrum. It uses archetypes at the coarsest level to rapidly emplot its meanings, but at the same time a large part of the point of good literature is to realize and interrogate a specific detailed phenomenology of a set of characters – that is, it deliberately preserves irreducible individuality. We can almost think of literature as generating synthetic individuality.