Competency in Navigating Arbitrary Spaces as an Invariant for Analyzing Cognition in Diverse Embodiments
The notion of action spaces shows how to connect goal-directed activity to notions of energy (e.g., of evolution toward an attractor). For biological systems at any scale, the relevant attractors are those that implement allostasis within the current environment. The variational free energy (VFE) principle formulates this requirement for allostasis in information-theoretic terms: living systems behave so as to maximize their ability to predict their own future states. When we identify VFE with uncertainty and hence with (the probability of) prediction error, the minima of VFE become the maxima of predictive success. Uncertainty, and hence VFE, is distinct from metabolic load. Hence, we can ask about the metabolic cost of achieving an increment of predictive ability. The metabolic cost of generating predictions through the use of some computational model of the environment emphasizes that moving through the search space really is a search. It takes effort and resources, including memory resources. As cells or other systems become stressed (e.g., by lack of sufficient free-energy resources), their generative models can be expected to deteriorate toward stochastic defaults, and hence their searches can be expected to deteriorate toward random walks. This has been observed at the cellular level [158] and is a commonplace observation in stressed organisms, including humans.
The act of navigating a state space, whatever its degrees of freedom, involves several fundamental components. First, there is the inverse problem of which effectors to activate to reach a preferred region of state space. It frames most aspects of survival as, fundamentally, a search with different degrees of capacity to look into the future. Second, it is greatly potentiated by the ability to maintain a record of the past (i.e., a reliable, readable memory). The central question faced by any system is what to do next among certain choices. Thus, it is important to begin to formalize a notion of decision making in a deterministic system. No finite agent can discover all of the causal influences that determine its own behavior, as global determinism logically disallows local determinism. Hence, global determinism assures “free will” from every (finite) local perspective [159]. The failure of theories modeling human decision making on what human agents can consciously report about their thought processes makes the need for a general theory evident [160,161]. The mathematical theory of active inference as Bayesian satisficing provides a scale-free framework for understanding this process. Given that spaces are an essential invariant for understanding biological adaptive activity, it is important to ask how these spaces originate.
Biological systems at all scales exist and maintain their integrity in active exchange with their environments. As first shown by Ashby [162], this exchange can be formalized as an exchange of information. From this formal perspective, organisms act so as to minimize the VFE of their interactions with their environments. While the FEP was originally formulated as a theory of brain function [163,164,165], it has since been applied to a wide range of biological systems and processes [27,142,166,167,168] and was recently shown to characterize any classical [169] or quantum [170] system that is sufficiently stable to be identifiable over macroscopic time. Allostasis is, in other words, not limited to biology; it is a general characteristic of all systems that resist the entropic forces of their environments long enough to be observed at multiple times. Indeed, the emergence of the structural complexity required to sustain allostasis can be seen as being driven by the environment as a means of producing entropy [171].
Maintaining allostasis is maintaining a distinction between “my states” and “my environment’s states,” where “my environment” here is everything other than me. A state is a collection of values of some degrees of freedom, or state variables. Position, temperature, viscosity, concentrations of various molecules, and electrical charge are all state variables of relevance to all organisms. The internal-external distinction can be formalized in terms of a Markov blanket (MB), a set of intermediate states that serves as an interface between inside and outside [27,172]. These interface states transfer information from outside to inside (i.e., implement perception) and from inside to outside (i.e., implement action). An organism and its environment share, by definition, the same MB; they merely “look at” different sides of it. Position states on a MB implement the “physical boundary” of an organism, with the cell membrane or the skin as examples. This boundary is, from the organism’s perspective, also the physical boundary of its environment. Most MB states encode values of variables other than the position (e.g., photon intensity and frequency (brightness, color, radiant temperature, etc.), air pressure (e.g., wind velocity and sound), or molecular concentrations (e.g., osmolarity, smell, and taste)).
As all information exchange between a system and its environment passes through the MB, an organism’s perceptions “of its environment” are, from a mechanistic perspective, data encoded by its environment on its MB (Figure 7). An organism’s actions “on its environment” are, similarly, data it encodes on its MB. Indeed, we can view an organism’s actions as its environment’s perceptions and vice versa. An organism has no access, even in principle, to the mechanisms by which its environment encodes data on its MB. This restriction on access is fully symmetrical; the organism’s environment has no access to how the organism encodes data on the MB. (Indeed, an organism is, by definition, “the environment” of its environment.) While such statements are sometimes considered “anti-realist” or “subjectivist” [173], they are just consequences of modeling physical interaction as information exchange [174,175]. Organisms such as humans that employ technologies to extend their perception and action capabilities are, effectively, extending their Markov blankets to encode the values of additional state variables.
When information flow is restricted by an MB, the task of minimizing the prediction error and hence minimizing the VFE becomes the task of predicting and then acting to regulate the future state of the MB. The good regulator theorem [176] requires any system capable of such regulation to be or to encode a generative model [27] of its environment’s actions on its MB. The state variables of this model are the state variables of the MB, the only state variables that can be either measured or predicted. The generative model encoded by an organism is thus the organism’s “theory” of its environment’s observable behavior (i.e., its environment’s actions on its MB) and includes, most importantly, its theory of how its environment will respond to each of its own actions on its MB.
It is important to emphasize that, as shown in [169,170], these considerations apply to all physical systems at all scales. While here, we will be concerned primarily with individual organisms, Markov blankets as system–environment interfaces and VFE minimization as an inferential mechanism characterize all systems identifiable as such over time, including macromolecules, biomolecular pathways, individual cells (whether free-living or components of multicellular organisms), organs and tissues, individual organisms, communities of organisms, ecosystems, and even larger structures. Indeed, the authors of [177] showed how to model the global climate system by minimizing the VFE across an MB. We can, therefore, consider MBs to be universal, scale-free structures and VFE minimization to be a universal, scale-free mechanism. Hence, MBs and VFE minimization are invariants that characterize all forms of behavior in all “spaces” occupied and explored by organisms.
We are now in position to define the spaces in which an organism operates. Suppose an organism’s MB encodes at most m distinct values of each of n distinct variables, and let N = nm. This number N is finite for any finite system (i.e., any system with finite energy resources and hence a finite measurement resolution). Any state of the MB can then be considered to be a vector in an N-dimensional vector space constructed by assigning a basis vector to each of the N variable-value combinations and adopting the standard notion of distance between vectors as a metric. Such vector spaces are called Hilbert spaces and are widely used in quantum theory. They can also be employed in classical physics. Any organism—indeed, any physical system, classical or quantum—can be considered to behave in the Hilbert space that characterizes its MB. A perception-action loop is, in this formalism, simply a mapping from an “input” vector representing the state of the MB at some instant t to an “output” vector representing the state of the MB at some later instant t + Δt. As this is a well-behaved map between vectors, it can be treated as linear independent of its implementation. It is this “hiding” of implementation details that renders MBs such useful theoretical tools. They can be thought of as defining application programming interfaces (APIs) around physical systems that specify the data structures that interactions must respect. The consequences of this are discussed further below.
For any biological system, the number N is enormous, and the complexity of a predictive (generative) model of an N-dimensional space increases combinatorially with N. Hence, organisms cannot be expected to implement full, predictive models of the Hilbert spaces of their MBs. Indeed, a model of the full Hilbert space of the MB is impossible in principle. The MB is, by definition, the sole interface in the joint system–environment state space between any system and its environment. Hence, some fraction of the states of the MB of any finite system must be allocated to free energy acquisition and waste heat removal [178]. The sector of the Hilbert space allocated to these thermodynamic functions is observationally inaccessible to the system; its sole function is a thermodynamic one. Hence, the MB of any finite system can be regarded as divided into at least three sectors, comprising sensory, active, and thermodynamic states as shown in Figure 8.
Because MB states cannot be modeled completely, organisms, including humans, instead implement partial models of sets of variables that have been observed to covary systematically. The positions of objects, for example, covary as an organism moves. A model that captures this covariance is a model of an ordinary 3D space. Concentrations of environmental chemicals also tend to covary; the space of chemical concentration gradients is the primary “space” in which chemotactic microbes operate [179] and is an important space for all organisms equipped with olfaction and taste. Organisms can, in general, be expected to optimize the use of their limited information-processing resources by limiting their generative models to just the principal components of their experience and segmenting these models into “spaces” spanned by covarying principal components.
Segmentation of an overall state space into predictively tractable subspaces is greatly facilitated by the fact that tractable subspaces tend to exhibit relatively simple symmetries. The most familiar are the translational, rotational, and relative motion symmetries of objects in a 3D space. Moving an object in a 3D space does not modify its properties or change its identity. These symmetries are described mathematically by the Galilean group in classical physics and by the Poincaré group when special relativity is taken into account. Hence, a generative model of spacetime is, at minimum, a representation of the Galilean or Poincaré groups. State variables that can be described by fields in spacetime (e.g., electric or magnetic fields) satisfy gauge symmetries, which prevent the state of the field from depending on how it is measured. Using a quantum theoretic framework, it can be proven that any state variables encoded on a Markov blanket must satisfy the gauge symmetries if they are represented as a field in spacetime [180]. See [143] for an application of gauge-theoretic ideas in neuroscience.
Symmetries create redundancy, and many different descriptions of a symmetric situation encode the same information. This redundancy enables data compression and coarse graining. It also makes any space characterized by symmetries an error-correcting code. Information that may be missing or ambiguous at one “location” in the space can be found at other locations [181]. Redundancy enables “babbling” as a strategy for discovering the symmetries of a space. The language and motor babbling of infants are a canonical example. Babbling can be considered a heuristic search strategy in which “random” actions are deployed to investigate the large-scale structure of a space, and more directed minor variations of “interesting” actions are used to investigate the local structure (for comparisons of human infant and developmental robotic implementations of babbling, see [182,183]). Such alternation between breadth-first and depth-first searching in a space is an ancient strategy of living systems, going back at least to the run-and-tumble behavior of chemotactic bacteria. The same strategy (in effect, babbling in a 3D space instead of in a linguistic space) has been shown to be very effective in robotics, enabling the building of adaptive robots that develop models of themselves and strategies to navigate the world de novo [184].
Problem spaces are defined by observers, as they make models to help explain, predict, and control other systems. Crucially, however, the system itself is also such an observer [185] and generates models of spaces to help guide activity. In humans, the personal past is such a space, as increasingly more detailed studies of the construction of episodic memories demonstrate [186,187]. A very fundamental way for even simple observers to generate the notion of spaces ab initio is from the commonalities between actions required to nullify changes in sensory experience. Actuations that result in predictable changes in sensory states can often be naturally represented as “movement” in a space [188]. As discussed above, an obvious and well-studied example is “babbling”—both vocal and motor—in human infants, a phenomenon also common in other animals [189] and increasingly employed in developmental robotics [182,190]. Such space construction is closely linked to the very basal capacity for homeostasis (keeping a sensory state in a constant range is just one step past keeping a specific variable, such as the pH level, in the right range), but this may involve feedback and monitoring of the result of one or more layers of processing past the raw sensor. This loop immediately provides the opportunity to scale intelligence via optimized user illusions (models) of spaces [181,183,191,192] because there are many levels of sophistication available to the overall project of keeping one measurable optimized by taking various actions. This scheme is not only about actuating muscle or ciliary motion to keep a constant relationship with a spot of light, for example; it works in other spaces, too. The barium-exposed planaria are looking for moves in transcriptional space that allow them to keep their normal physiological states.
Similarly, somatic tissue develops and monitors representations of its own anatomical layout, using bioelectrics in epithelia as a kind of “retina” that perceives the body structure [193,194] and is able to trigger movements in the morphospace to counter the induction of incorrect layouts in the large-scale morphospace (such as the fact that tails grafted onto inappropriate locations in amphibians will become remodeled into limbs, a structure more appropriate to the new location [195,196]). When brains developed, they retained and amplified the notion of modeling the self in anatomical space via the somatotopic homunculus [197].
Separating the overall (Hilbert) state space of the Markov blanket into predictively tractable components has the effect of breaking the overall prediction problem—the problem of minimizing VFE—into tractable components. As these component problems involve spaces with, in general, different symmetries, the most efficient methods of solving them will, in general, be different. They can each, in particular, be expected to involve data structures (“representations”) that encode the symmetries of the relevant space. These data structures are, in turn, encoded by the Markov blanket. Formally, they are the basis vectors of the corresponding Hilbert space. Maximizing efficiency (i.e., minimizing the resource requirements of information processing) requires that perception and action both employ the same data structures and hence respect the same symmetries. Hence, perception and action in any domain can always be viewed as acting via a particular, domain-specific component—a particular subspace with its own basis vectors—of the overall Markov blanket.
A perception–action module that imposes a particular data structure can be considered to define a reference frame, and when such a module is physically implemented by a finite system that consumes energy and dissipates heat, it becomes a quantum reference frame (QRF) [197,198] (see [170,199] for discussion in a biological context). The most familiar QRFs are artifacts, such as meter sticks or clocks, that we humans use to make external measurements. Employing such artifacts to measure distance and time, however, requires an internal sensory representation of distance and time. A person with no ability to sense duration, for example, could make no sense of a clock [200]. Hence, biologically implemented QRFs underlie the use of all artificial QRFs. Any pathway that employs a fixed (or only slowly varying) reference point (e.g., the midpoint of a sigmoid activity curve) to switch some behavior on or off can be considered a QRF. The use of the [CheY-P]/[CheY] concentration ratio to control the direction of flagellar motion in chemotactic bacteria provides an ancient example.
A perception–action module can compare perceptions and hence regulate actions only over the timeframe of its local memory. Maintaining a local memory requires energy. The [CheY-P]/[CheY] ratio, for example, is maintained by enzymatic activity and hence by metabolic activity. Selective pressure to minimize VFE is, therefore, selective pressure to expand the memory capacity (i.e., to allocate increased structural and energy resources to storing information about the consequences (in context) of past actions). As obtaining additional resources from the environment may require sensing and acting on the environment in new ways—from predation to social or economic exchange—VFE minimization can be expected, in general, to drive the development of new QRFs, with the elaboration of progressively more complex visual, auditory, and olfactory systems in lineages subject to different selection pressures as obvious examples. Increasing the information processing capability is, therefore, inevitably a positive feedback loop and hence effectively an arms race with selection pressures from the environment.
The Heisenberg uncertainty principle famously limits the simultaneous use or co-deployability of some pairs of QRFs (e.g., those for position and momentum) at high measurement resolutions. Interference between the measurements of degrees of freedom assumed a priori to be independent, generically termed context effects, can be generated even in classical systems [201] and can always be attributed to failures in commutativity (i.e., interference) between QRFs [202]. Competition for energetic resources between QRFs also limits co-deployability. Systems respond to limits on co-deployability by developing attention systems that prioritize both perceptions and actions. By serving as a resource allocation mechanism, attention itself becomes a resource.
When we perform experiments on a system, we are acting as part of that system’s environment. Our actions on the system and our measurements of its behavior depend on our QRFs and hence on the spaces in which we operate. Our inferred explanations of the system’s behavior become components of our generative models, which we test by testing their predictions.
How can we, from this position outside of the system’s Markov blanket, determine what QRFs the system is deploying and hence determine the spaces in which it operates? It is clear from the definition of a Markov blanket and a generic result of quantum information theory [178,203] that no such experimental determination can be made. The best that can be accomplished is an empirical model of the system’s QRFs and hence of its operating spaces, developed within the language imposed by the experimenter’s QRFs. Even biochemical pathways, from this strict perspective, are theoretical models based on evidence that may be limited or ascertainment-biased by the experimental procedures employed. The science of QRFs is, in other words, subject to the same fundamental limitations of any other science. It is greatly facilitated by building models of the system of interest as embedded in and communicating with its environment and then examining these models at multiple scales. The mammalian hippocampus, for example, functions in part as a spacetime QRF at the scale of the whole organism but functions as a pulse correlation generator at the scale of the local networks to which it supplies inputs [204].
From an operational perspective, probing a system’s QRFs is an exercise in reverse engineering, inferring a “design model” that meets the goals of functionality and efficiency from experiments that probe structure (to the extent that it is observationally accessible) and overt behavior. Inferring the representations and hence the data structures employed by an organism to process and act on information from its environment is, effectively, inferring the API of a computational system for which only the input-output behavior and external resource usage are initially known. While a recognizable hardware architecture can contribute useful information to this process, it places few if any constraints on the software architecture and hence on the structure of the API. The use of experimental methods modeled on those of cognitive psychology is the present state of the art for reverse engineering the functions implemented by deep learning systems following training [205]. We may increasingly expect the same to be the case in biology.
All complex biological systems are hierarchical; macromolecules are organized into larger-scale structures and pathways, which are organized into functioning cells, which form tissues, organs, and eventually whole organisms, which are then organized into societies and ecosystems, etc. A crucial aspect of this is that the hierarchy (modularity) is not simply structural. Each level contains its own agency, with agendas in various appropriate spaces. The robustness and plasticity of life may be due to the unique and powerful ways in which the lower levels’ activities (microstates) are harnessed toward the higher levels’ goals. Agents at higher levels (e.g., organs) deform the energy landscape of actions for the lower levels (e.g., cells or subcellular machinery), such as the example in Figure 3I. This enables the lower systems to “merely go down energy gradients”, or to perform their tasks with minimal cognitive capacity while at the same time serving the needs of the higher-level system, which has exerted energy (via rewards and other actions) to shape its parts’ geodesics to be compatible with its own goals. This is a very powerful aspect of multi-scale competency because the larger system does not need to micromanage the actions of the lower levels; once the geodesics are set, the system can depend on the lower levels to do what they do best: go down the energy gradient. The paths of least action in any space are implemented by the paths of least action (i.e., VFE-minimizing paths) in the overall manifold of the internal state probabilities. Recent work in Bayesian predictive processing shows how an agent’s information geometry is distorted by beliefs, a kind of gauge theory [143] that tightly links the notion of action in an arbitrary space to the cognitive state of the agent as an invested observer.
For example, the meta-cognitive level of attention can be seen as setting the precision (curvature of the free energy) for another part of the internal model. It is this set of multi-scale relationships, with parts deforming subparts’ action spaces toward goals in their own space, that distinguishes flat (single-level) systems that simply minimize energy (e.g., water flowing down a hill, which people do not think of as an action or decision) from ones that use the same kind of physics in a more obviously cognitive manner.
As noted above, systems at every level in such hierarchies can be described as performing active inference with the goal of minimizing environmental VFE, where the environment is everything other than the system. Systems at any hierarchical level can, in other words, be considered to deploy generative models of the behaviors of their environments to interpret what they perceive and to employ these same models to act on their environments in return. Biological hierarchies are, moreover, not just structural; they are also functional. How do actions or functions at one level affect the actions or functions at other levels? It is to this question that we now turn.
Consider an amoeboid cell. Actions in the macromolecular state spaces that define the genome, transcriptome, and proteome (e.g., expressing an actin gene) enable actions in the morphological space (e.g., pseudopod extension), as shown in Figure 9. The macromolecular actions are carried out by macromolecular complexes, in this case transcription, mRNA processing, and translation systems. The morphological actions that they enable are carried out by much larger-scale structures, in this case spatially organized associations between the cytoskeletal components and mitochondria. These organelle-scale actions in turn enable cellular scale actions such as environmental exploration and predation. Bottom-up enabling relations such as these have top-down counterparts; predation enables metabolism of prey components to yield usable free energy, which in turn enables macromolecular actions such as gene expression.
We have shown here how MBs and VFE minimization define mutually interdependent behavioral spaces at multiple scales and then organize behavior to maximize predictability, hence maximizing the probability of continuing allostasis within those spaces. This analysis leaves open, however, the question of how either MBs or VFE minimization are implemented in the vast variety of biological and increasingly hybrid biological and artificial systems to which we have experimental access. Hence, there remains a large number of further areas for conceptual development as well as empirical capabilities that should be investigated, which include the following.
While higher-level systems bend action spaces for lower-level subsystems, it can be predicted that the higher level no longer needs to operate in a very rugged space of microstates. Instead, evolution can search a coarse-grained space of interventions, which also includes changing the resource availability landscapes at both the lower and higher levels (e.g., inventing a mouth and a specialized digestive system). Computational models can be created to quantify the efficiency gains of evolutionary search in such multi-scale competency systems.
Links can be made to higher levels of cognitive activity and neuroscience. For example, yoga and biofeedback can be seen as ways for systems to forge new links between higher- and lower-level measurables. Gaining control over formerly autonomic system functions is akin to rerunning causal analysis functions on oneself to discover new axes in physiological spaces that the higher-level self did not previously have actuators for. Such processes clearly depend on interoception, a process for which active inference models are now well-developed [206,207], and being integrated with models of perception in a shared memory global workspace architecture [208].
More broadly, models of space traversal help flesh out a true continuum of agency, placing simple systems that only know how to “roll down a hill” on the same overall spectrum as psychological systems that minimize complex cognitive stress states. Concepts related to free energy help provide a single framework that is required to explain how complex minds emerge from “just physics” without magical discontinuities in evolution or development. The capacity to traverse a space without getting caught in local optima can be developed into a formal definition of IQ for a system in that space. This links naturally to the work in morphological computation and embodied cognition because body shape determines the IQ of traversing a 3D behavioral space. How does this extend into other spaces? Many fascinating conceptual links can be developed to work on embodied premotor cognition in math, causal reasoning, general planning, etc. [209,210,211].
How do cells, both native and after modification via synthetic biology tools, make internal models of their “body shape” in unconventional spaces, such as a transcriptional space? Cells in vitro can learn to control flight simulators [212], as can people with BCIs [213]. Brains can learn to control prosthetic limbs with new degrees of freedom [214]. What self- and world-modeling capacities are invariant across such problem spaces?