Active inference, morphogenesis, and computational psychiatry

4.1. Active inference formalism
30

In statistical physics, the ensuing dynamics are readily described in terms of density or ensemble dynamics, i.e., the evolution of the probability density p(x~), through the Fokker-Planck equation. We obtain the Fokker Planck equation from any Langevin equation using the conservation of probability mass:

31

where x~⋅p(x~) describes the probability current. This turns the Fokker-Planck equation into a continuity equation, which reads:

32

This is a partial differential equation that describes the time evolution of the probability density p(x~) under dissipative (first term) and conservative (second term) forces. At non-equilibrium steady-state, the density dynamics is just the solution to the Fokker Planck equation:

33

In other words, we can express a physical potential or a Lyapunov function L(x~) simply as the negative log probability of finding the system in any (generalized) state L(x~)=-lnp(x~). This is also known in information theory as the self-information of a state (also known as surprisal, or more simply surprise). In Bayesian statistics it is known as the negative log evidence.

34

This function can be bounded from above by the variational free energy function that is the foundation of the free energy principle underlying Bayesian inference.

35

Variational free energy is a function of internal states that allows one to associate the Lyapunov function with Bayesian model evidence and therefore characterize the system dynamics in terms of Bayesian inference and the implicit generative models. The key statistical mechanics tool employed to do this is to unpack the non-equilibrium steady-state flow of external, internal, and blanket states, called a Markov blanket partition.

36

A Markov partition, deriving from conditional independencies in Markov processes, which are implicit in the system's equations of motion or dynamics, separates all states x∈X into external e∈E, sensory s∈S, active a∈A, and internal states i∈I (with their generalized versions x~,ẽ,s~,ã, and i~, respectively), so that

37

The Markov blanket states separating external and internal states consist of S×A. Importantly, external and internal states depend only on the blanket states, with the further constraint that sensory states are not influenced by internal states and active states are not influenced by external states.

38

We note that the definition of a Markov blanket in a biological context is fairly intuitive, as robust literature demonstrates the ability of cells and many other aneural systems to measure aspects of their environment via specific sensors (Baluška and Levin, 2016). In fact, all biological systems can be analyzed in terms of sensory and internal states and the relationships between them (Rosen, 2012).

39

Following this Markov partition of states (and associated influences), we can decompose the flow f(x~) into fe(ẽ,s~,ã), fs(ẽ,s~,ã), fa(s~,ã,i~), and fi(s~,ã,i~).

40

We can derive a gradient descent for the Markov states using Equation (5) and its standard form expression for a flow f(x) subject to conservative and dissipative forces at non-equilibrium steady-state following (Yuan et al., 2014):

41

where Γ is the diffusion tensor defined as half the covariance of the dissipative random fluctuations, and a tensor Q describing friction satisfying ∇·Q∇L(x~)=0. The responses of active and internal states, to sensory stimuli under Markov partition m, therefore, become

42

Inserting the Lyapunov function from (c) into (a) and (b), gives us the resulting flow of active and sensory states as gradient descents on a log probability density:

43

Crucially, the autonomous states (i.e., states that do not depend upon external states: active and internal) of an agent depend upon the same quantity, which we have reduced to the log probability of finding the agent in a particular state; where the agent's states consist of the internal states and their Markov blanket.

44

Solving Equation (9) for the time evolution or flow f of active and internal states thus is equivalent to evaluating the gradients of the log probabilities above, corresponding to the Lagrangian of an open system. By minimizing the internal and active states of the partition instead of minimizing the Lyapunov or Lagrangian function (such as for a thermodynamic potential) as would be done in a classical physics approach, we can now replace the Lagrangian with a variational free energy functional of a probabilistic model of how a system thinks it should behave, as follows.

45

Using the above Markov blanket partition, we can now interpret internal states as parameterizing some arbitrary probability density q(ẽ) over external states. This allows us to express the Lagrangian or Lyapunov function as a free energy functional of beliefs, and implicitly a function of the internal states. We can express this variational free energy through the introduction of the Kullback-Leibler Divergence:

46

which is the expectation of the logarithmic difference between the probabilities p and q, where the expectation is taken using the probabilities p.

47

Thus, instead of taking the log density ln p(s~,ã,i~|m) above, we can now express a variational free energy F that corresponds to the logarithmic difference between the (variational) density or Bayesian beliefs about external states q(ẽ) and actual probability densities p(ẽ,s~,ã,i~|m) of all states under the Markov blanket m:

48

The first term is referred to as (Bayesian negative log) model evidence, or marginal likelihood, which denotes the likelihood that the sensory inputs were generated by a generative model implicit in the Markov blanket m. The second term is called relative entropy and works as to minimize the divergence between the variational and posterior density q(ẽ) and p(ẽ|s~,ã,i~), respectively. As a result, maximizing model evidence results in minimizing the free energy of the system, and because the divergence of the second term can never be less than zero, free energy is an upper bound on the negative log evidence. Using this expression, the flow of autonomous (i.e., active and internal) states becomes

49

Crucially, the gradient descent on variational free energy reduces the divergence in Equation (10) to its lower bound of zero (because the divergence cannot be less than zero). At that point, the gradients of the divergence in Equation (12) disappear and the dynamics reduce to the self-organization in Equation (9), which is what we aim to solve.

50

Now we can evaluate the variational free energy bound in Equation (11) in a straightforward way given a generative model; i.e., the joint probability over (generalized) external, internal, and blanket states. We can therefore associate the joint probability in Equation (11) with a likelihood; that is, the probability of a cell's states, given external states and a prior, in our case the prior probability of a cell's states (i.e., internal states and their Markov blanket). This means that q(ẽ) plays the role of a posterior density over hidden or external states under a particular Markov blanket or model (m). Importantly, this variational posterior is parameterized by internal states and we can talk about the internal states encoding beliefs about external states.

51

We now turn to the construction of this generative model, where we employ variational filtering as the method of quantification and minimization of a variational free energy, which places an upper bound on the dispersion of a particle's internal states and their Markov blanket (Friston, 2010; Buckley et al., 2017). This allows us to convert any process of self-organization into a gradient descent on a free energy landscape, where basins (minima) correspond to attractor states, or goal states—akin to the target morphology—as described next.

4.2. Active inference as a computational framework for morphogenesis
52

In this section, we introduce self organization to non-equilibrium steady-state under the lens of morphogenesis, using the variational principles described above. To that end, we simulated morphogenesis by specifying a generative model—and an implicit variational free energy function—and simulated self-organization by solving the equations of motion in Equation (12), as in Friston et al. (2015) and Kuchling et al. (2020).

53

Given a target morphology specified in terms of the location and differentiation of eight (undifferentiated) cells, we define a body morphology that the cells (collectively) have to reach, composed of a head, a body, and a tail (a single primary axis of positional identity, as exists for example in many metazoa). Migration and differentiation of each cell are key components of reaching a specific target morphology. In the active inference framework, it means that undifferentiated cells migrate and differentiate by minimizing free-energy. The dynamics of morphogenesis are mediated by chemotactic, biophysical, and electrochemical signals. Cell division is not taken into account in our model for the sake of simplicity. All cells are identical at the beginning of the simulation and they don't have any information on what kind of cell they are or where they are. Although they all have the same model, they are pluripotent and can differentiate into any kind of cell at the end.

54

In this multi-agent system, the active states of one cell (e.g., its secreted signaling molecules) are the external states to other cells (whose diffused concentrations it measures through it sensory states). Indeed, each cell has internal states and sensory states (of the Markov blanket) that correspond either to chemoreceptors of either extracellular and intracellular concentrations, while cell migration or the release of chemotactic signals are caused by the active states.

55

In terms of active inference, each cell is equipped with a generative model that encodes the beliefs it has about what chemotactic signals it should receive or express relative to its location in the target morphology.

56

We need to specify the generative model given by the probability density p(s~,ã,i~|m) of sensory states s, active states a, and internal states i, as well as the dynamics of the environment, determined through the flow fẽ and fs~ of external states e and sensory states s, respectively. This allows us to specify the requisite equations of motion for the system and its external states. Here, we will adopt a probabilistic nonlinear mapping with additive noise:

57

where the superscripts denote the first and second levels of our hierarchical model g. Gaussian assumptions about the random fluctuations or noise ω mean that we can write the requisite likelihood and priors as:

58

where N is the normal distribution, and Π(t) denotes the precision (or inverse variance) of the random fluctuations.

59

We then construct the approximate posterior density q(ẽ) introduced in Equation (10) using the associated Lagrangian or Lyapunov function