Wallace DC, 2013  ·  passages 0 to 29 of 58

Bioenergetics in human evolution and disease: implications for the origins of biological complexity and the missing genetic variation of common diseases

Abstract
0

Two major inconsistencies exist in the current neo-Darwinian evolutionary theory that random chromosomal mutations acted on by natural selection generate new species. First, natural selection does not require the evolution of ever increasing complexity, yet this is the hallmark of biology. Second, human chromosomal DNA sequence variation is predominantly either neutral or deleterious and is insufficient to provide the variation required for speciation or for predilection to common diseases. Complexity is explained by the continuous flow of energy through the biosphere that drives the accumulation of nucleic acids and information. Information then encodes complex forms. In animals, energy flow is primarily mediated by mitochondria whose maternally inherited mitochondrial DNA (mtDNA) codes for key genes for energy metabolism. In mammals, the mtDNA has a very high mutation rate, but the deleterious mutations are removed by an ovarian selection system. Hence, new mutations that subtly alter energy metabolism are continuously introduced into the species, permitting adaptation to regional differences in energy environments. Therefore, the most phenotypically significant gene variants arise in the mtDNA, are regional, and permit animals to occupy peripheral energy environments where rarer nuclear DNA (nDNA) variants can accumulate, leading to speciation. The neutralist–selectionist debate is then a consequence of mammals having two different evolutionary strategies: a fast mtDNA strategy for intra-specific radiation and a slow nDNA strategy for speciation. Furthermore, the missing genetic variation for common human diseases is primarily mtDNA variation plus regional nDNA variants, both of which have been missed by large, inter-population association studies.

Evolution and energetics
1

In the mid-nineteenth century, Charles Darwin and Alfred Russel Wallace proposed that species arose through natural selection, thus accounting for how one species can split into two similar species occupying adjacent niches. Darwin & Wallace [1,2] elaborated on this theory to explain many important biological questions. However, Wallace was concerned that natural selection was insufficient to explain the origins of consciousness and the human brain. At the time and even today, the primary debate is whether biology is a science that can be understood by natural principles. Hence, Wallace's concerns have been minimized by the natural scientists.

2

With the maturation of physical science theories–in particular, thermodynamics–an inconsistency arose. Thermodynamics taught that, in a closed system, entropy increases and complexity decays. Therefore, the stable existence of complex living organisms seemed to defy physics. This dilemma was partially resolved by Schrödinger, who pointed out that living organisms were not closed systems, but rather acquire energy from their environment as what he called ‘negative entropy’ [3]. The concept that the flow of energy through a system can generate and sustain complex structures is the domain of non-equilibrium thermodynamics that is now understood to be central to life, the reason why we eat and breathe [4–6].

3

Still, energy flow only explains the maintenance of complexity. It does not necessitate the development of complex forms. Resolution of this dilemma comes from information theory. The discovery of the structure of DNA revealed that biology is primarily about the storage and retrieval of information. Within a narrow thermal range, such as exists on Earth, the flow of energy through organic systems generates ordered structures. One of these structures is nucleic acids, which can accumulate information. Each year, the flow of energy through the biosphere from sunlight on the Earth's surface or from geothermal vents on the ocean floor provides the energy to generate more nucleic acids, and the more nucleic acids the more information, the more information the greater the complexity [7]. Originally, energy flow created nucleic acids directly, though inefficiently [6,8]. With the advent of intra-cellular DNA replication, however, nucleic acids accumulated rapidly through cell proliferation. Hence, the complexity of modern organisms is a consequence of four billion years of energy flow and the resulting information accumulation, making biological information stored energy. This implies that one of the more important actions of natural selection is enrichment for the more energy-efficient individuals among organisms attempting to exploit the same energy resource [7].

Evolution and Mendelian genetics
4

In the later part of the nineteenth century, Gregor Mendel outlined the rules of inheritance for sexually reproducing organisms. Each parent provides one copy of each gene to an offspring through fusion of the male and female gametes, generating a ‘diploid’ individual. At sexual maturity, the individual separates the two gene copies into his/her sex cells in preparation for conception of the next generation.

5

Subsequently, it was discovered that the behaviour of the genes corresponded to the behaviour of the chromosomes, and later that chromosomes packaged DNA. Since nucleic acids can replicate and in the process mutate, this led to the concept that mutations in the nDNA generate the variation that is acted on by natural selection to create organismal diversity. This neo-Darwinian synthesis implied that a significant proportion of the nDNA variation must be functional and, in the right environment, beneficial.

6

This concept stood until the 1960s, when molecular genetic studies on human nDNA revealed that much of the genetic variation that differentiated human populations is due to differences in the frequency of alleles common to both populations. Furthermore, the differences in allelic frequencies between populations could be explained primarily by statistical fluctuations [9–11]. This led Kimura to propose that virtually all extant chromosomal genetic variation was neutral, because the vast majority of functional mutations would be deleterious and removed by purifying selection [12]. This ‘neutralist’ hypothesis precipitated the neutralist–selectionist debate. If all intra-specific genetic variation was neutral, where was the functional variation that could permit individuals to adapt to environmental changes and ultimately give rise to new species?

7

Since before the time of Darwin and Wallace, species have been defined primarily by anatomical differences and all anatomical traits are coded by nDNA genes. Consequently, analysis of nDNA variation has been highly informative in understanding the progressive changes in anatomical traits that occur during speciation. By contrast, so much of human nDNA variation could be explained by stochastic processes that it soon became dogma that all intra-specific genetic differences were the result of stochastic processes such as genetic drift and founder effects.

8

However, humans are a single species and anatomy does not vary markedly within a species. Therefore, intra-specific variation must affect functions other than anatomical traits. Since energy is also fundamental to the life process, bioenergetic changes could be the source of intra-specific variation.

9

A new opportunity arose to find adaptive variation within the human nDNA as a consequence of the Human Genome Project. The screening of multiple human nDNAs permitted the identification of millions of single nucleotide polymorphisms (SNPs). These SNPs were then used to screen populations to identify chromosomal regions that are linked (in linkage disequilibrium) with loci that alter the risk of developing metabolic diseases. Major human metabolic phenotypes include diabetes and obesity and these clinical manifestations are directly related to individual responses to environmental differences in energy resources. In fact, many important environmental differences are related to the type and availability of calories and demands for use of those calories for tissue maintenance, physical work, combating infections, reproduction and to cope with environmental limitations such as oxygen deprivation and toxins. The aggregate of all of these factors will be referred to here as the energetic environment. Since the energetic environment within a species' niche can change, chromosomal locus variants that are beneficial at one time in one energetic environment might become deleterious in another. Therefore, chromosomal loci related to diabetes and obesity should provide insight into the genetics of energy metabolism.

10

Genome Wide Association Studies (GWAS) have identified 63 chromosomal loci associated with type 2 diabetes. However, all of these loci have very weak phenotypic effects, so in aggregate they account for only 5.7 per cent of the variance in disease susceptibility. Simulation studies have suggested that an additional 488 loci may contribute to type 2 diabetes risk, but again the aggregated effect of all such loci would still explain only 10.7 per cent of risk. Additional projections suggest that if it were technically possible to identify them, perhaps approximately 49 per cent of the risk variance might be explained by common variants with low phenotype effect [13]. Taking into account body mass index, additional insulin resistance loci have been identified [14], but the chromosomal variants found by GWAS still fall far short of accounting for the approximately 3.5-fold increased risk faced by first degree relatives of diabetes patients [15].

11

Since the GWAS study design relies on linked DNA polymorphisms to identify disease risk loci, it requires that one or more SNPs be stably linked to the functional locus. Because the observed loci have weak phenotypic effects, large numbers of samples have been required to achieve statistical significance. To obtain sufficient numbers, successful studies have combined the data from multiple geographically dispersed populations. This research design requires that the DNA polymorphism must have become associated with the functional locus early in human radiation so that the resulting linkage unit could become dispersed throughout global populations. Given that severely deleterious mutants would be eliminated by purifying selection, this research design necessitates that only the variants with the most modest phenotypic effects would survive long enough to remain associated within a single linkage group and become sufficiently dispersed throughout human populations to be detected by GWAS studies.

12

Since loci with large phenotypic effects would be rapidly eliminated by purifying selection, any extant high effect locus must be of recent origin. Having arisen recently, such large phenotypic effect loci must be confined to a single regional population and linked to a set of SNPs that are not associated with this functional locus anywhere else in the world. As a result, analysis of the relevant SNPs in a large multiple population study would dilute out any association between chromosomal SNPs and important regional phenotypic variants.

13

Such large phenotypic effect loci for diabetes and obesity loci have often been reported in regional association studies. One group of notable gene variants are those associated in the uncoupling protein (UCP) 1–3 genes [16–19]. However, these population-specific associations have often been dismissed as being ‘non-replicateable’ in other populations. But population-specific associations are precisely what would be expected for the most phenotypically significant gene variants.

14

One important type 2 diabetes locus is the peroxisome-proliferating-activated receptor γ (PPARγ) gene that has been associated with diabetes in GWAS studies [13,20,21]. However, a specific PPARγ variant, P121A, has also been associated with diabetes in specific populations and validated through family studies [22]. Hence, this locus encompasses both ancient, small-effect variants as well as recent, regional, large-effect variants. A G482S amino acid substitution in the functionally related PPARγ coactivator gene-1α (PGC-1α) gene has also been associated with metabolic alterations in regional populations, diabetes in the Danish [23] and altered lipid metabolism in the Pima Indians [24], but PGC-1α is not routinely detected by GWAS [13].

15

These observations suggest that there is a continuum of phenotypic effects among nDNA bioenergetic gene variants. The milder mutant phenotypes, which are less affected by purifying selection, may be retained for prolonged periods in the human population. Those that arose early in human radiation have remained with their original linkage group while being dispersed throughout the global population, and provide a signal in inter-population GWAS studies. More phenotypically significant variants have arisen throughout human history but have been acted on by purifying selection, eliminating them from the global population. Hence, these loci are not associated with widely dispersed chromosomal linkage groups. Some mutants, such as the PGC-1α G482S mutation, may be sufficiently adaptive to have arisen multiple independent times in different populations on different haplotypes, each new variant associated with a different linkage group rendering them undetectable by GWAS. To find these regional high impact variants it will be necessary to study regional populations for functional genetic mutations, presumably by whole-genome sequencing.

Evolution and energetics
16

PPARγ and PGC-1α are nuclear transcription factors that play a major role in regulating bioenergetics and particularly mitochondrial biogenesis. The UCPs regulate the coupling efficiency of the mitochondrial energy production system, oxidative phosphorylation (OXPHOS). Hence, the importance of these and multiple other loci identified by GWAS implicate mitochondrial bioenergetics in diabetes and obesity and, by extension, mitochondrial functional variation in human regional environmental adaptation. Consistent with this supposition, the mitochondria are estimated to generate about 90 per cent of the cellular energy in differentiated tissue cells and the mitochondrial genome encompasses in the order of one to two thousand nDNA genes and thousands of copies of the mtDNA. Hence, a large number of mitochondrial gene targets can be mutated and have significant effects on cellular bioenergetics.

17

The mitochondria are the product of a symbiosis between two micro-organisms that occurred about two billion years ago. The nature of the original partner organisms is actively debated, but the progenitor of the mitochondrion is thought to have been an α-protobacterium that harboured a complete OXPHOS system. Both of these organisms alone were limited in their complexity since a single bacterial cell can generate only enough energy to sustain about 10 000 genes [5,6]. However, when the host cell acquired multiple oxidative bacteria, the bacterial energy could be pooled to provide the required energy for adding more genes to the host cell's DNA to create more complex anatomical structures. There was a problem, however. The oxidative bacteria needed most of their energy to sustain themselves. This dilemma was resolved since bacteria readily exchange genes. By transferring structural genes from the oxidative bacteria to the host cell's DNA, the number of bacterial gene copies was reduced from thousands to two, the pair of homologues in the nDNA. This reduced the amount of DNA to be replicated and reduced the complexity of mtDNA transcriptional regulation, with significant savings in energy. The accumulation of nDNA also permitted the nDNA genes to radiate and address new functional genetic space [5,25]. Hence, there was a significant selective pressure to transfer most of the oxidative bacterial genes to the host cell's DNA, which ultimately became the eukaryotic cell nDNA. As the number of its genes declined, the oxidative bacterium became progressively more integrated into the host cell.

18

Surprisingly, the progressive transfer of genes from the mtDNA to the nDNA did not go to completion and all oxidative eukaryotic cells still retain a mtDNA. In humans, the mtDNA codes for 13 polypeptide genes plus the rRNA and tRNA genes for the mitochondrial, bacteria-like, protein synthesis system [26,27].

19

To sustain a complete bacterial biogenesis apparatus is very energetically expensive. Therefore, there must be a strong evolutionary advantage for retaining the mtDNA. While several hypotheses have been put forward for why the mtDNA has been retained [5,28], the fact that all of the polypeptide genes retained by the mtDNA are central to the mitochondrial energy-generating system OXPHOS provides one explanation [29].

20

All fungal and animal mtDNAs retain essentially the same set of OXPHOS polypeptide genes. In mammals, these include seven (ND1-3, 4L, 4-6) of the approximately 45 polypeptides of the electron transport chain (ETC) enzyme complex I, one (cytochrome b, cytb) of the 11 polypeptides of the ETC complex III, three (COI-III) of the 13 polypeptides of ETC complex IV (cytochrome c oxidase, COX) and two (ATP6 and 8) of the approximately 15 polypeptides of complex V, the ATP synthase.

21

Functionally, the mtDNA polypeptides are central to the electron and proton wiring system for mitochondrial energy production. In OXPHOS, reducing equivalents (electrons from reduced sources) derived from food flow from reduced to oxidized down the ETC that is embedded in the mitochondrial inner membrane. Starting with NADH, which is oxidized by complex I, and succinate by complex II, the electrons are transferred to coenzyme Q (CoQ), then to complex III, then cytocrome c, then to complex IV, and finally to oxygen to generate water. As the electrons traverse complexes I, III and IV, the energy released is used to pump protons from the mitochondrial matrix across the inner membrane to the inter-membrane space. This creates an electrochemical gradient that is acid and positive on the outside and alkaline and negative on the inside [30]. The resulting capacitance of about 0.2 V is the potential energy from which virtually all human biological processes are driven. Given that a human has in the order of 1017 mitochondrial capacitors, this is a great deal of potential energy, the vital force that animates our life. When breathing stops, the membrane potential collapses, energy transduction ceases and death ensues [27,31,32].

22

The potential energy stored in the mitochondrial capacitors can be used for many purposes: to take up Ca++ from the cytosol, modulate cellular REDOX status and reactive oxygen species (ROS) production, and transport proteins and substrates in and out of the mitochondrion. Particularly important, however, is that the mitochondrial inner membrane potential provides the motive force for driving complex V, the ATP synthase, to condense ADP and phosphate (Pi) to generate ATP [30]. ATP is the chemical energy carrier that is exported from the mitochondrion to the cytosol to energize cellular reactions and drive work.

23

Given the critical nature of the inner membrane potential, it is surprising that the efficiency by which the reducing equivalents from food calories are converted into ATP differs between different individuals and regional human populations. Some individuals are highly efficient at converting caloric energy into a proton gradient via the electron transport and by transforming the energy of the proton gradient into ATP. Consequently, these individuals need to burn the least number of calories for the required ATP. Since calories are a unit of heat, such ‘tightly coupled’ individuals generate the minimum heat for the ATP used. By contrast, some individuals are less efficient at converting reducing equivalents to membrane potential, and membrane potential to ATP. These ‘loosely coupled’ individuals burn more calories for the same amount of ATP and thus produce more core body heat per ATP used. This altered coupling efficiency can be modulated epigenetically, for example, by induction of uncoupling protein 1 (UCP1) in brown fat [33], or more stably by altering the sequence of the mtDNA and thus changing proton pumping efficiency [34,35].

24

In tropical and temperate environments, it is generally more advantageous to be more tightly coupled so as to produce the maximum ATP with the minimum heat. However, in the arctic the constraining factor is cold. Therefore, it can be beneficial to be less coupled so that more heat is generated to maintain core body temperature, provided sufficient dietary calories are available. That this type of adaptive variation is due to mtDNA changes has been supported by the demonstration that climate differences correlate with mtDNA rather than nDNA variation [36] and that the basal metabolic rate of Siberian populations is higher than that of more southern populations [37–39].

25

The importance of mtDNA variation in regional adaptation makes sense when it is realized that the mtDNA codes for the proteins that are central to the coupling of electron flow to proton pumping and thus ATP production. All four mitochondrial inner membrane complexes that incorporate mtDNA-coded polypeptides (I, III, IV and V) either generate or use the proton gradient. By contrast, complex II (succinate dehydrogenase), which transports electrons but does not pump protons, is composed of four nDNA polypeptides. Since the electrochemical gradient is a capacitor, the proton permeability of complexes I, III, IV and V must be balanced with each other. If any one of the complexes becomes leaky for protons, then the capacitor can short, which would be deleterious. Hence, the 13 polypeptides of the mtDNA represent an integrated electrical circuit in which each polypeptide must be functionally compatible with the other 12 mtDNA polypeptides of comparable coupling efficiency [27].

Evolution and mtDNA genetics
26

Given that the mtDNA polypeptides are highly variable in their energetic efficiency [40], then the random recombination of the 13 mtDNA polypeptides between different mtDNA lineages, for example, between mtDNAs that are more or less coupled, could erode energy efficiency. This dilemma is avoided by having these variable OXPHOS electron and proton transport genes linked together in a single non-recombining piece of DNA: the maternally inherited mtDNA. This forces all of the membrane potential elements to coevolve consistent with environmental factors such as climate and diet. Uniparental inheritance of the mtDNA is achieved by the concerted elimination of the paternal mtDNA at fertilization, thus blocking inter-individual mtDNA pairing and recombination [27]. Proof that absolute uniparental inheritance of the mtDNA is critical for animals has been obtained by artificially mixing two normal but different mtDNAs within the mouse female germline. The resulting ‘heteroplasmic’ mice manifested marked behavioural abnormalities and severe learning defects [41].

27

Because of strict maternal inheritance, the only way that the mtDNA sequence can change is by the sequential accumulation of mutations along radiating maternal lineages. Therefore, to a first approximation, the number of nucleotide differences between any two individuals is proportional to the time that they shared a common maternal ancestor. This unique feature of the mtDNA has permitted reconstruction of the ancient origins and migrations of women. By overlaying the mtDNA mutational tree, which shows the genetic affinities between indigenous peoples with the geographic location of the populations that harbour those mtDNAs, the progressive movement of humans was mapped (figure 1). The mtDNA proved to be a particularly powerful tool for studying human radiation, since the mtDNA sequence evolution rate was found to be much greater than that of nDNA-coded mitochondrial genes [44–46]. Indeed, the sequence evolution rate of the mtDNA has resulted in critical mtDNA changes corresponding remarkably well with the times of prominent transitional events in human geographical radiation. Figure 1.Diagram of the migratory history of the human mtDNA haplogroups. Homo sapiens mtDNAs arose in Africa about 130 000–200 000 years before the present (YBP), with the first African-specific haplogroup branch being L0, followed by the appearance in Africa of lineages L1, L2 and L3. In northeastern Africa, L3 gave rise to two new lineages M and N. Only M and N mtDNAs successfully left Africa about 65 000 YBP and colonized all of Eurasia and the Americas. The diverse array of mtDNA lineages that M and N spawned are clustered together as macrohaplogroups M and N. The founders of macrohaplogroup M moved out of Africa through India and along the Southeast Asian coast down along the Malaysian peninsula and into Australia, generating haplogroups Q and M42 around 48 000 YBP.

28

Subsequently, M moved north out of Southeast Asia to produce a diverse array of mtDNA lineages including haplogroups C, D, G and many other M haplogroup lineages. In northeast Asia, haplogroup C gave rise to haplogroup Z. The founders of macrohaplogroup N also move though Southeast Asia and into Australia, generating haplogroup S. In Asia, macrohaplogroup N mtDNAs also moved north to generate central Asian haplogroup A and Siberian haplogroup Y. In western Eurasia, macrohaplogroup N founders also moved north to spawn European haplogroups I, W, and X and in western Eurasia gave rise to sub-macrohaplogroup R. R moved west to produce the European haplogroups H, J, Uk, T, U, and V and also moved east to generate Australian haplogroup P and eastern Asian haplogroups F and B. By 20 000 YBP, mtDNA haplogroups C and D from M and A from N were enriched in northeastern Siberia and thus were positioned to migrate across the Bering land bridge (Beringia) to give rise to the first Native American populations, the Paleo-Indians. Haplogroups A, C, and D migrated throughout North America and on through Central American to radiate into South America. Haplogroup X, which is most prevalent in Europe but is also found in Mongolia, though not in Siberia, arrived in North America about 15 000 YBP, but remained in northern North America. Haplogroup B, which is not found in Siberia but is prevalent along the coast of Asia, arrived in North America about 12 000 to 15 000 YBP and moved through North and Central America and into South America, combining with A, C, D and X to generate the five dominant Paleo-Indian haplogroups (A − D + X). A subsequent migration of haplogroup A out of the Chukotka peninsula about 7000 to 9000 YBP gave rise to the Na-Déné (Athabaskins, Navajo, Apache, etc.).

29

Subsequent movement across the Bering Strait, primarily carrying haplogroups A and D after 6000 YBP, produced the Eskimo and Aleut populations. Most recently, eastern Asian haplogroup B migrated south along the Asian coast through Micronesia and out into the Pacific to colonize all of the Pacific islands. Ages of migrations are approximated using mtDNA sequence evolution rates determined by comparing regional archeological or physical anthropological data with corresponding mtDNA sequence diversity. Since selection may have limited the accumulation of diversity in certain contexts, ages for regional migrations can best be estimated from the diversity encompassed within an individual regional or continental lineage, since selection would have had its greatest effect in enriching for the founding mtDNA haplotype, after which mtDNA mutations would accumulate randomly and thus become clock-like [42,43] (reproduced from http://www.mitomap.org, with permission).