Towards a Physarum learning chip
The custom Petri-dish is far from ideal, as manufacture is time consuming and satisfactory growth is not always achieved. Multi-electrode arrays (MEA), also called micro-electrode arrays are devices which contain several recording electrodes, often in the space of several millimetres square; they usually have a central well for cell culture and surface adhesion. Typically MEAs are used for recording neuronal electrical activity at multiple sites within a small area; when neurons are excited, ionic currents are generated through the cell membrane which can be detected as a change in voltage between several of these electrodes. It is envisaged that instead of using the large custom Petri-dish, that a custom MEA could be used for P. polycehpalum network growth and subsequent measurement and stimulation. The size of the MEA would need to be larger than that of the neuronal MEAs, because one protoplasmic tube could cover the same area as the whole mass of neurons; this may require a custom MEA design with an area of several centimetres square and the possibility of agar as a supporting medium.
The associated response to negative and positive stimuli in the form of voltage application, can be regarded as a form of learning. Observing the stimuli-response correlation, it could be said that voltage stimulation in networks of protoplasmic tubes of Physarum polycephalum is equivalent to negative and positive reinforcement learning; connectivity in the network is directly correlated to applied stimulation. Abandonment of tubes is a behavior which is caused by conditioning. Tubes can be abandoned and then regrow to reform connections; this has been observed by the authors; it is possible that the learned state could be changed or reset by changing localised conditions such as light or temperature to initiate regrowth, or even reinforce the learning performed. One drawback to this method is the time taken to regrow tubes, typically a few hours, however this is a living system, much like the plasticity of neurons in the brain, rather than electronic solid state hardware which can be reset in a matter of milliseconds. The biggest limitation of this biological hybrid technology is the time taken for the system to learn; the stimulation time of several hours is currently a limiting factor. Future work will also be aimed at minimising this time.
These responses to negative and positive stimuli are more clearly visible in the modelling experiments. This is perhaps because adaptation of the model networks are not constrained by adhesion of the slime capsule in the Physarum plasmodium. There is avoidance of −ve stimuli in the model which slowly reverses after the stimuli are withdrawn. With very strong −ve stimuli the model networks withdraw completely from the site and do not return after withdrawal of the stimuli. This is suggestive of negative reinforcement learning. For +ve stimuli there is an aggregation of the model population towards the stimulated nodes during +ve stimuli. After the stimulus is withdrawn, however, the population withdraws from the stimulated nodes and is distributed around nearby nodes. On repeated application of +ve stimuli we found that the baseline population level of the stimulated node locations gradually increased after each stimulus period. This might correspond to a spatially implemented form of positive reinforcement learning.
It is difficult to directly relate the potential spatial learning effects of the Physarum plasmodium to neurally implemented learning. The amorphous nature of the plasmodium and the homogeneous nature of its constituent parts suggest that different learning mechanisms might be at work. The work of Reid59 suggests an environmentally mediated spatial learning and the work within this report suggests possible stimulus response effects and spatial aggregation effects that may also play a role. Nevertheless, recent advances in the understanding of electrical synapses as mediators of network connectivity during learning and memory60616263, and extension of these paradigms to non-neural tissues including bone and pancreas172021 suggest that important parallels may exist. Future work will compare the results of in vitro learning in cultured neural networks646566 with circuits composed of Physarum.
With more complex equipment, multiple tubes could be stimulated simultaneously, with negative and positive voltage stimulation; this could have the advantage of complex input-output processing. Much like the logic gates or geometric computation which can be performed by protoplasmic tubes, this system could enable the advancement of the first Physarum Chip; a processor whose processing power is derived from real biological components and their inherent learning ability with negative and positive reinforcement. If the network of protoplasmic connectivity could be trained and retrained with negative and positive reinforcement in the form of HAHF and LALF voltage stimulation, then the possibilities for computation are broad; geometric processing could be performed using inputs as weighted training functions, multiple input combination logic gates could be implemented, in a matter of hours, not including the time taken to produce the PhyChip. Protoplasmic tubes are self repairing and can occasionally reconnect or reconfigure themselves after trauma, so the possibility of an adaptive and reconfigurable network or protoplasmic tubes appears feasible from the findings presented here.
The mechanism of action for HAHF and LALF voltage stimulation is unknown. The most likely cause is that the voltage affects the trans-membrane voltage and therefore the voltage gated ion channels which control movement, environment sensing or decision making. It is possible that HAHF and LALF operate in different ways, for example it has been observed by the authors that applying high direct current voltages to a tube causes it to stop being conductive after a short time, possibly due to the tube being burnt with the high level of current. HAHF could simply be burning the stimulated tubes causing the organism to abandon this tube in the network and increase growth elsewhere in the network. It is however unclear why a LALF waveform would act in the opposite way and cause abandonment of non-stimulated tubes.
Our findings share many features with results reported from stimulating dissociated neuronal cell cultures grown on multi-electrode arrays. For example, Jimbo et al.67 reported how volleys of electrical stimulation can induce excitatory, inhibitory or neutral responses to subsequent electrical stimulation. This ability has subsequently been exploited for simple pattern recognition68.
For more complex scenarios, Bull et al.69 have suggested the use of machine learning techniques to control stimulation to induce required behaviour, using reinforcement learning in particular. The approach was shown able to control both cultured neurons and a non-linear chemical system. Future work with larger Physarum Chips will explore machine learning control of both the type and spacing of stimulation. This work has demonstrated an amorphous system which is capable of learning and adapting to negative and positive stimulation by rearranging network connectivity; this is inherently similar to the machine learning approach developed by Bull et al.69 using a neuronal culture grown on an MEA dish; stimulating the neuronal culture produced a “pathway-dependant plasticity”. The similarities between this and the results produced in this article are clear; positive and negative stimulation causes changes in network connectivity on a grid of electrodes which is classified as learning.
How to cite this article: Whiting, J. G. H. et al. Towards a Physarum learning chip. Sci. Rep. 6, 19948; doi: 10.1038/srep19948 (2016).