11 Meaningful Robot Learning
Cathrine Hasse
Robot learning ideally combines a robotic body with iterative self-adaptive machine learning. In 1988 the futurist and robot designer Hans Moravec predicted: “We are very near to the time when virtually no essential human function, physical or mental, will lack an artificial counterpart. The embodiment of this convergence of cultural developments will be the intelligent robot, a machine that can think and act as a human, however inhuman it may be in physical or mental detail” (1988, 2). However, it has proven to be more difficult than expected to combine an agile robot body with learning algorithms. In this memo, I argue that this difficulty has something to do with a lack of meaningful robot learning experiences involving perception and practice in actual environments.
From the beginning, robot learning entangled the behaviorist learning sciences, neurosciences, and cognitive computer sciences. All of these were already present in the 1940s and 1950s when one of the world’s first robot designers, William Grey Walter (1950), built robots that could move around and react to their environment. Since then, the field of robot learning has drawn on many sources of inspiration from the general learning sciences, which, basically, can be subdivided into three overall groups: behaviorist, cognitive, and social learning theories. Though many of the same theoretical frameworks have also inspired the computer sciences more generally, the field of robot learning is particularly fond of reinforcement learning, which inspired Walter’s robots. In this approach, “meaningful learning” is already present in that the robot is considered “teachable,” like animals (Walter 1951).
Yet the concept of agency is far from straightforward in this context. In robot learning, agency fluctuates between being externally programmed and internally emergent. Unlike human agency, which is deeply embedded in intentionality and meaning making, robotic agency is contingent upon the affordances of the environment and the constraints of algorithmic design. This calls into question the very notion of what it means for a robot to “act” or “learn” with agency.
Even if robot learning theory, to some extent, has included learning theories of meaning making, the concept of meaningful robot learning has yet to be explored. Here I suggest the field would benefit from looking into how humans perceive their environments in an open-ended process (e.g., Ingold 2000), which never ceases to engage humans and environments in meaningful ways. The term “meaningful” designates here an emotional and engaged correspondence with an environment (Hasse 2020).
Walter’s Tortoises
Walter was a British transdisciplinary cybernetician trained in neurophysiology. In 1949 he created what is considered the world’s first autonomous robots. He named his little tortoise-like creatures Machina speculatrix because “they illustrate particularly the exploratory, speculative behavior that is so characteristic of most animals” (Walter 1950, 43). Also in 1950, British computer scientist Alan Turing published a much more famous paper on how machines can be intelligent and think (Turing 1950). The different approaches taken in these two papers illustrate the difference between the algorithms needed in robot learning and the general development of algorithms in computer sciences. The main difference being that robotic machine learning, as Walter soon discovered, requires learning in and from an environment. The enacted body in this approach is not represented by semiotic markers on a computer screen (Hayles 1999, xiii) but moves about in physical space.
Walter draws this conclusion early. In order to perceive their environment, his tortoises had to learn. Walter followed his first paper describing his robots with a new paper published the following year detailing how he rebuilt the original speculatrix into two new robots, named Machina Docilis, which he translated as “teachable machine” (Walter 1951). He even gave them personal names, Elsie and Elmer, a tradition that has been carried on to this day in social robotics. The meaningfulness built into Elsie and Elmer was inspired by the Russian scientist Ivan Pavlov’s experiments with dogs and the behaviorist learning sciences prevalent in his days. Pavlov had noticed that when his assistants brought their lab dogs food they began to salivate. And if you rang a bell because it was feeding time they would begin to salivate as well even if no food was brought. This made Pavlov speculate that somehow the response of salivation could be triggered by a new learned behavior, where the bell came to represent the food for the dog. As summarized by Walter: “The basic event in this form of learning is that an unrelated stimulus, when repeatedly coupled with one that evokes a certain response, comes to acquire, the meaning of the original stimulus” (1951, 60). The dogs salivate when they hear the bell because food means something to them. Though Walter was completely aware that the robots were not finding any real meaning in the environment, his program CORA (Conditioned Reflex Analogue) experimented with Pavlov’s ideas. While Pavlov was conducting his research, American learning theorist Edward Thorndike was experimenting with teaching animals new behaviors. In an attempt to make learning theories more scientific and less introspective (where the researchers examine changes in their own learning), Thorndike devised a series of experiments to derive a set of “laws” regarding learning. The Law of Effect, for instance, emphasizes that we learn from consequences, either rewards or punishments, but if we want someone to learn something, reward is better than punishment. He also developed a Law of Exercise, which emphasizes that learning takes place over time when we practice, and its effect may disappear if we do not continue practicing (Thorndike 1905).
These “laws” and Pavlov’s insights became the cornerstone of the behaviorist learning sciences, which had a particularly strong influence on robot learning. Behaviorists like Burrhus F. Skinner and John B. Watson developed and refined the theories of conditioning and reinforcement learning. In the robot learning sciences these ideas were subsumed under the heading of “reinforcement learning,” to this day one of the most prevalent theoretical learning frameworks in robotics (Ramasubramanian 2019, 258).
In reinforcement learning (as inspired by behaviorist learning theory) a robot does not follow a predestined path. Reinforcement learning algorithms in robotics try to guide a robot’s agency through a kind of iterative self-learning based on punishment and reward. Like Walter’s first robots, the mechanical movement of the robot becomes a sign for the researcher of meaningful behavior. Note that this line of research is agnostic as to whether robots have a mind that extracts meaning from the environment. For the behaviorists, the mind is a black box that is best avoided; some even fiercely attacked the proponents of “mind,” who worked on ideas of learning as a cognitive process. In his defense of behaviorism over cognitive theories, Skinner, for instance, explicitly denounced the possibility of studying “meaning.” Skinner argued that in the case of Pavlov’s dog, it is Pavlov who assumes that the bell sound is meaningful to the dogs. In reality we do not know that the bell is signifying food for the dogs in some interior cognitive sense. Thus Skinner writes that “cognitive metaphor is based upon behavior in the real world. We store samples of material and retrieve and compare them with other samples. We compare them in the literal sense of putting them side by side to make differences more obvious. And we respond to different things in different ways. But that is all. The whole field of the processing of information can be reformulated as changes in the control exerted by stimuli” (1977, 7).
The “Mind” People
The “mind” people, on the other hand, were inspired by developments in the computer sciences that evolved simultaneously as the cognitive sciences (Boden 2006). Here the body was discarded as a relevant entity. Learning occurs in the closed circuit of the brain, which was aptly captured by “the computer,” conceived in terms of a brain as a symbolic information processor. For these researchers, social and cultural meaning making in real life environment was not part of the research. Building on cybernetics (and Turing’s initial insights) humans were seen as Cartesian mechanisms “that respond to their environments by trying to maintain homeostasis; the function of scientific language is exact specification; the bottleneck for creating intelligent machines lies in formulating problems exactly; and an information concept that privileges exactness over meaning is therefore more suitable to model construction than one that does not” (Hayles 1999, 67).
New symbol-processing computational algorithms also evolved from the learning sciences. In 1932, a psychology student named Donald Hebb wrote his Master’s thesis on how he imagined the basic learning processes worked in the human brain. This thesis was expanded upon in the book The Organization of Behavior in 1949. Later, the engineer Arthur Samuel took up some of these ideas and popularized the term “machine learning” in relation to a series of experimental chess programs (Samuel 1959). Machine learning infused new insights into general learning theories and vice versa. Hebb’s initial understanding of how neurons fired and connected with each other, over time, became a resource for machine learning in what came to be known as “neural networks” (Minsky 1954). Ever since the young Princeton student Marvin Minsky created the world’s first neural computer, the SNARC (Stochastic Neural Analog Reinforcement Computer) in 1954, the process of creating thinking machines (and learning robots) emerged alongside cognitive understandings of learning. Where the behaviorists had denied all references to the inner mind, the new cognitive sciences offered a way into understanding human inner life as a computational process (e.g., Simon and Newell 1971). The black box so efficiently slammed by the behaviorist would be opened and inside we would find a learning human brain functioning as an advanced computer.
In spite of the bitter fights between behaviorists and cognitivists, the new symbol-processing learning machines favored by the cognitivists also functioned on a kind of reinforcement system. In the new algorithmic system, machines “copied” the neurons and “fired” at each other. Although it was not a system of reward and punishment, paths between neurons were reinforced or weakened by number weights (Rumelhart and McClelland 1986). When two neurons fire together, their connection grows stronger (Hayles 1999, 261). Through this process, networks were built that could enhance certain connections and weaken others—which for instance could be seen as learning in parallel distributed processing systems (Rumelhart and McClelland 1986). That also changed how robot learning was envisioned from the 1950s onward. It eventually laid the ground for different types of machine learning where computers learn from data without explicit programming, with supervised learning using labeled data to make predictions and unsupervised learning finding patterns in unlabeled data.
Much of this innovation was spurred by Minsky’s move to MIT in 1958 and John McCarthy’s subsequent move from MIT to Stanford in 1963. As noted by Stuart Russell and Peter Norvig, authors of the influential computer science textbook Artificial Intelligence: A Modern Approach,1 there was a debate between Minsky and his colleague McCarthy. McCarthy created AI programming languages and claimed that by building on formal logic it was possible to create “programs with common sense” (McCarthy 1958), which used coded knowledge to look for solutions to defined problems. However, “McCarthy stressed representation and reasoning in formal logic, whereas Minsky was more interested in getting programs to work and eventually developed an anti-logic outlook” (Russell and Norvig 2010, 19). At Stanford, McCarthy inspired the Shakey robotics project at the Stanford Research Institute (SRI), which demonstrates a robotic integration of logical reasoning and physical activity. At MIT, Minsky, together with Seymour Papert, created the MA-3 Robotic Manipulator Arm.
The Cartesian Split and Social Learning
In all these developments we continue to find a Cartesian split between the mind and the body in robotics, as noted by philosopher Hubert Dreyfus (1972, 1992). The split divided the community of robot learning and the computer sciences, according to disciplinary distinctions articulated by Russell and Norvig. On one hand, we have behaviorist theories that focus on learning as initiating a changed rational or human-like behavior in a machine or a human. On the other hand, we have cognitive theories that see learning as initiating a changed rational or human-like thinking in a machine or a human (Russell and Norvig 2010, 2). The aspiration to combine the two has for a long time been difficult to obtain in robot learning.
This is mainly due to the problems arising from the robot body moving through an unknown environment and the difficulties in employing a formal logical system to the unforeseen events occurring when robots “go wild” (Bruun and Hasse 2016). Learning through agency is a highly complex and emergent process. It is much simpler to make systems work on computers, or robots work on formal logical systems, when they are enclosed in factories (i.e., houses built for robots). This is what the philosopher Luciano Floridi refers to as “ontological enveloping”: “The wheel is a good solution to moving only in an environment that includes good roads. Let us define as ‘ontological enveloping’ the process of adapting the environment to the agent in order to enhance the latter’s capacities of interaction” (1999, 214). While enveloped environments are designed for predictability and constraint, unenveloped environments unfold unpredictably and challenge the robot to continuously recalibrate in ways that attend to environments. Problems arise when robots encounter features of the environment that hold meaning for humans but not for robots (Sorensen et al. 2019). One of the main challenges in robotics has been how to connect the diverse types of machine learning processing algorithms used for supervised and unsupervised learning paradigms—of which most are developed primarily in the computer sciences—as the robot body attempts learning in real time in an environment.
The need to connect a robotic body with a meaningful world has inspired yet another link between the learning sciences and robotics: imitation learning or social learning inspired by the work of the psychologist Albert Bandura (1977). Robots and humans engage and learn from each other in two ways in social learning: Humans can learn to adjust their behavior to robots as if they were social beings (Xu 2023), and robots can learn to imitate human movements. For instance, virtual reality algorithms first record human movements and next let the robot repeat the movements. These multidimensional learning algorithms have made contemporary designs for human-robot cooperation possible (Zhang et al. 2024). Another example (for instance in the robots Pepper and Nao) finds a human lifting a robot arm where later the robot repeats the movement without human aid. In these new, more advanced examples of robotics, new types of supervised learning and reinforcement learning guide the robot to copy human behavior, making use of large language models, and both object and sound recognition technologies. The approach to learning, commonly known as imitation learning, can also be called learning by demonstration, apprenticeship learning, or programming by demonstration. Over the years there has been a consistent growth in this robot learning paradigm (Ravichandar et al. 2020, 298). Nevertheless, imitation learning still fails to address the challenge of making an environment more meaningful to robots.
Behaviorists focused on learning as a change in behavior. Gradually this approach was challenged by the cognitive sciences that connected the human mind and machine learning with the “computer as a mind” metaphor, where computer algorithms were said to mimic how the human brain functioned. Over time roboticists have realized that this Cartesian split did not work in practice. There has been an explosive development trying to connect the two learning paradigms with new types of social learning in robotics. Nevertheless, the algorithms have not been able to capture meaningful social learning as our bodies move through environments. In such movements, algorithms would not just have to process extremely quickly but also constantly connect to an emotionally freighted, sentient ecology. Whether this will become possible in robots has yet to be explored. Future developments, nevertheless, promise to blur further the line between cognition, body, and environment. New developments in for instance open-source platforms are enabling new levels of adaptive learning and dexterity and propose new ways to simulate learning in real-world environments for applications like household robotics. These developments point to a paradigm shift, where robots not only learn from the environment but also to some extent adapt in ways that further challenge the existing distinctions between human and artificial agency.
Note
1. The textbook has been reprinted many times but here I use the 2010 edition.
References
- Bandura, A. 1977. Social Learning Theory. Prentice Hall.
- Bateson, G. 1972 (1985). Steps to an Ecology of Mind. Chandler.
- Boden, M. A. 2006. Mind as Machine: A History of Cognitive Science. Vols. 1–2. Clarendon Press.
- Bruun, M. H., and C. Hasse. 2016. “Studying Robots in the Wild.” In What Social Robots Can and Should Do: Proceedings of Robophilosophy, edited by S. S. Andersen, J. Seibt, M. Norskov, and M. Norskov. IOS Press.
- Dreyfus, H. L. 1972. What Computers Can’t Do: A Critique of Artificial Reason. Harper and Row.
- Dreyfus, H. L. 1992. What Computers Still Can’t Do: A Critique of Artificial Reason. MIT Press.
- Floridi, L. 1999. Philosophy and Computing: An Introduction. Routledge.
- Gray, P. 2011. Psychology. 6th ed. Worth Publishers.
- Hasse, C. 2020. Posthumanist Learning: What Robots and Cyborgs Teach us About Being Ultra-social. Routledge.
- Hayles, N. K. 1999. How We Became Posthuman: Virtual Bodies in Cybernetics, Literature, and Informatics. University of Chicago Press.
- Hebb, D. O. 1949. The Organization of Behavior. Wiley.
- Hull, C. L. 1935. “The Conflicting Psychologies of Learning—A Way out.” Psychological Review 42 (6): 491–516. https://doi.org/10.1037/h0058665.
- Ingold, T. 2000. The Perception of the Environment: Essays on Livelihood, Dwelling and Skill. Routledge.
- Minsky, M. 1954. “Neural Nets and the Brain-Model Problem.” PhD thesis, Princeton University Press. Publication number 9438.
- Minsky, M., and S. Papert. 1969. Perceptrons. MIT Press.
- Moravec, H. P. 1988. Mind Children: The Future of Robot and Human Intelligence. Harvard University Press.
- Ramasubramanian, K., and A. Singh. 2019. “Machine Learning Theory and Practice.” In Machine Learning Using R: With Time Series and Industry-Based Use Cases in R, edited by K. Ramasubramanian and A. Singh. Apress. https://doi.org/10.1007/978-1-4842-4215-5_6.
- Ravichandar, H., A. S. Polydoros, S. Chernova, and A. Billard. 2020. “Recent Advances in Robot Learning from Demonstration.” Annual Review of Control, Robotics, and Autonomous Systems 3 (3): 297–330. https://doi.org/10.1146/annurev-control-100819-063206.
- Rumelhart, D. E., and J. L. McClelland, eds. 1986. Parallel Distributed Processing. MIT Press.
- Russell, S., and P. Norvig. 2010. Artificial Intelligence: A Modern Approach. 3rd ed. Prentice Hall.
- Samuel, A. L. 1959. “Some Studies in Machine Learning Using the Game of Checkers.” IBM Journal of Research and Development 3 (3): 210–29.
- Simon, H. A., and A. Newell. 1971. “Human Problem Solving: The State of the Theory in 1970.” American Psychologist 26 (2): 145–59. https://doi.org/10.1037/h0030806.
- Skinner, B. F. 1938. The Behavior of Organisms: An Experimental Analysis. Appleton-Century.
- Skinner, B. F. 1977. “Why I Am Not a Cognitive Psychologist.” Behaviorism 5 (2): 1–10.
- Thorndike, E. L. 1898. “Animal Intelligence: An Experimental Study of the Associative Processes in Animals.” Psychological Monographs: General and Applied 2 (4): i–109.
- Thorndike, E. L. 1905. The Elements of Psychology. A. G. Seiler.
- Turing, A. M. 1950. “Computing Machinery and Intelligence.” Mind LIX 236:433–60.
- Walter, W. G. 1950. “An Imitation of Life.” Scientific American Magazine 182 (5): 42–45.
- Walter, W. G. 1951. “A Machine That Learns.” Scientific American Magazine 185 (2): 60–64.
- Xu, K. 2023. “A Mini Imitation Game: How Individuals Model Social Robots via Behavioral Outcomes and Social Roles.” Telematics and Informatics 78:101950. https://doi.org/10.1016/j.tele.2023.101950.
- Zhang, X., Y. Wang, C. Li, A. Fahmy, and J. Sienz. 2024. “Innovative Multi-dimensional Learning Algorithm and Experiment Design for Human-Robot Cooperation.” Applied Mathematical Modelling 127 (3): 730–51.