4 Neuro-symbolic Algorithms and the Infant Mind
Elizabeth de Freitas
Classical debates in epistemology and philosophy of mind often reference infant behavior as they argue the merits of nativism versus empiricism, where the former assumes innate cognitive faculties and the latter assumes a relatively blank slate. This debate corresponds quite accurately to disagreements among AI theorists, where advocates of neural network models are usually cast as empiricists, and advocates of more traditional symbolic AI are cast as nativists. Glossing these debates with these terms overlooks much of the nuance, but helps us map the theoretical stakes. Despite the impressive ability to simulate human language-use or generate new images, many thinkers criticize amped-up empiricist deep learning models because they are inefficient learners when compared to humans, requiring enormous amounts of data and energy, and forming rigid concepts that lack the cognitive suppleness and adaptive learning one finds in children (Lake et al. 2015; Levine 2023). Deep neural network AI models are thus castigated as poor learners because of the excessive amount of data they require, and because they cannot pivot easily after learning from mistakes.
Critics of current AI models frequently cite studies of child behavior and the capacity of infants to learn quickly after only one or few encounters (“one-shot agile learning”) to support the claim that a priori cognitive symbolic structures are the foundation of human and machine intelligence. Noam Chomsky and Gary Marcus, as well as computer scientists like François Chollet, Brenden Lake, and Josh Tenenbaum, refer to studies of childhood developmental psychology to support their preference for a machine intelligence that would supplement neural nets with more traditional symbolic AI—an approach often referred to as “neuro-symbolic” (see Lake et al. 2015; Ellis et al. 2020). Indeed, Chollet (2019) offered a challenge to the computer science community, to find a rapprochement between deep learning and symbolic approaches. Chollet (2019) takes up the theories of cognitive psychologist Elizabeth Spelke at the Harvard Center for Brains, Minds and Machines who claims that the aim of AI should be to “reverse engineer the infant mind and create smart machines” (Spelke 2020). In this memorandum, I show how this ongoing debate in AI research rests on a particular image of the infant mind, elaborated by Seymour Papert and Marvin Minsky at the MIT Artificial Intelligence Lab in the 1960s.
The Influence of Jean Piaget
One of the founders of AI, Seymour Papert (1928–2015), worked with Jean Piaget (1896–1980) at the University of Geneva, between 1958 and 1962, and was one of Piaget’s protégés, before codirecting the Artificial Intelligence Lab at MIT. Piaget’s theories of childhood development and “genetic epistemology” strongly influenced the newly emerging Cognitive Science of the 1960s and the burgeoning field of symbolic AI. At the heart of Piaget’s cognitive psychological approach was both an experimental methodology and a developmental ideology that stipulated the order and age at which children (can/should) develop specific conceptual skills (Piaget 1954). Papert brought Piaget’s experimental methodologies and ideas about learning to his work in the AI lab at MIT. Marvin Minsky and Papert’s 1971 Progress Report on Artificial Intelligence stated that their studies had “become closely bound to the study of development of intelligence in children” (2). In their report, Minsky and Papert critique the early artificial neural nets (i.e., early connectionist efforts in the DL tradition) of the 1950s and ’60s, noting that “the descriptive powers of these quasi-linear learning schemes have such peculiar and crippling limitations that they can be used only in special ways” (31). In other words, cranking through empirical data could only support limited kinds of learning, and what was needed was a more nativist approach (a symbolic AI model). The criticism was quite harsh, although they admitted that sometimes a “weighted decision” and “incremental adaptation” is better than nothing. Still, and foreshadowing what would come, they worried that simply layering neural nets only masks the problem of what they called the “terminal” learning trajectory of neural nets (31).
Minsky and Papert’s 1971 report reveals the image of AI that was emerging in the MIT lab at that historical moment. The report supported the kind of situated or embodied nativism found in Piaget whose idea of “mental schema” was modeled on the Kantian “schemata.” These schemata or mental schema were said to be the innate structures in children’s minds that were used to process and make sense of perception. Papert and Minsky’s Piaget-inspired constructionist (and ultimately Kantian) image of the infant mind substantially influenced computer science as it pursued symbolic AI for the next fifty years, rather than develop AI architectures predicated on neural nets. In fact, it was Minsky’s and Papert’s experimental psychological approach with infants and children that displaced the connectionist neural net models of AI developed by McCulloch, Pitts, Rosenblatt, and other AI researchers in the 1970s. It wasn’t until the twenty-first century that neural net methods would return and dominate the field of computer science.
To be fair, Minsky and Papert embraced a grounded and embodied approach to learning, situating their study of children’s reasoning in experimental design of simple tasks, creating learning environments where knowledge was said to be constructed rather than instructed. They stated, “our main experimental subject worlds, namely the ‘blocks world’ robotics environment and the children’s story environment, are better suited to these studies than are the puzzle, game, and theorem-proving environments that became traditional in the early years of AI research” (1).
My aim in this memo is not to critique Papert’s constructionist pedagogy nor his aim to make mathematics and science education more imaginative and “concrete” with technologically rich environments. But informed by posthuman and eco-cognitive perspectives, I do hope to show how his constructionism ultimately sprang from a nativist epistemology, where supportive learning environments were said to elicit children’s “spontaneous” logic or geometry (Papert 1991, 2). Borrowing from Piaget’s experiments, Minsky and Papert (1971) utilized a series of tasks designed to address and elicit specific cognitive skills in children, especially those skills related to visual perception and spatial sense, as these were considered skills that might be programmable in machines. Many of the experiments involved spatial analogies and attempts to discern separate objects within a scene, always starting with the prompt of asking the child to produce a “symbolic description” and thereby communicate and represent their knowledge about the situation, their process for problem solving, and any methods of computation. Their aim was to study these child-produced descriptions for how they revealed what Minsky and Papert called “heterarchical control structures” within the child’s mind. These structures were said to function as cognitive “problem-solving programs” that involved situated heuristics responsive to specific prompts. In effect, the heterarchical control structures were assumed to be symbolic and computational a priori programs innate to the infant mind and corresponding (à la Kant) to the material world. Minsky and Papert understood the child as an agile agent with built-in reasoning programs. They stated as much when they argued that “such an agent must know to a greater or lesser extent how to plan, produce, test, edit and adapt procedures. In short, it must know a lot about computational processes. We are not saying that an intelligent machine, or person, must have such knowledge available at the level of overt statements or consciousness, but we maintain that the equivalent of such knowledge must be represented in an effective way somewhere in the system” (1971, 2).
Built-In Capacities to “See” Structure
Importantly, Minsky and Papert (1971) acknowledged that children’s initial accounts were typically inexact and failed to offer formal rules. Consider, for instance, tasks that involve spatial analogical reasoning, and the request that the child select the appropriate choice to make the analogy “A is to B, as C is to which of 1, 2, 3, 4, 5?” (see Figure 4.1). The children’s initial justifications for their choices were vague, and they failed to show how their analogical reasoning could become a rule that generalized to all similar cases. To address this inadequacy, Minsky and Papert asked the children for descriptions of what they saw, rather than for justifications of their choice. The emphasis on description—rather than justification—became crucial for Minsky and Papert because they believed that the process of generating a description would more effectively attend to the common structure that masks the tacit rule. In other words, they argued that a child’s “mini-theory” can be articulated only if the child is first asked to “make up a description” and then asked to change that description so that it describes a new example (and thus the child generates an analogical rule for generalizing). The experiments with children showed how analogical reasoning rested on the habit of formulating a verbal description of a state of affairs, and then transferring that description to a new state of affairs.
Figure 4.1 “A is to B as C is to which one of these?” Analogical thinking, Minsky and Papert (1971).
Figure Description
This images presentes a multiple choice visual analogy question. The question prompt reads, “A is to B as C is to which one of [the following]”. The top row has images labeled A, B, and C. The bottom row has five possible choices: images labeled 1, 2, 3, 4, 5. The A and B images indicate a spatial relationship between a large circle, small circle, and small square. These shapes change positions between A and B. The C image shows a large triangle, a small circle, and a small square. The 5 possible responses show circles, squares, and triangles in various arrangements, one of which represents the correct transformation to match the analogy.
Minsky and Papert argued that machine intelligence should be crafted to mimic the infant mind. Rather than develop a connectionist approach that incrementally adapts through various kinds of network reasoning, Minsky and Papert sought a more qualitative image of learning, one in which learning occurred after only a few encounters, and where the qualitative leap to structural learning was characterized by innate programs. In other words, the use of language in the descriptions is simply the best medium to get to the more fundamental logical structure—that being the computable operation or program. For Minsky and Papert, programs are not merely strings of symbols but must include “primitive” symbols that represent some selection of the features and their relationship in a situation. Notice that a description is always already a selection, a procedural (program-like) engagement. They claim that “the description is itself a MODEL—not merely a name” (4, caps in original).
Figure 4.2 Child’s drawing of a cube (left) versus adult “iconic” drawing of a cube (right).
Figure Description
The child draws an “open” cube showing the equal sides in a flat display; the adult conventional cube uses perspective drawing.
To what extent do the ideas of Minsky and Papert provide an image of the infant mind that is neuro-symbolic? They assume that the children’s thoughts are “an abstract data structure” that represents all kinds of features, relations, procedures, and other information. Children’s accounts are thus used here as templates for the symbolic (logical) process presumed to undergird analogical spatial reasoning. Minor pedagogical nudges might have occurred with the tasks, and perhaps there was a kind of updating and neural network learning elicited in the Piagetian experiments, but it’s not clear from how these are presented.
After all, Minsky and Papert dismissed the rival AI empiricist approach that emphasized neural network connectionist approaches. Instead, they claimed that their research illustrated how children’s descriptions were not oriented to perceptual cues, and in fact were often at odds with these perceptual cues. This was a crucial observation in their theory of the infant mind. For instance, when asked to draw a cube (Figure 4.2), the children were said to ignore the perceptual cues and attend to important structural features, whereas the adults attended to the perceptual signal, and produced iconic representations of the cube (“iconic imitation”). Minsky and Papert thus argued that children override their senses and use highly schematic descriptions to represent geometric information. In other words, infant minds will consistently override perceptual cues to “describe” the logical and causal structure. In the case of a cube, for instance, the adult draws for resemblance, while a child draws to capture key relationships (i.e., perpendicularity, number of faces). It was these “formalist tendencies” in children’s drawings which lead Minsky and Papert to endorse a symbolic approach to AI, and not one that engages “more directly . . . with the optical image” (7). Although they acknowledged that children can be trained to draw and describe with “quantitative accuracy,” they argue that “the symbolic mode is the more normal manner of performance” (8). This finding was precisely what supported their belief in a symbolic AI, against the “empiricist” AI models that emphasized neural net incremental training on perceptual data.
The Mind as Program
To bolster the symbolic approach to AI, children’s concepts and heuristics are here cast as program-like procedures that often run in parallel. Minsky and Papert’s image of the infant mind as a “problem solver” is not so much hierarchical, in which a unified control center manages various parts and subassemblies. Instead, they proposed a “heterarchy” where computational processes overlap. They state: “we do not mean to suggest that our child had in his mind anything like the graphical image of his drawing, but rather that he has a structural network of properties, features, and relations of aspects of the cube, and that what he drew matches this structure better than does the adult’s more iconic picture” (6). This “structural network” is essentially a set of built-in cognitive programs, arranged in a heterarchy, but not a free-ranging connectionist network model.
They see their approach as an advancement from “abominable” behavioristic models because the activity of description was interactive (child, block) and the child was given more agency than in the usual behaviorist experiment. These ideas will feed learning theories for decades, fueling the learning trajectory theories of the 1990s. Minsky (1986) will go on to characterize the mind as replete with “agents,” stating “minds are what brains do,” while Papert (1980) will advocate for computer literacy and constructionist learning theories in education policy. Note that Papert’s connectionist pedagogy was contrasted with “instructionism” and banking models of education (Papert 1991); in other words, he did not consider his “nativist” theory of internal structural but flexible programs to be at odds with his pedagogy that was focused on creating alternative learning environments where children might creatively develop STEM literacy through concrete and material activities (Papert 1980; Turkle and Papert 1990). His emphasis on “epistemological pluralism” was essentially a pedagogical position, exploring the distinctive ways that different contexts shape learning—this was an extremely important insight about pedagogy that, however, kept his Kantian ontological assumptions unquestioned. In fact, his Kantian epistemology was well suited to such pedagogy, because such an epistemology assumes a convenient correlation between the presumed cognitive faculties of the human mind and the material conditions of the world (Meillassoux 2008).
In conclusion, one can identify several tacit philosophical assumptions about minds and language in the approach of Minsky and Papert. First, they assumed an unproblematic transparent translation between thought and language in their emphasis on descriptions. Despite their equivocation, that by “description” they did not mean “verbal description” but rather “an abstract data structure in which are represented features, relations, functions, references to processes, and other information” (9), their experimental methodology ultimately rested on the child’s language account of their thoughts. In other words, each experiment involved a request that the child explain what they were thinking, or describe their actions, or justify in words the choice they had made, or draw a diagram to capture the analogy, and so on. Setting aside all the complex ways in which thought is not adequately captured by signification or conscious communication, Minsky and Papert went forth with a naive realist approach to sign-making: “to find the secret, one has merely to ask any child to justify his choice” (3). Second, knowledge was assumed to be programs that captured structural and logical relationships, conditioned by innate symbolic capacities rather than networked connectionist models. Minsky and Papert were adamant that their approach to knowledge and learning pointed to inherent cognitive faculties, which I have argued aligns with the Kantian tradition. As mentioned in the opening remarks, this position is epistemically opposed to deep learning AI, where incremental deep training models are said to learn without built-in programs (Bengio 2009).
As the limitations of deep learning AI become increasingly obvious, we find scientists turning back to these early symbolic methods (Levine 2023). In the words of contemporary computer scientists advocating for an alternative neuro-symbolic approach, we hear claims that the conceptual and material world is composed of “programs” (Ellis et al. 2020). And we note the continued influence on AI theory of cognitive scientist Elizabeth Spelke (2017) who argues that there are innate “core knowledge” structures that determine learning progressions. Spelke, who was a keynote speaker at the 2024 annual meeting of the Association for the Advancement of Artificial Intelligence, argues that there is a set of six cognitive domains of which infants seem to have intuitive (i.e., innate) knowledge. These are (1) Objectness, (2) Number (counting), (3) Geometry and space, (4) Agentness, (5) Place, and (6) Social relations. It is this Piaget-Vygotsky-inspired perspective that continues to stand as the foundation of most psychometric benchmarks for spatial and logical reasoning in children and machines. Attempts to build machine intelligence as neuro-symbolic draw explicitly from this work, including the 2019 proposal ARC (Abstraction and Reasoning Corpus) that cites Spelke’s cognitive framework of “core knowledge” as a basis for shaping future machine intelligence (Chollet 2019). Such benchmarks and other related tests have been designed to show the poor learning skills of neural net models—including those using transformer algorithms—compared to the agile and rapid learning found in children. Despite the success of various “empiricist” machine learning algorithms, further enhanced with generative strategies, the nativist theories of the infant mind persist in AI debates today.
References
- Bengio, Y. 2009. “Learning Deep Architectures for AI.” Foundations and Trends in Machine Learning 2 (1): 1–127.
- Chollet, F. 2019. “On the Measure of Intelligence.” arXiv.org.
- Ellis, K., C. Wong, M. Nye, M. Sable-Meyer, L. Cary, L. Morales, L. Hweitt, A. Solar-Lezema, and J. B. Tenenbaum. 2020. “DreamCoder: Growing Generalizable Interpretable Knowledge with Wake-Sleep Bayesian Program Learning.” arXiv.org. https://arxiv.org/abs/2006.08381.
- Lake, B. M., R. Salakhutdinov, J. B. Tenenbaum. 2015. “Human-Level Concept Learning Through Probabilistic Program Induction.” Science 350 (6266).
- Levine, E. V. 2023. “Cargo Cult AI.” Communications of the ACM 66 (9): 46–51.
- Magnani, L., ed. 2023. Handbook of Abductive Cognition. Springer Verlag.
- Meillassoux, Q. 2008. “After Finitude: An Essay on the Necessity of Contingency.” A & C Publisher.
- Minsky, M. 1986. The Society of Mind. Simon and Schuster.
- Minsky, M., and S. Papert. 1971. Progress Report on Artificial Intelligence. MIT Press.
- Papert, S. 1980. Mindstorms: Children, Computers, and Powerful Ideas. Basic Books.
- Papert, S. 1991. “Situating Constructionism.” In Constructionism: Research Reports and Essays, edited by I. Harel and S. Papert. Ablex.
- Piaget, J. 1954. The Construction of Reality in the Child. Translated by M. Cook. Basic Books.
- Spelke, E. 2017. “Core Knowledge, Language and Number.” Language, Learning and Development.
- Spelke, E. S. 2020. “What Infants Know (and Don’t Know). CVPR Workshops on Minds vs. Machines: How Far Are We from the Mind of a Toddler?” Talk available at https://www.youtube.com/watch?v=r-F8Dlp-POk.
- Turkle, S., and S. Papert. 1990. “Epistemological Pluralism: Styles and Voices Within the Computer Culture.” Signs 16 (1): 128–57.