26 Machine Learning and Its Operational Diagrams
Goda Klumbytė
The epistemology of machine learning has often been described in pragmatist terms, juxtaposing it with classical statistical or more experimentalist approaches. As Wheeler writes, “For a pragmatist, . . . knowledge is a means of control, not a special state of mind. Uncertainty is to be exploited rather than extinguished, and epistemic notions ought to be derived from the roles that they play in inquiry rather than the other way around” (2017, 2). Approaching machine learning systems as assemblages and investigating their operations through identifying and analyzing their underlying diagrams opens the line of inquiry toward the diagrammatic epistemology of such systems. Such epistemology brings together pragmatist approaches to knowledge as tethered to means and ends of knowledge making, with the idea that abstract processes, such as diagrams, have profound material and materializing effects. This allows us to approach machine learning as an abstract-concrete technology that organizes phenomena and knowledge about phenomena in an abstract manner while also materializing and inscribing its (historically and technologically contingent) modes of operation into the learning process. This memo aims to outline a proposition for diagrammatic epistemology of machine learning, whereby machine learning systems are seen as assemblages, whose emergence and functioning is piloted by operational diagrams.
Diagrams and Assemblages
In his book Machine Learners: Archaeology of a Data Practice (2017), Adrian Mackenzie describes machine learning as “a diagrammatic practice in which different semiotic forms—lines, numbers, symbols, operators, patches of color, words, images, marks such as dots, crosses, ticks, and arrowheads—are constantly connected, substituted, embedded, or created from existing diagrams” (49). Following this logic, each specific instance of a trained machine learning system—that is, a system that has been trained on data and calibrated to perform a specific task of inference—could be addressed as a particular assemblage of human and nonhuman elements, including specific data, algorithm(s), infrastructures, as well as processes of labor (engineering, labeling, designing, training), epistemic premises, and disciplinary techniques.
For Deleuze and Guattari, assemblages are agentive arrangements (agencement), while diagrams are expressions and enactments of relations that produce other diagrams and assemblages (Deleuze and Guattari 1987; Livesey 2010). Diagrams are pragmatic in that they describe the function of an assemblage, and they are generative in that in their abstract quality they express, enact, and engender forces and relations. Delanda calls them “mechanisms of variation” (Delanda 2016) because they are amenable to experimentation. Functionally, they can enact both forces of territorialization and deterritorialization (Deleuze and Guattari 1987). As territorializing tools of control, they allow for capture and value extraction. For instance, Guattari argues that capitalism draws power from its diagrammatic operations: Capitalism inscribes itself operationally into various spheres of life as a principle of organization and functionality (Guattari 2009). At the same time, diagrams, being abstract and generative, also have a potential to open up lines of flight for unexpected, qualitatively new and different realities to emerge (Knoespel 2001). In this dual quality, diagrams mark a point of intersection between invention and control (Knoespel 2001).
Because of their generative, productive capacity, diagrams are “abstract machines” that are historically situated (Deleuze and Guattari 1987; Knoespel 2001). Jeremy Bentham’s Panopticon structure analyzed by Foucault (Foucault 1991, 1995) is an example of a historically situated disciplinary diagram that can be transposed from prison to classroom, hospital, and other spaces (Butler et al. 2014). This historical and cultural situatedness allows us to analyze the genealogy of various diagrammatic processes and their transpositions.
Figure 26.1. Linear regression scatterplot diagram showing random data points and their linear regression line.
In this memo, I address diagrams as abstract processes that find their expression in other forms, such as visual diagrams, capturing and mobilizing quantitative relations that have become part of the machine learning algorithmic milieu (see Figures 26.1 and 26.2). Expressions of diagrams, as well as their logic and historical genealogy, can be studied in order to understand their specific piloting functions in the formation of learning assemblages. I argue that, as historically situated machines, diagrams prescribe certain “ways of seeing” (Haraway 1988) and organizing phenomena, materializing these ways of seeing into concrete sociotechnical systems. In the case of machine learning, they do so by piloting the learning process and thereby acting as pedagogical instructions for machine learning, which I call operational diagrams.
Algorithms as Operational Diagrams
Machine learning algorithms—that is, the specific architectures and learning functions that are used to construct a specific model—are piloted by operational diagrams that express abstract-concrete relations between different elements in the learning assemblage. For example, linear regression, k-nearest neighbor, support vector machines, and neural networks are all such operational diagrams that, through the learning process, construct the model as a learning outcome. Approaching algorithms as operational diagrams allows us to ask what kind of perspectives they introduce and what kind of materializations they produce through the learning process—in other words, their prescriptive conditions and effects as well as their situatedness.
Let us investigate two examples of an operational diagram: linear regression and k-nearest neighbor (k-NN). Linear regression is a way of modeling a relationship between two or more variables through finding the line of best fit, that is, a line that expresses the scalar change of the dependent variable in relation to the change of the independent variable. The “best fit” is defined by choosing specific mathematic criteria to calculate the distance between the actual variable values and the regression line , so as to minimize the distance (Figure 26.1).
As a statistical method, linear regression was formulated by a nineteenth-century polymath Francis Galton based on his study of genetics and heredity. He noticed that various features in populations (such as parent/child size differences in sweet pea plants and height differences in human parents and children) tended to the mean—they exhibited a “regression towards mediocrity” (Galton 1886). Galton’s interest in heredity and normality was part of the emergent field of population statistics, and directly linked to Adolphe Quetelet’s concept of the “average man” (2014 [1842]), to which he added an understanding of variance and deviation as features that can be manipulated. Coining the term “eugenics,” Galton connected the science of statistics to a social program of population improvement by pushing the average toward the right-hand side of the normal distribution. This was to be done by encouraging procreation of those with “best qualities” and discouraging reproduction of “worst qualities,” where “best” and “worst” were defined through reference to class, dis/ability, and race (Thomson 2010; Turda 2010).
Linear regression as a method is entangled with this history of normalization and the casting of difference as either pejorative or as advantage and thus rendering it always already in relation to the norm. As an operational diagram of machine learning, linear regression first relates data to each other through pairing and renders them as stable categorical concepts (variables). Then, as the function of best fit is constructed, the values of variables acquire meaning in relation to a fictional line of best fit. Through this ideal construct, linear regression materializes the concept of a norm expressed as a stable, calculable relation against which difference is plotted, predicted, and managed. In machine learning this norm, contrary to the “moral norm” of Quetelet or even the “mediocre norm” of Galton, is rendered as an expression of a standardized, calculable, and to some extent regularized—and thus regulatable, governable—relation. In effect, the operational diagram of linear regression pilots the machine learning assemblage not toward addressing difference in itself or internal differentiations of phenomena, but rather toward the normal/nonnormal dynamic. Learning here becomes an intervention that produces and enacts normalized relations, piloted by abstract yet historically situated sets of instructions (diagrams).
Definition and production of difference is also a feature of another classical statistical method, widely used in machine learning—the so-called proximity-based estimator k-NN. Here, the value of the query will be defined by comparing it to 𝑘 number of its closest neighbors—depending on where the so-called “neighborhood line” is drawn (i.e., what value of k is chosen), the query will be compared within the values in that neighbourhood and categorised (or “voted” to belong to a specific category) according to the majority principle (Figure 26.2).
The logic of K-NN can be described as the logic of homophily (“birds of a feather flock together”). Chun (2018, 2021) has traced the concept of homophily to Lazarsfeld and Merton’s research into mixed-race neighborhoods in the US in the 1950s, which posited segregation as the norm and naturalized the idea that people form friendships based on similarity (Lazarsfeld and Merton 1954). According to Chun, homophily, as one of the core principles of network science, connects it to histories of organizing difference and sameness through tools of social engineering, such as segregation, reservations, and camps, not only positing homogeneity and similarity as the norm but also exacerbating differences between the clusters.
Figure 26.2 Diagram of a k-NN scatter plot with two categories (triangles and squares). The value of the point labeled “query” will be determined by defining the hyperparameter k. If k=3 (i.e., three closest neighbors are considered), the query will be categorized as a triangle. If k=9 (i.e., nine closest neighbors are considered), the query will be categorized as a square.
As an operational diagram, k-NN inscribes sameness and difference through spatialized partition, not simply “discovering” clusters in data but actively performing segmentation and control of class divisions and categorical difference. This is exemplified discursively in the language of “outliers,” “infiltrators,” “smoothing out the boundaries” that patterns contemporary k-NN papers, as well as in the use cases that k-NN and homophily analysis renders itself to, which has included analysis of judicial decisions as early as 1970s (see Mackaay and Robillard 1974) and predictive policing (Kaufmann et al. 2019). Again, the process of learning here is rendered as an active production of patterns through clustering, which is informed by the genealogy of such homophilic “way of seeing” and the discursive practices enmeshed within this genealogy. Notably, both examples discussed here operate through distance measures, thereby lending them to a diagrammatic analysis.
Diagrammatic Epistemology and Lines of Flight
The above examples of operational diagrams show that they act as pedagogical instructions for the learning that machines do. By iteratively acting on other elements of the assemblage, such as data and the relations in the specific application domain, operational diagrams enact and materialize historicized ways of seeing and doing. This merits further investigations of what kind of learning is machine learning and how does it compare, contrast, and interact with human forms of learning. For instance, Reigeluth and Castelle (2021) ask precisely what kind of learning is performed by machines and argue that machine learning would benefit from a more social theory of learning, addressing machine learning beyond individual training techniques, and rather as cultural activity in a sociotechnical milieu. Machine learning assemblages also could be further approached from the perspective of learning theories as “pedagogical devices,” as theorized by Basil Bernstein and successors (Sadovnik 1991; Singh 2002): devices that establish principles for the pedagogization of knowledge, by which knowledge becomes communicable. Addressing machine learning as a diagrammatic process emphasizes the relation between abstraction and materialization in learning processes, and draws attention to genealogical, historical conditions within which various diagrams come to dominate particular cultural practices.
Diagrammatic operations within machine learning introduce their logic into larger learning assemblages, establishing and modulating relations (between data, phenomena, humans, contexts, etc.), rendering them available for exploration, experimentation, as well as value and knowledge extraction. They also fix such relations in place through the process of learning—a trained model technically only reflects temporally situated relations yet it is used as a stable pattern for future action, often with preemptive effects (Rouvroy 2020).
Diagrammatic epistemology points to the nonrepresentational function of diagrams, revealing their generative and constructive activity. Operational diagrams, however, are not to be understood as pure “abstract machines” in the Deleuzeo-Guattarian sense—they are rather abstract-concrete patterns of operation that, through choreographing and reassembling, can become abstract machines. While linear regression and k-NN as concrete algorithms are axiomatic and representational (rather than inventive), they can and do become germinal of new relations to emerge within an assemblage and across different machine learning assemblages.1 Diagrammatics can also enable forms of modulation, translation, and transposition across diverse phenomena, which is evident in the variety of data and application domains deterritorialized by machine learning techniques. This allows us to address diagrammatics as a political space of both control and intervention (de Freitas 2014).
To address machine learning epistemology as diagrammatic allows one to recognize the principles of operation that pilot the larger learning assemblage, and to trace the ways in which learning is a worlding process that involves linking, connecting, and relating phenomena. Since diagrams as abstract-concrete objects can move across different life domains, they can be intervened into and modulated through other abstract-concrete tools, such as critical concepts (Klumbytė et al. 2022). This opens the possibility for more interdisciplinary, egalitarian, and situated engagement with machine learning, not simply as concrete technological implementations but as diagrams for worlding otherwise.
Note
1. Thank you to the editors and anonymous reader for their helpful comments and feedback on this point.
References
- Butler, N., E. Jeanes, and B. Otto. 2014. “Diagrammatics of Organization.” Ephemera: Theory & Politics in Organization 14 (2): 167–75. https://ephemerajournal.org/contribution/diagrammatics-organization.
- Chun, W. H. K. 2018. “Queerying Homophily Muster der Netzwerkanalyse.” Zeitschrift Für Medienwissenschaften 10 (18–1): 131–48. https://doi.org/10.14361/zfmw-2018-0112.
- Chun, W. H. K. 2021. “The Space Between Us: Network Gaps, Racism, and the Possibilities of Living in/Difference.” Catalyst: Feminism, Theory, Technoscience 7 (2). https://doi.org/10.28968/cftt.v7i2.34903.
- de Freitas, E. 2014. “Diagramming the Classroom as Topological Assemblage.” In Deleuze & Guattari, Politics and Education: For a People-yet-to-Come, edited by M. Carlin and J. J. Wallin. Bloomsbury.
- Delanda, M. 2016. Assemblage Theory: Speculative Realism. Edinburgh University Press.
- Deleuze, G., and F. Guattari. 1987. A Thousand Plateaus: Capitalism and Schizophrenia. University of Minnesota Press.
- Foucault, M. 1991. “Governmentality.” In The Foucault Effect: Studies in Governmentality, edited by G. Burchell, C. Gordon, and P. Miller. University of Chicago Press.
- Foucault, M. 1995. Discipline and Punish: The Birth of the Prison. Vintage Books.
- Galton, F. 1886. “Regression Towards Mediocrity in Hereditary Stature.” Journal of the Anthropological Institute of Great Britain and Ireland 15:246–63.
- Guattari, F. 2009. “Capital as the Integral of Social Relations.” In Soft Subversions: Texts and Interviews 1977–1985, edited by Sylvère Lotringer. Translated by Chet Wiener and Emily Wittman. Semiotext(e).
- Haraway, D. 1988. “Situated Knowledges: The Science Question in Feminism and the Privilege of Partial Perspective.” Feminist Studies 14 (3): 575–99. https://doi.org/10.2307/3178066.
- Kaufmann, M., S. Egbert, and M. Leese. 2019. “Predictive Policing and the Politics of Patterns.” The British Journal of Criminology 59 (3): 674–92. https://doi.org/10.1093/bjc/azy060.
- Klumbytė, G., C. Draude, and A. S. Taylor. 2022. “Critical Tools for Machine Learning: Working with Intersectional Critical Concepts in Machine Learning Systems Design.” In 2022 ACM Conference on Fairness, Accountability, and Transparency. ACM. https://doi.org/10.1145/3531146.3533207.
- Knoespel, K. J. 2001. “Diagrams as Piloting Devices in the Philosophy of Gilles Deleuze.” Théorie—Littérature—Enseignement 2001 (19): 145–65.
- Lazarsfeld, P. F., and R. K. Merton. 1954. “Friendship as a Social Process: A Substantive and Methodological Analysis.” In Freedom and Control in Modern Society, edited by M. Berger, T. Abel, and C. H. Page. Van Nostrand.
- Livesey, G. 2010. “Assemblage.” In The Deleuze Dictionary. Rev. ed. Edited by A. Parr. Edinburgh University Press.
- Mackaay, E., and P. Robillard. 1974. “Predicting Judicial Decisions: The Nearest Neighbour Rule and Visual Representation of Case Patterns.” In Datenverarbeitung im Recht (DVR): Band 3, Heft 3/4. De Gruyter.
- Mackenzie, A. 2017. Machine Learners: Archaeology of a Data Practice. MIT Press.
- Quetelet, L. A. J. 2014 (1842). A Treatise on Man and the Development of His Faculties. Cambridge University Press.
- Reigeluth, T., and M. Castelle. 2021. “What Kind of Learning Is Machine Learning?” In The Cultural Life of Machine Learning, edited by J. Roberge and M. Castelle. Springer International Publishing.
- Rouvroy, A. 2020. “Algorithmic Governmentality and the Death of Politics.” Green European Journal. https://www.greeneuropeanjournal.eu/algorithmic-governmentality-and-the-death-of-politics/.
- Sadovnik, A. R. 1991. “Basil Bernstein’s Theory of Pedagogic Practice: A Structuralist Approach.” Sociology of Education 64 (1): 48. https://doi.org/10.2307/2112891.
- Singh, P. 2002. “Pedagogising Knowledge: Bernstein’s Theory of the Pedagogic Device.” British Journal of Sociology of Education 23 (4): 571–82. https://doi.org/10.1080/0142569022000038422.
- Thomson, M. 2010. “Disability, Psychiatry, and Eugenics.” In The Oxford Handbook of the History of Eugenics, edited by A. Bashford and P. Levine. Oxford University Press.
- Turda, M. 2010. “Race, Science, and Eugenics in the Twentieth Century.” In The Oxford Handbook of the History of Eugenics, edited by A. Bashford and P. Levine. Oxford University Press.
- Wheeler, G. 2017. “Machine Epistemology and Big Data.” In The Routledge Companion to Philosophy of Social Science, edited by L. MacIntyre and A. Rosenberg. Routledge.