27 Algorithmic Creativity, Deception, and Delirium
Elizabeth de Freitas
Generative algorithms seem to have introduced a new paradigm into the field of AI. Braidotti and Fuller (2019) suggest that generative machine learning algorithms be considered posthuman precisely because they engage in “processes of abstraction that may in turn generate grounds of operation that are outside of the original conditions abstracted from, and thus producing novelty” (9). They suggest that machine abstraction is more than the subtraction of “nuisance features,” and more than invariance seeking, and involves an act of creative composition whereby the new emerges. Buckner (2018, 2020) similarly argues that machine learning algorithms abstract in several philosophically important ways, and might mimic the capacity of human imagining, as conceived within the Humean tradition of empiricism. In a kindred effort, Fazi (2018, 2019) seeks a computational aesthetics or an aesthetics of discreteness that can think with the potentiality of digital processes, attending to how they are generative and capable of creative abstraction. This memorandum focuses on the algorithmic learning that is performed by Generative Adversarial Networks (GAN), which are machine learning frameworks designed by Ian Goodfellow in 2014. They are historically significant, and offer an interesting case study. To what extent do these algorithms exhibit a robust kind of learning, insofar as they work through creative abstraction and hypothesis generation? My aim here is to explore the extent to which GANs leverage the power of the false and fabulation, as a way of learning how to create the new.
GANs involve a particular process of machine learning in which two adversarial neural nets participate in a competitive zero-sum game, gradually training each other until they converge toward a shared knowledge (Goodfellow et al. 2014). GANs have been used in both science and art, supporting experimental practices, exploring hypothetical spaces, and testing simulations. The majority of GAN applications have been in image processing, but the approach has been applied in other areas. Malicious applications have led to concerns about GAN “deepfakes,” and legal responses have included the 2020 California law that controls the use of this kind of technology, named “human image synthesis technologies.” Terms such as “fabulation,” “hallucination,” and “dream” often appear in papers about GANs. We might justifiably curtail the use of GANs because they are often used to do harm, and their criminal application contributes to our distrust of algorithms. Moreover, like other deep learning networks, GANs seem to attend to features that are inscrutable to humans, fueling debates about algorithmic explainability.
In this memorandum I argue that generative algorithms, as embodied in GAN, perform a kind of creative practice whereby they engender the new, simulating “totalities that are not given in nature” (Deleuze 1991 [1953], 86). In addition, I show how such algorithms reveal a delirium interior to machine thought, and that this delirium mirrors an essential aspect of human learning. My argument is that GANs learn according to the same polemic that Deleuze (1991) sees at the heart of empiricism; there is a madness or delirium interior to thought, fueled by the imagination and its constant correction from perception. According to Deleuze, delirium marks the underlying polemic by which learning occurs. I argue that this same polemic is embodied in the GAN learning architecture.
Simulation
Chun (2021) refers to GAN composites when discussing current use of artificial neural nets; and reminds us that “averaging” or normalizing methods of linear, logistic, and other forms of regression always perform a dangerous erasure of complex diversity. She references GAN-produced photographic portraits known as the “non-persons” at the website This Person Does Not Exist (https://thispersondoesnotexist.com/) (Figure 27.1). We gaze upon these faces with dismay at their uncanny resemblance to an actual “someone.” We are engaged and bemused precisely because these GAN-generated faces seem so real but are yet fabulations. We are impressed that they are “fakes” created by an algorithm all too skilled at occupying the hypothetical space of the plausibly human. It’s as though these non-persons stare back at us without permission. They are abstractions that have trespassed into accurate resemblances of the particular—the result of creative mimesis and algorithmic simulation. These images remind us of previous infringements on the individual face, and the sociocultural dangers of abstraction. The eugenicist Francis Galton’s “composite portrait” photographs were presented as both the average and the ideal type, used to typecast the criminal, the working class, the ethnic, racial, and sexual category. Galton’s composite faces served as racist and sexist forms of control and oppression, sorting society hierarchically into types and kinds, according to the averaging of large data sets.
Figure 27.1 www.thispersondoesnotexist.com.
The GAN image differs from the Galton average face, however, partially because the process does not treat the training data in terms of facial features—nose, ears, eyes, cheekbones, etc.—but as merely a distribution of pixels. This matters, because in Galton’s case, the eyes were used as an anchoring feature—in other words, the eyes were a reference point or anchor around which one could align the various photographs; the materiality of the photographic plate then allowed features (brows, chins, noses, etc.) to aggregate as though chemically merging. In the case of the GAN, the training images are not faces in any human sense, since the GAN does not perceive facial features in the data, and has no concept of facial features or face as we understand it. Rather, the GAN scans the training image for mere pixel grayscale distributions. The GAN doesn’t come ready handed with a concept of eye, or ear, but must learn the concept of face from the enormous data set, and must also learn how a face is composed of nose, eye, cheekbone, etc. Neural net processes like this are data hungry and inefficient—they need an excessive amount of examples and nonexamples in order to generalize, even with heuristic searches. Their reliance on expansive data resources—unlike humans who can generalize after only a few examples and nonexamples—is directly linked to their need to abstract from first pixels; the pixel operates here like a floating percept, processed by the GAN without it knowing the concept of face, as it recursively revisits these pixel arrays, until eventually a steady-state distribution emerges that creates a stable and adequate image of a face. But how is that adequacy evaluated? When does the iteration end?
Chun (2021) emphasizes the use of logistic regression in neural net machine learning, as evidence that such algorithms are always subtracting features when forming an abstraction or generalization. However, this overlooks an important distinction between classical averaging techniques (“logistic regression to the mean”) and the manner by which a GAN “composite” is achieved. It would be wrong to suggest that logistic regression, embedded into the nodes of an artificial neural network, is what characterizes the algorithmic power and methods of abstraction we find in a GAN. Logistic regression and related methods have been used to spread data across a probability distribution for centuries. This is not new. What makes deep neural networks (DNN) powerful and dangerous goes well beyond this procedure. GANs, for instance, may use logistic regression at every node in the network, but what is actually significant and possibly unexplainable is the iterative flow of work accomplished across the complex compositional network. How does the specificity of algorithmic construction make these GAN-made portraits something else altogether, something that might be deemed a creative rather than a reductive abstraction? To what extent is creative abstraction entailed in this kind of image composition, as part of their learning process?
Deception and Adversarial Games
The painting entitled Edmond de Belamy (Figure 27.2)1 more than adequately captures the European painterly portrait of a certain era. Produced by the art collective Obvious, and selling for $433,000 in 2018, at Christie’s first auction of an AI-generated art, this painting was composed by a GAN that was trained on fifteen thousand portraits from various time periods. The painting is signed at the bottom right corner, with the GAN loss formula, which encapsulates the distinctive process of GAN learning, where D and G are functions that seek to maximize or minimize the loss equation:
We can unpack this equation, and reveal some of the learning theories built into the GAN architecture, which operates according to a split identity—generator (G) and discriminator (D)—obtaining its objective through dueling adversarial activity, whereby the generator and the discriminator learn as a collective unit. The zero-sum game structure involves the generator creating “fake” data and trying to trick the discriminator into thinking the fake data is one of the “real” data from the training set. The loss function captures this dueling action, where the generator tries to minimize and the discriminator tries to maximize the loss. The key point is that the GAN learns from an internal discernment process, which involves creating nonexamples and trying to fool an internal discriminator. Initially, the generator is bad at fooling the discriminator, creating nonexamples that stray too far from the training data, which are thus easily detected as false by the discriminator. With each iteration, the generator gets better at creating plausible nonexamples, until it is able to “fool” the discriminator, whose skills at discernment, however, have also improved. Ultimately, the GAN is able to generate hypothetical data (images, text, sounds, and other digital modes) that are eerily similar to the real data, as in the case of Edmond de Belamy and painted European portraits.
Figure 27.2 Edmond de Belamy.
Machine learning in such cases involves the usual neural net iterative correction, but also generative trickery and deception. Deception figures prominently in the success of the GAN; indeed, the entire learning process depends on the GAN’s willingness to play a game of deception. As a learning process, this adversarial approach is not so much a predator-prey model, as a symbiotic collective effort of fictioning or imagining otherwise from within the algorithmic architecture—in the end, the output is the result of D and G converging. The generator-discriminator have complementary needs that are served by their entanglement—the generator-forger improves its skills through repeated correction, and the discriminator-detective comes to expect better and better forgeries, until the two agree that the image meets the standards of the fine art of portraiture.
As a machine learning process, GANs gain skill through what is termed “unsupervised learning.” GANs train on given data but they also generate plausible images, and train on both the training set and these generated images. Computer scientists call this generative process of formulating nonexamples to test and correct algorithms a matter of “fiction,” “fabulation,” “dreaming,” and “hallucination” (Ellis et al. 2020). In other words, they suggest through the use of this language that the machines are operating according to a process of imagining otherwise. The dueling structure of fiction and correction (hallucination-seeing, dreaming-waking, etc.) is part of many generative algorithmic architectures, and not solely the purview of GANs. I am focusing on GANs precisely because they were a foundational and game-changing invention, using the built-in dueling structure of generator-discriminator, where the generator acts like the human imagination in extending and creating fictions, while the discriminator guesses whether the work is fiction or not, whether it is to be trusted as real, whether it aligns with expectations, and so on. In the case of face data, the first “guesses” of what might be considered a plausible face are absurd, barely face-like or “close” to the concept, generated as they are from a random flux of pixels, selected without conceptual guidance. This noisy generator—proffering “useless hypotheses” plucked from the vast hypothesis space—is precisely what is needed for the machine learning to be robust. Gradually, the generator learns from the process of judgment and generates hypothetical faces that are more face-like and capable of fooling the discriminator. These generated images are creative abstractions insofar as they capture something “about faces” while not presenting any currently existing particular face.
Learning Through the Absurd
GANs generate novelty through the introduction of a noise vector that randomly selects some of the pixels, gradually learning, through trial and error, to transform the noise vector in such a way that the generated image mimics the original data set with adequate nuance. This is an impressively inefficient way of generating plausible hypotheses, but it is also what guarantees that the learning is not overfit to the training set, and it is also one of the ways in which speculation, conjecture, and abduction enter the process of machine learning. This noise vector is crucial for the learning and the creative abstraction to emerge; without the introduction of randomness, the neural nets would be overfitted to the original data set. In other words, they would be merely deterministic models that aimed at accuracy through massive training sets. In the GAN, the generator is key; the learning occurs through a recursive movement of creation and correction, leveraging the noise vector as a creative force.
It’s essential that we understand how generative machine learning algorithms operate according to the mathematics of the continuous (i.e., topological space, distance metrics, gradient descent, etc.). In other words, the mathematics of continuous functions and real numbers is at the heart of this particular approach to AI, unlike classical symbolic (or discrete) AI models. What matters in a neural net is how value is correlated with weight, and how these weights are distributed and revised across the network; this is machine learning within a vast undulating multidimensional hypothesis space, where minimizing and maximizing is a process of navigating the hills and troughs of weighted significance. The algorithm modifies weights so as to reduce overall error, thereby inducing a refined error function which emerges from the process. Others have rightly raised concerns about bias that is built into our databases, but this memo aims to dig deeper and expose the technical being of the GAN (Simondon 2017). We can better understand the dangers of automated reason when we understand how generative algorithms simulate or diverge from our own techniques of speculation, grounded in the material labor of the imagination. In other words, human reason also involves a speculative reach and reliance on the power of imagining otherwise. In what ways is it different?
GAN learning is achieved through a polemic of fictional reconciliations between principles of correction and invented images. Learning is characterized through this polemic between a generator that composes monstrous combinations of past givens and a discriminator that manages the repetition of a habit. The shared objective of reducing their disagreement, in the name of an emergent fiction, affirms the delirium interior to machine thought. We might dismiss my defense of GAN learning, by countering with the claim that tedious methods of recursive correction seem too controlling, and the search for error minimums seems a form of incremental bottom-dredging. But let’s not forget that our own learning processes often entail this kind of methodical work. The generative method of the random noise vector, which is absolutely foundational for the GAN to function and generate hypotheses, introduces a recurring accidental element, so that the absurd and the fanciful enter the process. This might seem a rather unsatisfying form of imagining otherwise, not a very sophisticated creative abstraction—but there is often mundane material labor entailed in learning processes. Human learning involves a deep dive into the unknown of counterfactuals, as well as an affirmation of the absurd and the accident, and a compositional practice of disjunctive synthesis. Learning requires this kind of play. And indeed we find all these habits in the GANs, in their traipsing across the topological contortions of their error landscape, and their gregarious tendency to fabulate and propose the absurd hypothesis.
Finally, we judge neural networks as woefully inefficient learners, in relying on obscene amounts of computing power and data, but through this extravagance and excess the GAN learning seems to embody the same sort of delirium interior to all human thought. By mobilizing chance and randomness, and brazenly mutating that which is given, by partially differentiating and continuously modulating a vast interleaving network of composed functions, the generative algorithm gives itself over to hypothesis generation. This kind of algorithmic sense-making involves fabulation and the power of the false, a creative counterfactual speculation that seems similar, in many respects, to the role of the imagination.
Note
1. “Belamy” sounds like “Bel ami,” which is a cute reference to “Goodfellow,” the inventor of GANs.
References
- Braidotti, R., and M. Fuller. 2019. “The Posthumanities in an Era of Unexpected Consequences.” Theory, Culture & Society 36 (6): 3–29.
- Buckner, C. 2018. “Empiricism Without Magic: Transformational Abstraction in Deep Convolutional Neural Networks.” Synthese 195 (12): 5339–72.
- Buckner, C. 2019. “Deep Learning: A Philosophical Introduction.” Philosophy Compass 14 (50). https://doi.org/10.1111/phc3.12625.
- Chun, W. 2021. “Authenticating Figures: Algorithms and the New Politics of Recognition.” Keynote presented at the New Materialism Informatics Conference. https://www.uni-kassel.de/forschung/en/iteg/veranstaltungen/nmi-2021.
- Deleuze, G. 1991 (1953). Empiricism and Subjectivity: An Essay on Hume’s Theory of Human Nature. Translated by C. V. Boundas. Columbia University Press.
- Ellis, K., L. Wong, M. Nye, M. Sablé-Meyer, L. Cary, L. A. Pozo, L. Hewitt, et al. 2023. “DreamCoder: Growing Generalizable Interpretable Knowledge with Wake-Sleep Bayesian Program Learning.” Philosophical Transactions of the Royal Society. https://doi.org/10.1098/rsta.2022.0050.
- Fazi, B. 2019. “Digital Aesthetics: The Discrete and the Continuous.” Theory, Culture & Society 36 (1): 3–26.
- Fazi, B. 2018. Contingent Computation: Abstraction, Experience, and Indeterminacy in Computational Aesthetics. Rowman & Littlefield.
- Goodfellow, I. J., J. Pouget-Abadie, M. Mirza, B. Xu, D. WardeFarley, et al. 2014. “Generative Adversarial Networks.” arXiv.org.