Skip to main content

Learning Under Algorithmic Conditions: 12 Bioinformatic Algorithms and Educational Genomics

Learning Under Algorithmic Conditions
12 Bioinformatic Algorithms and Educational Genomics
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomeLearning Under Algorithmic Conditions
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. Cover
  2. Half Title Page
  3. Title Page
  4. Copyright Page
  5. Contents
  6. Introduction
  7. Part 1. Imitation, Thought, and Reason
    1. 1. Technics and Text: Guided by Gilbert Simondon
    2. 2. Deviation Games: Desire and the Pedagogy of Thought
    3. 3. Number Sense in Large Language Models
  8. Part 2. Bodies, Brains, and Common Sense
    1. 4. Neuro-symbolic Algorithms and the Infant Mind
    2. 5. The Problem of Algorithmic Commonsense Learning
    3. 6. Learning on the Neuromorphic Circuit
  9. Part 3. Curriculum, Control, and Computation
    1. 7. Who Controls the Curriculum for AI? The Limits of Participatory Design for Educational AI
    2. 8. Learning to Program
    3. 9. Computational Thinking and Software Studies
  10. Part 4. Mysticism, Robots, and Genetic Algorithms
    1. 10. Machine Learning Ecologies and Self-Organization
    2. 11. Meaningful Robot Learning
    3. 12. Bioinformatic Algorithms and Educational Genomics
  11. Part 5. Viral Affect and School Interfaces
    1. 13. The Urban Public School as Cybernetic Apparatus
    2. 14. Algorithms and Immediacy
    3. 15. Responsible AI and Learning to Language
  12. Part 6. The Onto-Epistemology of Colonial Instrumental Reason
    1. 16. Machining Coloniality and Learning Otherwise
    2. 17. Noisy Compression and Colonial Violence
    3. 18. Instrumentalizing Colonial Reason
  13. Part 7. Life and the Limits of Computation
    1. 19. Learning in the New Dispersed Prime Time
    2. 20. Machine Learning and the Digital Archiving of Death
    3. 21. Thinking Softly with Incomputability
  14. Part 8. Multimodal Learning with Unruly Tools
    1. 22. Learning by Co-constructing with Stupid (but Useful) Generative AI
    2. 23. Digital Technologies and Perceptual Curation
    3. 24. Technosocial Scotomas in the Algorithmic Age
  15. Part 9. The Disruptive Technical Being of Generative AI
    1. 25. Prompt Battles and the Conundrums of Logos
    2. 26. Machine Learning and Its Operational Diagrams
    3. 27. Algorithmic Creativity, Deception, and Delirium
  16. Acknowledgments
  17. Contributors

12 Bioinformatic Algorithms and Educational Genomics

Ben Williamson

Educational Genomics

The use of genetic “big data” to address research questions about educational outcomes and associated behaviors has proliferated since 2010, resuscitating historical scientific thinking about genetics and intelligence in education (Martschenko et al. 2019). The development of bioinformatics technologies—which incorporate computer algorithms to analyze biological big data—has reanimated attempts to discover the genetic basis of IQ, cognitive ability, and learning outcomes in educational discourse (Williamson et al. 2024). Characterized as a “genomic revolution for education research and policy” (Morris et al. 2022, 1), researchers in fields of social and behavioral genomics have begun using bioinformatics to explore the “genetic influences” on educational outcomes (Harden 2021) and predict academic achievement from DNA (Selzam et al. 2017). These “educational genomics” studies collectively claim that the common “genetic architecture” of learning outcomes has been made legible through the analysis of vast genomic datasets (Chen et al. 2024). The genetic architecture of learning outcomes has been mapped using advanced technologies, such as bioinformatics instruments and genomic databases for genotyping and analyzing DNA, leading to claims of the “discovery of genes associated with human phenotypes such as educational attainment” (Madole and Harden 2023, 1).

This memo offers a critical examination of “EduYears,” a measure of educational attainment that, scientists claim, is influenced by a complex genetic architecture of “polygenic” associations (Lee et al. 2018). In particular, it reviews the specific bioinformatic devices and algorithmic formulas that produce a “polygenic score” for EduYears in order to raise questions about concomitant ideas concerning human learning.

EduYears, educational genomics researchers claim, is influenced by miniscule molecular variations in a genetic architecture of thousands of interacting variants, albeit with considerable environmental and cultural confounding (Gusev 2024a). It is often viewed as a proxy measure for cognitive development and as “evidence of the supposed biological reality of IQ” (Gusev 2024b, n.p.). Educational genomics programs are therefore concerned with tracing the genetic architecture of learning outcomes, illuminating the interior biological interactions that constitute cognition and intelligence, and using polygenic score calculations to “genomically predict” educational outcomes from genetic data (Kawakami et al. 2024). Polygenic scores are summary statistics of all the weighted genetic variants—biomarkers known as single nucleotide polymorphisms or SNPs—that affect an outcome like EduYears (Burt 2024). These calculative devices are increasingly positioned as generating “policy-relevant” insights into the genetic heritability of learning outcomes, which could be used for various forms of genetic testing, screening and intervention (Asbury et al. 2022). As devices with claimed policy relevance, polygenic scores reveal how learning and cognition are being reconceived as interior biological processes through molecular genomics methods and bioinformatics technologies.

However, polygenic scores do not simply reveal the heritability of cognition or intelligence. Instead, they produce a bioinformatic proxy measure of the outcomes of learning. Polygenic scores “do not ‘exist’ in the same way” as other measurable biological processes, but only as “algorithms” and “models” that might “create” a “new biological reality” distinct from the biospecimens from which they are derived (Janssens 2019, 147–48). As an algorithm or model, the EduYears polygenic score acts as a proxy for educational attainment, which is in turn a proxy for biologically interior processes of cognition often captured and quantified as IQ. Computational genetic sciences do not reveal a biological body at all, but a bioinformational substitute—a statistical, calculable body characterized in terms of networks, patterns and codes, or “homo algorithmicus” (Kotliar and Grosglik 2023).

Making Polygenic Scores Through Algorithmic and Bioinformatic Infrastructures

Polygenic scores are artifacts of a complex infrastructure of statistical genetics, molecular genomic databases, bioinformatics instruments, and the algorithms that orchestrate them, which together function to make the outcomes of learning appear legible at a molecular level of analysis (Williamson et al. 2024). The “EduYears polygenic score” is a product of a series of “big data” genomics studies of educational attainment conducted by the Social Science Genetic Association Consortium, an international research infrastructure established 2011 to apply genomic methods to social science and economics research. It has catalyzed the recent international growth of educational genomics research. The SSGAC’s educational genomics studies have, it is claimed, identified those “genes involved in brain-development processes and neuron-to-neuron communication” (Lee et al. 2018, 1112), which also “predict a range of cognitive phenotypes” and educational outcomes such as test scores and high school grades (Okbay et al. 2022, 438). An EduYears polygenic score was constructed to predict educational outcomes from genetic data among independent samples, generating promissory enthusiasm for other educational genomics studies.

Polygenic scores are presented by social and behavioral genomics scientists as statistical formulas, or “a single quantitative variable that summarizes an individual’s genetic predisposition to a trait” (Mills and Tropf 2020, 558). However, they are better understood as artifacts of bioinformatics instruments and infrastructures, and the “architectural-algorithmic and organisational-work practices” involved in generating, processing, organizing and using genomic data (Mackenzie 2003, 318). Bioinformatics techniques “promise the detection of hitherto unknown features and properties of biological molecules through computer database searching and comparison” (2003, 327).

The biological realism of educational genomics research treats the bioinformatic infrastructure underpinning EduYears studies as neutral, unbiased tools of discovery. However, the bioinformatization of biological investigation has significantly influenced the character of scientific knowledge, making it possible “not only to aggregate and compare data, but to parse, rearrange, and manipulate them in a variety of complex ways that reveal hidden and surprising patterns” (Stevens 2013, 65). Genomic insights are generated through “data journeys” involving encounters with algorithms embedded in genotyping platforms, genome-wide scanning robots, microarray chips, and an array of biostatistical practices and bioinformatics software, all of which can variously influence the results and conclusions of genomic studies (Leonelli 2014). Rather than objectively discovering the biological interiority of student cognition and learning processes, then, EduYears findings are mediated and shaped by the design and constraints of the scientific infrastructure at every step of processing.

Among the bioinformatics applications through which polygenic scores are assembled are microarrays, or SNP chips, used for genotyping individuals from DNA samples—such as the millions of samples included in EduYears and related studies (Malanchini et al. 2020). Microarrays do not merely “discover” but rather “produce” human genetic variation by applying techniques from the computing and data sciences, such as pattern recognition, data mining, machine learning, computational algorithms, robotics, and automation (Kragh-Furbo et al. 2016). The design, statistical power, algorithmic parameters and technical constraints of the chips have “locked in” the primary “conceptual frameworks humans should use to consider, organize, work with, and ultimately act on genetic differences” (Rajagopalan and Fujimura 2018, 862). The microarrayed data can then be entered into biobanks. The SSGAC’s samples are primarily obtained via legal data sharing contracts with the UK Biobank and 23andMe, two of the largest repositories of genetic data. This selection introduces sample biases since biobanks are overrepresented by healthy, well-educated, and wealthier-than-average individuals of European ancestry (Roberts and Rollins 2020).

As the SSGAC EduYears studies specify, the genotyped biobank data are then “pruned” using “clumping algorithms” so that only the “lead SNPs” most strongly associated with an outcome remain (Lee et al. 2018). These steps of microarraying, clustering, and mapping the genomic associations with learning outcomes take place through large-sample genome-wide association study (GWAS) methodologies that involve scanning the human genome for SNPs associated with a particular outcome or trait (Kawakami et al. 2024). GWASs analyze each SNP individually and aggregate them into summary statistics that are later used to construct the polygenic scores. The SSGAC compiles the results of such studies through a meta-GWAS methodology that involves combining the biobank data with dozens of other genetic cohort studies. It functions as a bioinformatic search engine by scanning multiple databases for patterns among millions of SNPs, then surfacing relevant results as the basis for its EduYears calculations. Indeed, GWAS methodologies are built on many of the data-scientific and computational pattern-matching algorithms developed for searching the web (Stevens 2016).

While GWAS methods only produce summary statistics rather than identifying specific causal mechanisms, “bioinformatics annotation” technologies, it is claimed, make it possible to “locate specific cells, tissues, and organs where relevant genetic variants are expressed” and “gain deeper insights into the biology of complex behavioral outcomes” like educational attainment (Madole and Harden 2023, 13). Using bioinformatics annotation, the SSGAC prioritized genes associated with EduYears that “encode proteins that carry out neurophysiological functions” (Lee et al. 2018, 1114). The polygenic architecture of learning outcomes can therefore be made legible as codes for gene clusters and graphical displays of their anatomical location, annotated with known biological functions at the individual SNP level, all orchestrated by the constitutive algorithms of bioinformatics. But this does not guarantee causal-mechanistic explanation of the pathways from genotyped somatic substance to phenotypic outcomes like IQ, academic achievement and educational attainment (Matthews and Turkheimer 2022).

Following this data journey through bioinformatics microarrays, biobanks and bioannotation instruments, the SSGAC constructed its EduYears polygenic score with the software packages LDpred and PLINK. LDpred is a computationally intensive bioinformatics method utilizing probabilistic Bayesian algorithms for “deriving polygenic scores based on summary statistics and a matrix of correlation between genetic variants” (Privé et al. 2020, 5424). Mobilizing “stratification” and clumping algorithms, PLINK makes it possible to process “large data sets comprising hundreds of thousands of markers genotyped for thousands of individuals” and identify “polygenic effects” (Purcell et al. 2007, 559). There are, however, multiple ways to create polygenic scores, each involving different assumptions and goals, measurement instruments, technical adjustments, calculation methods, and algorithm parameter specifications, which can introduce technical biases into the results (Burt 2024).

Polygenic scoring software and its constitutive algorithms therefore mediate and produce certain configurations of “biological reality,” represented as a single quantitative “signal” of polygenic influence on social outcomes, rather than discovering causal-mechanistic biological insights from bioinformation (Burt 2023). The data used for educational genomics studies have to journey through complex “logistical infrastructures” of storage, processing, comparison and distribution, encountering algorithmic techniques and instruments that change the biological objects under investigation into reinvented forms with different potentialities for scientific knowing and understanding (Mackenzie et al. 2016).

Molecularized Learning Under Bioinformatic-Algorithmic Conditions

The bioinformatics apparatus of polygenic prediction and scoring in educational genomics produces a molecularized epistemology of educational outcomes. Human learning, specifically, is maintained as a series of interior biological processes that correlate with phenotypical traits and outcomes. This epistemology of learning is based on mapping and making legible the thousands of minute SNP biomarkers that constitute the genetic architecture of educational attainment. The EduYears polygenic score functions as a proxy for cognitive ability at the molecular level, and is further correlated with educational outcomes such as test scores, standardized grade achievements, and IQ measures (Harden 2021). This biomarkerization of learning has enabled some researchers and startup companies even to propose polygenic IQ tests (Smart 2023), genetic talent-testing for children (Au 2022), or early years genetic screen-and-intervene programs using polygenic scores (Asbury et al. 2022).

Learning under bioinformatic-algorithmic conditions is therefore understood within educational genomics research as an interior process that is surveyable at the molecular level through genomic scanning and polygenic calculation of minute genetic differences—seemingly realizing the historical aims of behavioral geneticists to uncover the biological substrates of intelligence, cognition, and educational outcomes. Far from being unmediated measurements of interior biological substances, however, polygenic scores represent an epistemological shift that is “linked to a technological shift associated with the widespread use of computers” (Stevens 2013, 67). A “data-centric” epistemology in genomics privileges biological data mining and algorithmic analysis as a correlational, unbiased and objective “discovery” method (Leonelli 2014). Through this shift, the genetic architecture of human learning is claimed to be discoverable by data mining millions of bioinformational data points recovered from vast biobanks, though every algorithmic step of the data journey taken to produce polygenic scores mediates the analyses and the results. Just as the genomics “laboratory has become a kind of factory for the creation of new forms of molecular life” (Rose 2007, 13), the educational genomics laboratory has become a factory for fabricating new models of learning at the molecular level. This molecularization of learning imputes educational outcomes to genetic causes, although polygenic scores may be almost entirely confounded by social and environment factors, throwing their salience as a source of policy or practice intervention into serious doubt (Govindaraju and Goldstein 2025).

The polygenic body produced by educational genomics is an algorithmic artifact of bioinformatics instruments and infrastructures, legible only from scanning, clumping, and stratifying genetic data in computers. In this sense, as algorithmic artifacts of data journeys through bioinformatics infrastructures, polygenic scores are performative of distinctive biological realities, separated from the bodies and behaviors they purport to predict by cascades of bioinformatic algorithms and practices. They fabricate and format human subjects in terms of an informational epistemology that links genetic codes to computer codes (Koopman 2020). Polygenic scores are assigned to the informationalized body of “homo algorithmicus” (Kotliar and Grosglik 2023) as a proxy of the embodied subjects of education. Educational genomics has reanimated historical aims to uncover the genetic substrates of intelligence by deploying a cascade of bioinformatic-algorithmic operations, and calculating polygenic scores for diagnosis, screening, and intervention into human learning itself.

Acknowledgment

This research was supported by a Research Project Grant awarded by the Leverhulme Trust (grant number: RPG-2020-395).

References

  • Asbury, K., T. McBride, and R. Bawn. 2022. “Can Genomic Research Make a Useful Contribution to Social Policy?” Royal Society Open Science 9:220873220873: http://doi.org/10.1098/rsos.220873.
  • Au, L. 2022. “Testing the Talented Child: Direct-to-Consumer Genetic Talent Tests in China.” Public Understanding of Science 31 (2): 195–210.
  • Burt, C. H. 2023. “Challenging the Utility of Polygenic Scores for Social Science: Environmental Confounding, Downward Causation, and Unknown Biology.” Behavioral and Brain Sciences 46:e207, 1–19.
  • Burt, C. H. 2024. “Polygenic Indices (aka Polygenic Scores) in Social Science: A Guide for Interpretation and Evaluation.” Sociological Methodology, https://doi.org/10.1177/00811750241236482.
  • Chen, T. T., J. Kim, M. Lam, et al. 2024. “Shared Genetic Architectures of Educational Attainment in East Asian and European Populations.” Nature Human Behaviour 8:562–75.
  • Govindaraju, D. R., and A. M. Goldstein. 2025. “The Elusive Associations of Nucleotides with Human Success: Evolutionary Genetics in Education and Social Policies.” Evolution: Education and Outreach 18:4. https://doi.org/10.1186/s12052-025-00218-3.
  • Gusev, A. 2024a. “The Heritability of Educational Attainment.” GusevLab: http://gusevlab.org/projects/hsq/#h.a6jctcodj87b.
  • Gusev, A. 2024b. “The Heritability of IQ Test Performance I: What Does IQ Measure?” GusevLab, http://gusevlab.org/projects/hsq/#h.u5i4y14hya4j.
  • Harden, K. P. 2021. The Genetic Lottery: Why DNA Matters for Social Equality. Princeton University Press.
  • Janssens, A.C.J.W. 2019. “Validity of Polygenic Risk Scores: Are We Measuring What We Think We Are?” Human Molecular Genetics 28 (R2): R143–50.
  • Kawakami, K., F. Procopio, K. Rimfeld, et al. 2024. “Exploring the Genetic Prediction of Academic Underachievement and Overachievement.” npj Science of Learning 9:39. https://doi.org/10.1038/s41539-024-00251-9.
  • Koopman, C. 2020. “Coding the Self: The Infopolitics and Biopolitics of Genetic Sciences.” Hastings Report 50 (3): 6–14.
  • Kotliar, D. M., and R. Grosglik. 2023. “On the Contesting Conceptualisation of the Human Body: Between ‘Homo-Microbis’ and ‘Homo-Algorithmicus.’” Body and Society 29 (3): 81–108.
  • Kragh-Furbo, M., A. Mackenzie, M. Mort, and C. Roberts. 2016. “Do Biosensors Biomedicalize? Sites of Negotiation in DNA-Based Biosensing Data Practices.” In Quantified: Biosensing Technologies in Everyday Life, edited by D. Nafus. MIT Press.
  • Lee, J. J., R. Wedow, A. Okbay, et al. 2018. “Gene Discovery and Polygenic Prediction from a Genome-Wide Association Study of Educational Attainment in 1.1 million Individuals.” Nature Genetics 50:1112–21.
  • Leonelli, S. 2014. “What Difference Does Quantity Make? On the Epistemology of Big Data in Biology.” Big Data & Society 1 (1). https://doi.org/10.1177/2053951714534395.
  • Mackenzie, A. 2003. “Bringing Sequences to Life: How Bioinformatics Corporealizes Sequence Data.” New Genetics and Society 22 (3): 315–32.
  • Mackenzie, A., R. McNally, R. Mills, and S. Sharples. 2016. “Post-Archival Genomics and the Bulk Logistics of DNA Sequences.” BioSocieties 11:82–105.
  • Madole, J. W., and K. P. Harden. 2023. “Building Causal Knowledge in Behavior Genetics.” Behavioral and Brain Sciences 46:e182, 1–57.
  • Malanchini, M., K. Rimfield, A. G. Allegrini, S. J. Ritchie, and R. Plomin. 2020. “Cognitive Ability and Education: How Behavioural Genetic Research Has Advanced Our Knowledge and Understanding of Their Association.” Neuroscience and Biobehavioral Reviews 111:229–45.
  • Martschenko, D., S. Trejo, and B. W. Domingue. 2019. “Genetics and Education: Recent Developments in the Context of an Ugly History and an Uncertain Future.” AERA Open 5 (1): 1–15.
  • Matthews, L. J., and E. Turkheimer. 2022. “Three Legs of the Missing Heritability Problem.” Studies in History and Philosophy of Science 93:183–91.
  • Mills, M.C., and F. C. Tropf. 2020. “Sociology, Genetics, and the Coming of Age of Sociogenomics.” Annual Review of Sociology 46:553–81.
  • Morris, T. T., S. von Hinke, L. Pike, N. R. Ingram, G. Davey Smith, M. R. Munafò, and N. M. Davies. 2022. “Implications of the Genomic Revolution for Education Research and Policy.” British Educational Research Journal. https://doi.org/10.1002/berj.3784.
  • Okbay, A., Y. Wu, N. Wang, et al. 2022. “Polygenic Prediction of Educational Attainment Within and Between Families from Genome-Wide Association Analyses in 3 Million Individuals.” Nature Genetics 54:437–49.
  • Privé, F., J. Arbel, and B. J. Vilhjálmsson. 2020. “LDpred2: Better, Faster, Stronger.” Bioinformatics 36 (22–23): 5424–31.
  • Purcell, S. et al. 2007. “PLINK: A Tool Set for Whole Genome Association and Population-Based Linkage Analyses.” American Journal of Human Genetics 81:559–75.
  • Rajagopalan, R.M., and J. H. Fujimura. 2018. “Variations on a Chip: Technologies of Difference in Human Genetics Research.” Journal of the History of Biology 51:841–73.
  • Roberts, D., and O. Rollins. 2020. “Why Sociology Matters to Race and Biosocial Science.” Annual Review of Sociology 46:195–214.
  • Rose, N. 2007. The Politics of Life Itself: Biomedicine, Power, and Subjectivity in the Twenty-First Century. Princeton University Press.
  • Selzam, S., E. Krapohl, S. von Stumm, et al. 2017. “Predicting Educational Achievement from DNA.” Molecular Psychiatry 22:267–72.
  • Smart, A. 2023, 27 October. “From a Fledgling Genetic Science, A Murky Market for Prediction.” Undark. https://undark.org/2023/10/27/consumer-genetic-testing-science/.
  • Stevens, H. 2013. Life Out of Sequence: A Data-Driven History of Bioinformatics. University of Chicago Press.
  • Stevens, H. 2016. “Hadooping the Genome: The Impact of Big Data Tools on Biology.” BioSocieties 11:352–71.
  • Williamson, B., D. Kotouza, M. Pickersgill, and J. Pykett. 2024. “Infrastructuring Educational Genomics: Associations, Architectures, and Apparatuses.” Postdigital Science and Education. https://doi.org/10.1007/s42438-023-00451-3.

Annotate

Next Chapter
Part 5 Viral Affect and School Interfaces
PreviousNext
The University of Minnesota Press gratefully acknowledges the generous assistance provided for the publication of this book by the University of British Columbia, Columbia University, and Adelphi University.

Chapter 1 contains portions previously published, in modified form, from Elizabeth de Freitas, “Fragile Books and Machine Readers: Trans/in/dividual Reading Tactics in a Complex Technical Milieu,” International Journal of Qualitative Studies in Education 37, no. 6 (2024): 1655–65; reprinted by permission of the publisher (Taylor & Francis Ltd, https://www.tandfonline.com). Portions of chapter 5 were previously published in a different form in Carolyn Pedwell, “The Intuitive and the Counter-intuitive: AI and the Affective Ideologies of Common Sense,” New Formations 112 (2024): 70–93.

Copyright 2026 by the Regents of the University of Minnesota

Learning Under Algorithmic Conditions is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License (CC BY-NC-ND 4.0), https://creativecommons.org/licenses/by-nc-nd/4.0/.
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org