Notes
12 Transforming the Genome into a Clinical Resource
DNA, Data, and Algorithms in Medicine
Ramya M. Rajagopalan
1. Introduction
In 2001, the International Human Genome Sequencing Consortium, a multi-country collective of thousands of scientists that had for over a decade conducted a $3 billion effort to elucidate the sequence of nucleotides in a human genome, published its draft findings in the journal Nature (Lander et al. 2001). As many have argued, the Human Genome Project (HGP) modeled how molecular biology could be fashioned into a “big science” (Kevles 1997). It claimed, almost heretically at the time, that molecular biology could be hypothesis-free or “discovery driven”—that is, it could sidestep the aim of most basic science, to answer a specific question(s) or resolve a specific hypothesis. Rather, it posited that the molecular sciences could equally proceed via an alternative model of research, one that aimed primarily to be generative of data and tools, at unprecedented scales. Further, such activities could themselves be understood as having inestimable scientific value, because their material outputs could serve as the catalyst for driving the formation of hundreds or thousands of hypotheses in areas that spanned the entire field of the molecular health sciences. But even if the HGP represented in many ways a project that was self-reflexively and decidedly not motivated by a circumscribed hypothesis framed in concrete terms, there was nevertheless an unambiguous, if veiled, hypothesis solidly undergirding the project, and the HGP benefited from the scientific legitimacy that this hypothesis conferred. The hypothesis guiding the project was a belief in the future, medical value of the genome.
The architects of the HGP posited that knowing the genome sequence in its entirety would prove of invaluable medical importance, justifying the expense required to deduce it. In opposition, some within the scientific community and the general public pushed back against this hypothesis, questioning the value of investing billions of dollars in elucidating something as esoteric as the sequence of DNA, which seemed at the time only remotely relevant to medicine and of dubious practical significance in the clinic. As an article in the New York Times put it at the time, “Critics of the program quarrel with its medical and scientific claims. They say that the best approach to understanding human disease is not to thwack away randomly at the thick forest of human DNA . . . but to study one specific disease at a time, as scientists traditionally have done” (Angier 1990). Nonetheless, the HGP was inaugurated with this very applied promise that genome sequencing would help researchers address basic questions of human biology, inform understanding of disease, and fundamentally accelerate medical advances in diagnosis and therapy. This narrative imbued virtually all the publications produced by the HGP. For example, a public-facing booklet describing the project claimed, “The impact that will be felt in medicine and health care alone, once we identify all human genes, is inestimable . . . if we are ever to uncover the mysteries of carcinogenesis, if we are ever to know how biochemistry contributes to mental illness and dementia, if we ever hope to really understand the processes of growth and development, we must first have a detailed map of the genetic landscape” (US DOE 1996). A decade after the project launched, in the draft genome publication, the HGP’s authors concluded, “The scientific work will have profound long-term consequences for medicine, leading to the elucidation of the underlying molecular mechanisms of disease and thereby facilitating the design in many cases of rational diagnostics and therapeutics targeted at those mechanisms” (International Human Genome Project Consortium 2001).
This chapter seeks to unravel how these promises were pursued, that is, how “the genome” was transformed into a clinical resource. How did DNA come to symbolize (conceptually and semantically) and epitomize (through transformations in clinical research practice) a key repository of medically relevant data? Though in retrospect these transformations may seem unsurprising, even inevitable, given the motivation for the project, a substantial amount of effort and investment, institutional and scientific realignment, and professional transformations had to take place to begin to realize these promises. Indeed, the medical intentions of the HGP architects were tempered from the outset with an understanding that their prognostications could be years in the making. As they went on to caution in the draft publication, “We must set realistic expectations that the most important benefits will not be reaped overnight” (International Human Genome Project Consortium 2001). This caution offered a recognition that a significant amount of institutional and scientific labor would be required to render massive amounts of genomic information into forms that might be usefully applied in clinical settings. Two decades later, much of this rendering remains in the research and development stage, fueled and furthered by the work of a powerful set of interests and actors.
This chapter will track how medical orientations within human genomics have developed and shifted in the last two decades, exploring how the genome and the massive data streams it inspired have been framed as key clinical resources, against a framing of “genomic medicine” and later “precision medicine” as emerging arenas of significance for clinical care. I draw on ethnographic fieldwork conducted at National Institutes of Health (NIH)-supported clinical research laboratories, a population-based DNA biobank at a US medical research clinic, and scientific conferences and meetings, as well as documentary analysis of key publications and archival material at the National Human Genome Research Institute (NHGRI) History of Genomics program and the NHGRI Oral History Collection, to discuss some of the scientific projects, institutional realignments, and professional transformations through which the genome has increasingly been made legible to clinicians and clinical practice. By unraveling how the language of the genome sequence was recast and transmuted into the language and practices of the clinic, the chapter demonstrates the importance of attending to the many ways in which arguments about genome function have been yoked to arguments about the genome’s health relevance.
Tracing how the genome came to be framed as a clinical resource requires understanding how genomic medicine was built, piece by piece, in the early 2000s, from a distant goal of the HGP to a ubiquitous presence in clinical research and in discourse about the health applications of genome information. Several elements of these activities to assemble and package genome information into clinical tools merit attention. The first is how researchers have framed and circumscribed what counts as health-relevant information. This has taken shape primarily in two forms: the electronic health record (EHR), a composite and multilayered data source including detailed, individualized information on medical history and disease status; and genome sequence data, seen as containing individualized information on an individual’s unique disease and disease-risk status. The second element of building genomic medicine that I will give some attention to is the brisk refocusing of clinical care around computation and algorithms. I will discuss how both the EHR and the genome sequence, cast as forms of health data, have facilitated and been a product of the flow of computation and algorithms into medicine. Both are sculpted and reworked into computationally tractable information, through a cascade of activities, from the way they are collected, to how they are algorithmically analyzed, manipulated, stored, digitized, and deployed. Rapid increases in the digital resources and computing power available in the last decade, and in the case of DNA, the technology available for genotyping, sequencing and analyzing DNA, have accelerated the clinical embrace of computation and algorithms. Algorithms in particular have come to be viewed as essential clinical tools for deriving insights from DNA sequence to guide or improve clinical decisions.
Finally, I will discuss the role of the NHGRI-funded research consortia that have sought to enact and hasten these transitions, part of a suite of initiatives that formed the backbone of the NHGRI’s genomic medicine program since the early 2000s. These consortia have structured and organized lines of research by fusing data to clinical needs, in attempts to carve a specialized role and relevance for genomics in clinical decision-making. They have done so through the participation of health-care provider–driven research biobanks. The chapter will track the infrastructural investments in large-scale biobanking for genomics research by which large patient cohorts for genomics research have been built, detailing the ways in which these biobanks became data engines fueling the project of precision medicine.
2. Genomic Healing: Medical Promises of the Human Genome Project
Up until the HGP, much of the work of genetics in the clinic centered around the medical genetics and genetic counseling professions, which had been guiding patients through complex decisions around genetic risks for debilitating diseases for decades (Paul 1995; Stern 2012). The HGP made concrete the aspirations of human genetics to have broader clinical significance than the narrow focus on rare familial disorders that fell under the patient care work performed by medical geneticists and genetic counselors. With the completion of the HGP in the early 2000s, these aspirations crystallized in the rhetoric of terms and concepts like “bench to bedside” and “translational” medicine, as well as “precision” medicine. As this chapter will discuss, these terms took on new life as fundamental shapers of research and funding trajectories, materializing into the umbrella priorities that have guided medical genomics research during the past two decades.
The concept of “precision” or “personalized” medicine, coined during the HGP, began to capture research interest and funding almost as soon as the HGP concluded (Prainsack 2017). Precision medicine became a scientific paradigm of its own, a buzzword for a set of interests that were buttressed by significant institutional backing and resources. In this way, precision medicine became an enabling engine for the goal of making the basic research question of the HGP (to elucidate the genome sequence) of concrete value to medical practice—even if it was never precisely defined what exactly precision meant, or how it might differ from earlier eras when medical practitioners already provided individualized care to their patients. Instead, researchers had an idea of what they wanted precision to mean: the elucidation of genome function and the relevance of those functions to individuals’ health, and the use of this genome information to customize therapeutic care for the patient. In algorithmic terms, this might be framed as something like, individually unique DNA sequence in, and tailored patient care out. The operative metaphors of the HGP, that DNA was the “code of life,” as unique as a fingerprint, helped yoke DNA sequence steadfastly to the notion of precision. Although genome functions turned out to be not so easy to decipher, let alone connect to specific health outcomes (Callaway 2017), over the years the hopes for quick answers to the significance of particular tracts of genome sequence have given way to more tempered expectations about the insights that can, and maybe cannot, be derived from the genome. Nevertheless, as an aspirational goal, the idea of precision spurred investment in tools to build a clinical understanding of the genome: specifically, gene variants and their roles in disease and drug metabolism.
To that end, in the years immediately following the completion of the HGP, genomics research focused on building tools for Genome-wide Association Studies (GWAS), which advanced the view that genetic variation was important for understanding humans’ differential susceptibilities to diseases and drug therapies (Rajagopalan and Fujimura 2018). GWAS were designed to explore the space of genetic variation among humans to uncover genetic loci involved in common complex disease like heart disease, auto-immune disease, and metabolic diseases. Over eighty million variants have been identified to date in human genomes, the vast majority of which remain of uncertain significance for human health and disease (Rehm 2017). Partly owing to this, and partly to the fact that for variants with some evidence of a link to disease, the risk factors are typically quite small, there has been debate about whether and to what degree GWAS have succeeded in increasing understanding of common disease (Callaway 2017). Nevertheless, GWAS played key roles in pointing to the ways that data gleaned from genome-wide scans of the DNA might be coaxed into furnishing information that might then be retooled to improve on the diagnostic and treatment journeys experienced by patients, whether by improving patient outcomes or hastening the speed with which decisions around diagnosis or treatment might be made.
The search for genomic variants that might have direct effects on therapeutic outcomes, such as drug metabolism or tolerance, an area known as pharmacogenomics, became a leading edge in efforts to bring genomics approaches to clinical care. Early in the 2000s, researchers viewed pharmacogenomics as one of the genome’s more promising avenues of entry into clinical care. As geneticists noted in 2003, “Pharmacogenetics seeks to reduce the variation in how people respond to medicines by tailoring therapy to individual genetic make-up. It seems increasingly likely that investment in this field might be the most effective strategy for rapidly delivering the public health benefits that are promised by the Human Genome Project and related endeavours” (Goldstein et al. 2003). Thus, pharmacogenomics was promoted early on as a favorable avenue to pursue, to bring genomics to bear on public health concerns.
At the same time, by the end of the HGP, the generation of massive amounts of data saw genomics become increasingly entwined with computation (Stevens 2013). Genomic research in the 1990s, modeled by the HGP, focused on generating and archiving both “raw” and “finished” DNA sequence, in both physical and computational form. Physical media for genomic preservation took the form of bacterial artificial chromosome and cosmid libraries, bacterial and yeast chromosomes containing large stretches of the human chromosomes, propagated as needed in living cells for analysis. These research tools generated by the HGP also critically drew biology deeper into the analytic capacities of digital computation, not for the first time in molecular biology (Strasser 2009; November 2012) but certainly on much larger global scales than previous eras. These efforts greatly accelerated and magnified the data streams funneling into the NIH’s first international publicly and freely accessible database for archiving genetic sequences, Genbank, established in 1982, but also the many other NIH-hosted databases that spun out from the work of the HGP.
3. (Bio)banking on the Future: Building Repositories of Biospecimens for Genomic Medicine
The HGP’s architects had operated under the assumption that each individual’s genome held “clues” to that individual’s health status, and the work of the HGP proceeded as though this were self-evident. Post-HGP, the genome was increasingly framed as a forensic tool one could excavate, pointing to patterns in DNA sequence variation that might make some individuals more susceptible than others to disease (Rajagopalan and Fujimura 2018). Medical geneticists were seen as the detectives, and DNA variants were the clues that could solve the whodunit of complex diseases; they were the canaries in the coalmine, molecular sentinels marking individuals’ supposedly innate predispositions to disease.
But the HGP had produced a singular composite genome, which was meant to represent everybody and yet be produced from nobody in particular. There was virtually no documented variation in the human genome sequence it generated. The HGP had sequenced DNA from tens of anonymous donors, although later admitted that most of the sequence, about 70 percent, was generated by analyzing the DNA extracted from a single donor, an unnamed male in Buffalo known only as RP11 (Zhang 2018). A single individual’s genome could not shed light on the genetic variation that researchers proposed were relevant to health. To move genomics into the clinic, researchers aimed first to establish the institutional scaffolds that would pave the way to innovating the bioscientific tools needed to generate insights from genome sequence variation on a mass scale, and to do so in ways that were proximal and responsive to clinical needs.
A central institutional scaffold for this work was the DNA biobank, an innovation emerging directly from the perceived need, generated by the HGP, for large patient cohorts to power studies of genome variation and health. Especially in countries with nationalized or more federally coordinated health systems, governments were among the first to embrace building DNA biobanks, committing taxpayer resources to do so. deCODE genetics, a private company in Iceland, was an early pioneer in these efforts, convincing the Icelandic government in the late 1990s to invest in a large nationalized biobank comprising data and biospecimens donated from over one-third of the country’s residents. Soon after, the National Health Service in the United Kingdom launched the UK Biobank. Similar efforts took place in Estonia, China, Canada, Singapore and the Middle East. Governments sometimes contracted with private biotech companies, and in countries with nationalized health systems they worked directly with hospitals and clinics to recruit members of the general population to participate in these biobanks and donate blood, tissue, and medical health histories to research.
In the United States, by the early 2000s, researchers at prominent biomedical research centers, including those directly affiliated with hospitals and clinics, as well as federal funders like the NIH, began to promote large-scale biobanking, framing large patient cohorts as an essential resource for advancing a new translational research agenda around genomic and precision medicine (Precision Medicine Initiative Working Group 2015). In response, academic health centers, insurers, and private health networks were among the first to establish large DNA biobanks representing their patient populations, some collaborative, many aided by specific RFPs and targeted funding solicitations announced by the NIH. Some of the sites that built DNA biobanks included large regional research clinics and academic medical centers, such as Vanderbilt University Medical Center’s “BioVu” biobank, the Cleveland Clinic’s “CC-BioR” biorepository, the Personalized Medicine Research Project biobank at the Marshfield Clinic Research Institute, and others. In many ways, research capacities affiliated with large academic and US-based teaching hospitals and population health clinics were well positioned to establish the first DNA biobanks in the United States focused on genomic materials and data, by leveraging existing infrastructural and administrative systems that serviced, supported, and delivered health care to large local and regional clinical populations to establish biorepositories of patient-donated samples. By the mid-2000s, the construction of research biobanks in the United States was well underway at several university-affiliated health centers on both coasts and in the Midwest United States.
The research clinic was a key site for the development of such biobanks, and precision medicine tools more broadly, owing to several features that facilitated the traffic of samples, data, and findings between the laboratory “bench” and the patient “bedside.” Genomics studies typically employ large numbers of samples. At research hospitals, researchers had immediate access to and familiarity with a large group of practicing physicians. Their professional relationships with these doctors, many of whom were themselves investigators on biobank-driven research projects, facilitated enrollment of study participants into biobanks. The project of translating genomics to medicine has thus been enhanced by the co-location of molecular research facilities with clinics and patient donors, allowing the flow and sharing of patient samples, data, and findings among fast-assembling collaborative teams of doctors and researchers.
Biobanks in the genome era, since the completion of the HGP, have typically cataloged and archived the biomaterials of hundreds of thousands of patients, including blood, saliva, DNA, and other types of biological specimens, warehousing the sequencing and genotyping data and biological measurements gleaned from each sample. Often, these biobanks were sited at university-affiliated health centers or large regional clinics with a strong research track record, a long history of excellence in both clinical research and patient care, and large patient populations. Indeed, many academic medical centers and research clinics have a long-standing history of successfully enrolling patients into clinical studies and trials because they have worked to foster relationships of trust and ongoing dialogue with the communities that they serve, so that clinical needs drive research rather than the other way around. For example, the Marshfield Clinic PMRP biobank regularly published and sent out newsletters to all biobank participants, with updates on research being conducted on biobank samples (McCarty et al. 2008), and Vanderbilt’s BioVu consented patients at phlebotomy visits, periodically offering an opportunity to opt out of research (Roden et al. 2008). The research clinic co-locates professional interests in clinical studies, facilitating the design and testing of new ways to diagnose and treat patients, with the patient-facing daily activities of clinical care. This aids recruitment efforts among prospective participants who may be willing to donate their biospecimens and data to clinical research studies, and patients receiving health care may also be receptive to being more readily recruited into clinical trials to help test the impact of new clinical tests or interventions on patient outcomes.
Private health systems such as Geisinger Health and Kaiser Permanente have also established their own biobanks, often in collaboration with academic genomics researchers. Commercial and biotech interests have also invested significant resources in constructing large repositories of DNA samples collected from patient donors. As genomics studies in the 1990s and later GWAS in the early 2000s generated findings suggesting the medical significance of DNA sequence variants, private biotech interests in the United States began to see business opportunities in selling consumers their own genetic information, and this led to the formation of many direct-to-consumer (DTC) genetic testing companies between 2008 and 2010, including Navigenics and the Google-backed company 23andMe (which, as competing companies struggled to generate profits, eventually came to dominate the industry). These companies have never made secret the fact that their value derives from the stockpile of consumer genetic data they have, which has been key to the formation of research collaborations with pharmaceutical companies and academic interests keen to mine the genomic data for medically useful information.
What was different about genome era biobanking, compared to biobanks of blood, tissues, and other biospecimens predating the HGP, was the sheer scope and scale of the efforts, both in terms of the number of participants involved but also in terms of the data generated. Although blood had been collected since the 1970s from patient donors undergoing chromosomal and genetic testing for specific Mendelian diseases, by the early 2000s, blood sampling and DNA collection became increasingly systematized in genomics research, morphing into massive population-scale repositories. Indeed, some of these biobanks emerged from longitudinal cohort studies in epidemiology and public health that had tracked individuals’ health for decades, which were retooled to the goal of identifying genetic variants associated with medical conditions (Fujimura and Rajagopalan 2020).
More recently, the US government established a national biobank of its own. With strong bipartisan support in Congress, in his 2015 State of the Union Address, President Barack Obama announced a $215 million commitment to precision medicine, declaring it a public health priority. “I want the country that eliminated polio and mapped the human genome to lead a new era of medicine—one that delivers the right treatment at the right time. Tonight, I’m launching a new Precision Medicine Initiative to bring us closer to curing diseases like cancer and diabetes—and to give all of us access to the personalized information we need to keep ourselves and our families healthier.” The Precision Medicine Initiative (now known as “All of US”) is centrally coordinated through the NIH. As the organizing node for the initiative, the NIH works with grantees from other genomics programs, which include medical centers involved in genomics studies that have built biobanks under other NIH funding streams. The aim is to recruit one million research participants into a longitudinal study of health and disease, with data sources including health histories, biomarker assays, genome sequencing, and data from wearable technologies and sensors that capture information on participants’ bodily functions and vital signs in real time.
All of these biobanks have aimed to study the nexus of genome sequence and personal health histories across large cohorts. They have employed a framework of enrollment and recruitment that asks patients to contribute to the project of building personalized or precision medicine, bringing public benefits to all through the donations of individuals. Participants are typically told their contribution could help build medical knowledge toward improving the overall health of their own and other communities, a goal that is seen to require massive datasets of intimate patient data.
The infrastructures that support DNA biobanks leverage existing strengths and resources at clinics, including longstanding clinical expertise in organizing and cataloging biological samples, and informatics expertise in using the EHR as a research tool. Participants donate blood, often filling out lengthy diet and physical activity questionnaires to supplement information on clinical encounters stored in the medical record (Fujimura and Rajagopalan 2020). Recruitment literature promotes biobanks as projects that generate substantial public benefits, asserting that participants’ contributions can help build medical knowledge toward improving the overall health of their own and other communities.
The research infrastructure subtended by biobanks is co-located with real-world patient populations at health centers. This clinic-patient-research lab nexus maps an inquiry- and discovery-oriented landscape that facilitates the pursuit of the translational promises of the genome. While early “bench-to-bedside” research often took place in basic science contexts, at arm’s length from clinical encounters, biobanks and the research agendas they support are located epistemologically and materially adjacent to the clinics they are intended to serve. This proximity grants them a “translational legitimacy” that was often lacking in earlier waves of research intended to be “translational,” and that legitimacy has also aided in securing long-term sustainability through the attention and financial support of NIH and other federal health research funding.
4. The Genome as Health Sentinel: Biobanks and Population-Scale DNA Collection and Analysis
DNA biobanks in the United States in the early 2000s were established as research tools to facilitate clinical studies. Architects made a focus on DNA a central part of these efforts, and biobanks served as important conduits by which the genome slowly gained clinical significance and made its way into clinical thinking. The significance of DNA to routine functions in clinical care, such as diagnosis and treatment of complex diseases, was not a foregone conclusion; rather, a significant amount of labor was involved in rendering DNA into forms useful and legible to clinical practitioners. In addition, DNA biobanks have benefited from and leveraged the significant gains in the technology available for high-throughput DNA studies over the past two decades, such as high-throughput genotyping and sequencing platforms. Genotyping technologies were first marketed commercially just as the first population-based DNA biobanks were getting underway in the early 2000s, while technologies for large-scale sequencing, like Illumina’s “next generation sequencing” platforms which would allow whole genome or exome sequencing, became available in the 2010s (Rajagopalan and Fujimura 2018).
In the United States, there was also a concerted push during the 1990s, across many (but not all) health systems, to digitize patients’ medical histories through EHRs. This push to consolidate health-relevant data into digital formats reshaped clinical duties and responsibilities, as medical practitioners were increasingly expected to use EHRs to capture information on medical events including diagnoses, procedures, medications, clinical notes, billing codes, and radiology, lab, and clinical observations. Not only could data flow from the professional settings of the clinic into the EHR, but increasingly, patients may contribute to improving population health by directly donating their data to research, or participate in monitoring their own bodies and vital signs, through wearables, body sensors, and health-tracking devices that collect and stream data directly and continuously into their own EHR. In this way, personal health surveillance has become an unabashedly data-generating activity, and a conscious public good. In turn, precision medicine, powered by genomics, encompasses the idea of “continuous medicine,” by which health outcomes could be improved by avoiding adverse outcomes through the predictive, pre-emptive powers ascribed to DNA and data from health-tracking devices.
Regardless of the domains in which these biobanks were constructed, academic, federal, or private, all aimed to actualize “big data,” as a means of tailoring and “personalizing” (and by implication improving) health care. They sought big data by triangulating patients’ biomarkers, genotypes, and genome sequences, with the personal health data and medical histories of large cohorts of patients. In so doing, DNA biobanks became simultaneously producers, warehousers, and consumers of biomedical “big data,” which has come to embody method, tool, and goal in clinical genomics research. By collecting patients’ DNA, genotyping or sequencing it, and linking it to EHRs, early biobank architects hoped to identify genetic loci that may be involved in diseases, with the aim of feeding any validated genetic findings back into patient EHRs across the health system, to inform future point of care decisions. The relations that have been built between DNA biobanks and EHRs in clinical research have sparked new workflows and new data analytic processes in clinical research. These are transforming how diseases are defined, profiled, diagnosed, and quantified for clinical trials and patient care—often using biomarker or genomic data. On the one hand, as much as DNA biobanks are research tools, they are also seen as important sources of data for intervening in clinical treatment decisions. On the other hand, as much as EHRs and patient records more generally have been used by clinicians and administrators as tools for streamlining and enhancing clinical care, they are also seen as a data resource for clinical genomics research.
5. Data in Sight: Viewing Diagnostic and Therapeutic Decision-Making Through the Lens of Genomic Data
American biobanking projects have accelerated the volume and speed of data accretion, accumulation, and aggregation around patients. They have done so partly through federally funded research consortia, whose goal has been to render genomes and genomic data visible to the daily practices of clinical care. The NHGRI supported a number of institutional investments in research consortia that pursued investigations into the clinical applications of data on genome sequence and genome variation. These represent a set of initiatives and attempts to realize a new phase of medicine, variously described as “translational,” “personalized” or “precision,” or “genomic” medicine, by attempting to install the genome and its variants at the center of clinical decision-making. The eMERGE network and the IGNITE program are illustrative and discussed below.
Other overarching initiatives similarly encouraged researchers to think more broadly and collectively about issues relevant to attempts to move genomics into clinical contexts. For example, from 2011 on, the Division of Genomic Medicine at NHGRI began hosting themed annual meetings numbered with roman numerals in a series entitled Genomic Medicine, bringing together researchers in clinical genomics. Researchers shared presentations on current research projects, and experiences with scientific and institutional challenges they were working through, as they tried to mold genome information into forms that hospital administrators, insurers, doctors, and patients could see as valuable. Each meeting emphasized different aspects of translating genome information to clinical applications, and the obstacles encountered, scientific, technical, and administrative. Uniting these initiatives, consortia, and meetings, was a focus on identifying genome variants that might have significance for specific diseases, conditions, or clinical outcomes, and developing these into genetic tests that could be done with greater speed, efficiency, and accuracy in clinical contexts, preferably at the point of care, to inform diagnosis and treatment options for patients (Manolio et al. 2017). Two of these initiatives in particular, eMERGE and IGNITE, illustrate how an explicit focus on developing clinical utility and validity of genome information has been a crucial goal of genomic medicine efforts.
5.1. eMERGE: Electronic Medical Records and Genomics Consortium
In 2007, as part of its Division of Genomic Medicine, the NHGRI initiated the establishment of a consortium of research groups called the eMERGE (Electronic MEdical Records and GEnomics) network, with the goal of developing methods for “using the electronic medical record as a tool for genomics research” (Gottesman et al. 2013). eMERGE member sites sought to generate a deeper understanding of a range of clinical conditions by analyzing patients’ health histories (as captured in EHRs), alongside their genomic risks (gleaned from the genomic variation each harbored in their DNA, archived in the biobanks established at each participating consortium site). eMERGE focused predominantly on high-throughput GWAS analysis as the means of identifying potentially medically relevant DNA variants from genome sequence. Thus, the eMERGE network and the research it pursued followed close on the heels of the first GWAS studies, as part of a medical genomics frame for thinking about and understanding disease and clinical care, from the vantage of genetic variation.
eMERGE initially comprised a consortium consisting of five member institutions: Group Health Cooperative/University of Washington; Marshfield Clinic; Mayo Clinic; Northwestern University; and Vanderbilt University. The network was intended “to provide support for investigative groups affiliated with existing biorepositories to develop . . . methods and procedures for genome-wide studies in participants with phenotypes . . . derived from EMR”1 (Gottesman et al. 2013). Assembling several working groups focused in areas that included clinical annotation of genomic variants and their potential functions; integration into the EHR, tracking the impact of genomic results on health-care utilization and outcomes for patients and families, and ethical, legal, and social implications (ELSI) of their work, eMERGE proceeded through three phases. The main aim of eMERGE Phase I was to demonstrate that EHRs could be used to describe the phenotypes that would qualify individual biobank participants as eligible for inclusion in case-control GWAS studies for a particular medical condition. eMERGE sites also developed approaches to address ethical questions attending the individually identifiable medical data warehoused in the electronic medical record (EMR) systems, including the de-identification of data and samples, protection of individuals’ confidentiality across the network while data and samples were shared, and strategies for dealing with the return of results.
Beginning in 2011, additional institutions joined eMERGE, including Geisinger Clinic and Mount Sinai School of Medicine and later, Children’s Hospital of Philadelphia, Cincinnati Childrens’ Hospital, and Boston Children’s Hospital. Most sites had genotyped and sequenced the DNA of 3,000 to 9,000 of their biobank participants; Vanderbilt had genotyped over 27,000 and Geisinger had genotyped over 61,000 biobank participants. Thus, the network in aggregate had collected data on over 100,000 DNA samples (eMERGE Network 2017). Previously, Phase I eMERGE researchers had used EMRs as a data source to help improve the way they conducted GWAS and uncover possible genotype-phenotype associations. In Phase II, there was a much stronger emphasis on applying these findings to real-world test cases in the clinic. Accordingly, Phase II eMERGE researchers designed pilot studies investigating the potential for clinical application. They used GWAS findings to inform the incorporation of genotype information back into the EMRs of participants in their biobanks. With the genome firmly ensconced in the medical record, their studies asked whether and how genome sequence information could inform (and potentially improve) clinical decision-making. Such decisions included pharmacogenetic questions like whether a given medication was likely to help a given patient and if so at what dose, or more general medical questions such as a patient’s risk for a disease condition. Whereas eMERGE Phase I sought to capitalize on the rich data of EMRs to “link genes to diseases” (Zimmer 2013), eMERGE Phase II sought to make strides toward the “possibility of comprehensive genomic information being the norm in clinical care” (Gottesman et al. 2013). In 2015, eMERGE leaders transitioned the consortium into Phase III, continuing earlier efforts from Phases I and II but giving increasing attention to elucidating the clinical significance of rare variants in one hundred genes deemed clinically relevant. In a sense, the successive phases of eMERGE sought increasingly greater degrees of translation into the clinic of its work and findings. These efforts were shepherded under a guiding principle of “implementation,” whereby finding ways to actualize genetics findings in the clinic, through clinical decision-making, became the primary goal.
5.2. IGNITE—Implementing Genomics in Practice Network Consortium
Key moments in doctor-patient encounters include decisions around diagnoses, and decisions around treatment. For many conditions, these decisions follow a fairly specific decision tree, established long before the advent of the genome era. Given that turnover in clinical standards of care is largely dependent on demonstrable evidence of a new intervention’s efficacy, tracking how genomic findings have been shepherded along the road to clinical relevance requires some attention to the ways that clinical researchers have framed ideas of clinical validity and utility around the genome. Put another way, medical genomics researchers have focused their efforts on answering the question, how can health practitioners “implement” the genome, accurately and in cost-effective ways that improve patient outcomes (thus satisfying both the clinical validity and clinical utility of genome information)?
One of the overarching concerns of eMERGE was to assess whether the incorporation of genomic information into the medical record, and more importantly, into clinical decision-making, could be done cost-effectively, an important consideration if genomic medicine tools were to be broadly embraced by insurers and providers. eMERGE Phases II and III in particular emphasized research to evaluate cost-effectiveness strategies to render the clinical implementation of genomics more palatable to insurers and hospital administrators. These efforts were enshrined in NHGRI’s Strategic Plan, “Charting a Course for Genomic Medicine from Base Pairs to Bedside,” released in 2011, which emphasized the need for identifying and pursuing the most promising pathways by which genomic medicine could be approved for clinical deployment (Green 2011). For example, one of the main focal areas for eMERGE’s implementation efforts was identifying and validating pharmacogenes, or genes involved in drug tolerance or drug metabolism in the body. The motivating idea was that DNA variation at these genes might assist clinical decisions about which drug therapy to use for a given patient, or suggest a tolerable dose range to prescribe to individuals with certain variants in those genes, in order to maximize drug efficacy and minimize adverse drug reactions. To that end, eMERGE II sequenced eighty-four candidate pharmacogenes in nine thousand biobank participants across participating consortia sites, and established an ongoing investigation to see “what happens when large-scale genetically tailored drug prescription is implemented across several large medical centers . . . [representing] a first step toward a vision of incorporating large-scale sequence information into the flow of routine health care” (Rasmussen-Torvik et al. 2014; Bush et al. 2016).
In addition to eMERGE, another NHGRI-initiated effort that sought to clinically actualize genomic medicine was IGNITE, a funding program within NHGRI that awarded its first set of grants in 2013 and 2014. While eMERGE’s member institutions represented relatively well-resourced settings in which genomics research and clinical trials of genomic interventions might readily fit within existing infrastructures for clinical research, IGNITE had a slightly different emphasis, which was to focus on “real-world health care delivery.” Successfully funded projects were pitched by academic health centers that had partnered with health-care providers with little prior exposure to genomics tools or approaches, including health centers in underserved areas, family practices, and military and VA hospitals. In this way, IGNITE filled a niche that eMERGE could not address, by seeking to demonstrate the clinical validity and utility of genome information at “real-world” sites (Ginsburg et al. 2021). Such sites might be less or under-resourced, but the hope was that if genomic data could make a demonstrable and cost-effective difference to patient outcomes in such contexts, the findings would be compelling to insurers, payers, and hospital administrators while rendering the economic, resource, and labor price tag of clinical genomics interventions a less imposing barrier to widespread embrace of genomic interventions. This call for hard evidence of the clinical utility of the genome, from the early planning stages of IGNITE, was especially important for prioritized efforts to extend the reach of genomics to underrepresented populations and address socioeconomic disparities in access to genomics interventions. The first phase of IGNITE funded several pilot demonstration projects. The second phase funded “pragmatic” clinical trials, which aimed to test the efficacy of genomic interventions in socioeconomically diverse clinical settings.
The IGNITE program is especially revelatory of how genomic data has been framed and developed for use in the clinic. While clinical trials have yet to convincingly establish medical relevance for most genomic variants (Rasmussen-Torvik et al. 2014), medical relevance is being sought in real time through targeted projects that establish relevance in particular contexts. Just as precision medicine is being framed as uniquely tailored to the individual, it is also being uniquely tailored to particular sites.
Through their work, both eMERGE and IGNITE programs attempted to instantiate predictive capacities within DNA sequence for use in medicine. They attempted to make a case for the clinical utility and clinical validity of genetic information, especially around the elevated risk for disease or adverse drug reactions that they attributed to certain DNA variants. Their work aimed to stabilize the inclusion of genetic information in the medical record as a matter of routine clinical practice. IGNITE, more than eMERGE, tested the utility of genomic information in challenging settings. Whereas eMERGE conducted trials in what came to be seen as “ideal” clinical settings receptive to genomics, with significant infrastructural and professional capacities for processing genomic materials and data, IGNITE sites were chosen to be more reflective of “real-world” health-care settings where such resources may not be available, and where partnerships with better-equipped sites (or outsourcing of genome-associated activities, such as genotyping, to commercial outfits) would need to be operative.
5.3. Clinical Sequencing Exploratory Research (CSER) and Clinical Genomics Resource (ClinGen)
Two other programs at NHGRI have been important in promoting the diffusion of genomics into clinical work. The first is CSER, a consortium initiated in 2010 to evaluate the introduction of whole genome sequencing methods at the point of care in traditionally underserved and underrepresented clinical populations. The program’s several sites, primarily affiliated with academic medical centers as with eMERGE, sought to define and generate evidence of the utility of clinical genome sequencing. In its Phase II, initiated in 2017, CSER was renamed to Clinical Sequencing Evidence-Generating Research, to better reflect its aims. CSER pursued a strong focus on ensuring that ELSI issues related to clinical sequencing received equal attention, and that stakeholder input (including from patients and providers) was solicited in the assessment of its work and best practices. Participating sites in CSER included Kaiser Permanente Northwest, Baylor College of Medicine, UNC-Chapel Hill, Icahn School of Medicine at Mount Sinai Hospital, University of California San Francisco, HudsonAlpha Institute for Biotechnology, and UW-Seattle.
ClinGen, a resource initiated in 2013 and funded primarily by NHGRI, is a central clearinghouse for evidence on the relationships between gene variants and diseases. It brings together clinical and research experts to develop standard processes for assessing DNA variants and their links to disease, and it oversees standardized, federated databases of variants for clinical and research use. It aims to “develop and implement standards to support clinical annotation and interpretation of genes and variants” and “develop data standards, software infrastructure and computational approaches to enable curation at scale and facilitate integration into healthcare delivery” (Rehm et al. 2015).
The programs discussed above all furthered the project of precision medicine by developing the means and mechanisms for durably yoking genomic data to clinical needs, at scientific, technical, institutional, and regulatory levels. Involving academic medical centers with a long tradition of excellence in both clinical research and clinical care strategically positioned these programs within key urban patient populations where the chance of successful acceptance and integration of genomic findings was reasonably high. In aggregate, the NHGRI’s suite of programs dedicated to genomic medicine have established strong institutional frameworks for translating the language of genomics into forms that are legible to clinical contexts, and that address pressing and vexing problems in clinical care as a means of defining the ways that genomics can uniquely contribute to clinical decision-making.
6. Genomes and Algorithms at Work: Restructuring Clinical Labor Around DNA
The thrust of NHGRI’s programs post-HGP have focused on implementing the medical promises of the genome. Through these efforts, big-data collection, analytics, and cloud computing have become pervasive in medicine. What transformations in work practice, both in the clinic and in research, have taken place and continue to take place to facilitate these transitions?
One arena of transformation has taken place in the realignments and synchronization of research and clinical priorities. While the eMERGE and IGNITE research programs illustrate the continuing R&D character of contemporary attempts to instantiate genomics in clinical practice, they nevertheless have also modeled what a clinical experience oriented around DNA might look like if deployed. Thus, these programs blur the interface between research and clinical practice; clinical practice is part of research, and vice versa.
This can be seen for example in how clinical research workflows have been transformed in the genome era. The technologies for populating data in EHRs have changed dramatically. Data in medicine used to be limited to a few vital signs taken at the point of care, primarily physician generated; through the increasing prevalence of wearables, patients now actively participate in streaming their own data into their EHR. And, with entire genomes potentially entering the EHR, the sheer scope and scale of data in medicine has radically shifted. In order to triangulate across disparate data sources, such as the EHR and individual’s genotyping or sequencing data, these massive datasets first have to be generated, cleaned, and prepped for analysis, using new tools and methods that must be innovated to handle such volumes of data at scale. Such functions are typically carried out by teams of informatics specialists affiliated with DNA biobanks. Thus, the EHR and genome data are both a product of, and a conduit for, an increasing computational presence in medical research, which spills over into medicine through the kinds of implementation demonstration projects that eMERGE and IGNITE have supported. The EHR and the genome are both molded and shaped through these computational activities, which archive, organize, and analyze the data and render them into forms that are recognizable to clinicians, such as key nodes in decision trees that help health professionals make diagnostic or treatment decisions.
These developments are furthered by the processes that are embedding algorithms in clinical decisions, particularly around the interpretation of genomic variants. The translation of genomics from bench to bedside has been mediated through computation. It is the algorithms designed for use on genomic data that are tasked with mediating the interpretive dynamics of this translation. Researchers have designed and used algorithms to generate risk estimates, which summarize the risk of incurring various outcomes associated with having a particular DNA variant, such as an elevated risk for a particular disease, or an elevated risk for an adverse reaction to a drug, or to being non-responsive to a particular therapy. Researchers also design and use algorithms that could help decide whether to prescribe a given drug to a given patient, for medications like warfarin (an anti-clotting drug) or Herceptin (a chemotherapy), whose turnover in the body is thought to be mediated by particular genetic variants in drug-metabolizing genes. Algorithms therefore play a key role in genomic implementation efforts in that they act as translators of genomic information into medically actionable findings.
Participants in research trials, especially at sites in the eMERGE and IGNITE networks, have been among the first to have their genomic variants assayed and their genotypes and in some cases genome sequences deposited and accessible to their health-care providers in their EHR. For example, eMERGE sites have deployed pharmacogenomics testing for genome variants associated with drug responses, for thousands of biobank participants in the United States. It remains to be seen the extent to which prescribing physicians will make use of these data, but the availability of genome data and genome-interpreting algorithms are already beginning to shift the modalities of clinical decision-making as the collection and analysis of patients’ sequence information is slowly being integrated into clinical workflows, for example, through eMERGE’s pharmacogene sequencing project.
Millions more patients now have access to their genomes online as customers of DTC genetic testing companies like 23andMe. Expanded algorithmic capacities, the wheelhouse of Silicon Valley’s tech companies, has led many established and startup commercial interests to try to use data and computation to intervene in, and possibly disrupt, health care. For example, companies are using artificial intelligence (AI) and machine learning to develop “real-time” clinical decision support systems that could incorporate an assessment of genome variants to help doctors make diagnostic or treatment recommendations. Such systems, they argue, could potentially sift through exponentially large data streams and detect early warning signs of disease in patients, at home or at the point of care, and help doctors more rapidly apply preventive screening and therapy. At IBM, programmers tried to use the Watson supercomputer’s AI to determine optimal cancer treatments and assign patients to clinical trials, though with mixed success (Chen 2018). Google developed an AI system to diagnose breast cancer tumors and another, DeepVariant, to sift through genome data and identify mutations (Knight 2017).
Algorithms being developed by clinical researchers and by tech companies seek to introduce AI and machine-learning approaches into point-of-care diagnosis and treatment. These algorithmic approaches portend an increasing reliance on computational and algorithmic thinking to make key decisions about patients’ lives. But algorithms do not act on their own. Algorithms play directly with data, but in seeking to find predictive links between DNA sequence variation and medically important outcomes, their coders must, like all scientists, make some assumptions and some presuppositions to coax the algorithm to do work that will be seen as legible and meaningful in a clinical context. Algorithmic computation is always interpretive; human coders seek to align “raw” data with real-world contingencies and needs, and in so doing they craft and create analyses that are embedded in the code but also reflect layers of interpretative assumptions that take hold prior to and in interaction with the data.
One example of this is seen in pharmacogenomics algorithms, which seek to predict a more tailored dose of a particular drug, like warfarin, so that it is more compatible with the suspected drug-metabolizing effects of an individual’s genetic variants in key drug-metabolizing genes. In defining such algorithms, clinical bioinformaticians make assumptions about the role of other potentially intervening factors beyond genetics, such as age, sex, and race, especially how they might influence or mediate a prediction based on genetics alone. This adds a layer of human interpretative work to the seemingly “objective” algorithm, and can introduce social biases into the equation (Kahn 2012).
This chapter has attempted to chart the movements, material and epistemic, that are hastening the transformations in clinical care needed to accommodate genomic data. These transformations are taking place even as an understanding of the vast majority of the genome’s clinical significance remains in its infancy. The NIH has been a key player in rendering the genome as “data,” both in the sense of aggregations of large-scale bits of information that can be acted on digitally by computers and algorithmic tools, and in the sense of the material configurations and professional alignments across research and clinical settings that have built the complex web of work practices that give meaning and significance to these aggregations. Such data has been promoted as being of “medical” consequence, but this is not innate or given to the data itself, or to DNA sequence; it is realized and actualized through the interpretive work that human-coded algorithms perform, and that human researchers of DNA variation enact through their work.
7. Conclusion
Though genetic testing is not new to the clinic, the era of genomic medicine portends significant re-orientations in the practices of health care, and especially in the kinds of predictions of disease risk or drug response that have proven more ambiguous than many examples of Mendelian inheritance of disease-proximal genes that are more tightly linked to disease. If widely implemented, such risk predictions would entail significant reorientations in the practices of health care.
This chapter has asked two questions: How is genome information framed and developed for use in clinical decisions? And what transformations in work structure and practice, in research and clinical practice, have taken place to facilitate this transition? From its inception, the architects of the HGP premised the project on a medical imperative, and they proceeded by building into the project an abiding orientation to improving human health, culminating in the draft genome in 2001. Tracing the ways that the genome has been imbued with meaning over the past twenty years illustrates how rendering the genome as sequences and data has allowed researchers to attempt to make medicine algorithmic. As this chapter has discussed, the representation of DNA as sequence opened up intellectual and institutional spaces through which the layers of work built around the project of elucidating DNA themselves became conduits for big data into biomedicine. Put another way, big data became the cargo that DNA and all the activities organized around it have carried into medicine and into patient-proximal activities in the clinic.
As this chapter has shown, the siting of biobanks at US biomedical research clinics, where bench science labs sit adjacent to sites of patient care, has created key institutional sites at which these practices have taken shape. Biobanks today are large repositories linking patients’ health histories and current health data to their biological specimens and DNA samples. They create durable links between electronic medical histories, DNA sequence, and risk estimates for health and disease. In this way, biobanks serve as infrastructures that hold together particular health-relevant modalities, bringing them into new epistemic alignment and interaction with each other. Clinicians and researchers, through their work, brought patients, blood, DNA, genomics technologies, information technology, algorithms, and biostatistical architectures into new constellations of work and knowledge production. The establishment of large biobanks and patient data repositories, and the crafting of new tools for algorithmically sifting through these datasets, were aligned with and oriented to the goal of translating genomic findings to clinical decision-making practice.
The efforts described in this chapter illustrate how the research clinic as an institution in American biomedicine became a key site at which the envisioned promises of precision medicine are being put to the test. All of these efforts have been aimed at rendering genomic information into useable forms for health care in clinical settings. Though demonstrations of clinical utility have been slow, genomics continues to be a driver of medical research especially at large academic hospitals and a key recipient of federal and private investment. Hospitals continue to invest resources in high-throughput genome technologies in the belief that the data they generate in copious volumes will soon be clinically meaningful. A prominent impetus for these efforts has been the strengthening of institutional linkages offered by the construction of biobanks, which durably cement (while blurring away distinctions between) biomedical research and clinical care.
The NIH has been a central organizing node and source of financial support for the project of building the sites and tools of precision medicine. A number of programs and consortia convened under the aegis of the NHGRI and supported by NIH funding and resources explicitly took on the project of translating genomics to the clinic. By bringing together researchers and clinicians to advance the project of precision medicine, they have helped establish the material, epistemological, and professional conduits along which the knowledge and findings labeled as precision medicine flow. They have also evangelized the project of precision or genomic medicine, raising its visibility and enrolling new actors and supporters into the networks through which it is being enacted.
Amid these developments, precision medicine has proven to be part conceptual, part research agenda, and part professional movement, a “buzzword science” that nevertheless powerfully marshals resources and labor under its umbrella. Furthermore, as this chapter has shown, precision medicine has served as a growing platform for innovating new ways of doing cross-regional, multi-site science, and a conduit for generating, analyzing, and applying large-scale datasets comprised of individual-specific data types for assessing health susceptibilities. “Implementing the genome” has meant finding ways of plugging genomic data into the existing structures of everyday medicine, discerning open needs or questions and, in some cases, constructing new ones within daily health care, where the explanatory power ascribed to DNA sequence is seen to be a bulwark for caulking or buttressing against these cracks and gaps.
Precision medicine, however, belies the uncertainties that persist despite the best efforts of researchers to pin down the precise meaning and significance of DNA sequence and DNA variation to individuals’ current and future health. The clinical recommendations that attend genotypes are virtually always probabilistic, statistical estimates, never stable or firm. The term precision, then, could even be said to be misleading, since DNA sequence, or at least scientists’ current understandings of its clinical significance, introduces just as much uncertainty and probability into clinical treatment evaluations and calculations as it strives to inject an aura of exactitude and certainty.
Despite, or perhaps because of, this, health systems continue to expand the infrastructures for collecting, assessing, and triangulating across massive and disparate datatypes, in the belief that more data will help deliver more efficient, precise, and standardized patient care.
Note
1. The terms electronic medical record (EMR) and electronic health record (EHR) have been used interchangeably, and I use both in this chapter, depending on context and usage in cited sources.
References
- Angier, N. 1990. “Great 15-Year Project to Decipher Genes Stirs Opposition.” New York Times, June 5.
- Bush, W. S., D. R. Crosslin, A. Owusu-Obeng, J. Wallace, B. Almoguera, M. A. Basford, et al. 2016. “Genetic Variation Among 82 Pharmacogenes: The PGRNseq Data from the eMERGE Network.” Clinical Pharmacology & Therapeutics 100 (2): 160–69.
- Callaway, E. 2017. “New Concerns Raised over Value of Genome-Wide Disease Studies.” Nature 546 (7659): 463.
- Chen, A. 2018. “IBM’s Watson Gave Unsafe Recommendations for Treating Cancer.” The Verge, July 26.
- Electronic Medical Records and Genomics (eMERGE) Network. 2017. Accessed August 25, 2017. https://www.genome.gov/27540473/electronic-medical-records-and-genomics-emerge-network/.
- Fujimura, J., and R. M. Rajagopalan. 2020. “Race, Ethnicity, Ancestry, and Genomics in Hawai‘i: Discourses and Practices.” Historical Studies in the Natural Sciences 50 (5): 596–623.
- Ginsburg G. S., L. H. Cavallari, H. Chakraborty, R. M. Cooper-DeHoff, P. R. Dexter, M. T. Eadon, et al. 2021. “Establishing the Value of Genomics in Medicine: The IGNITE Pragmatic Trials Network.” Genetics in Medicine 23 (7): 1185–91.
- Goldstein, D. B., S. K. Tate, and S. M. Sisodiya. 2003. “Pharmacogenetics Goes Genomic.” Nature Reviews Genetics 4:937–47.
- Gottesman O., H. Kuivaniemi, G. Tromp, W. A. Faucett, R. Li, T. A. Manolio, et al. 2013. “The Electronic Medical Records and Genomics (eMERGE) Network: Past, Present, and Future,” Nature 15:761–71.
- Green, E. D. 2011. “Charting a Course for Genomic Medicine from Base Pairs to Bedside.” Nature 470:204–13.
- International Human Genome Project Consortium. 2001. “Initial Sequencing and Analysis of the Human Genome.” Nature 409:860–921.
- Kahn, J. 2012. “The Troubling Persistence of Race in Pharmacogenomics.” Journal of Law, Medicine & Ethics 40 (4): 873–85.
- Kevles, D. J. 1997. “Big Science and Big Politics in the United States: Reflections on the Death of the SSC and the Life of the Human Genome Project.” Historical Studies in the Physical and Biological Sciences 27 (2): 269–97.
- Knight, W. 2017. “Google Has Released an AI Tool That Makes Sense of Your Genome.” MIT Technology Review, December 4.
- Lander, E. S., L. M. Linton, B. Birren, et al. 2001. “Initial Sequencing and Analysis of the Human Genome.” Nature 409:860–921.
- Manolio T., D. M. Fowler, L. M. Starita, et al. 2017. “Bedside Back to Bench: Building Bridges Between Basic and Clinical Genomic Research,” Cell 169 (1): 6–12.
- McCarty, C. A., D. Chapman-Stone, T. Derfus, P. F. Giampietro, and N. Frost. 2008. “Community Consultation and Communication for a Population-Based DNA Biobank: The Marshfield Clinic Personalized Medicine Research Project.” American Journal of Medical Genetics 146A:3026–33.
- November, J. 2012. Biomedical Computing: Digitizing Life in the United States. Johns Hopkins University Press.
- Paul, D. 1995. Controlling Human Heredity: 1865 to the Present. Humanities Books.
- Prainsack, B. 2017. Personalized Medicine: Empowered Patients in the 21st Century? NYU Press.
- Precision Medicine Initiative Working Group. 2015 “The Precision Medicine Initiative Cohort Program—Building a Research Foundation for 21st Century Medicine.” Precision Medicine Initiative (PMI) Working Group Report to the Advisory Committee to the Director, NIH.
- Rajagopalan, R. M., and J. H. Fujimura. 2018. “Variations on a Chip: Technologies of Difference in Human Genetics Research.” Journal of the History of Biology 51 (4): 841–73.
- Rasmussen-Torvik, L. J., S. C. Stallings, A. S. Gordon, et al. 2014. “Design and Anticipated Outcomes of the eMERGE-PGx Project: A Multicenter Pilot for Preemptive Pharmacogenomics in Electronic Health Record Systems.” Clinical Pharmacology & Therapeutics 96 (4): 482–89.
- Rehm, H. L. 2017. “A New Era in the Interpretation of Human Genomic Variation.” Genetics in Medicine 19:1092–95.
- Rehm, H. L., J. S. Berg, L. D. Brooks, et al. 2015. “ClinGen—the Clinical Genome Resource.” New England Journal of Medicine 372 (23): 2235–42.
- Roden D. M., J. M. Pulley, M. A. Basford, et al. 2008. “Development of a Large-Scale De-identified DNA Biobank to Enable Personalized Medicine.” Clinical Pharmacology & Therapeutics 84 (3): 362–69.
- Stern, A. M. 2012. Telling Genes: The Story of Genetic Counseling in America. Johns Hopkins University Press.
- Stevens, H. 2013. Life Out of Sequence: A Data-Driven History of Bioinformatics. Johns Hopkins University Press.
- Strasser, B. 2009. “Collecting, Comparing, and Computing Sequences: The Making of Margaret O. Dayhoff’s Atlas of Protein Sequence and Structure, 1954–1965.” Journal of the History of Biology 43 (4): 623–60.
- US Department of Energy and the Human Genome Project. 1996. To Know Ourselves. Edited by Douglas Vaughan. https://doe-humangenomeproject.ornl.gov/wp-content/uploads/2022/08/tko.pdf.
- Zhang, S. 2018. “300 Million Letters of DNA Are Missing from the Human Genome.” The Atlantic, November 28.
- Zimmer, C. 2013. “Linking Genes to Diseases by Sifting Through Electronic Medical Records.” New York Times, November 28.