Notes
9 Technological Change Driving Scientific Questions
Genomic Sequencing as a Case Study
Adam L. Felsenfeld and Kris A. Wetterstrand
1. Introduction
It is not controversial to consider that scientific advances have often been driven by technical advances that allowed new observations to be made. These observations have then become the basis for generating questions that can be addressed by experiment and further observation, and/or observation of consistent association between phenomena. The International Human Genome Project (IHGP), and the era of large genome projects that followed—a period beginning in the late 1990s and continuing through the present time—provides a unique lens with which to examine the relationship between technological advance and scientific advance, for several reasons. First, this time period is relatively brief and unusually productive. Second, technological advances that directly affected the quantity and quality of the data generated occurred in rapid succession and were widely adopted, so the temporal relationships are easy to establish. Also, the nature of the data provided direct proxies—cost per base pair or cost per genome, and overall throughput—for technical progress, making the technical progressions easy to measure. Third, although the literature in genomics is now large, in recent memory it was still possible for one person to be aware of most of the significant advances,1 and to speak with most members of the core large-scale genome sequencing community. As a result of these circumstances, it has been possible for us to track the history of cost decreases in genome sequencing in a reasonably consistent way, between the completion of the draft Human Genome Project (HGP) in 2001, and the present. This information has been widely disseminated to the point of being iconic, and is reproduced from the National Human Genome Research Institute (NHGRI) website (Wetterstrand 2023) as the starting point for this discussion (Figure 9.1).
We note that some important sequencing projects sequenced exomes (also other forms of targeted sequencing) rather than whole genomes—these will be noted in the text. Exome sequencing provides only data from coding and nearby regions (a few information-rich percent of the sequence) and has consistently cost between one-third and one-seventh of a whole genome sequence since the advent of exome sequencing.
A more detailed discussion of costs, including cost components, is detailed elsewhere in this volume (see chap. 3).
Figure 9.1. Cost per sequencing a human genome between 2001 and 2022. Data were taken from quarterly retrospective assessments of genome sequencing costs from grantees of the National Institutes of Health (NIH), funded by the National Human Genome Research Institute (NHGRI) Genome Sequencing Program (GSP), and its predecessors [see chapter 3]. Costs here are tracked from the period directly after the initial publications of the draft human genome sequence (Lander et al. 2001) estimated to have cost on the order of $3 billion, spent across multinational efforts, and including the costs of pilot efforts. Costs include all those allocated to all steps from sample preparation through release of data (i.e., sample preparation, library preparation, genome sequencing, data processing, and deposition, including amortized equipment costs where applicable and allocable indirect costs).
Figure Description
A graph showing the same NHGRI genome sequencing cost curve data in Figure 3.1. The first data point begins at a cost of roughly100 million USD per genome in August 2001. This part of the graph contains consistent decreasing cost data representing Sanger-based (dideoxy chain termination) capillary sequencing until January 2008 where second generation or “next-generation” sequencing by synthesis platforms come online at NHGRI sequencing centers. There is an exponential drop from this point until 2012 or so, where cost decreases are smaller each reporting period. There is one further inflection point where high throughput next generation sequencing (branded as Illumina X Ten) platforms come online and drive costs to roughly 1000 USD per genome. In addition to the cost curve data, a straight line shows the inverse of “Moore’s Law.” That line descends less steeply than the curve data, dropping from 100 million dollars to 100 thousand dollars.
Date x-axis | Cost per Genome y-axis | Moore’s law y-axis |
|---|---|---|
September 2001 | $95,263,072 | $95,263,072 |
March 2002 | $70,175,437 | $80,106,376 |
September 2002 | $61,448,422 | $67,361,164 |
March 2003 | $53,751,684 | $61,770,460 |
October 2003 | $40,157,554 | $56,643,761 |
January 2004 | $28,780,376 | $43,678,311 |
April 2004 | $20,442,576 | $40,053,188 |
July 2004 | $19,934,346 | $36,728,935 |
October 2004 | $18,519,312 | $33,680,582 |
January 2005 | $17,534,970 | $30,885,230 |
April 2005 | $16,159,699 | $28,321,881 |
July 2005 | $16,180,224 | $25,971,279 |
October 2005 | $13,801,124 | $23,815,768 |
January 2006 | $12,585,659 | $21,839,156 |
April 2006 | $11,732,535 | $20,026,594 |
July 2006 | $11,455,315 | $18,364,468 |
October 2006 | $10,474,556 | $16,840,291 |
January 2007 | $9,408,739 | $15,442,615 |
April 2007 | $9,047,003 | $14,160,940 |
July 2007 | $8,927,342 | $12,985,640 |
October 2007 | $7,147,571 | $11,907,884 |
January 2008 | $3,063,820 | $10,919,578 |
April 2008 | $1,352,982 | $10,013,297 |
July 2008 | $752,080 | $9,182,234 |
October 2008 | $342,502 | $8,420,146 |
January 2009 | $232,735 | $7,721,307 |
April 2009 | $154,714 | $7,080,470 |
July 2009 | $108,065 | $6,492,820 |
October 2009 | $70,333 | $5,953,942 |
January 2010 | $46,774 | $5,459,789 |
April 2010 | $31,512 | $5,006,648 |
July 2010 | $31,125 | $4,591,117 |
October 2010 | $29,092 | $4,210,073 |
January 2011 | $20,963 | $3,860,654 |
April 2011 | $16,712 | $3,540,235 |
July 2011 | $10,497 | $3,246,410 |
October 2011 | $7,743 | $2,976,971 |
January 2012 | $7,666 | $2,729,894 |
April 2012 | $5,901 | $2,503,324 |
July 2012 | $5,985 | $2,295,558 |
October 2012 | $6,618 | $2,105,036 |
January 2013 | $5,671 | $1,930,327 |
April 2013 | $6,618 | $1,770,118 |
July 2013 | $5,550 | $1,623,205 |
October 2013 | $5,096 | $1,488,485 |
January 2014 | $4,008 | $1,364,947 |
April 2014 | $4,920 | $1,251,662 |
July 2014 | $4,905 | $1,147,779 |
October 2014 | $5,731 | $1,052,518 |
January 2015 | $3,970 | $965,163 |
April 2015 | $4,211 | $885,059 |
July 2015 | $1,363 | $811,602 |
October 2015 | $1,245 | $744,243 |
May 2016 | $1,176 | $625,831 |
August 2016 | $1,508 | $573,890 |
November 2016 | $1,356 | $526,259 |
February 2017 | $1,015 | $482,582 |
May 2017 | $1,333 | $442,529 |
August 2017 | $1,134 | $405,801 |
November 2017 | $1,844 | $372,121 |
February 2018 | $1,232 | $341,237 |
May 2018 | $1,463 | $312,916 |
August 2018 | $1,467 | $286,945 |
November 2018 | $1,392 | $263,130 |
February 2019 | $989 | $241,291 |
May 2019 | $606 | $221,265 |
August 2019 | $942 | $202,901 |
November 2019 | $695 | $186,061 |
February 2020 | $645 | $170,618 |
May 2020 | $702 | $156,458 |
August 2020 | $689 | $143,472 |
November 2020 | $512 | $131,565 |
February 2021 | $851 | $120,645 |
May 2021 | $454 | $110,632 |
August 2021 | $562 | $101,450 |
November 2021 | $552 | $93,030 |
February 2022 | $525 | $85,309 |
May 2022 | $525 | $78,229 |
The salient feature of Figure 9.1 is that the costs of generating genome sequence data have plunged six orders of magnitude between the generation of the first genome sequence and the present time. Although platform technology improvement was in the main responsible for driving large cost decreases (see chap. 6)—and has been even more dominant since about 2007—it was not the only factor. Significant incremental cost decreases were realized by inventiveness within large genome sequencing centers, including taking advantages of efficiencies of scale, and through incremental technology development to streamline the use of the platforms (e.g., through improved sample preparation, or more efficient reagent use). In addition, once a high-quality human genome reference was made available in 2001, it became possible to obtain a new good-quality genome by “re-sequencing,” an obsolete term that refers to the ability to assemble genome sequence reads from a new individual with the aid of a reference sequence scaffold obtained from the same species (see chap. 3).
Before proceeding, it is important to note that this perspective is written based on the personal experience of the authors working with the National Institutes of Health (NIH)/NHGRI component of the IHGP and that program’s successors spanning the period from 1997 to the present. While this experience is direct, rich, and encompasses the early childhood through the early adolescence of the genomic age, it is liable to miss important viewpoints developed outside the IHGP and NHGRI. Moreover, we focus on genomic sequencing for the sake of obtaining direct information about genomes; a more extensive discussion may be worthwhile for other sequencing modalities (e.g., applications of sequencing that are assays for gene expression or chromatin state, etc.).2
2. Cost and Questions
The main aim of this paper is to draw attention to some prominent examples where it is easy to see the relationship between genome sequencing costs and throughput, the kinds of new questions that could be asked and answered, and some of the insights they provided. Note that the time periods and cost ranges for these examples are overlapping, because that is the way that the science developed. The range of dates provided for each example are meant to indicate the time and cost when the type of project described became technically and practically possible, and the time period over which the most rapid new progress was made, and during which the listed insights and questions developed. Importantly, the costs in the section headings are costs per human genome as listed at the time. However, several of the scientific questions mentioned in each section could be explored, indeed, had to be explored, using less sequencing data than required for a full human genome. Most of the types of projects discussed below continued beyond the indicated dates, even if the specific project mentioned ended.
2.1. $109 per Genome; 2001 and Prior. The IHGP and Pilot Genomes
In order to have this discussion, we need a reasonably consistent measure of costs over time for comparison. However, it is problematic to start from the IHGP itself, which is estimated to have cost over $2.7 billion in 1991 dollars overall. But this estimate depends on what specific activities are included. For example, the pilot efforts for the IHGP included substantial funding for sequencing genomes of E. coli, S. cerevisiae, C. elegans, and D. melanogaster (note that NHGRI was not the sole funding source for some of these). In addition, a substantial amount was spent by NIH early on for HGP pilot efforts starting in 1996 (National Human Genome Research Institute 1995) and extended in 1998 (National Human Genome Research Institute 1998a). Moreover, the non-US contributions are difficult to estimate. These costs are not consistently included in overall IHGP cost estimates. More to the point, for the purposes of this article, the $2.7 billion figure is not useful as a point of comparison.
The consistent cost accounting (per base pair, and later per genome, reflected in Figure 9.1) did not begin until after this period.3 At the start of this time period, much of the sequencing was still being done on slab gels, but by the end capillary sequencing had enabled the cheaper higher-throughput production of higher quality data that ultimately was used for generating the vast majority of data for the project.
While the major rationale for sequencing the human genome was to understand how our genes influence our health, this was a long-term goal that would mostly not start to be realized until years later (though perhaps it is more accurate to say that we had a few examples before the availability of a human reference sequence, growing to many thousands through the application of some of the examples below, most of which are still poorly understood or characterized). The more proximate questions were more fundamental, and in the context of this article, rather poorly formed simply because we then knew so little about genomes. These are questions like:
- What is the overall sequence-level structure of human and other genomes?
- How can we identify genes within the sequence?
- How many genes are there?
Note that, along with these biological questions, other sorts of important questions were raised and addressed from the start of the IHGP to the present time, such as “How can we improve our ability to sequence and assemble genomes?” and “How will the new scientific information affect human society and what are the ethical, legal, and social implications?” These important questions are not addressed in this article.
The major insights for the first human genome assemblies were described in two papers (Lander et al. 2001; Venter et al. 2001). Between the two publications, these include the approximate number of genes (though the initial IHGP paper [Lander et al. 2001] somewhat overestimated this number), and the highly uneven distribution of almost any base-level feature that could be assessed (GC content, genes, transposable elements, etc.). Hints were seen about the origins of some genes; and there was an initial characterization of the number and distribution of transposable elements, repeat structures, mutation rates, and beginning information about human genetic variation at genome scale.
From our experience, most of these questions were not asked before the first human assembly was available, as high-level drivers at the NIH of the project in advance. Instead, at least for the IHGP effort, many of these questions were being pursued by individual investigators interested in individual topics who were more or less associated with the IHGP (the “Genome Analysis Group”), and who were brought together late in the project with the IHGP genome sequencing centers to draft the main themes of the 2001 paper (Lander et al. 2001).
Although it should be obvious, it cannot be overstated that the first human genome assemblies, and assemblies for other species, were prerequisites for all the work described below.
2.2. $108–107 per Genome: ~2002–2005. The Advent of Large-Scale Human Variation Studies
The International HapMap Project (HapMap) was initiated and published results at this time (National Human Genome Research Institute 2002; International HapMap 2005). Notably, the cost point per genome was still prohibitive for this effort, even given the very compelling rationale to understand human variation, a key step in developing the ability to make correlations between variation and disease or other phenotype at any scale. Consider that, in order to identify enough common (5 percent or greater) variants in several populations, over 250 genomes were sampled—at this time period, this might have cost upwards of $2.5B. Fortunately, a work-around was developed. Instead of sequencing multiple human genomes to 7X capillary coverage (either clone-based or whole genome shotgun) as was done with the initial assemblies, targeted portions of the genome were captured, sequenced, and assembled to somewhat lower quality to identify variants, which were then genotyped on a larger number of samples representing multiple populations. A fuller description of HapMap is included elsewhere in this volume (see chapter 5). Note that even with this work-around, and even given the high level of justification for this project, the lowered cost per genome was still enabling, as the methods used still required a lot of genome sequencing.
Although HapMap was primarily motivated as a tool for enabling association studies between variants and phenotypes, the project enabled a number of questions including:
- What is the amount and structure of human genomic variation?
- What features related to genetic variation are conserved, and different, within and among populations?
- Addressing these led to the following insights (International HapMap 2005):
- There are ~three million genomic differences between any two individual haplotypes
- There is a block-like structure of linkage disequilibrium making it possible to discern a correlation between variants (i.e., if one is found, we can understand the likelihood that a nearby variant will also be present, and the result can be imputed into cases where there is no direct observation of the coincidence between variants).
- This information can be used to design and analyze genome association studies.
The insights from the data were new, but the questions above are not—population geneticists had long been working on them, using progressively better methods including phenotypes, enzyme isoforms, and genetic markers such as restriction fragment length polymorphisms. But the existence of more-or-less comprehensive sequence-level data from multiple populations is surely a major stimulus. The insights from these studies in turn were suggested to enable more refined or even new questions, for example, what is the detailed landscape of recombination across the human genome and how did it shape genome structure? And what regions/variants across the entire human genome bear the signatures of recent natural selection?
Variation data discovered in HapMap and its successor (1000 Genomes Project [1000G]; see below) were used to develop genotyping chips. Initially, these were able to assay tens of thousands of variants at a time; currently, they can assay nearly two million variants for about $50 per sample (Illumina). These enabled Genome-Wide Association Studies (GWAS), which have now identified strong statistical associations between over seventy thousand genomic variants and human phenotypes (Buniello et al. 2019).
2.3. Acceleration of Cost Decreases
Starting in about 2006–2007, sequencing costs started to drop even more rapidly, due to the introduction of so-called “second-generation” instruments—Solexa and SOLiD used in the NHGRI programs, and others used elsewhere (Heather and Chain 2016). Capillary instruments were phased out from NHGRI data production pipelines over time (but are elsewhere still used for some lower-throughput, high-quality applications, including assessing individual clinical samples). This change occurred during, or prior to, the lifetimes of the next several examples.
2.4. $107–104 per Genome: ~2002–2010. Comparative Genomics
Sequencing of the genomes of multiple eukaryotic species started with the pilot IHGP efforts on four species (baker’s yeast, C. elegans, and D. melanogaster) and has only accelerated through the present. Prior to 2002, NHGRI released separate funding announcements for mouse (National Human Genome Research Institute 1998b) and rat (National Human Genome Research Institute 2000). A Zebrafish sequence was begun at this time, under other auspices. Although useful for comparative genomics studies, the main rationale for these choices was to add critical value to these research models; the importance of the model justified the high cost in ~2002 (many tens of millions of dollars for a high-quality clone-based reference genome assembly). A Funding Opportunity Announcement (FOA) for a second fruit fly, D. pseudoobscura, was published in 2001 (National Human Genome Research Institute 2001)—instead, the justification was almost solely for the purposes of comparative genomics analysis. Fortunately, Drosophila genomes are roughly one-twentieth the size of most vertebrate genomes, so the cost was relatively low for the time—the funding announcement anticipated a $5 million cost.
Due to decreases in cost during this time period, NHGRI changed the way it organized genome sequencing efforts. At the time, in order to realize continued cost decreases, we still needed to fund large, stable centers to achieve efficiencies of scale. At the same time, many of the organismal sequencing projects were too small to fill the capacity of a large center for a reasonable time period corresponding to a single grant award (three to five years). In addition, it was clear that the applications for large-scale sequencing capacity would change rapidly, and that the ideas for the new applications would increasingly come from investigators who were not expert at large-scale sequencing, or even necessarily genomicists, but rather investigators with important questions but no data. This circumstance cannot be accommodated by the usual approach of “one project, one grant.” As a result, NHGRI issued an FOA for generic large-scale sequencing capacity (National Human Genome Research Institute 2003). A group of outside scientific investigators was recruited to advise NHGRI on new sequencing targets based on community proposals; several of those proposed multiple species selected to address a scientific question, including some of those listed below in this section.
Costs at the beginning of this period were still very high (multiple tens of millions of dollars for a state-of-the-art assembly of a mammalian genome). Initial choices of large genomes were still mostly for research models (macaque, chimp, dog, chicken) or for species with smaller genomes (fungi, sea urchin). Many of the mammalian genomes were initially sequenced at very low coverage (2X in capillary sequencing coverage) because even at this very low quality, some comparative analyses could be done. (These genomes were “topped up” when costs dropped.) At the end of this period (National Human Genome Research Institute 2010a), reasonably high-quality genomes—often large vertebrate genomes—were being produced with whole genome shotgun data, assembled on its own, or on a clone-based scaffold. Ultimately, over 250 projects were sequenced4—some were important biomedical models, some were human pathogens and disease vectors, some displayed uniquely informative biology.
But the most interesting in this context are the genomes produced for comparative analysis, because these raised the most interesting new questions in one or more areas:
- Direct comparison at the sequence level to annotate vertebrate genomes.
- Understanding rates of change in genomes; gene loss and gain in clades; selective forces acting on genomes; speciation
- Understanding the evolutionary origins of genes and other genomic features
- Understanding the genes underlying major evolutionary transitions
- Improving/resolving phylogenetic relationships
Each of these areas raised multiple new questions. Of course, some of the important questions were not new, but could only previously be explored with limited data on genomic regions over small numbers of species.
Some new questions generated by the ability to access comparative genomics data include:
- How far can we pursue the paradigm of identifying functional elements in genomes by comparing them and looking for conserved or rapidly changing sequences in order to annotate the human genome? A more refined question along these lines was, How many different species, at what amount of evolutionary separation, need to be sequenced and compared in order to identify all conserved regions that are over an arbitrary small size? This is the key question for understanding how to annotate all the conserved regions of the human genome. Using comparative genomics was, in the first decade of the century, the most powerful way to identify putative functional regions of the genome, and was one foundation for the ENCODE effort (National Human Genome Research Institute) discussed elsewhere in this volume (see chap. 4).
- Can we map the genomic differences observed between organisms to the anatomical and physiological differences that we see between them? For example, if we compare genomes of multicellular species with those of the living descendants of their unicellular relatives, can we identify the genes underlying multicellularity (Ruiz-Trillo et al. 2007)? Note that the transition from unicellular to multicellular life occurred multiple times in at least three independent lineages, plants, fungi, and animals—are these the same, or different? How easy or hard is it to become multicellular? Or, can we discern the genes for making a flagellum by comparing genomes from organisms that have them, to those that lack them (Li et al. 2003)?
- Can we detect signatures of rapid change of, and selection on, genes within a specific lineage? And how does that relate to phenotype differences?
A significant conceptual foundation that drove the design of comparative genomics studies was the idea that it would be possible to compare genomes and see significant genomic differences and similarities that correlated with phenotype. This concept was also at the foundation of later efforts, including sequencing to understand cancer, and sequencing to understand inherited diseases.5
2.5. $106–104 per Genome: ~2008–2010. 1000 Genomes (1000G)
The drop in costs after 2007 enabled another critical human variation effort, the 1000G (Abecasis et al. 2010).6 1000G began in 2008, by which time costs were in the range of $250–500K per genome, much lower than costs when HapMap was initiated. Even at this cost point, the project was too expensive to be practical if the approach was to produce high-quality (now 30X coverage) data from all one thousand genomes from individuals from multiple populations (which eventually was adjusted to two thousand). Therefore, the project began with testing multiple work-arounds: using lower whole genome coverage (7X using the then-new platforms), and sequencing targeting only the exomes (whole exome sequencing), less than 5 percent of the genome, but containing the regions coding for proteins and immediately adjacent sequence. This was approximately one-third the cost of full 30X coverage.
1000G extended and pursued some of the same questions that motivated and that were pointed out by HapMap, including addressing differences between populations, and characterized signatures of selection across the genome and differences in these regions among different populations7 (Abecasis et al. 2010). It also suggested questions about the rate of mutation and the number of gene-damaging variants in the average person (Auton et al. 2015).
2.6. $106–103 per Genome: 2008–2014. Sequencing Tumor Genomes
The Cancer Genome Atlas (TCGA), a joint effort between NHGRI and the National Cancer Institute,8 officially began at the end of 2005, but genome sequence data production did not start in earnest until nearly two years later/cheaper.9 Although consortium papers were still being published on the data until very recently, the data production for that project was completed in 2014. Even given the very compelling nature of the project, costs were too high to make it practical to produce whole genome sequence for every sample. Moreover, for each cancer, both tumor and normal tissue were sequenced for up to hundreds of individuals. For these reasons, TCGA began with sequencing targeted sets of genes, and eventually progressed to whole exomes (using second-generation sequencing platforms), at the time about one-fifth or less the cost of whole genome data.
One technical feature of the newer platforms was that so much data could be gathered from each tumor that it was possible to detect cancer mutations in mixed populations of cells—this is important because many tumor samples from biopsies are a mix of normal and mutated cells, with the number of sequence reads from the mutated genes being approximately representative of the number of copies of those genes10 in that sample. It became possible to discern genomic changes in tumors as they evolved over time, including detection of different cancer cell lineages within a tumor.
Early in this time frame (2008), the first use of whole genome sequencing to characterize the cancer of an individual patient was published (Ley et al. 2008).
A good account of the beginnings of, and rationale for, genomic characterization of cancers was published in 2009 (Mardis et al. 2009).
The impact of sequencing tumor genomes, in terms of the questions and insights about cancer biology that it led to, not to mention the clinical implications, is hard to overstate. It had of course long been known that changes in genomes—ranging from large chromosomal rearrangements to individual base-pair changes in specific genes—were associated with and in many cases were sufficient to cause the development of cancers. The ability to comprehensively identify genomic changes in a large number of tumors from many cancer types led to many insights, including the discovery of nearly three hundred “driver genes” for cancer, with the observation that some correlate with tissue of origin, and others with cell type of origin; also, some drivers underly only one cancer type, while others are present in multiple cancer types. This work also made it clear that sequencing individual tumors could help identify specific drugs that were likely to be useful for treatment of the patient (Bailey et al. 2018, and see other references summarizing the work [Atlas, n.d.]). Sequencing tumor genomes is now so inexpensive, and so informative for individual patients, that it is increasingly a routine component of cancer treatment in the clinic.
The ability to obtain tumor sequence data on this scale enabled significant new questions to be asked, and even where these questions were not completely new, allowed them to be addressed:
- How do cancer genome drivers assort with cancer types (tissue of origin)?
- Given that not all instances of a specific cancer type have the same drivers, what is the actual distribution of drivers? Is this a small set of genes for any cancer type, or are there a large number of different genetic changes that cause the same cancer type? How common or rare are each, in any cancer type? Do these all fall into the same molecular genetic pathways; is there another way to understand what they have in common?
- Can we determine the origins and causes of specific cancers based on the mutations we see? For example, does lung cancer caused by smoking have a different mutational signature than non-smoking-related lung cancer?
- What are the changes in populations of cancer cells within a tumor over time that accompany key features of cancer progression, including metastasis to distant sites? How do the genomes of different populations of tumor cells change in response to treatment and acquisition of drug resistance?
- What are the implications for the new findings for disease prognosis and treatment both of different cancer types, and individual tumors?
2.7. $104 –103 per Genome: 2011–2019. Common and Mendelian Disease Sequencing; Clinical Genomes Beyond Cancer
As early as 2005, our programs, and others, undertook what were then called “medical sequencing projects” by NHGRI. These were in several areas: sequencing genes and surrounding regions that were already known to be involved in disease, to look for new disease alleles in those genes (“allelic spectrum”); looking for disease-causing genomic variants in regions that had previously been genetically mapped, but for which no specific gene or variant had been identified; and a version of that, sequencing exomes in cohorts with an X-linked disease where no causal gene was known. These designs reflect the cost limitations at the time, focusing on small regions of the genome (known disease genes, mapped loci, and exomes for just one chromosome).
As costs approached $1000 per genome, it became possible to pursue genome sequencing studies on inherited diseases in a more comprehensive way. Mendelian diseases, by definition, are caused by variants of very strong effect.11 In 2010, the first paper was published demonstrating that whole exome sequencing comparing exomes of a small number of related and unrelated affected individuals, with those of unaffecteds, could reveal a variant and implicate a gene causing a Mendelian disorder (Ng et al. 2010). Following this, NHGRI issued an FOA for Mendelian Genomics centers (National Human Genome Research Institute 2010b), which would follow this model, and develop others, to identify the causes of Mendelian diseases more broadly, perhaps comprehensively (that is, understand the genetic basis for every Mendelian disease). This approach has been highly successful at the costs obtainable during this time, in large part because a relatively small number of samples is needed for each different disease (because the point is to look for variants of very strong effect, which are easier to find because they will almost always be present in affected individuals but almost never present in their unaffected relatives), and it was logical to look only in exomes (which were at this point one-third the cost of whole genomes) because variants that have a very strong effect are most likely to be in gene coding regions.
A number of questions came out of this work, including:
- How many Mendelian conditions are there? Is there one (or more) for every gene in the human genome?
- How often do different alleles in a single gene cause the same phenotype vs. a different one?
- How do variants in different genes lead to the same phenotype?
- Can an individual have more than one Mendelian condition, or a condition that is caused by more than one strong-effect variant? How often does this happen?
- How can we explain the rare instances where someone has one copy of a dominant allele, or two copies of a recessive one, but is not affected? Is this as rare as we initially thought?
- If we look carefully, can we see phenotypes in people with one copy of a recessive disease allele?
- How do some of these questions lead us to reconsider a simple Mendelian model for some of these conditions?
Again, most of these questions are not entirely new—genetic studies in model systems have revealed interactions between strong-effect mutations, and variation in penetrance and expressivity of mutations has long been known. Nonetheless, the ability to understand these phenomena in humans, tied to specifying gene variants and disease phenotypes, is new.
During this time period, use of sequencing in the clinic became possible and was pursued by NHGRI (National Human Genome Research Institute 2010a; Biesecker 2012) and by others. This was not just because costs were permissive, but also it became increasingly possible to interpret the genetic variants identified by sequencing individuals due to accumulating work from “discovery” projects that linked variants with disease phenotypes in large numbers of study subjects.
An approximately $1000 genome also allowed sequencing to be used for the investigation of common diseases with a complex genetic basis (e.g., coronary artery disease, type 2 diabetes, epilepsy, schizophrenia). GWAS were, before this time, identifying relatively common variants associated with disease phenotypes,12 and it soon became evident that the variants they were finding did not account for all of the incidence of disease in a population (Manolio et al. 2009). Moreover, GWAS assayed the more common variants (currently ~1 percent frequency in the population or greater, but of somewhat higher frequency initially)—a variant that was not represented in the assay could not be scored directly in a new sample. Genomic sequencing approaches, though up to 20X more expensive, had the potential to identify more genomic variants associated with diseases. Several common disease sequencing programs started at this time, at NIH notably the NHGRI Centers for Common Disease Genomics and the National Heart Lung and Blood Institutes TOPMed program (National Heart Institute 2015). Some of these were exome-based, and some were whole-genome based.13 These two ongoing NIH projects aim to produce, in combination, 300,000 whole genomes and 250,000 exomes, covering multiple common disease phenotypes. Many insights have arisen from the combination of GWAS and genome sequence data in common disease studies, including:
- It is possible to use genome sequencing to find variants associated with disease, including ones missed by GWAS. But it is much harder to conclude which associated variants are actually responsible for the effect.
- Most of the variants underlying a common disease in a population are relatively common and individually of weak effect in terms of disease risk. The great majority are in non-coding regions. Some, however, are rare and may be of stronger effect. Outside of the protein-coding regions, it is a challenge to link any particular variant in a region associated with a phenotype to a specific gene.
- The variants underlying an individual common disease phenotype are often numerous (dozens to hundreds of genes/alleles involved, different combinations in different individuals). The characteristics of variant number and type, population frequency, and effect size can be considered as elements of the genetic architecture of common disease. Different common diseases have different architectures. Assumptions about disease architecture influence how best to find the variants.
- Because common disease sequencing studies are often looking both for common variants and rare variants across a range of effect sizes, very large numbers of samples are needed to achieve statistical significance (Zuk et al. 2014). However, this is only really daunting if the goal is to find all of the variants underlying a complex disease. If the goal is instead to find some of the more common, or larger effect variants that underlie the disease in order to gain a toehold of understanding, then there are clever designs that require fewer samples (Timpson et al. 2018).
Some questions that arose from these insights include:
- What is the range of different common disease architectures? What forces shape them—for example, how do natural selection and recent population demographics shape genetic architecture?
- How can we better interpret or validate associations, especially in non-coding regions?
- How can we connect thousands of known common disease non-coding variants to the genes they act on?
- What study designs are best for which diseases?
- Can we find every genetic variant underlying every common disease?
- What can we learn from these studies about basic disease biology, including learning about sets of genes and biological processes that underlie any particular phenotype; can we use gene and variant commonalities between different diseases to find unsuspected connections between different diseases? Can we use the information to make inferences about environmental influences on common disease?
- How can we translate knowledge of common disease variants into clinically useful information, including estimations of risk, prognosis, disease course, and interventions?
As with other topics presented here, some of these questions are not completely new, but most are newly formed or refined versions of older questions.
Another consequence of very cheap genomes during this period, along with the explosion of the science of genomics, is that there was a sharp increase in the number of different groups that were sequencing for many different purposes, especially in clinical labs for cancer, newborn diseases and birth defects, diagnosis of Mendelian diseases, and biobanks, but also in direct-to-consumer services. These are only the biomedical applications, and do not include use in agriculture, infectious diseases, microbiome studies (human and environmental), hominin evolution, and many others. By the end of this period, the NHGRI outlay for genome sequencing dropped significantly, both in absolute dollars and even more as a proportion of output worldwide (see chap. 3).
2.8. A Future Below $103 per Genome
It is already possible to predict at least some of what still lower costs will enable—even if it is hard to predict specific questions that will arise as a result. For example, it is easy to predict that clinical sequencing will become more and more routine, and not just for cancer and Mendelian diseases, but for a host of clinically useful information, including response to drugs, and increasingly, polygenic risk for common disease.
The groundwork has already been laid for this at the current cost through programs such as Electronic Medical Records and Genomes network14 and the All of Us program.15 For common disease studies, the setting will likely increasingly move away from individual disease sequencing studies. Instead, very cheap sequencing will allow health management systems and biobanks to gather sequence data as an adjunct to the rest of the data and materials they have (electronic health records, biosamples) for the millions of patient samples that they collect in the course of routine care. These data could be aggregated and provided to researchers as a substrate for large enough analyses to power studies for many common diseases, with the advantage that in some cases one could follow subjects over time. The cost of sequence data will likely not be a limiting factor, but rather the hindrance will increasingly be the ability to aggregate or federate the data, which is limited by technical data issues and the ability to ensure that the use of the data is consistent with the consent of the participants.
It is also evident that short-read data will not be the only useful sequencing product. Years of experience with the shortcomings of cheap data stimulated the development of technologies that produced much higher quality data, including so-called “long read” technologies (e.g., Pacific Biosciences, Oxford Nanopore) and synthetic long-read and optical mapping methods that could produce long scaffolds on which the rest of the genome data could be assembled. This high-quality data is expensive—today, a high-quality genome assembled from this type of data would cost on the order of $20,000,16 but some of these modalities are less expensive than others and can yield intermediate results. These data types make it possible to assemble substantially more contiguous genomes; they also can recover phasing information—that is, information about whether multiple nearby variants exist on the same chromosome homolog/haplotype. This is important, for example, in studying interactions between gene regulatory variants as they act on a nearby gene. In addition, long reads have a potential to consistently and reliably resolve structural variation, which is relatively invisible to sort read methods—in fact, as we learn more about structural variants, they are turning out to have effects that are greater than would be expected simply based on their number.
3. Lessons and Ideas
For genome sequencing, the temporal relationships between technological advance, reductions in the cost of data generation, and the kinds of scientific questions that could be asked—though not necessarily answered yet—are unusually clear. As such, they are a good foundation for more thorough consideration of a number of ideas about how we should think about scientific discovery, or about how the organization of bioscience is changing around the kinds of questions that are its subject. Many of the questions that arose in genomics were not the ones that most prominently motivated the HGP in the first place, and often were not fully anticipated. In particular, new questions were stimulated by cost decreases that allowed sufficient data to be produced to reveal previously unknown features of genomes—for example, amount, type, and genome-wide distribution of variation, especially structural variation and revelation of the detailed haplotype structure of the human genome. In addition, new capabilities, such as the ability to do comparative analysis on whole vertebrate genomes, stimulated new biological questions and the analysis methods needed to answer them. Some of these questions were not completely original with the availability of the advance, or had been constructed in some form, especially the population genetics questions that had previously been addressed with genetic markers, and aspects of the comparative genomics questions that had been pursued with single-locus studies. But the data to fully realize the question were not previously available or were not sufficient to ask the question comprehensively across the whole genome and across populations.
It is probably an oversimplification to say that any particular advance is essentially dependent on cost; other factors are at play. Data quality is also important. For example, assessment of large-scale structural variation was technically possible using capillary and clone-based sequencing methods at least as early as 2004, but the cost was prohibitive at that time. It actually became less technically possible with the advent of short-read platforms, because although the data was cheaper, it was not fully possible to use short reads to detect the larger structural variations. With the availability of higher-quality long-read data platforms, structural variation is becoming both more tractable and practical.
3.1. Discovery Science vs. Hypothesis-Driven Science
At the start of the HGP, it was common to hear the molecular biology and genetics communities cast genomics as “discovery science” in contrast to “hypothesis-driven science.” This was sometimes not meant to be a compliment to genomics,17 although it should be said that this idea has older origins, reflected in the quote attributed to Rutherford that all science is either physics or stamp collecting.18 And older origins still—though more neutral—in the distinction between inductive and deductive elements in science. The rapid cycles described here of technology change, ability to produce more data, and resulting new questions considerably blur these distinctions—the way genomics has been done in effect erases the temporal condition that makes it possible to make neat distinction between discovery and hypothesis-driven science. The notion that genomics blurs this distinction was explicitly advanced by Kell and Oliver (2004): “We argue here that data- and technology-driven programmes are not alternatives to hypothesis-led studies in scientific knowledge discovery but are complementary and iterative partners with them.” We extend their discussion to the exemplar of the NHGRI Genome Sequencing Program (GSP) and our understanding of the GSP from our personal vantage point. The description of events here confirms those views and adds context and granularity to them, based on our personal direct experience over essentially the whole of the genomic era to the present time.
The previous paragraph starts to raise a couple of ideas whose fuller explorations go beyond the scope of this article. First, above, the implicit relating of the word hypothesis to the word question, used throughout most of this article, needs explanation. At least, in a practical sense, questions lead to hypotheses, for scientists. It is unclear that hypotheses can exist in a meaningful sense without corresponding questions. Second, and related, is the status of inductive versus deductive or experimental approaches, which is worth examining further in the context of genomic science, because it is so heavily data- and association-driven, and because it marks a recent rapid change in the context of biological sciences from the hypothesis-driven approach, possibly illuminating ideas about what it means to understand something in genomics (Mazzocchi 2015).
Although we believe that the last two decades of recent progress in genome sequencing have not been exceptional in terms of the fundamental structure of discovery, it is clear that practical considerations have significantly changed how researchers are organized to do the science.19 Large groups of people with highly specialized expertise were required to develop the means, and then produce, the data, requiring a concentration of resources over a period of years. Due to size and complexity, additional layers of management were acquired. It took time for the data to be delivered to the community, and for the community to learn to use the data. As sequencing gets more affordable and consistent (e.g., as a result of fee-for service data generation, or as a result of miniaturization of the platform), these large dedicated centers will no longer be needed to produce data for many purposes; datasets that were once considered large will become routine; and the superficially separate phases of data generation and hypothesis testing will re-integrate.20 At the same time, large, useful high-quality datasets will increasingly be seen as basic resources, rather than as a separate kind of scientific activity (“discovery”).
This does not mean that large dedicated centers for data production will go away. As is already happening, they will move on to problems that can only be answered with data from millions of genomes. By all current appearances, this will require another reorganization around the necessities of making available very large numbers of well-annotated and phenotyped samples; and organizing, processing, and serving the aggregated or federated data.
Finally, it is worth noting that the increasingly widespread use of genome sequencing and its increasingly routine incorporation into biomedical research and clinical application—including the current indications that the lines between the two will be erased—means that NHGRI is less and less likely to directly fund efforts that are essentially large genome sequencing centers. This means that NHGRI will likely lose its ability to do cost accounting for genome sequencing.
Notes
The authors are employees of the National Human Genome Research Institute.
1. A Google Scholar search for “Whole Genome Sequencing” used in the title or abstract of a publication indicates 321 new publications in 2000; 1040 in 2005; 3,400 in 2010; and 25,400 in 2018.
2. We note that the same technical advances that made genome sequencing per se very cheap—instruments that produce billions of sequencing reads per machine run—also made possible the “counting applications” that are at the heart of, for example, quantitative gene expression assays, chromatin accessibility assays, chromatin structure assays, etc.
3. Several efforts were made in 1998–1999 to measure and consistently report cost and quality; quality metrics were established and assessed prior to the publication of Lander et al. (2001). The current set of cost metrics used by NHGRI was initially proposed at a meeting of the NHGRI-funded genome sequencing grantees in 1998, but the current consistent cost reporting was not established until 2001.
4. NHGRI maintains a list of all comparative genomics projects and all supporting community proposals; see https://www.genome.gov/25521739/comparative-genome-evolution#al-2.
5. Regarding comparative approaches, it is generally easier to see differences between two closely related genomes (e.g., between two cells from the same individual, or between parents and a child). It is much harder to interpret differences between two unrelated people because there are roughly three million differences between any two people, and almost all of them, in the best case, will have no bearing on the differences in phenotype that one is interested in.
6. 1000 Genomes Project, https://www.internationalgenome.org/home.
7. A version of the 1000 G project was proposed and discussed at NHGRI for more than a year (in other words, soon after the HapMap publication) before the decision was made to initiate it. Initial objections were made about the fact that the samples would not have phenotype data associated with them. In addition, the value of the information, considering the cost, was initially not considered to be high enough by some program advisors. The information from the 1000 G samples have proven to be extremely valuable over time: in 2019, the 1000 G samples were sequenced again to update the data to the latest platform standard. The cost was about $3 million for two thousand genomes at full coverage (about 30X).
8. See https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga.
9. It is important to state that genome sequence was only one data type produced in TCGA or other cancer genome studies. Other data types included array data to detect structural variation, RNA expression data, and others. Critically, the most efficient genome sequence modality of the time was largely unable to reveal several types of structural variants that are known to be a key feature of cancer genomes.
10. The number of copies could be representative of the proportion of cancer cells within the tumor, and/or may be representative of genomic duplications of the gene locus within the cancer cell lineage.
11. Mendelian traits, including diseases such as cystic fibrosis or Duchenne muscular dystrophy, follow the genetics that most are taught in secondary school: for a recessive trait, those who inherit (or acquire) two “bad” copies of the gene will almost always have the phenotype. Some variants/mutations can be dominant, that is, only one bad copy is needed.
12. Identifying variants associated with diseases does not usually identify the actual gene responsible. Many GWAS “hits” are actually outside of genes and non-coding variants are not always readily connectable to their genes due to a lack of basic knowledge. The ENCODE dataset can help with this, but the state of the art is still far from where it needs to be.
13. The basic concept, as with comparative genomics, tumor sequencing, and Mendelians, was to sequence samples from individuals with one phenotype (a common disease in this case) and compare with the sequence from another (unaffected). A difference here is that, for common disease, comparisons had to be done across unrelated individuals; because there are so many background sequence differences between any two individuals, the comparisons are noisy, requiring large numbers of samples.
15. All of Us program, https://allofus.nih.gov/.
16. Although this cost seems to be high, please note that to get this quality of assembly cost on the order of $1 million only five years ago and is on its own cost curve.
17. This is the direct experience of one of the authors on several occasions, including an assertion that high-throughput discovery efforts would make postdocs and graduate students “stupid” because all they would do was collect data but never analyze it.
18. However see https://www.patreon.com/posts/did-rutherford-23413944 for a different interpretation—that Rutherford actually meant to insult theoretical approaches, rather than data gathering.
19. This reorganization was not entirely compatible with the standard academic model for research—principal investigators and their trainees working on well-defined and self-contained projects—which is probably partly why it created friction.
20. It is interesting to see that some mid-scale bioscience projects are essentially replicating the “genomic” scheme of creating cycles of “observe-hypothesize—perturb/test-observe etc.” as an integrated method. See https://www.broadinstitute.org/center-cell-circuits.
References
- Abecasis, G. R., D. Altshuler, A. Auton, et al. 2010. “A Map of Human Genome Variation from Population-Scale Sequencing.” Nature 467 (7319): 1061–73.
- Atlas, T. C. G. n.d. “TCGA Research Network Publications.” https://www.cancer.gov/about-nci/organization/ccg/research/structural-genomics/tcga/publications.
- Auton, A., L. D. Brooks, R. M. Durbin, et al. 2015. “A Global Reference for Human Genetic Variation.” Nature 526 (7571): 68–74.
- Bailey, M. H., C. Tokheim, E. Porta-Pardo, et al. 2018. “Comprehensive Characterization of Cancer Driver Genes and Mutations.” Cell 173 (2): 371–85.e318.
- Biesecker, L. G. 2012. “Opportunities and Challenges for the Integration of Massively Parallel Genomic Sequencing into Clinical Practice: Lessons from the ClinSeq Project.” Genetics in Medicine 14 (4): 393–98.
- Buniello, A., J. A. L. MacArthur, M. Cerezo, et al. 2019. “The NHGRI-EBI GWAS Catalog of Published Genome-Wide Association Studies, Targeted Arrays and Summary Statistics 2019.” Nucleic Acids Research 47 (D1): D1005–12.
- Heather, J. M., and B. Chain. 2016. “The Sequence of Sequencers: The History of Sequencing DNA.” Genomics 107 (1): 1–8.
- Illumina. n.d. “Understanding Our Genetic Variation.” Accessed December 23, 2025. https://emea.illumina.com/techniques/microarrays/human-genotyping.html.
- International HapMap, C. 2005. “A Haplotype Map of the Human Genome.” Nature 437 (7063): 1299–320.
- Kell, D. B., and S. G. Oliver. 2004. “Here Is the Evidence, Now What Is the Hypothesis? The Complementary Roles of Inductive and Hypothesis-Driven Science in the Post-Genomic Era.” BioEssays 26 (1): 99–105.
- Lander, E. S., L. M. Linton, B. Birren, et al. 2001. “Initial Sequencing and Analysis of the Human Genome.” Nature 409 (6822): 860–921.
- Ley, T. J., E. R. Mardis, L. Ding, et al. 2008. “DNA Sequencing of a Cytogenetically Normal Acute Myeloid Leukaemia Genome.” Nature 456 (7218): 66–72.
- Li, J. B., S. Lin, H. Jia, et al. 2003. “Analysis of Chlamydomonas Reinhardtii Genome Structure Using Large-Scale Sequencing of Regions on Linkage Groups I and III.” Journal of Eukaryotic Microbiology 50 (3): 145–55.
- Manolio, T. A., F. S. Collins, N. J. Cox, et al. 2009. “Finding the Missing Heritability of Complex Diseases.” Nature 461 (7265): 747–53.
- Mardis, E. R., L. Ding, D. J. Dooling, et al. 2009. “Recurring Mutations Found by Sequencing an Acute Myeloid Leukemia Genome.” New England Journal of Medicine 361 (11): 1058–66.
- Mazzocchi, F. 2015. “Could Big Data Be the End of Theory in Science? A Few Remarks on the Epistemology of Data-Driven Science.” EMBO Reports 16 (10): 1250–55.
- National Heart Institute. 2015. “NHLBI Trans-Omics for Precision Medicine (TOPMed).” https://www.nhlbiwgs.org/.
- National Human Genome Research Institute. 2024. “ENCODE Funded Programs and Projects.” https://www.genome.gov/Funded-Programs-Projects/ENCODE-Project-ENCyclopedia-Of-DNA-Elements.
- National Human Genome Research Institute. 2024. “NHGRI History and Timeline of Events.” https://www.genome.gov/about-nhgri/Brief-History-Timeline.
- National Human Genome Research Institute. 1995. “Pilot Projects for Sequencing of the Human Genome.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-95-005.html.
- National Human Genome Research Institute. 1998a. “Human Genome Sequencing Projects Receive Third Year of Funding.” https://www.genome.gov/10000602/1998-release-human-genome-sequencing-projects-funding.
- National Human Genome Research Institute. 1998b. “Network for Large-Scale Sequencing of the Mouse Genome.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-99-001.html.
- National Human Genome Research Institute. 2000. “Network for Large-Scale Sequencing of the Rat Genome.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-00-002.html.
- National Human Genome Research Institute. 2001. “Comparative Genomic Sequencing of Drosophila DNA.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-01-003.html.
- National Human Genome Research Institute. 2002. “Large-Scale Genotyping for the Haplotype Map of the Human Genome.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-02-005.html.
- National Human Genome Research Institute. 2003. “Large-Scale Sequencing Capacity.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-03-002.html.
- National Human Genome Research Institute. 2010a. “Clinical Sequencing Exploratory Research (RFA-HG-10–017).” https://grants.nih.gov/grants/guide/rfa-files/rfa-hg-10-017.html.
- National Human Genome Research Institute. 2010b. “Mendelian Disorders Genome Centers (RFA-HG-10–016).” https://grants.nih.gov/grants/guide/rfa-files/rfa-hg-10-016.html.
- Ng, S. B., K. J. Buckingham, C. Lee, et al. 2010. “Exome Sequencing Identifies the Cause of a Mendelian Disorder.” Nature Genetics 42 (1): 30–35.
- Ruiz-Trillo, I., G. Burger, P. W. Holland, et al. 2007. “The Origins of Multicellularity: A Multi-Taxon Genome Initiative.” Trends in Genetics 23 (3): 113–18.
- Timpson, N. J., C. M. T. Greenwood, N. Soranzo, D. J. Lawson, and J. B. Richards. 2018. “Genetic Architecture: The Shape of the Genetic Contribution to Human Traits and Disease.” Nature Reviews Genetics 19 (2): 110–24.
- Venter, J. C., M. D. Adams, E. W. Myers, et al. 2001. “The Sequence of the Human Genome.” Science 291 (5507): 1304–51.
- Wetterstrand, K. 2023. “DNA Sequencing Costs: Data from the NHGRI Genome Sequencing Program (GSP).” https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data.
- Zuk, O., S. F. Schaffner, K. Samocha, et al. 2014. “Searching for Missing Heritability: Designing Rare Variant Association Studies.” Proceedings of the National Academy of Sciences USA 111 (4): E455–64.