Notes
11 “Variations on a Theme”
A History of Errors and Polymorphisms in the Human Genome Project and Beyond
Christopher R. Donohue
Although a great deal has been written concerning the Human Genome Project (HGP) and genomics, there remain considerable gaps in the understanding of historical HGP efforts and of the development of genomic science after the “completion” of the HGP in 2003. There have been various popular accounts of the HGP (see Shreeve 2007; Ferry and Sulston 2010) as well as numerous and trenchant discussions of the ethics and the epistemic consequences of the HGP and genomics (such as Sarkar 1998; Roberts 2011). As importantly, neither the history of the HGP nor of genomics has addressed variation as a technical and conceptual problem, especially in the context of sequencing and of assembly.
Such an omission is somewhat surprising given the importance of polymorphism to the history of biology and genetics. One only need look at the energy that E. B. Ford gave the topic many years ago (Ford 1945). Douglas Futuyma observed that “The theory of population genetics has expanded beyond its preoccupation with polymorphism” with population geneticists studying the “evolution of gene families,” “neutral molecular evolution,” “sexual selection,” and other topics (Futuyma 1988, 217). As is well known and has been widely described, among the foundations of the evolutionary synthesis was the “the demonstration of the existence of ample genetic variation in nature” (Pigliucci 2007, 2744). Critiques and extensions of that evolutionary synthesis have strongly and, sometimes fruitfully, suggested that though the genotypic variation in nature be vast, such variation is still insufficient to explain all phenotypic variation, thus leading to such questions as ones surrounding so-called “missing heritability” (Manolio et al. 2009). Nevertheless, even with this complex and rich history, variation as such has never been approached as a technical, historical, and conceptual problem in the context of the HGP and genomics.
In this chapter, I will narrate how the distinction between a polymorphism and an error emerged as one of the key concerns in the sequencing of the genome. I will detail how rather than being a formal standard, with deep vetting and consideration, the sequencing error metric emerged rather flexibly from both the technological and organizational knowledge of the time. It was adopted, though not universally assented to, because such a standard was cost-effective and achievable. As importantly, this metric reflected a specific knowledge of the nature and extent of variation in the genome, which, because of issues such as read length, focused on SNPs (single nucleotide polymorphisms) rather than more difficult to ascertain features in the genome, such as structural variation. I conclude with how the decisions around such standards, though pragmatic and flexible, still have real contemporary legacies.
Lisa Gannett has underscored that the HGP could be rightly criticized because it treated either some or all human genomic variation as a departure from the “norm” (Gannett 2003). A version of Gannett’s worry was expressed by others within the genomics grantee community. Evan Eichler worried that investigators would normalize “the human genome to an a priori expectation of allelic variation” where bioinformatics tools would disregard all SNP reads above a certain threshold or have difficulty distinguishing these reads from errors (Eichler et al. 2004). Thus, the question for the genomics grantee community during the HGP was not merely quantifying the amount of variation in the first read of a human genome assembly, but also how to distinguish it from those errors, which came about through the process of sequence and assembly.
This required developing an error standard for finished sequence deposited into GenBank (the central repository for sequence during the duration of the HGP). The development of this standard was neither sudden nor easy, and it depended upon historical, contingent, and scientific understandings of what constituted an error and what constituted variation or polymorphism (which were used interchangeably by the genomics grantee community and staff) for the duration of the HGP. Thus, the error problem especially was acknowledged very early on in the history of the HGP, and discussions of the distinctions between errors and variations/polymorphisms continue to this day regarding the reference sequence and the difficulties attending variant calling, assembly errors, and the continual difficulties surrounding structural variation.
Discussions of polymorphisms were also contingent, because such was the amount of sequence generated that significant amounts of polymorphism in sequence data were not considered a serious issue in practice (although it was always realized in theory) until about 1996. This was because prior to that year there was a relatively small amount of finished sequence deposited in GenBank. By 1998, about 116MB was finished, with 70MB had been produced in the preceding year (NHGRI Internal History Archive. Scanned/Box List 15/5 Bermuda-II).
As Guyer, Peterson, Wetterstrand, and Jordan in this volume underscore, the HGP, being a novel endeavor, had to develop standards using the expertise of its program officers and from the wider grantee community. Like many standards, these began with so-called back of the envelope calculations based upon a few studies, which then developed into a shared consensus within the genomics grantee community. This shared consensus over the difference between an error and a polymorphism, and the sequence quality and contiguity, were not agreed upon by the entire community, and as the archival record illustrates, generated considerable discussion.
Eventually, two members of the grantee community, Maynard Olson and Phil Green, argued in an article entitled “A ‘Quality-First’ Credo for the Human Genome Project,” which outlined how this metric was insufficient. They underscored that “even at a 10−4 error rate, on the order of 10% of the discrepancies with the reference sequence that will arise during intraspecies comparisons will be due to errors in the reference data” (Olson and Green 1998). They pointed to other features of successful genome assembly, such as contiguity and assembly checking, which would be as critical (if not more so) for the kinds of functional and comparative studies that would need to be done in the future in order to capitalize on the initial work and insights provided by the sequence.
I will argue in the last part of the chapter that Olsen and Green have predicted many of the issues now apparent with the reference sequence, which has remained, until recently, riddled with assembly errors and where, until recently, there were immense difficulties with rare and structural variants (SVs). This can be in part traced back to the initial assumptions about error rate and polymorphism in the middle to late 1990s. As importantly, the present difficulties with genome reference can also be traced back to perfectly reasonable, pragmatic assumptions about the nature of variation in the genome and the degree of sequence accuracy that could be achieved for a specific cost.
That these data quality standards took six years or more (if one counts early efforts by the Department of Energy [DOE] and by the human genetics community) to define illustrates a fundamental reality about the HGP. That is, that the quantity of sequence produced prior to that point was likely insufficient for a standard to exist at all. This reality also underscores that the last three years of the project were years of successful and at scale sequencing, which could not necessarily have been predicted by previous years. The first years of the HPG were defined by mapping technology development and by fundamental foundational efforts to determine grantsmanship and funding mechanisms.
In the “Mapping and Sequencing the Human Genome” National Research Council Report (1988), the authors noted that “The rate of DNA sequence determination is also limited by the fact that all techniques currently use polyacrylamide gels that resolve no more than 250 to 500 nucleotides at a time.” The report went on to underscore that at that time (in 1988) “The accuracy of DNA sequencing has not yet been firmly established.” The report continued that “A careful and experienced laboratory probably achieves an accuracy of about one error in every 5,000 nucleotides (0.02 percent error rate) in the finished DNA sequence, but this degree of precision requires careful attention to virtually every nucleotide in the sequence.” But such attention to the error rate “inevitably slows the sequencing rate.” They continued, “it will be difficult to hold the error rate to this level in a large-scale nucleotide sequencing project” (National Research Council Report 1988). Thus, although mapping was slow and cumbersome, there was also a real worry among the grantee community that not only would sequencing not scale, it would not scale accurately, even with dense and accurate maps.
Although the HGP began at the National Institutes of Health in 1990 (with efforts at the DOE having begun some years before), only in the mid-1990s was it then possible to ascertain sequence and sequence at scale. Key efforts in this regard were the sequencing technology development Requests for Applications (RFAs) and the sequencing pilot program managed by Jane Peterson. As Jeff Schloss underscores in his contribution to this volume, among the significant efforts in the mid-1990s were those that encouraged applications on “Improved Electrophoretic DNA Sequencing Technology.” The “pilot projects,” began in 1995 and designed to cover three fiscal years, sought to capitalize and extend upon that technology and others like it. The pilot programs gave centers an opportunity to balance between developing sequencing technology and implementing that sequencing technology. The RFA for the pilot projects underscored that “In the past year, there has been significant progress in developing the capability for large-scale DNA sequencing. Several laboratories have generated at least one, and as much as ten, megabases of DNA sequence each.” The RFA went on to describe the key question that remained was “how to manage large laboratory production efforts devoted to rapid data production.” Nonetheless, “sequencing technology and sequencing capability must still be improved considerably to achieve the DNA sequencing capability necessary to determine the sequence of the three billion base pairs of human DNA within the time and cost goals for the HGP” (see NHGRI Internal History Archive. Scanned Box 1272, folder 001).
To ascertain how accurate this newly produced sequence was, the genomics community used software known as Phred/Phrap. Phil Green’s laboratory developed Phred (as well as other bioinformatics tools), which allowed for better assembly and annotation of finished genome sequence in 1996 (though the paper they authored on the package was published in 1998) (see Ewing et al. 1998). This tool would assign each sequence base an accuracy score and allow for a consistent across-center metric. Oftentimes, sequencing groups would choose to manually re-sequence raw data that had a low “Phred score,” while depositing the rest into GenBank. As importantly, sequencing groups would also often manually check raw data to ensure that the Phred score matched the real ascertained sequence data (this is because in a technical sense Phred scores are error probabilities that needed to be matched against error rates).
Another key moment where the consensus on the error standard came together was the months surrounding the Second International Strategy meeting on Human Genome Sequencing, which was also known as the meeting responsible for developing the Bermuda Principles or Accords. There has been considerable amount of discussion by Ferry and Sulston (2010) and Maxson-Jones and colleagues (Maxson Jones et al. 2018) (see also chap. 2, this volume) of the data release principles emerging from the Bermuda meetings, but far less discussion of the data quality standards and their significance for the HGP.
It was here that there first appeared an explicit and public acknowledgment of sequence quality and the defined 1/10,000 error rate. The Bermuda meeting report underscored that, “the nucleotide error rate should be 1 error in 10,000 bases or less for most sequence.” It also underscored, with implications for future issues with the reference sequence, that, “The agreed long-term goal is no gaps, recognizing that this is not yet routine.” It was also agreed that sequencing centers needed to describe what methods they used to ascertain errors in their sequence and that their sequence data needed to be annotated in a way to refer to this. The above standards and the above considerations needed a significant amount of sequence to be produced (and this was only just coming into being), as it required for there to be some amount of consensus sequence to be checked against the raw sequence data.
The Bermuda meeting was also among one of the first explicit discussions of polymorphism in the context of sequence error discussions, as well as in the context of the HGP. This was the very brief mention that polymorphisms could be distinguished by virtue of deep coverage where there was a high number of unique reads for each portion of reassembled sequence.
Discussions of polymorphism and errors continued to gain in frequency through the late 1990s after the Bermuda meeting, where there were conversations around relationship between errors and polymorphisms in the context of addressing errors in sequencing in scale and a related effort in polymorphism discovery. Indeed, discussions about errors and polymorphisms need to be considered as part of an explicit effort to understand the nature of polymorphism in the human genome (which in the mid- to late-1990s was thought of as SNPs) because of the available technology and of the cost of using that technology. At the same time, error rates used assumed an SNP density and distribution in the genome, which became more refined as the HGP sped to completion in the late 1990s (NHGRI Internal History Archive. Scanned/Francis Collins Papers, Box 7118, folder 012).
In 1996, Dan Drell, a scientist at the DOE, observed that since the “normal rate of human polymorphism is about 1/1000,” that would mean that any accuracy standard to direct errors would have to be 1/10,000. Drell detailed that circa 1996 standard for accuracy, which was about ten times better than the expected polymorphism rate, would be enough. However, Drell and others’ assumption concerning polymorphism and this accuracy measure underscored that polymorphism and variation would only be usefully found within coding regions of the genome, which subsequent research has shown to be false and that such non-coding variation is actually important for health and disease (see Ward and Kellis 2012). Drell also assumed that SNP variation would be constant across the genome (see NHGRI Internal History Archive. Scanned Box 024, file 008).
Drell’s analysis also hinted that the HGP had little direct experience of structural variation and its role in the genome. He noted, too, that the presence of structural variation 1 kb or larger in size would also cause significant gaps in the sequence because it was unresolvable with current technology, underscoring that this was a “variation on a theme.” Accordingly, the data quality standard policy developed in 1997 stated that sequence error standards mandated that sequence could contain an error in no more than 1 out of 10,000 bases and that the sequence be contiguous with no gaps (with a minimum contig length of 30 kb) (see NHGRI Internal History Archive. Scanned Box 1272, folder 005).
However, there was fierce disagreement with this standard, which some grantees, namely Maynard Olson and Phil Green, thought minimal at best. Phil Green, in an email to Jeff Schloss, underscored that the 1/10,000 standard was problematic, and would more likely need other stringent controls, since according to Green since “sequencing reads generally include substantial stretches of low quality data towards the ends of the reads, which have a high error rate but are just accurate enough to align against other reads” (email from Phil Green to Jeffery Schloss, 1/14/1997). It turns out, Green underscored, that since there tended to be low-quality data at the end of reads, these reads could have error rates of from 10 percent to 20 percent. This would not necessary be picked up by a Phred or Phrap quality score.
Maynard Olson, who was copied on this correspondence between Schloss and Green, underscored that the 1/10,000 standard was a “minimal one.” Olson had concerns that because there was no “gold standard sequence” with which to compare it to and only the sequence that centers had themselves produced, this would lead to serious issues, as without this “gold standard sequence” there would be no way to properly distinguish between known sequence and “real world data sets of poorer quality.” Using the current standard, Olson underscored that even with the 1/10,000 standard, which could (in theory) distinguish between a polymorphism and an error, it would nonetheless be difficult to achieve in practice, since “much sequence that is currently going into the data base fails the test and people are doing the easy things first” (NHGRI email archive Maynard Olson to Jeffery Schloss, 1/15/1997).
This was precisely the position that Olson and Green took in their paper, “A ‘Quality-First’ Credo for the Human Genome Project,” published in 1998. Their concerns, I would argue, were prescient especially for future discussions of issues with the reference sequence. In their article, Olson and Green noted that at the current error rate the emerging “reference sequence” would not be “complete” or without “error.” An ideal reference sequence was one that would “reflect real biological effects or limitations on future experimental or theoretical methods—not errors in the reference data.” Because of the issues with end-of-sequence reads and the emphasis on base-pair accuracy, rather than other metrics, there remained the possibility that “on the order of 10% of the discrepancies with the reference sequence that will arise during intraspecies comparisons will be due to errors in the reference data.”
What sequencing centers failed to realize, for both authors, was that the very nature of the HGP itself, producing tiny fragments of sequence that would then be aligned computationally was prone to introduce errors, which made extracting biological insight or comparative analysis from the sequence immensely difficult. Because of the size of the sequence fragments (100 kb) there were bound to be many unresolved gaps. As importantly, at 100-kb length with numerous gaps, it would be difficult to ascertain genes or coding regions. Thus, the authors noted, “Gaps not only leave users of the data uncertain as to the precise genomic origins and biological content of particular sequences, but they also make the sequences difficult to validate” (Olson and Green 1998, 414). And it is precisely these gaps and the relative error laden-ness of the sequence that has caused difficulties (until very recently) with ascertaining the significance of novel variation and its connection to disease.
Later in 1997, program officers began a broad effort at the National Human Genome Research Institute (NHGRI) to assess the relative quality of sequence being produced by the centers. As part of this effort, there were very specific instructions of how to specifically address errors and determining polymorphism. Centers were instructed to cease work and have a technician manually check the raw data if there were more than five errors per ten thousand bases. It would then be noted in the sequence submitted to the center for the “sequence quality assessment exercise” whether they were indeed base-calling errors, which could then be eliminated by resequencing. Many times, however, single base discrepancies could not be resolved and were often described at this stage as errors in “otherwise good data.” On a more system-wide level, centers were implored to try to resolve gaps and ambiguities, but the NHGRI understood that many of these gaps and ambiguities could not be resolved (and indeed many remained even after the “finishing” of the HGP in 2003). See NHGRI Internal History Archive. Scanned Box 1272, file 005.
As initiatives to ascertain the quality of sequence produced at (for 1997) relative scale were underway, there was an effort to specifically address the issue of polymorphism. This was a result of centers having by 1997 some actual contact with sequence, which validated some initial assumptions of earlier studies. The “Workshop on Human DNA Sequence Variation,” which convened at the end of March and early April 1997, underscored that DNA sequencing could provide direct insights into the nature and significance of human variation in the genome. The initial results of genomic sequence analysis indicate “that, on average, there is one variant nucleotide (nt) per 1000 nt (one kilobase, or kb) screened.” This confirmed the studies done in the 1980s using restriction fragment-length polymorphisms. While the number of studies the workshop cited were small (and the initial sequencing validation based on limited sequence available), and the data was not exhaustive, the NHGRI was prepared to move forward with more aggressive ascertainment of polymorphism, underscoring that as more DNA was sequenced and technology improved, the distinctions between errors and polymorphisms would become more refined. Even at this early stage there was clear evidence in sequencing data that SNPs possessed a block-like (haplotype) structure where contiguous segments of DNA had been inherited without recombination. The report went on to speculate that “As the functional information is contained within these blocks, association studies can be used to correlate haplotypes with specific phenotypes.” Here then were the roots of Genome Wide Association Studies and the International HapMap Project, as detailed by Boyer, McEwan, and Brooks in this volume (see NHGRI Internal History Archive. Scanned/Box 2025, folder 35 1998 Planning-003).
Errors and polymorphisms and the distinctions between them continued to generate animated discussions in the grantee community as more sequence was generated in the late 1990s. With the founding of Celera in May 1998 and the subsequent acceleration of the public side of the HGP, the grantee community adhered to the base pair quality standard while learning more about how to identify and to describe polymorphism across the genome. By the late 1990s, program officers at the NHGRI and the grantee community had developed sufficient technology and expertise (using array technology developed by Affymetrix; see interview with Steve Fodor 2019) to not only identify polymorphism (SNPs), but also to begin to empirically verify polymorphism and its distribution in the genome.
The efforts in 1997 and 1998 to ascertain variation led to the development of the “SNP Consortium.” This was an ambitious public-private partnership that sought to develop methods of resolving SNP in diverse individuals (twenty-four individuals were part of the initial study). While this effort was apart from the initial sequencing of the genome that was occurring at the same time, many of the same research groups worked on the same projects (including the Sanger Centre). Among the questions that these early discussions of ascertaining SNP variation attempted to answer were to what degree were there “patterns of variation” across the genome, how much linkage disequilibrium existed (and whether that was constant across populations), and how to understand the connection between SNPs and biological function. This included not only defining sampling procedures and ethical considerations but also assessing how and whether analytical tools needed to be developed in order to better access SNP variation across the genome. Many of the discussions pertained to how to scale the ascertainment of variation through genotyping technology in the same way as sequence (see NHGRI Internal History Archive. Scanned Francis Collins Files, Box 7080, folder 003).
In the HGP proper, the same sequencing groups began in 1999 to discuss with significant detail the distribution of variation in the genome itself. Having much more sequence (as chapter 9 in this volume also details) and much more familiarity with sequencing technology and analysis allowed for a better understanding of variation and therefore better discrimination between polymorphism and error (although even with these improvements much checking still had to be done manually). Error checkers were able to distinguish between errors and polymorphism, for example, on the X chromosome, because of observed differences in the rate of polymorphism versus autosomal chromosomes. This meant that if an ambiguity was observed, it was more likely to be an error (or some other issue with the sequence data quality). A detailed understanding of genome biology was then critical because that knowledge allowed for a working account of which regions had higher degrees of polymorphism. Here, comparative studies, such as rat and mouse sequence, were of significance for helping make this determination.
Error checkers also used smaller labs’ studies of gene-rich regions to ensure that what was computationally determined to be an error or a polymorphism using Phred scores was in reality one or the other. This problem became particularly interesting when the HGP confronted sequencing those regions of the genome that were highly repetitive. Many of these regions, even within the euchromatic portion, remained only partially sequenced and understood until relatively recently (e.g., centromeric tandem repeats; see Hartley and O’Neill 2019). Simple sequence repeats, or where certain DNA motifs are repeated, were studied for their apparent lack of polymorphism (and whether this lack of polymorphism was due to issues with sequence read length or the actual absence of polymorphism in those segments of the genome).
Therefore, as public project raced toward the completion of the draft sequence in 2000, bioinformatics programs (such as the “Golden Path”) became much more able to ascertain the difference between error and polymorphism. As importantly, sequencing centers increased the number of reads (up to tenfold coverage), as this gave them a better understanding of under what conditions there would be single base-pair rearrangements, which occurred in the normal stitching together of clone fragments. Last, using standardized data from other labs as well as those members of the HGP, bioinformaticians were able to disregard stretches of the genome that were known to be highly repetitive in order to focus on those small segments that were full of gaps, and which really contained ambiguities and errors (NHGRI Internal History Archive. Scanned/Box 20, 2015 Bermuda folder).
However, the above processes were extremely difficult until there was an actual standard reference set of data to compare subsequent sequencing to, and the reference genome in many ways introduced its own suit of complexities and problems. Prior to the completion of the HGP, there was no standard dataset for the entirety of the human genome. Rather, there were just high-quality sequences, which could be used to model and refine tools and methods, as well as many fine-grained studies that pointed to specific features of genome biology, which allowed for bioinformaticians to make good informal estimations of various genome features. Thus, much of the knowledge around errors and polymorphisms was still ambiguous. Indeed, because distribution of errors across a specific genome sequence was always non-uniform, it was difficult to accurately distinguish polymorphism from error for a number of years after the project (He et al. 2013).
Another issue that was not discussed in the race to the “finished” sequence was the presence of structural variation. SVs, which were known to have some role in the genetic architecture of disease, could not be measured using late-1990s sequencing technology. Structural variations, such as insertions or deletions as well as translocations, were commonly referred to as structural if they were greater than whatever short-read sequencing technology was available at the time, although most in the grantee community understood SVs to be larger than 1 kb in size. Discussions of polymorphism and error were limited not only by the lack of a standard reference, or the ability to consistently close gaps for contiguous sequence, but also due to a conceptually and scientifically truncated conception of variation itself, which focused primarily, though not exclusively, on SNPs, where the lab of Evan Eichler and that of a few others were notable exceptions. Indeed, the focus on SNPs grew out of sequencing technology, which could capture base-pair changes. And accuracy itself was privileged on sequencing technology and bioinformatics with single base-pair resolution. This was technology that could not, in most cases, due to its short-read length, capture insertions and deletions of more significant size (and certainly not without a reference sequence to compare data).
This changed after 2003 with the “completion” of the human genome sequence and the development of a standard reference for sequence comparison, where individual groups could submit alternative configurations as annotations to the existing reference (such as Ref Seq). However, Olson and Green’s worries proved true. Because of the adherence to that specific 1/10,000 base-pair accuracy standard, and because of numerous discussions (and lack of real agreement between sequence centers) over the optimal sequence fragment length, the finished euchromatic portion of the genome possessed a number of significant gaps and errors. By 2014, Evan Eichler’s laboratory had resolved a number of these gaps and ambiguities but noted that “aspects of its structural variation remain poorly understood ten years after its completion.” And while they were able to close numerous gaps, especially in areas of the genome with high numbers of repeats using longer more contiguous reads, ambiguities still remained (Chaisson et al. 2015).
This posed a number of problems for the scientific community in the era after the HGP. In particular, the clinical genomics community began to notice that although there had been many, many improvements to the original 2003 genome, with continual “builds” throughout the 2000s and 2010s, there were enough gaps and errors that oftentimes it was difficult to gain any information about a variant unless it was well represented and annotated within the reference sequence. SVs, far from being responsible for solely significant phenotypic changes as had been assumed, are also important even for more common diseases, being responsible for protein changes, among other features (Ballouz et al. 2019). Indeed, researcher after researcher underscored the difficulties with all sorts of analysis as resulting not merely from the sequence gaps, but how the sequence itself was assembled through relatively short, relatively error-laden reads.
As importantly, the democratization of sequence due to significant cost reductions has shown that individuals exhibit a great deal more variation than has been anticipated, and the idea that a single standard reference sequence could represent the variational plurality of humanity appears a bit dated, and certainly not appreciative of the wealth of individual variation that has been discovered. The idea of a single reference as built up on mostly short-read consensus fragments underscores the worries of both Eichler and Gannett that there were assumptions being attached to the level of variation in the human genome sequence, and that this was a normative or a computational assumption (in the case of Eichler), which was not well represented by the data.
Thus, the sequencing error standard was agreed upon by most of the sequencing community but was viewed with significant caution by others. This standard had developed from the understanding of a particular kind of sequence resolution and assembly for a specific amount of time, effort, and funds. Such was the result of a few studies and back-of-the-envelope calculations about how prevalent polymorphism should be in the genome, adding another order of magnitude to distinguish polymorphism from error.
These practical suggestions were adopted when there was not a great deal of sequence, and with the existence of only a few effective tools (such as Phred and Phrap). This meant that any consensus would be tested and retested and refined as more sequence became available. And indeed, toward the end of the Genome Project, distinguishing polymorphism due to a better understanding of genome biology became apparent. Nonetheless, there was almost immediate dissent from Olson and Green that the error standard was, at most, a minimum standard. This meant that though there was a standard that the genomics community followed, this standard was not universally assented too.
In 2000, a draft sequence was produced, which yielded a wealth of biological information. This was followed by a sequence produced in 2003, which resolved many of the difficulties (errors and gaps) but also left many. Subsequent builds of this now “reference sequence” continue to demonstrate instead the particular difficulties of working with short-read technology, which left many gaps, and under the assumption that SNPs that would be ascertainable by short-read technology as well as distinguishable from errors. This resulted in several difficulties with SVs, which as being larger in structure were not able to be traced by mid- to late-1990s sequencing technology. If the Olson-Green standard had been followed (with 1/100,000 base-pair accuracy and larger fragments), then perhaps the project would have become prohibitively expensive, and it certainly would not have finished on time. The sequence, nevertheless, would have been more complete and more accurate than it is now, with accuracy, fidelity, and contiguity being pursued over efficiency.
Archival Sources
- Steve Fodor Interview. Originally conducted March 2, 2019. Available online at https://www.genome.gov/leadership-initiatives/History-of-Genomics-Program/HGP-30th-Anniversary/Jjy2Jzq1zUY.
- NHGRI Internal History Archive. Scanned/Box List 15/5 Bermuda-II.
- NHGRI Internal History Archive. Scanned Box 1272, folder 001.
- NHGRI Internal History Archive. Scanned/Francis Collins Papers, Box 7118, folder 012.
- NHGRI Internal History Archive. Scanned Box 024, file 008.
- NHGRI Internal History Archive. Scanned Box 1272, folder 005.
- NHGRI Internal History Archive. Scanned Box 2025, folder 35 1998 Planning-003.
- NHGRI Internal History Archive. Scanned Francis Collins Files, Box 7080, folder 003.
- NHGRI Internal History Archive. Scanned Box 20, 2015 Bermuda folder.
- NHGRI Internal Archive. Correspondence from Phil Green to Jeffery Schloss 1/14/1997.
- NHGRI Internal Archive. Correspondence from Maynard Olson to Jeffery Schloss 1/15/1997.
References
- Ballouz S., A. Dobin, J. A. Gillis. 2019. “Is It Time to Change the Reference Genome?” Genome Biology 20 (1): 159.
- Chaisson, M., J. Huddleston, M. Dennis, et al. 2015. “Resolving the Complexity of the Human Genome Using Single-Molecule Sequencing.” Nature 517:608–11.
- Eichler, E. E., R. A. Clark, and X. She. 2004. “An Assessment of the Sequence Gaps: Unfinished Business in a Finished Human Genome.” Nature Reviews Genetics 5 (5): 345–54.
- Ewing, B., H. LaDeana, M. C. Wendel, et al. (1998). “Base-Calling of Automated Sequencer Traces Using Phred. I. Accuracy Assessment.” Genome Research 8 (3): 175–85.
- Ferry, G., and J. Sulston. 2010. The Common Thread. Transworld.
- Ford, E. B. 1945. “Polymorphism.” Biological Reviews 20 (2): 73–88.
- Futuyma, D. J. 1988. “Sturm und Drang and the Evolutionary Synthesis.” Evolution 42 (2): 217–26.
- Gannett, L. 2003. “The Normal Genome in Twentieth-Century Evolutionary Thought.” Studies in History and Philosophy of Science Part C: Studies in History and Philosophy of Biological and Biomedical Sciences 34 (1): 143–85.
- Hartley, G., and R. J. O’Neill. 2019. “Centromere Repeats: Hidden Gems of the Genome.” Genes 10 (3): 223.
- He, Z., X. Li, S. Ling, et al. 2013. “Estimating DNA Polymorphism from Next Generation Sequencing Data with High Error Rate by Dual Sequencing Applications.” BMC Genomics 14: 535–43.
- Manolio, T. A., F. S. Collins, N. J. Cox, et al. 2009. “Finding the Missing Heritability of Complex Diseases.” Nature 461 (7265): 747–53.
- Maxson Jones, K., R. A. Ankeny, and R. Cook-Deegan. 2018. “The Bermuda Triangle: The Pragmatics, Policies, and Principles for Data Sharing in the History of the Human Genome Project.” Journal of the History of Biology 51 (4): 693–805.
- National Research Council Report. 1988. “Mapping and Sequencing the Human Genome.” https://www.nap.edu/catalog/1097/mapping-and-sequencing-the-human-genome.
- Olson, M., and P. Green. 1998. “A ‘Quality-First’ Credo for the Human Genome Project.” Genome Research 8 (5): 414–15.
- Pigliucci, M. 2007. “Do We Need an Extended Evolutionary Synthesis?” Evolution: International Journal of Organic Evolution 61 (12): 2743–49.
- Roberts, D. 2011. Fatal Invention: How Science, Politics, and Big Business Re-Create Race in the Twenty-First Century. The New Press.
- Sarkar, S. 1998. Genetics and Reductionism. Cambridge University Press.
- Shreeve, J. 2007. The Genome War: How Craig Venter Tried to Capture the Code of Life and Save the World. Random House Publishing Group.
- Ward, L. D., and M. Kellis. 2012. “Interpreting Non-Coding Variation in Complex Disease Genetics.” Nature Biotechnology 30 (11): 1095.