Skip to main content

Perspectives on the Human Genome Project and Genomics: 5 NHGRI Genetic Variation Program

Perspectives on the Human Genome Project and Genomics
5 NHGRI Genetic Variation Program
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomePerspectives on the Human Genome Project and Genomics
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. Cover
  2. Half Title Page
  3. Series List
  4. Title Page
  5. Copyright Page
  6. Contents
  7. Preface
  8. List of Abbreviations
  9. Introduction: Complexity, Contingency, and Controversy in Genomics
  10. Part 1. Producing the Genome
    1. 1. Challenges in the Early Years of the Human Genome Project at the National Institutes of Health: A Personal Retrospective
    2. 2. Unsung Contributors to the Human Genome Project: NIH Staff and Advisors
    3. 3. The NHGRI Genome Sequencing Cost Curve: An Indicator of Scientific Progress
    4. 4. History of the Encyclopedia of DNA Elements (ENCODE) Project
    5. 5. NHGRI Genetic Variation Program
    6. 6. Genome Technology Development Grants for the Human Genome Project and Beyond
  11. Part 2. Contextualizing the Genome
    1. 7. The Nature of Genomic Publishing
    2. 8. Europe and the Genome: An Overlooked Strategy for a Translational Genomics
    3. 9. Technological Change Driving Scientific Questions: Genomic Sequencing as a Case Study
    4. 10. Addressing Ethical, Legal, and Social Implications (ELSI): Navigating Ongoing Productive Tensions
    5. 11. “Variations on a Theme”: A History of Errors and Polymorphisms in the Human Genome Project and Beyond
    6. 12. Transforming the Genome into a Clinical Resource: DNA, Data, and Algorithms in Medicine
  12. Part 3. Interpreting the Genome
    1. 13. The Difference Genomics Makes: Characterizing Human Differences After the Human Genome Project
    2. 14. The Trouble with Being “Socially Responsible”: Science, GWAS, and Sexual Orientation
    3. 15. Epigenetics in Public Health: Comments on the “From Cells to Society” Approach
    4. 16. When Eugenic Enhancement Meets the Myth of Genetic Reductionism
    5. 17. modENCODE and the Elaboration of Functional Genomic Methodology
    6. 18. The Cancer Genome Atlas Project: Data-Driven, Hypothesis-Driven, or Something In-Between?
    7. 19. Large-Scale Biology: Philosophical, Historical, and Computational Perspectives
  13. Contributors
  14. Index

5 NHGRI Genetic Variation Program

Lisa D. Brooks

This paper is about my experience as a National Human Genome Research Institute (NHGRI) Program Officer overseeing the Genetic Variation Program and the major scientific issues that the program addressed. My training was in classical population genetics, a field with a rich theoretical foundation and, in the late 1990s, an increasing amount of information on protein variation in genes in many species. I worked with Marcus Feldman at Stanford as an undergraduate, with Richard Lewontin at Harvard for my PhD, and with Bruce Weir at NCSU as a postdoc. I was a faculty member at Brown University and a program officer at National Science Foundation (NSF) before moving to NHGRI in 1997. As an NHGRI Program Officer, I managed large projects in human genetic variation and a portfolio of grants in population genetics and statistical genetics. I managed the production of resources that vastly increased the amount of information on human genetic variation and enabled this information to be applied to studies of human traits and diseases. I worked closely with Jean McEwen of the Ethical, Legal, and Social Implications (ELSI) Program at NHGRI to deal with the ethical issues related to human variation and population sampling, especially processes for community consultation.

The human sequencing for the Human Genome Project focused on finding one sequence across the genome. The DNA libraries used for the sequencing came from an ancestrally diverse set of anonymous volunteers for this project. The reference sequence is a mosaic of chromosome regions from several people; in any region it is from one person, but data from about twenty-eight people made up the initial sequence. As the sequencing improved, data from many of the smaller libraries were replaced, leading to the reference human sequence coming from nine people. When it became possible to identify ancestries based on genome-wide sets of variants, analysis showed that the largest proportion, about 70 percent, came from an African American donor, the next largest proportion came from a donor with East Asian ancestry, and the remaining seven donors had European ancestry (Green et al. 2010).

Of course, the reference human genome sequence has been hugely useful for human biology, showing the locations of genes, regulatory elements, and heterochromatic regions. However, studying the genetic contribution to phenotypes and disease risk requires understanding the genetic variation where the genomes of people differ. I set up the Genetic Variation Program at NHGRI in 1997 to develop the resources that would be needed for these sorts of studies.

1. The State of Human Variation Knowledge at the End of the 1990s

By the end of the 1990s, there was much information on polymorphism in specific human genes. The overall pattern of human variation had been established: most common variants were in all populations (e.g., ABO), although the allele frequencies could differ among populations (Barbujani et al. 1997). This variation came from the original human population in Africa. As these people migrated to the rest of Africa and the rest of the world, they carried much of the genetic variation with them, although not all. This resulted in African populations, with long histories of relatively large population sizes, having more variation than non-African populations. As people moved out of Africa and as populations in and out of Africa underwent population bottlenecks, they lost many variants, especially rare ones. With the dispersal of humans around the world, mutations occurred in those populations, leading to variants found only in local regions.

2. The Polymorphism Discovery Resource (PDR)

Most of the common genetic variants could be found by looking in any population. For finding less common variants, though, many populations of diverse ancestries would be needed. Although much genetic research is disproportionately done in populations of European ancestry, NHGRI wanted to include samples from a diversity of populations to ensure that all populations would benefit from the research. Given the many variants shared among populations, no population would be frozen out of the benefits from these studies if they did not participate. However, rare variants in a population group would be better identified if that group or a related one were included. Since much ELSI discussion was still needed, Francis Collins decided that for this resource the population that each sample came from would not be named.

To address the need for a diverse set of samples, NHGRI set up the PDR (Collins et al. 1998). This set included 450 samples from African Americans, Mexican Americans, Americans with European ancestry, Americans with ancestry from several countries in East and South Asia, and American Indians. People with ancestry from more than one region were included. No phenotypic, medical, or ethnicity information was included for each sample. Samples from the first population groups were relatively easy to obtain; however, including samples from American Indians took more effort. After consultation with experts at the National Institutes of Health (NIH) on working with American Indians for genetic research, a group of NHGRI staff including Francis Collins attended a meeting with the leaders of the National Congress of American Indians in Washington, DC. There we discussed the goals, benefits, and risks of the project, and the processes and concerns of the American Indians. They confirmed that each tribe is sovereign and can make its own decision about participating in such a study. I also attended an Indian Health Service research meeting for more insight.

I attended a meeting in October 1998 on Genetic Research and Native Peoples: Colonialism through Biopiracy. Although the participants were not particularly interested in the goals of this study, they appreciated that NIH staff attended. They made clear that with all five hundred or so treaties between Indian tribes and the US government having been broken, there was little trust in the US government. Any research would need to be well justified and discussed with the tribes prior to any permissions. The most heartbreaking moment for me was when we talked with an elderly Indian man; he thought that diabetes had been solved for white people, but the cure was being withheld from Indians.

Based on the concerns of the American Indians, researchers who used the data at the National Center for Biotechnology Information (NCBI) or the cell lines at the nonprofit Coriell Institute for Medical Research had to agree to not try to identify the person or the population that each sample came from. Although identifying the general population group that a sample came from seemed out of reach in 1998, now it is easy to do. Thus, the PDR has been closed and will no longer distribute samples. Many other samples are available for which the donors provided consent for sharing the data and samples and labeling the populations.

3. Genetic Variation Across the Genome

Improvements in the technology for finding variants led to the discovery of many more. The most common variants in the genome are called single nucleotide polymorphisms (SNPs). This refers to a specific site in the genome where people differ. For example, the site may have both A and C bases, with people having genotypes AA, AC, or CC. Most SNPs are biallelic, but three or even four bases at a site may exist in a population. Sites may also differ by having deletions or duplications of one or up to many bases; these are called insertion/deletions (indels) when small (fewer than about fifty bases) or structural variants (SVs) when larger. SNP variants are easier to detect than other types of variants and were the most studied initially. However, both types may contribute to differences among people in phenotypes and disease risk; the large size of SVs means that more of the genome is variable for indels and SVs than for SNPs.

When multiple variants in a chromosome region are studied, the haplotypes are important descriptions of the data. A haplotype is the set of variants in a region along a specific chromosome. Aside from the X and Y chromosomes in males, people have two copies of each chromosome, and so have two haplotypes for each region of a chromosome. A haplotype could describe a chromosome from one end to the other, but usually smaller regions are studied.

Linkage disequilibrium (LD) refers to how associated the variants are in haplotypes. If there are N biallelic variants in a chromosome region, for example, there could be up to 2N haplotypes in that region. If all possible haplotypes exist at the frequencies of the product of the allele frequencies, then LD is zero and the variants are not associated with each other. If fewer haplotypes exist or their frequencies are not simply the product of the allele frequencies, then the alleles are associated with each other. This means that information about the alleles at one or a few variant sites on a haplotype in a chromosome region can provide information about which other variants are on that haplotype in that region.

When SNPs were studied more comprehensively in specific regions of the genome, the genome was found to be “blocky”: instead of all possible haplotypes being seen for a set of variants in a chromosome region, in much of the genome only a handful of haplotypes were found (Daly et al. 2001; Patil et al. 2001; Reich et al. 2001; Gabriel et al. 2002). This makes sense based on a limited set of haplotypes in the original human population, the human population bottlenecks, and the relatively short time for recombination to occur in small regions that would scramble the variants to make many haplotypes.

In a block region with only a few haplotypes, all the SNPs are in high LD with each other. This means that knowing a few variants in a region provides accurate information about the rest of the variants on the haplotypes in that region. Where recombination hotspots occur, there are many haplotypes across those hotspot borders, until the next block region with few haplotypes.

4. The International HapMap Project

These findings were the basis for the International HapMap Project. Since genotyping was expensive, especially for hundreds of thousands of SNPs, the most informative way of doing genotyping would be to genotype many SNPs in regions of low LD, and to genotype few SNPs in block regions with high LD where those SNPs contained information on the haplotypes and other SNPs in those regions.

Although the block structure of the genome was found to be an oversimplification, the idea was still sound that identifying regions needing more or fewer SNPs would allow the design of efficient and comprehensive assays across the genome. This would set the stage for genome-wide association studies (GWAS), by developing a haplotype map showing the patterns of haplotypes and LD across the genome in a several populations. Once these patterns were found, SNPs could be chosen to tag each region. There was nothing special about the tag SNPs; any of a set of SNPs could be chosen that would statistically represent a region. The set of tag SNPs could be placed on an array, which could be used to assay the genomes of many people with or without a disease or trait. Any SNPs found to be significantly more common in people with the disease or trait would be statistically associated with the trait, along with SNPs in LD with them in those regions. Tag SNPs thus tagged regions associated with a trait or disease. The LD allowed a subset of SNPs to be used to assay the genome, but within a tagged region all the SNPs would be in high LD with each other and associated with the disease or trait. Thus, identifying the causal variants, as compared simply to the statistically associated variants, would require more work, such as functional studies, but only in regions found to be associated with the disease.

Another justification for the HapMap Project was the Common Disease/Common Variant hypothesis, that common variants (minor allele frequency more than 5 percent in a population) are the major genetic contribution to the risk for many common complex diseases (Reich and Lander 2001). GWAS studies, especially small ones, would be successful only if the variants associated with a disease were common.

Based on the papers showing the block structure of the genome, researchers urged NIH to start a haplotype map project. NHGRI convened phone calls with international partners, including the Wellcome Trust in the UK and researchers in Canada, Japan, and China. In July of 2001, we had an initial meeting in Washington, DC (Brooks 2001). The scientific topics discussed included patterns of human variation, similarities and differences in these patterns among populations, and the goals and elements of a haplotype map project. Another major topic of discussion was whether it would be ethical to identify which populations the individual samples came from; with input from the ELSI participants we decided that it would be.

A project about human genetic variation clearly had important ELSI considerations in its design and implementation. Jean McEwen at NHGRI oversaw the ethical and sampling parts of the Project (chapter 10, this volume). The Project set up an ELSI group from the beginning, with international participation (International HapMap 2004). The first major question was what framework for the sampling (population vs. grid vs. other) would meet the scientific needs and ethical standards of the Project. Since haplotypes and LD are population-level phenomena, and there are ethically sound ways to study populations, it made sense to study populations.

The ELSI group discussed criteria for which populations to include. This Project aimed to support medical studies. This was in contrast to the Human Genome Diversity Project (HGDP), which had an anthropological focus on studying many specific small Indigenous human populations, to study questions such as the relatedness among these populations (Cavalli-Sforza 2005). The HapMap Project needed an ancestrally diverse set of population in order to include many variants and haplotypes, to provide a resource that would be useful for all populations. However, no specific populations had to be included, so the Project decided not to include vulnerable or small populations.

NHGRI staff consulted with a group of American Indian geneticists and doctors about whether it would be appropriate to include an American Indian population in the HapMap Project. This group understood the goals and benefits of the Project. Since only a few populations would be included, and all populations would benefit from the Project, they thought there was no need to include an American Indian population in the HapMap, so none was included.

The populations included were large populations without sharp boundaries for membership. Two populations had been suggested from each ancestral group, to make it clear that a population did not represent an entire ancestral group and to minimize the possibility that people who used the resource would inappropriately conflate ancestral geography with popular conceptions of race. However, given the cost of genotyping, the Project did not have the capacity to study two populations for each ancestral group. Once researchers in both Japan and China were involved and both governments provided funding, the Project decided to in clude samples from both of those countries.

A total of 270 samples were genotyped, mostly from mother-father-adult child trios. Although using trios resulted in fewer genomes being studied, they had the advantage of helping define haplotypes and allowing data quality assessment. The Centre d’Etude du Polymorphisme Humain (CEPH) samples, from US residents with European ancestry, already had been extensively studied for the human genetic map; it made sense to use samples from thirty CEPH trios to tie the HapMap data to this body of research. Charles Rotimi’s group collected thirty trio samples from the Yoruba population in Nigeria. Yusuke Nakamura’s group collected forty-five individual Japanese samples, and Henry Yang’s group collected forty-five individual Chinese samples; since both of these populations were included, the Project decided that the most information would come from individual samples rather than trios.

The criterion used for most populations was that a sampled individual had to have at least three of four grandparents from the population being studied. Most participants had four of four grandparents from that population, but the complexity of real and large populations was not a problem.

No phenotypic or medical information was collected about the sample donors. A project with only a few hundred samples would have limited power to show associations of variants with the phenotypes anyway, and not including phenotypic or medical information also helped to decrease concerns about individual donor privacy.

The ELSI group decided that community consultation would be needed for each community that might participate in the Project. The HGDP had met resistance from the Indigenous communities that it wanted to sample. The HapMap Project had a medical focus so that different types of populations were to be included. However, the HGDP issues underlined the need for careful community consultation, including taking no for an answer, as occurred with a population that was consulted in a similar manner for the subsequent 1000 Genomes Project. The ELSI group developed guidelines on how to do this consultation and a template for obtaining individual consent (Rotimi et al. 2007). The donors gave broad consent to open release of their data and to future use of the data and samples, including variation, gene expression, and proteomic studies. A Community Advisory Group was established in each community. As a demonstration of respect, the Project promised to keep the communities informed about how their samples and the HapMap resource were being used. Now, fifteen years after the Project was completed, that promise is still being kept; the Coriell Institute for Medical Research, which distributes the samples, gives quarterly summaries to the sample collection PIs to share with interested members of the communities where the samples were collected.

One topic that the ELSI Group and each community considered was how to name each community that provided the samples. The sampling was not done in exactly the same way for each population, which was fine for the medical-resource goals of the project. The ELSI group developed a statement about naming the communities, saying that they should be labeled in a way that is accurate, neither overly general nor too specific, and acceptable to members of the sampled communities (Coriell, n.d.).

The Project had several working groups (ELSI, genotyping, analysis), with weekly group phone calls and Steering Committee calls, and twice-yearly meetings. Progress was monitored frequently for sample collection, genotyping, and analysis. Ad hoc groups addressed specific needs (SNP production, IP issues, paper writing). The Data Coordination Center, jointly run by the European Bioinformatics Institute and NCBI, was crucial for monitoring progress, cleaning the data, and distributing the data. The Project had “bake-offs” to compare approaches for genotyping and data analysis, which allowed the Project to develop standard approaches. The Project developed standards for the field for representing variation data, such as the formats SAM (Sequence Alignment/Map), BAM (a binary form of SAM), and CRAM (a compressed form).

The initial plan was to type 600,000 SNPs across the genome (National Human Genome Research Institute 2002; International HapMap 2003). Competition among the genotyping companies resulted in large reductions in the cost of genotyping, allowing more SNPs and samples to be typed. The genotyping companies Illumina and Affymetrix joined the consortium under the same conditions as the academic groups, including providing SNP data publicly, not using participation in the Project for unfair advantage, and working cooperatively with the other groups. The genotyping company Perlegen Sciences developed cheaper technology, joined the consortium under the same conditions, and produced much of the phase II data (National Human Genome Research Institute 2004). The phase I paper included data on 1 million SNPs; the phase II paper included data on 3.1 million SNPs (International HapMap Consortium 2005, 2007a). The Phase 3 HapMap was based on the four initial HapMap populations plus an additional seven populations, to encompass more human variation. It provided data on 1.6 million SNPs in 1,184 samples (International HapMap 3 Consortium 2010).

A major concern was intellectual property and public data access. The funders and researchers in the Project were committed to open data release for the maximum public good. The legal advice we received, however, was that any group could add another SNP to the data already released and then patent those haplotypes in a way that could exclude public data access. Based on this advice, the Project reluctantly instituted a click-wrap data use agreement; data users had to promise that they would not exclude others from using the data. The downside of such an agreement is that the data could not be released publicly or included in other public genome-wide datasets. In 2004, Perlegen Sciences released a haplotype map with data on 1.6 million SNPs in seventy-one samples from three populations (Hinds et al. 2005). Based on the public release of this dataset and of the HapMap Project’s phase I dataset, the Project decided that there was much less risk of IP problems with the HapMap dataset and made all HapMap data public (Spencer 2004). There were no subsequent IP problems.

The HapMap data showed that the studied populations had similar patterns for the location of regions of higher LD and of recombination hotspots, with less LD in the Yoruba population than in the other populations. This meant that tag SNPs developed for the Yoruba population would also work for the other populations. Because of the lower LD, the Yoruba population would never be as well tagged as the other populations for a specific number of tag SNPs on an array. This was a problem when the first arrays with just a hundred thousand SNPs were developed but stopped being a problem with arrays of millions of SNPs. The LD patterns meant that untyped SNPs in high LD with typed ones could be inferred accurately, called imputation. The more tag SNPs used, the better the imputation. This was especially true for the more common SNPs.

The HapMap data allowed arrays to be developed to assay the genome efficiently and comprehensively in an unbiased manner. This meant that the entire genome could be studied to assess associations with diseases and traits, rather than looking only at genes that researchers already knew to be related to those diseases. Many GWAS studies have been done (Collins et al. 1998; Buniello et al. 2019); a major result is that many genomic regions found to be associated with and then some genes causally implicated in the disease process were not predicted by the field (Manolio et al. 2008; Buniello et al. 2019). Most GWAS-associated regions are not protein-coding regions, pointing to the importance of variants in regulatory regions.

Another major result is that the genomic basis for disease is more complicated than predicted when GWAS studies were starting. Many variants contribute to risk for most diseases, mostly with small effects. This has been called a failure of GWAS; however, it reflects the reality of genetic contributions to disease. Complex diseases are actually complex!

5. The 1000 Genomes Project

With the success of GWAS studies in finding genomic regions associated with disease, and the cost of sequencing having dropped considerably, a consensus emerged in 2007 that it would be valuable to sequence a diverse set of samples and release the data publicly (International HapMap Consortium 2007b). As with the HapMap, any such study would be too small to find associations between variants and diseases, so no phenotypic or clinical data were collected or distributed for this Project. This would also help protect participant privacy, with the sequence data being made public. The 1000 Genomes dataset would allow GWAS researchers to see most of the variants in any disease-associated region without having to take the time and expense to sequence their own samples. Researchers using other genomic datasets, such as ENCODE, would be able to see the variants in the functional genomic elements or other genomic features.

Common variants in a sample can be found by genotyping. The major added information from sequencing is finding rare variants, especially ones not already known and on genotyping arrays. Long sequencing reads can also show SVs.

The cost of sequencing was still too high to sequence each sample deeply. The Project decided to use shared haplotype information across samples to provide more accuracy for the genotype of each sample. Variants had to be seen in at least two samples to be called, so the variants seen are highly likely to be real and not cell line artifacts. Having at least fifty samples per population would allow LD patterns to be established by population. Studying samples from a set of related populations would find rare alleles better than the same number of samples from one population, since genetic drift in separate populations would increase the frequency of many rare variants in at least one such population.

Despite a two-year hiatus between the end of the HapMap Project and the beginning of the 1000 Genomes Project, the projects had many of the same participants, with some new centers doing the sequencing. These included the US sequencing centers at the Broad Institute, Baylor College of Medicine, and Washington University; centers in the UK at the Sanger Institute, in China at BGI, and in Germany at the Max Planck Institute for Human Genetics; and Illumina. The 1000 Genomes Project used the same structure and Data Coordination Center as the HapMap Project, adding some working groups such as for structural variation, exomes, functional interpretation, and chromosome Y. Thus, many of the interactions and processes had already been worked out.

Many research groups worked with communities for the sampling. A similar, but more streamlined, process of community engagement as for the HapMap was used for the additional populations in the 1000 Genomes Project. The same broad consent form was used. The Project studied 2,504 samples from twenty-six populations, from five major geographic ancestral groups. The Project included the four original HapMap populations, five of the additional HapMap 3 populations, and additional populations.

A large group of researchers analyzed the data. The Project started with a pilot phase to test the methods while the additional samples were being collected. Overall, the Project found eighty-eight million SNPs, small and large insertions and deletions, copy-number variants (CNVs), and other types of complex variants (1000 Genomes Project Consortium 2010, 2012, 2015a, 2015b). The samples were resequenced in 2019 to the field’s current sequencing depth and data quality, making the data even more useful (Byrska-Bishop et al. 2022). This sequencing found 117 million variant sites.

This is the largest dataset of human sequence data that is publicly available. Despite the number of samples studied seeming smaller each year as disease sequencing studies get larger, pretty much all human genetic studies use the data. They show variants and haplotypes in regions, and show what sequence data look like for “sanity checks” in other sequencing studies. As a central resource for the field, they become more valuable as other studies are done on the same samples. All the samples are available as cell lines, so researchers have used them to study cellular phenotypes such as drug reactions and RNA expression; the sequence data allow GWAS studies of these phenotypes (Gamazon et al. 2011; Stranger et al. 2012).

6. The Uses of Genomic Variation Data

The HapMap was developed to allow GWAS studies for diseases and traits, and the 1000 Genomes Project was developed to follow up those disease studies. These datasets have been widely used for those reasons. A natural follow-on has been the interest in personal genomics for clinical use. However, although the clinical effects of many variants are known (e.g., ClinGen),1 the clinical interpretations of the majority of variants are unknown (Rehm et al. 2018). Even diploidy is a problem; the effects of most variants in both heterozygous and homozygous forms are not well known. Incomplete penetrance also makes it hard to interpret variants. Since complex diseases are affected by multiple variants as well as interactions with the environment, predicting and then understanding genomic risk will require integrating information on variants across the genome. Large sample sizes are needed to find the contributions of any variants, common or rare, with small added risk, and to find rare variants that contribute to risk. However, there are limits on the ability to make predictions, based on the vastness of the “environment” and random effects. Expanded polygenic risk scores, incorporating much genetic and non-genetic information, will be important to provide as much information as possible about the risk of diseases.

The interest in genetic variation data that surprised Jean McEwen and me was the explosion of ancestry testing that occurred several years after these projects ended. Real populations have lots of complexity, historical layering, and fuzzy boundaries, but there are signals of ancestry locations, admixture, and historical population sizes. With enough datasets from many populations, ancestry testing can provide information, as long as one remembers that past populations have also had plenty of migration and admixture.

GWAS using LD information with genotype or sequence data is useful for finding regions associated with a disease, but functional information of many types will be needed to figure out which variants contribute to a phenotype or disease risk; causality is difficult to prove. Just as human genome sequencing started with the reference sequence and expanded to study genetic variation among many people, the ENCODE Project started with finding functional genomic elements and has expanded to study how variation in those elements affects how they function (chapter 4, this volume). Large-scale functional studies, including genome-editing methods, will be needed to understand how variants function differently to lead to different phenotypes and disease risk (Zhao et al. 2018). This information will be needed for interpreting variants in individual genomes and for figuring out interventions in the disease process.

Note

  1. 1. ClinGen, https://clinicalgenome.org/.

References

  • Barbujani, G., A. Magagni, E. Minch, and L. L. Cavalli-Sforza. 1997. “An Apportionment of Human DNA Diversity.” Proceedings of the National Academy of Sciences USA 94 (9): 4516–19.
  • Brooks, L. D. 2001. Developing a Haplotype Map of the Human Genome for Finding Genes Related to Health and Disease. NHGRI.
  • Buniello, A., J. A. L. MacArthur, M. Cerezo, et al. 2019. “The NHGRI-EBI GWAS Catalog of Published Genome-Wide Association Studies, Targeted Arrays and Summary Statistics 2019.” Nucleic Acids Research 47 (D1): D1005–12.
  • Byrska-Bishop, M., U. S. Evani, X. Zhao, et al. 2022. “High-Coverage Whole-Genome Sequencing of the Expanded 1000 Genomes Project Cohort Including 602 Trios.” Cell 185:3426–40.
  • Cavalli-Sforza, L. L. 2005. “The Human Genome Diversity Project: Past, Present and Future.” Nature Reviews Genetics 6 (4): 333–40.
  • Collins, F. S., L. D. Brooks, and A. Chakravarti. 1998. “A DNA Polymorphism Discovery Resource for Research on Human Genetic Variation.” Genome Research 8 (12): 1229–31.
  • Coriell. n.d. “Guidelines for Referring to Populations.” Accessed November 12, 2025. https://www.coriell.org/1/NHGRI/About/Guidelines-for-Referring-to-Populations.
  • Daly, M. J., J. D. Rioux, S. F. Schaffner, T. J. Hudson, and E. S. Lander. 2001. “High-Resolution Haplotype Structure in the Human Genome.” Nature Genetics 29 (2): 229–32.
  • Gabriel, S. B., S. F. Schaffner, H. Nguyen, et al. 2002. “The Structure of Haplotype Blocks in the Human Genome.” Science 296 (5576): 2225–29.
  • Gamazon, E. R., R. S. Huang, M. E. Dolan, and N. J. Cox. 2011. “Copy Number Polymorphisms and Anticancer Pharmacogenomics.” Genome Biology 12 (5): R46.
  • Green, R. E., J. Krause, A. W. Briggs, et al. 2010. “A Draft Sequence of the Neandertal Genome.” Science 328 (5979): 710–22.
  • Hinds, D. A., L. L. Stuve, G. B. Nilsen, et al. 2005. “Whole-Genome Patterns of Common DNA Variation in Three Human Populations.” Science 307 (5712): 1072–79.
  • International HapMap Consortium. 2003. “The International HapMap Project.” Nature 426 (6968): 789–96.
  • International HapMap Consortium. 2004. “Integrating Ethics and Science in the International HapMap Project.” Nature Reviews Genetics 5 (6): 467–75.
  • International HapMap Consortium. 2005. “A Haplotype Map of the Human Genome.” Nature 437 (7063): 1299–320.
  • International HapMap Consortium. 2007a. “A Second Generation Human Haplotype Map of Over 3.1 Million SNPs.” Nature 449 (7164): 851–61.
  • International HapMap Consortium. 2007b. Meeting Report: A Workshop to Plan a Deep Catalog of Human Genetic Variation. https://www.internationalgenome.org/sites/1000genomes.org/files/docs/1000Genomes-MeetingReport.pdf.
  • International HapMap 3 Consortium. 2010. “Integrating Common and Rare Genetic Variation in Diverse Human Populations.” Nature 467 (7311): 52–58.
  • Manolio, T. A., L. D. Brooks, and F. S. Collins. 2008. “A HapMap Harvest of Insights into the Genetics of Common Disease.” Journal of Clinical Investigation 118 (5): 1590–605.
  • National Human Genome Research Institute. 2002. “[HG-02–005] Large-Scale Genotyping for the Haplotype Map of the Human Genome.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-02-005.html.
  • National Human Genome Research Institute. 2004. “[HG-04–005] Additional Genotyping for the Human Haplotype Map.” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-04-005.html.
  • 1000 Genomes Project Consortium. 2010. “A Map of Human Genome Variation from Population-Scale Sequencing.” Nature 467 (7319): 1061–73.
  • 1000 Genomes Project Consortium. 2012. “An Integrated Map of Genetic Variation from 1,092 Human Genomes.” Nature 491 (7422): 56–65.
  • 1000 Genomes Project Consortium.2015a. “A Global Reference for Human Genetic Variation.” Nature 526 (7571): 68–74.
  • 1000 Genomes Project Consortium. 2015b. “An Integrated Map of Structural Variation in 2,504 Human Genomes.” Nature 526 (7571): 75–81.
  • Patil, N., A. J. Berno, D. A. Hinds, et al. 2001. “Blocks of Limited Haplotype Diversity Revealed by High-Resolution Scanning of Human Chromosome 21.” Science 294 (5547): 1719–23.
  • Rehm, H. L., J. S. Berg, and S. E. Plon. 2018. “ClinGen and ClinVar—Enabling Genomics in Precision Medicine.” Human Mutation 39 (11): 1473–75.
  • Reich, D. E., M. Cargill, S. Bolk, et al. 2001. “Linkage Disequilibrium in the Human Genome.” Nature 411 (6834): 199–204.
  • Reich, D. E., and E. S. Lander. 2001. “On the Allelic Spectrum of Human Disease.” Trends in Genetics 17 (9): 502–10.
  • Rotimi, C., M. Leppert, I. Matsuda, et al. 2007. “Community Engagement and Informed Consent in the International HapMap Project.” Community Genetics 10 (3): 186–98.
  • Spencer, G. 2004. International HapMap Consortium Widens Data Access. NIH News Release.
  • Stranger, B. E., S. B. Montgomery, A. S. Dimas, et al. 2012. “Patterns of cis Regulatory Variation in Diverse Human Populations.” PLoS Genetics 8 (4): e1002639.
  • Zhao, J., F. Cheng, P. Jia, N. Cox, J. C. Denny, and Z. Zhao. 2018. “An Integrative Functional Genomics Framework for Effective Identification of Novel Regulatory Variants in Genome-Phenome Studies.” Genome Medicine 10 (1): 7.

Annotate

Next Chapter
6 Genome Technology Development Grants for the Human Genome Project and Beyond
PreviousNext
Copyright 2026 by the Regents of the University of Minnesota

All rights reserved.
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org