Notes
4 History of the Encyclopedia of DNA Elements (ENCODE) Project
Elise A. Feingold
1. Background
The Human Genome Project (HGP) was guided by a series of strategic plans that set out specific goals to be accomplished in five-year time frames. As the completion of the HGP was anticipated in the early 2000s, the National Human Genome Research Institute (NHGRI) initiated a strategic planning process to consider the needs and opportunities that would follow this milestone. To inform this process, NHGRI held a series of workshops on a wide range of topics (National Human Genome Research Institute 2012c) to obtain input from the scientific research community. The resulting plan—A Vision for the Future of Genomics Research—was published in April 2003 (Collins et al. 2003).
One of these workshops—The Comprehensive Extraction of Biological Information from Genomic Sequence—was held in July 2002 with the goal of exploring ways to find all genes and other sequence-based genomic features in the human genome. The outcome of that workshop was the establishment of the Encyclopedia of DNA Elements (ENCODE) Program in 2003. The expectation was that ENCODE could provide a deeper understanding of the human genome and disease and contribute to disease prevention and treatment.
In 2001, the ability to identify protein-coding genes was reasonably good, but still limited. There were two main ways in which genes were identified. Computational tools, aided by the genetic code that had been determined in the 1960s (Nirenberg and Matthaei 1961; Leder et al. 1963), could be used to scan DNA sequence to find open reading frames, that is, regions of the genome that could code for proteins. But these approaches were hampered by the fact that most protein-coding genes are not coded in the genome in one contiguous sequence; rather, they are present in shorter segments (exons) that are interrupted by typically longer segments (introns) that do not code for proteins. When DNA is transcribed into RNA, these introns are spliced out, and often exons are also differentially spliced, resulting in what is termed RNA isoforms. As a result, the fine structure of genes was very challenging to determine in silico. Experimental methods, primarily through sequencing of cDNA libraries, had its own set of limitations, including finding genes expressed at low abundance, in rare cells, in restricted time in development, or in response to environmental perturbations. In addition, most DNA sequencing (until recently) had been done with short sequence reads, which made in silico assembly into specific isoforms very challenging.
In the early 2000s, our knowledge about other sequence-based features such as non-coding genes and regions that regulate gene expression was even more limited, and there was no known “regulatory” code analogous to the genetic code. Gene regulation was known to be directed by sequence features such as promoters and enhancers, but much information, particularly regarding enhancers, was limited to well-studied genomic regions and experimental methods to identify these regulatory sequences generally targeted one or a limited number at a time.
As the first draft of the human genome was being completed, there was a significant debate about how many protein-coding genes it encoded. Numbers ranged from 30,000—120,000 (Liang et al. 2000; Lander et al. 2001; Venter et al. 2001; Wright et al. 2001). The International Human Genome Sequencing Consortium, of which NHGRI was a part, estimated that the number of genes was on the lower end of that range (Lander et al. 2001). Following discussions on this topic, including at a brainstorming meeting in Santa Cruz on the future of genomics in August 2001, emerged the idea for a pilot project to resolve this question by systematically testing 1 percent of the human genome using all available approaches to predict and validate gene models; this information could then be used to extrapolate a firmer estimate for the entire genome (Eric Lander, Francis Collins, personal communication, 2001). Following subsequent informal discussions, this idea expanded to include an evaluation of all sequence-based features of the human genome including those in non-coding regions.
In July 2002, NHGRI held a planning workshop for this pilot project—The Comprehensive Extraction of Biological Information from Genomic Sequence organized by Eric Green (then Chief of Genome Technology Branch and Director of the NIH Intramural Sequencing Center) and Mark Guyer (then Director of Extramural Research and Assistant Director for Scientific Coordination), both at NHGRI. They convened approximately thirty scientific researchers as well as NHGRI staff to consider a draft proposal for a pilot project to analyze 1 percent of the human genome to identify all functional elements within that ~30Mb of DNA, which was provided to stimulate discussion. The workshop had a three-part structure: The first question posed to the workshop participants was whether such a pilot project was worthy and feasible to pursue. This was followed by a discussion of what such a pilot could look like and then a final session devoted to how it could be implemented. Unlike many of the other Strategic Planning Workshops, which were fairly wide open to a broad range of ideas from the community, this one was more of an organizational meeting, as the idea for what a project could look like was already fairly well developed.
The proposed pilot project was enthusiastically endorsed by the workshop participants, and major discussion points were articulated in the workshop summary (National Human Genome Research Institute 2002a). This summary laid out a series of key points that became the foundation for the research network that ultimately became the ENCODE Project. These key points included (1) an initial list of functional elements to be identified; (2) an initial list of potential technologies and strategies that could be used as well as the recognition of the importance of close interactions between computational and experimental approaches; (3) a suggested process for selecting the target regions that would comprise the 1 percent (~30 Mb) of the human gene based on different genome features; (4) criteria for participation in the research network that tried to balance being as open as possible while at the same time recognizing the scientific contributions of the participants; and (5) organizational issues including committees to manage and advise the project.
Based on the strong support for the proposed pilot project, NHGRI immediately moved forward to develop plans to fund it through the issuance of Requests for Applications (RFAs). The first step was to write a Concept “Pilot Project for the Exhaustive Determination of Functional Elements in the Human Genome” for discussion at the September 2002 National Advisory Council for Human Genome Research (NACHGR) that laid out a proposed scope for the initiative (National Human Genome Research Institute 2002b). This Concept described plans to initiate the pilot project essentially as described in the workshop report through the support of research projects solicited by an RFA. It did not specify a target budget or duration, which are typically included; rather, the Concept indicated that these would be determined based on the availability of funds and the actual scope of the project. The Concept envisioned that the major activities would be supported by the Cooperative Agreement funding mechanism that allows for substantial programmatic input (see below). In addition to the pilot project RFA, a second RFA focused on technology development was also proposed in the Concept, as it was recognized at that time that new technologies for high-throughput methods to identify functional elements would be needed to meet the long-term goal of a “comprehensive encyclopedia” of all sequence features in the human genome (National Human Genome Research Institute 2002b). The Council approved this Concept and also suggested consideration of a similar project in model organisms (see modENCODE below).
1.1. Encyclopedia of DNA Elements (ENCODE)
Immediately following approval of the Concept, NHGRI Program staff working to implement the pilot project considered what to name it. It was clear that the name of the workshop (Comprehensive Extraction of Biological Information from Genomic Sequence) and the name used in the Concept (Exhaustive Determination of Functional Elements in the Human Genome) were both awkward and too long. Several names were put forward by the group: ENCODE (ENCyclopedia Of DNA Elements); CAFÉ (Comprehensive Atlas of Functional Elements); DECIFER (DEtermining the Complete Inventory of Functional Elements in a “Region” with the notation that a better “R” word was desirable); and FEATURE (Functional Element ATlas for the FutURE). Although there was a passing concern that it was too close to “deCODE,” the name ENCODE was unanimously selected by program staff and Francis Collins.1 This was particularly gratifying to me since I came up with the name! Since ENCODE was implemented, several of other groups have used variations of “ENCODE” to name related projects.2
Two RFAs to support ENCODE were released in February 2003. The first, RFA HG-03–003 Determination of All Functional Elements in Human DNA (National Human Genome Research Institute 2003a), supported the pilot project, and the second RFA HG-03–004 Technologies to Find Functional Elements in Genome DNA (National Human Genome Research Institute 2003c) was the first of a series of calls for technology development projects. Shortly after these RFAs were released, NHGRI held the ENCODE Pilot Project Launch Meeting (National Human Genome Research Institute 2003b) to officially launch the project and to provide information to potential applicants. Following the standard NIH review process, awards were made at the end of September 2003.
1.1.1. Pilot Project
Based on peer review, programmatic and scientific balance, and input from the NACHGR, NHGRI funded eight awards for pilot projects (National Human Genome Research Institute 2012b). Although all funded awards were made to researchers outside of the NIH, several NHGRI intramural groups were actively involved in the ENCODE Pilot Project, most notably one led by Eric Green. His group’s role was to select homologous bacterial artificial chromosome clones from DNA libraries constructed from the DNA of a wide range of non-human species and sequence them to identify highly conserved regions that were expected to be functionally important by virtue of being evolutionarily conserved.
While the methods and technologies used in the different phases of ENCODE evolved, the basic premise of most of the experimental projects was to apply high-throughput assays to measure biochemical activities associated with specific regions of the genome that had previously been mechanistically associated with particular classes of functional genomic elements (see Figure 4.1). For example, projects mapped binding sites of transcription factors and histone modifications that were associated with promoters and enhancers. However, it is important to point out that this approach identifies genomic regions that have the potential to have functional activity, and additional work is needed to determine if, when, and where these regions have functions in particular biological settings (see below). Unlike the genome sequence, which is nearly identical in every cell in an organism, most genes and non-coding sequences that regulate the genome are active only in specific biological contexts.
Several key steps were needed to get the Consortium launched. An organizing committee was convened prior to the funding of the Pilot Project awards. One of its main tasks was to select target sequences comprising the 1 percent (~30Mb) to be analyzed. The committee decided on a mixed strategy. Half of the sequences would be selected manually based on the existence of substantial prior knowledge about genes and/or other functional regions and for which substantial comparative genomic data was available. Fourteen target regions ranging in size from 0.5 to 2 Mb (totaling ~15Mb) were selected. The other half of the ~30Mb were thirty 0.5Mb regions that were selected using a “stratified random-sampling strategy based on gene density and level of non-exonic conservation” to include regions with a range of gene and functional element density (National Human Genome Research Institute 2012a).
Figure 4.1. ENCODE Assays and Data Types. Schematic representation of the high-throughput biological assays that are employed to generate biochemical data mechanistically associated with specific functional elements. These data are computationally analyzed to identify candidate regulatory elements and transcripts used to populate the ENCODE Encyclopedia. Each technology, written in a box within the figure, points to the biological region or material that it works upon.
Figure Description
Figure 4.1. A strand of DNA is unwound from chromatin, which extends from the q-arm of a chromosome. The strand of DNA is wrapped around histones on the left side of the figure and is progressively unwound towards the right, revealing a double helix. Relevant sites are annotated for polymerase binding sites, methylation sites, and those sites which serve as long range regulatory elements, promotors and genes. The bottom of the figure includes a line diagram representing double-stranded DNA and the points on the DNA where genomic technologies can be used to interrogate functional elements.
As described in more detail below, tight coordination and project management enabled by frequent communication had been an essential feature in the success of the HGP and ENCODE followed this successful model. For example, monthly ENCODE teleconference calls were held, fittingly, on Fridays at 11:00 a.m. ET, which previously had been the timeslot for the HGP sequencing teleconference calls. This time of day was chosen to accommodate participants across a wide range of time zones, including GMT (4:00 p.m.) and PT (8:00 a.m.). Also, discussed in more detail below, was the early recognition for the need to have a centralized repository to store, display, and make data available.
Since ENCODE was a new project, the Consortium decided to write a “Marker Paper”—a brief introduction to the project including a description of its goals—that was published in Science in October 2004 (ENCODE Consortium 2004). It laid out the research strategy, including target sequences, organization, operating principles, plans for data management, and long-term goals. The writing of a Marker Paper was in keeping with NHGRI’s practice of informing the community about what resources they could anticipate having access to and striving for transparency and inclusion (Strausberg et al. 1999; ENCODE Consortium 2004).
The pilot project concluded with series of papers including a capstone publication (Birney et al. 2007) and twenty-eight companion papers in the Genome Research ENCODE Special Issue (ENCODE Consortium 2007b). As the first major set of publications from ENCODE, this served to bring the community’s attention to the project at a time when the next phase of ENCODE was being planned and modENCODE (see below) was just being initiated. The process of writing and publishing a Consortium-wide paper along with companion papers, however, was met with a number of challenges that presaged challenges in future ENCODE and modENCODE publications (see below). For example, the Consortium initially submitted five independent papers on different themes, including transcription and gene regulation that were of varying strength. Based on the reviewers’ and editors’ feedback, these five papers were condensed into the capstone paper.
Some of the key findings and surprises included the discovery that most of the ~30Mb of the human genome studied in the pilot is transcribed into RNA, that about half of the potential functional elements identified do not appear to be located in evolutionarily conserved regions, and that regulatory elements are found within and downstream of genes more frequently than previously thought. Most importantly, the pilot project demonstrated the feasibility of its approach to functionally annotate the human genome.
1.1.2. ENCODE 2
Based on the success of the pilot project, NHGRI decided to fund another phase of ENCODE (ENCODE 2) that continued support of pilot projects focused on the selected target regions as well as supported new projects that scaled to whole-genome analysis. The plan was endorsed by the NACHGR in May 2006 (National Human Genome Research Institute 2006d), and three RFAs were subsequently released and applications funded (National Human Genome Research Institute 2018a): (1) To support data production: Creating the Encyclopedia of DNA Elements (ENCODE) in the Human Genome (U01 and U54) (National Human Genome Research Institute 2006a); (2) To support data management: A Data Coordination Center for the Encyclopedia of DNA Elements (ENCODE) Project (U41) (National Human Genome Research Institute 2006b); and (3) To support data analysis: A Data Analysis Center for the Encyclopedia of DNA Elements (ENCODE) Project (U01) (National Human Genome Research Institute 2007a). Two additional RFAs were funded to support additional technology development.
Major technological advances in DNA sequencing at this time enabled the adaptation of previously used assays reliant on microarray readouts to sequencing readouts as well as the development of new sequence-based approaches to map biochemical activities; the advances increased the feasibility of whole-genome analyses. (See Technology Development section below.) However, in these early days of implementing the new technologies, uncertainties for success in making the leap from 1 percent to the entire genome were of concern. To mitigate these uncertainties, NHGRI supported several groups working on similar approaches and made provisions to move funds between groups in the later years if the need arose, although in the end that was not necessary.
Many of the observations that had been made in the pilot project were confirmed and expanded upon in this second phase of ENCODE. ENCODE 2 provided a richer and more complex view of the human genome, identifying a vast number of sequences with potential regulatory activity that control when and where genes are expressed. For example, experimental evidence was found that regulatory elements can work from a considerable distance more often than had been previously realized. Many disease-associated genetic variants were found to be located within the non-coding regulatory regions that ENCODE identified, which was surprising at the time, as it was expected that most would be located within protein-coding regions and amplified the importance of studying non-coding sequences in a systematic way. By studying more than one hundred different cell types, additional observations could be made about which functional elements are expressed in which cell types, enabling the study of understanding why different types of cells have different characteristics and functions—and this only increased as ENCODE expanded the repertoire of cell types studied in ensuing years. What became evident was the power of genome-wide functional data to enable the derivation of generalizable observations on genome structure and function that heretofore was based on a limited number of genes or genomic regions.
In September 2012, ENCODE published these findings in a paper reporting on integrated analyses of ENCODE data in Nature (ENCODE Consortium 2012b) along with ~thirty companion papers in Nature and other journals (ENCODE Consortium 2012a). As part of this collection were several publishing innovations developed with Nature, including “threads,” themes that ran through multiple publications. In a remarkable demonstration of collaboration, all publications, whether they were from a journal within the Nature family or other journals (e.g., Cell, Science, J. Biol Chem.), were published simultaneously and linked to this Nature site.
With the availability of ENCODE data and tools now publicly available and an increased awareness of the project based on these Consortium publications, the utility of ENCODE as a community resource was starting to be more fully realized. In fact, it was shortly after this time that community publications using ENCODE data and tools started its exponential growth (see below and Figure 4.2).
There was some negative response to these publications, however, that largely revolved around two main points. The first was the perennial concern expressed by some in the biomedical community that “big consortium science” was a waste of resources and that the funds devoted to it would be better spent on individual investigator-directed research. This argument was not dissimilar to objections raised by the HGP and other large consortia. As discussed below, it is NHGRI’s view that there is great value in developing resources for the community that are generated in a coordinated, standardized, and efficient manner rather than being done in a distributed fashion such that comparisons across projects are not nearly as feasible.
The second main objection was the use of the term “functional” when describing the regions of genomic DNA that had been identified using the high-throughput assays employed by ENCODE researchers that primarily measured biochemical activity and, in fact, had not demonstrated definitive functional activity. This objection was fueled by the statement that the data could allow “us to assign biochemical functions for 80% of the genome, in particular outside of the well-studied protein-coding regions,” which was misinterpreted by some to mean that 80 percent has a definitive biological role. This led to a debate in the scientific community in 2013 (Doolittle 2013; Eddy 2013; Graur et al. 2013; Brunet and Doolittle 2014) about what it means to ascribe function to any genomic region.
In writing the paper, the Consortium members were aware of the sensitivities on describing ENCODE’s findings with respect to function and had tried to address it directly by stating upfront: “Operationally, we define a functional element as a discrete genome segment that encodes a defined product (for example, protein or non-coding RNA) or displays a reproducible biochemical signature (for example, protein binding, or a specific chromatin structure)” and otherwise providing qualifiers. However, our attempts were not sufficiently clear to avoid some confusion and criticism. While some negative reception, differences of scientific opinion, or objections to a scientific approach or conclusion are not unexpected, I was surprised by the intensity of the reaction by several individuals and how social media can amplify the voices of a few (which in hindsight is not as surprising).
Figure 4.2. (a) Consortium and Community Publications. Number of publications (x-axis) over time (y-axis) identified by NHGRI to have used ENCODE data and other resources. As of 2019, papers published by the consortium exceeded 1,000. Papers published by the ENCODE user community neared 2,300. (b) Diseases Studied in Community Publications. Frequency of specific diseases studied publications from the ENCODE user community demonstrating the broad translational value of the ENCODE data and resources. Disease categories range from cancer to trauma.
Figure Description
The first component for Figure 4.2, labeled a, is a graph, described by the following table, showing the number of papers which make use of ENCODE data. These papers come from the consortium and from the community of non-funded users. The Y-axis includes the number of publications and the X-axis shows the year from 2007–2019. ENCODE Consortium-supported publications are indicated in gray and publications by others are indicated in black. The ENCODE consortium papers grow linearly to over 1,000 in 2019 and the ENCODE community papers grow sigmoidally to around 2,300.
Date x-axis | ENCODE Consortium Publications y-axis | ENCODE Community Publications y-axis |
|---|---|---|
November 2007 | 125 | 7 |
February 2008 | 138 | 7 |
May 2008 | 145 | 10 |
August 2008 | 152 | 15 |
November 2008 | 159 | 16 |
February 2009 | 172 | 20 |
May 2009 | 179 | 22 |
August 2009 | 184 | 27 |
November 2009 | 203 | 32 |
February 2010 | 218 | 32 |
May 2010 | 233 | 38 |
August 2010 | 246 | 47 |
November 2010 | 261 | 59 |
February 2011 | 275 | 75 |
May 2011 | 288 | 96 |
August 2011 | 306 | 120 |
November 2011 | 329 | 161 |
February 2012 | 348 | 190 |
May 2012 | 356 | 211 |
August 2012 | 402 | 250 |
November 2012 | 421 | 289 |
February 2013 | 440 | 337 |
May 2013 | 456 | 429 |
August 2013 | 481 | 532 |
November 2013 | 530 | 646 |
February 2014 | 553 | 742 |
May 2014 | 584 | 863 |
August 2014 | 605 | 938 |
November 2014 | 649 | 1047 |
March 2015 | 691 | 1175 |
May 2015 | 724 | 1330 |
August 2015 | 749 | 1431 |
November 2015 | 782 | 1569 |
February 2016 | 806 | 1665 |
May 2016 | 836 | 1754 |
August 2016 | 855 | 1833 |
November 2016 | 879 | 1962 |
February 2017 | 907 | 2020 |
May 2017 | 931 | 2071 |
August 2017 | 954 | 2114 |
November 2017 | 981 | 2173 |
February 2018 | 1003 | 2193 |
May 2018 | 1031 | 2241 |
August 2018 | 1046 | 2289 |
The second component of Figure 4.2, labeled b, is a pie chart, whose values are included in the following table, showing the diversity of the topics covered by papers which make use of the ENCODE data, organized by disease categories studied by the ENCODE user community. The three largest disease categories are cancer at 36%, autoimmunity/allergy at 12 % and human genetics at 12%. The other, smaller disease categories include: neurological/psychiatric, cardiovascular, multiple diseases, metabolism, infectious disease, musculoskeletal, ophthalmology, aging, developmental, respiratory, reproductive/dimorphism, hematology, dermatological and trauma, and range from 10% to less than 1%.
Disease Category Studied by ENCODE | Percent of Total |
|---|---|
Cancer | 36% |
Autoimmunity/Allergy | 12% |
Human Genetics | 12% |
Neurological/Psychiatric | 10% |
Cardiovascular | 8% |
Multiple Diseases | 4% |
Metabolism | 4% |
Infectious Disease | 3% |
Musculo-skeletal | 3% |
Ophthalmology | 2% |
Aging | 2% |
Developmental | 1% |
Respiratory | 1% |
Reproductive/Dimorphism | 1% |
Hematology | 1% |
Dermatological | 1% |
Trauma | 0% |
A group of ENCODE 2 participants addressed these concerns in a paper (and follow-up letter) outlining the challenges of ascribing function to specific regions of the genome, including the lack of consensus on how to define function particularly across different scientific disciplines (Kellis et al. 2014a, 2014b). This paper discussed the strengths and limitations of the three approaches taken to defining function—experimental, sequence conservation, and genetic—describing how each provides complementary lines of evidence for function and all are needed to identify functional DNA sequences, determine their boundaries, and ascertain their biological roles at the molecular, cellular, and organismal levels.
ENCODE has made a sustained effort by to be transparent in presenting ENCODE data, analyses, resources and in communicating inferences, assumptions, and conclusions. For example, as discussed below, ENCODE compiled a “Registry of candidate cis-Regulatory Elements (cCREs)” using data from ENCODE and the RoadMap Epigenomics Project.3 In addition, a visualization tool “SCREEN” was developed to enable exploration of cCREs, how they are classified, and the supporting data (Moore et al. 2020).4
1.1.3. modENCODE and mouse ENCODE
When NHGRI Program staff presented the Concept for the ENCODE Pilot Project, the question was raised during Council discussion about why we were planning to start with the study the human genome and not first, as had been the successful strategy of the HGP, start with studying5 model organism genomes. This idea was supported by others on the Council—as noted in the Council minutes from September 2002 (National Human Genome Research Institute 2002b): “The Council approved the concept, but noted that there would also be considerable merit in a comparable program focused on model organisms.” There was also enthusiasm for such an endeavor among NHGRI staff and from that important discussion grew the modENCODE project (for model organism ENCODE).
In considering which organisms to include in modENCODE, program staff reached out to several key model organism communities for which the completed sequence was available: S. cerevisiae (yeast), C. elegans (worm), and D. melanogaster (fruit fly). Each community considered the opportunities based on the state of genomics at the time and provided NHGRI with White Papers for what such projects could look like. With community input from these White Papers, NHGRI developed two RFAs to support the modENCODE effort. These RFAs were issued in March 2006: one for data production—Identification of All Functional Elements in Selected Model Organism Genomes (National Human Genome Research Institute 2006e); and one for a Data Coordination Center (National Human Genome Research Institute 2006c) to support the generation and management of functional data from C. elegans and D. melanogaster. Based on the advanced state of genomics of S. cerevisiae and the path described for further development in the White Paper, NHGRI decided to not include this organism in modENCODE. modENCODE awards were made in April 2007, five months before ENCODE 2 was funded (National Human Genome Research Institute 2014). The modENCODE Consortium also published a White Paper outlining the goals and objectives for the project (Celniker et al. 2009).
The modENCODE Consortium members both collectively and individually published a number of papers reporting on their results. Of note are two sets of Consortium papers: The first set of papers was published online in December 2010, primarily focused on analyses within one species or the other (Gerstein et al. 2010; Ikegami et al. 2010; Ledford 2010; modENCODE Consortium et al. 2010; Allen et al. 2011; Berezikov et al. 2011; Brooks et al. 2011; Cherbas et al. 2011; Chung et al. 2011; Eaton et al. 2011; Ercan et al. 2011; Graveley et al. 2011; Hoskins et al. 2011; Kharchenko et al. 2011; Liu et al. 2011; Lu et al. 2011; modENCODE Consortium 2011; Niu et al. 2011; Nordman et al. 2011; Riddle et al. 2011; Spencer et al. 2011). The second set, published in August 2014, included both species-specific analyses as well as three major papers reporting on comparative analyses across species (Araya et al. 2014; Boyle et al. 2014; Brown et al. 2014; Gerstein et al. 2014; Ho et al. 2014). The former papers provided a detailed view about the functional components of each genome, and the latter set of papers described ways in which gene regulation in the fly, worm, and human genomes follow similar overall principles yet differ in the specific encoding and execution of their functions.
1.1.4. Impact of ARRA Funding
In 2009, NIH received significant additional funding through the American Recovery and Reinvestment Act and NIH Institutes/Centers solicited supplements and new applications (National Institutes of Health 2009).6 With this unexpected opportunity, NHGRI was able to support a Data Analysis Center focused specifically on modENCODE, which greatly enhanced the ability to analyze modENCODE data (Johnson 2011). Supplements were also made to existing modENCODE and ENCODE projects, primarily to augment data production.
In addition, three new awards were made to expand ENCODE to analyze the mouse genome, an important, well-studied model organism whose genome is more closely related to the human than fly or worm. The three groups were joined by a fourth group supported by an investigator-initiated R01 award to form the mouse ENCODE Consortium (National Human Genome Research Institute 2016e). An initial description of the resource was published in 2012 (mouseENCODE, et al. 2012). This group subsequently published 10 papers in 2014 (mouseENCODE Consortium 2014). These papers describe the data generated by mouse ENCODE as well as comparative analyses between mouse and human that identified both conserved sequence features as well as transcribed and regulatory regions that show significant divergence (Yue et al. 2014). Following mouse ENCODE, both mouse and human genomes were studied in subsequent phases of ENCODE.
1.1.5. ENCODE 3
It was anticipated that only a limited number of biological samples (mostly cell lines) would be studied by the end of ENCODE 2, and so another phase of ENCODE—ENCODE 3 (National Human Genome Research Institute 2018d)—was funded with the aim of greatly expanding the biological space to be interrogated. One RFA (National Human Genome Research Institute 2011c) supported functional element mapping in both mouse and human. A second RFA (National Human Genome Research Institute 2011b) supported a combined Data Analysis and Coordination Center with the goal that these two activities be closely coordinated. A third RFA (National Human Genome Research Institute 2011a) was a new effort to expand the computational expertise brought to bear on ENCODE data with the goal of strengthening the utility of the ENCODE resource for the research community. As in ENCODE 2, a suite of RFAs were funded in ENCODE 3 to support additional technology development (National Human Genome Research Institute 2011e–g).
During ENCODE 3, emphasis was placed on taking ENCODE data, as well as RoadMap Epigenomics data and creating a series of annotations that comprise the “Encyclopedia.”7 The Encyclopedia includes ground-level annotations of individual or related datatypes and integrative level annotations. The main product of the integrative level annotations is the registry of candidate cis-regulatory elements (cCREs) as well as a web-based tool—Search Candidate cis-Regulatory Elements by ENCODE (SCREEN)—to connect to other annotations and raw ENCODE data. The Encyclopedia is a “living” entity that will evolve and be updated, enabled by additional data (e.g., from Functional Characterization Centers [FCCs], see below) and analyses.
While the ENCODE 3 Consortium members published papers throughout the funding period, a major body of work was published in a coordinated set of papers in 2020 featuring a flagship paper and twenty-nine additional peer-reviewed papers (Nature 2020). The flagship paper provided a detailed view of the updated catalog of candidate functional elements ENCODE had compiled, including the millions of cCREs found in the human and mouse genomes, as well as using SCREEN to navigate ENCODE 3 data (Moore et al. 2020). The additional papers featured more targeted investigations into mouse epigenetics and gene expression, human gene expression and RNA regulation, and computational analysis and tools. Finally, the ENCODE 3 PIs wrote a Perspective on ENCODE to provide a brief history of ENCODE and high-level summary of ENCODE data production, genome annotation, and publications over the years (Snyder et al. 2020).
1.1.6. ENCODE 4
A fourth phase of ENCODE was funded in 2017 after receiving strong enthusiasm from the research community for continuing mapping and computational analyses as well as moving beyond cataloging candidate functional regions to functionally validating them in specific biological (National Human Genome Research Institute 2015b). As a result of these recommendations and approval of the Concept by the NACHGR (National Human Genome Research Institute 2015a), NHGRI released another series of RFAs and supported five interrelated activities: (1) continued support of mapping candidate functional elements in human and mouse (National Human Genome Research Institute 2016a); (2) new functional characterization efforts to implement a number of different approaches to test candidate elements in different biological contexts (National Human Genome Research Institute 2016b); (3) continued support of computational groups to bring additional expertise to the analysis of the ENCODE resource (National Human Genome Research Institute 2016b); (4) a Data Coordinating Center (National Human Genome Research Institute 2016d); to closely interact with (5) a Data Analysis Center (National Human Genome Research Institute 2016c). ENCODE 4 was managed similarly to previous phases of ENCODE and resulted in a greatly expanded number of mapping datasets, functional characterization of ENCODE data, and updated versions of the Encyclopedia. ENCODE 4 (National Human Genome Research Institute 2020) ran from 2017 to 2022.
Funding for ENCODE 4 ended in 2022 and another series of publications are anticipated. NHGRI developed plans to sustain the ENCODE resource for years to come, as it does not plan to fund an additional phase of ENCODE. While considerable effort is still needed to fully annotate and understand the human and mouse genomes, NHGRI’s recent strategic vision for how genomics can be applied for improving human health underscored how much genomics has evolved over recent years, warranting fresh approaches and perspectives to address these important questions (Green et al. 2020). Of note are two new efforts underway at NHGRI that address genome function in novel ways: (1) Impact of Genomic Variation on Function was funded in 2021 to systematically understand how genomic variation effects genome function and phenotype (National Human Genome Research Institute 2022b) and (2) Molecular Phenotypes of Null Alleles in Cells (MorPhiC)—Phase 1 is expected to be funded in 2022 with the long-term goal of creating a catalog of molecular and cellular phenotypes for null alleles for every human gene (National Human Genome Research Institute 2022c).
2. Hallmarks of ENCODE
A number of hallmarks of ENCODE contributed to its success. Many of these were envisioned early in the development of the project—in the planning meeting pilot project proposal, concepts, RFAs and Marker Paper (ENCODE Consortium 2004)—and served as a strong foundation that was built upon over the ensuing years. These features are described in detail below.
2.1. Consortium Organization and Membership
The ENCODE Pilot Project comprised a group of investigators who were expected to work together in a highly collaborative manner to evaluate computational and experimental approaches for their potential for high-throughput identification of a range of functional elements at genome scale. This style of Consortium member interactions has continued through all phases of ENCODE and is considered one of its strengths because of the value of providing multiple scientific points of view, discussing and evaluating approaches as well as facilitating coordinated production, analysis, and distribution of data.
The basic infrastructure supporting the Consortium was largely similar throughout each phase of ENCODE, although it evolved as additional activities were added (e.g., Computational Analysis and Functional Characterizations Centers). With the exception of the technology development awards, all ENCODE-funded projects were supported through the Cooperative Agreement mechanism (National Institutes of Health 2020), which allows for substantial staff involvement. Cooperative Agreements are defined by the relationship between any awardee and NIH staff but are typically used in research consortia; they provide the means for active participation and flexibility to ensure that the goals of individual projects as well as the Consortium as a whole are met. For mapping projects, annual milestones were set, and progress reports were submitted on a regular basis, typically quarterly, following the practice of the HGP.
Frequent communication by a series of means was a necessary feature of ENCODE. In addition to monthly Principal Investigator (PI) and Consortium teleconference calls there were frequent calls of various working groups focused on particular topics, for example, issues related to specific data types and data analysis. Multiple means of electronic communication were employed beyond email, including use of a common wiki site, Google documents, and Slack channels. However, there was no substitution for face-to-face meetings to facilitate communication and build relationships. Annual Consortium-wide meetings as well as periodic PI and analysis-focused meetings were held in person, except during the SARS-CoV-2 pandemic.
In addition to project oversight by NHGRI staff, NHGRI convened an External Consultants Panel (National Human Genome Research Institute 2018b). This group of outside researchers provided feedback on the progress and scientific direction of the program, as described in the RFAs (e.g., National Human Genome Research Institute 2016a–e).
It was also recognized early on that there would be immense benefit in opening the Consortium to any academic, government, or private-sector researcher interested in participating, to expand the expertise needed to address this enormous challenge. In each phase of ENCODE, interested investigators not funded by ENCODE, but who agreed to abide by criteria of participation (National Human Genome Research Institute 2018a) (e.g., to make a significant scientific contribution and to release data in accordance with the Consortium’s data release policy), were brought into ENCODE and participated in Consortium activities. Allowing for broader participation than what NHGRI could support through ENCODE enriched the Consortium and its resources by providing additional analytical expertise, data types, and perspectives.
2.1.1. Interactions with Other Genomics Projects
The advantage of collaborating and synergizing with other, related, genomics projects such as the Mammalian Gene Collection (Gerhard et al. 2004) and International HapMap Project (International HapMap Consortium 2003) was also recognized to be important during the ENCODE Pilot Project. Additional collaborations were established during subsequent phases of ENCODE with several NIH Common Fund Programs, for example, RoadMap Epigenomics, Genotype-Tissue Expression (GTEx),8 and 4D Nucleome Programs,9 as well as the International Human Epigenome Consortium,10 the latter of which ENCODE was a member. The ENCODE Portal11 now hosts the RoadMap Epigenomics Program’s data where they are provided to the research community in the same processed form that is used for ENCODE data, ensuring interoperability and consistency between the Projects’ data.
2.1.2. Technology Development
NHGRI has a strong and sustained history of supporting technology development since the initial days of the HGP. When ENCODE was being planned, it was recognized that to reach the long-term goal of developing a comprehensive annotation of the human genome it would be critical to stimulate new and improved approaches, methods, and technologies to have a sufficient set of tools available for high-throughput, genome-wide identification of functional elements. To fulfill this need, NHGRI supported four rounds of funding for computational and experimental tool technology development. The first two rounds were funded during the pilot project in 2003 and 2004 and supported a total of twelve projects focused on a wide range of biochemical features of DNA including regions that were potentially functional based on their accessibility to DNA cleavage, long-range interactions between regions, and regions where regulatory proteins bind (National Human Genome Research Institute Institute 2012a). At the time of funding ENCODE 2, NHGRI supported a third round of technology development projects (National Human Genome Research Institute Institute 2018c), and the fourth was funded in 2012 at the same time as ENCODE 3 awards (National Human Genome Research Institute Institute 2018d). By the time that ENCODE 4 was being planned, NHGRI recognized that many of the technologies that were or could be developed for ENCODE had the potential for broad applicability and that methods being developed by a more general technology development call (which NHGRI was expanding) could be applied to ENCODE. Therefore, the decision was made to broaden the general call for technology development and put these efforts together under a single umbrella (National Human Genome Research Institute Institute 2018e).
In large-scale genomics projects, there is always a tension between maintaining use of existing technologies in data production throughout the course of a phase of the project versus incorporating significant technological improvements to existing methods or implementing completely new technologies that were thought to have significant benefits. Generally speaking, NHGRI was supportive of data producers implementing technical improvements to existing technologies, but it was often not feasible to switch to completely new methods during any given phase of ENCODE. However, technological advances achieved by ENCODE Technology Development grantees or by the broader research community have been incorporated into each new phase of ENCODE.
Several technologies developed with ENCODE support were sufficiently robust to be implemented in subsequent data production activities. For example, a method for whole genome bisulfite sequencing developed by Dr. Joseph Ecker’s group at the Salk Institute was used in data production of DNA methylation data in ENCODE 3. Similarly, a research team led by Dr. Jay Shendure at the University of Washington and Dr. Nadav Ahitav at the University of California, San Francisco were supported to develop a massively parallel reporter assay that was then implemented in an FCC under ENCODE 4.
Perhaps the most profound technological advance that took place near the end of the pilot project that completely changed the face of ENCODE (and well beyond) was the availability of “next generation” sequencing platforms that allowed for massively parallel sequencing, especially that provided by Illumina (see chapter 6). By 2007, shortly before the funding of ENCODE 2, sequencing machines were commercially available for researchers to implement readouts of ENCODE-style assays at an unprecedented scale and at a greatly increased specificity than the microarray-based platforms used during the pilot. Nearly all assays used in ENCODE since the pilot have employed these short-read sequencing platforms, which initially involved the successful development and adaptation of methods by ENCODE investigators. The more recently available long-read sequencing platforms from, for example, PacBio,12 were also incorporated into specific applications where the tradeoff between read length and data quality/cost was worthwhile, for example, sequencing of RNA transcripts to identify different isoforms.
CRISPR-Cas9 technologies that can introduce pre-determined and precise changes in DNA sequence made it possible to perturb sequences within candidate functional elements at high resolution and throughput (Catarino and Stark 2018). This advance supported the concept of establishing FCCs in ENCODE 4, in which a number of these centers used CRISPR-Cas9-based genome editing approaches.
2.1.3. Importance of Working on Common Targets/Samples/Reagents
The utility of having common sets of reagents such as cell lines and antibodies to facilitate comparison between different technological platforms and approaches was also identified as a key feature of Consortium interactions from the beginning, a practice that continued through ENCODE 4. In addition to serving as a means of evaluating methods and data quality, it has the tremendous advantage of enabling integration of similar as well as different data types, strengthening the ability of ENCODE to richly annotate the human and mouse genomes (see above).
2.1.4. Data Management
Overall, data management is a critically important, time-consuming and expensive component of any large genomics project, including ENCODE. Meeting the requirements of data management and analysis that the diverse data types generated by ENCODE were considered and planned for from the start of the project. The Marker Paper described the need to establish an ENCODE Portal that would serve to index the data and allow for user queries of both data and metadata. From 2003 to 2012, the primary location for community access to sequence-based ENCODE data was the UCSC Genome Browser13 with data also available at GEO.14 As the needs of ENCODE evolved, a portal and platform that could flexibly handle the volume and diversity of data was established in 2012 at the ENCODE Data Coordination Center at Stanford University, while continuing to make data available at GEO.15 This portal now hosts all ENCODE data as well as data from related projects, including modENCODE, mouse ENCODE, and RoadMap Epigenomics. It also hosts data analysis products that provide ground level as well as integrative level annotations that all contribute to the compiled encyclopedia (see above). As of December 28, 2021, 19,803 experiments were publicly available on the ENCODE portal (ENCODE Consortium 2000) with more than 14,000 generated by ENCODE. The portal also serves as a centralized source of information on ENCODE for both Consortium members and outside users, including access to publications using ENCODE data (see below), assay standards, file formats, software tools, uniform data processing pipelines, and the data use policy.
2.1.5. Data Standards
The need to develop data standards for all data types was also identified in the pilot project and continued to be a major effort as assays were improved or new assays were incorporated into ENCODE. The availability of data standards, quality metrics, experimental guidelines, and data processing pipelines ensured that high (and known) quality data were available (ENCODE Consortium 2019b). In addition, capturing detailed and well-curated metadata in a searchable form was critical to making data readily accessible to the research community. The development and use of common ontologies and controlled vocabulary both within ENCODE and shared between other related projects, such as IHEC, has enhanced reproducibility, transparency, and interoperability with other projects (ENCODE Consortium 2019c). Of the many advances in data storage and management, one key development that ENCODE has taken advantage of is the advent of cloud computing. The ENCODE DCC stores ENCODE data and metadata on the cloud through Amazon Web Services. It also makes its uniform data processing pipelines available for community use through DNAnexus (ENCODE Consortium 2019d).
2.1.6. Data Release and Accessibility
From the start, ENCODE was designated as a “community resource project” as defined in a report on genomics data sharing that came out of a meeting held in Fort Lauderdale in 2003 (Wellcome Trust 2003) (see chapter 2). The ENCODE pilot project data release policy (ENCODE Consortium 2007a) was based on the expectations laid out in that report, specifically (1) data producers will make data publicly available as soon as they are verified to be of high quality/reproducible; (2) data users should cite the source of the data by citing the Marker Paper and acknowledging the ENCODE Consortium and the data producers; and (3) data users should voluntarily afford the data producers the first opportunity to publish on the data and analyses of these data. Subsequently, ENCODE 2 and modENCODE established a joint data release policy (ENCODE Consortium 2009). The main change in this policy was to elaborate point # 3 above in more detail. It strived to be more specific about the expectations for use and publication by both data producers and data users. Data producers were encouraged to publish on their own data in a timely manner (i.e., submission of manuscript for publication within nine months of data release). Within this nine-month time frame, data users could use the data in their own analyses but were asked not to publish or publicly present on unpublished ENCODE data. After the nine-month publication moratorium (or earlier if the data had been published), data users were free to publish or otherwise report on their use of ENCODE data but were asked to acknowledge the Consortium and data producers.
In ENCODE 3 and 4, the decision was made to remove the nine-month publication moratorium, and data users are free to use data as they are released (ENCODE Consortium 2014). As before, data users were asked to reference the datasets used and acknowledge the Consortium and data producers.
2.1.7. Unrestricted Data Access
A fundamental tenet of ENCODE has been to create as broadly accessible a resource as possible. Because ENCODE serves a wide range of users, this means availability of even raw, unprocessed data in unrestricted (open access) databases such as GEO since requesting and gaining access to controlled access data in databases, for example, dbGaP, can take considerable time and effort. As ENCODE moved beyond using long-established cell lines for which a substantial amount of genomic data was already in public databases, it transitioned to using new primary cell and tissue biosamples that were explicitly consented for unrestricted data sharing (thereby allowing for general research use). The first big step in this direction came during ENCODE 3 through a collaboration with the GTEx project, named “EN-Tex,” in which approximately thirty tissue samples from each of four donors were obtained using the GTEx consent form onto which an addendum had been added that explicitly allowed for unrestricted data sharing (National Disease Research Interchange 2014). In ENCODE 4, the expectation was that all biosamples used would have been obtained using consent explicitly allowing for unrestricted data sharing. (Exceptions were allowed only in limited instances for the use of some historically used cell lines, e.g., for assays not previously used in ENCODE.) This expectation preceded the new NHGRI requirement for explicit consent for future research uses and broad data sharing that started in early 2020 (Ganguly 2019).
2.1.8. Open-Source Software
In addition to providing freely accessible data, starting with the pilot project the expectation was set forth in the RFAs for open-source software sharing. The most recent, and detailed, plan for software and analysis release was established in ENCODE 4 (ENCODE Consortium 2019d).
2.1.9. Publications and Community Outreach
Another feature of ENCODE through the years has been the publication of Consortium-wide papers as well as papers by individual participants. In keeping with the goal of serving as a community resource, the Consortium-wide papers have largely been published under Open Access agreements.
All ENCODE participants were free to publish on their own work when ready for publication, and most papers from ENCODE participants were published in this manner. However, there was always an inherent tension between meeting the needs of individual researchers to publish on their own timeline and the needs of the overall Consortium to publish as a group. While individual opinions within the Consortium have differed, the general sentiment, particularly early on, was that there was significant value in publishing Consortium-wide papers and other related papers from individuals or small groups in paper “packages,” as they served to raise the visibility of the project by notifying the broader research community about the content and utility of the resource. However, from my viewpoint, putting together Consortium-wide and associated papers came at a significant cost as well. The papers, particularly those that were Consortium-wide, always took significantly longer to write (one to two-plus years) than anticipated. Reaching consensus about topics for which there is a wide range of scientific opinion can be very challenging and time-consuming, as can coordinating analyses and companion papers to avoid unnecessary duplication and overlap. Inevitable delays often resulted in some papers being held up for long periods of time, which had the potential to negatively impact some Consortium members, especially those more junior in their careers. Further, with a project running as long as ENCODE has, with participation of some of the same researchers in multiple phases, inevitably led to some stressed relationships and publication preparation “fatigue” that have amplified these challenges. Nonetheless, publication by individuals, small groups, and the Consortium as a whole has been an important aspect of outreach to the research community as well as well-deserved recognition for work completed.
As further outreach to the community, ENCODE hosted tutorials and Users Meetings to inform the community about the available resources and how they can be used to advance the users’ own research.
2.2. Impact of ENCODE
One way to measure the impact of ENCODE has been by tracking publications using ENCODE data, tools, and other resources. Program staff tracked publications by members of the ENCODE, mouse ENCODE, and modENCODE programs, which is relatively straightforward by tracking publications that cite ENCODE grant numbers as the funding source. Less straightforward has been tracking the use of ENCODE resources by the research community because of several factors, including ENCODE’s non-unique name, lack of citation of ENCODE DCC accession numbers, and access to ENCODE data via other routes, such as GEO or the UCSC browser, where it may not even be transparent to the user that the data was generated by ENCODE. With these caveats, NHGRI has identified and collated more than 1,000 publications from NHGRI-supported ENCODE awardees and more than 2,200 additional publications from the research community (National Human Genome Research Institute 2022) (see Figure 4.2a). These community publications have been further classified as being focused on basic biology, tool development, or human disease; approximately one thousand of these were categorized as focused on human disease research, attesting to the translational value of the resource. While the potential for ENCODE to enable discoveries relevant to human disease and medicine was a motivating factor for the project, I did not fully anticipate the degree and speed to which the resource would be applied to disease research and therapeutics.
A striking and personally exciting therapeutic application of the ENCODE resource has been its contribution to identifying genomic targets for gene therapy for both sickle cell disease (SCD) and beta-thalassemia, which are both caused by mutations impacting the expression of the beta-globin gene. The beta-globin gene is expressed after birth, while the two related gamma globin genes are expressed in fetal development. The gamma and beta globin proteins that result from the translation of these genes each form heterodimers with the alpha globin subunit to make fetal (HbF) and adult (HbA) hemoglobin, respectively. It has been long known that some people, for genetic reasons, have higher levels of HbF persisting in adulthood and that those individuals who also have beta-thalassemia or SCD fair better clinically. For decades, there has been keen interest in understanding how to prevent gamma globin gene expression from turning off (or learning how to turn it back on) with the hope of ameliorating the symptoms of these diseases (Nienhuis and Stamatoyannopoulos 1978; Vinjamur et al. 2018). In fact, I have been fascinated with this question since the early 1980s when I first pursued understanding beta and gamma globin gene expression by the detailed study of the DNA of an individual with Hereditary Persistence of Fetal Hemoglobin as a graduate student (Forget et al. 1983; Feingold and Forget 1989).
Recent advances have been made in understanding how the gamma-globin genes are turned off after birth. GWAS studies showed that SNPs associated with an increase in HbF mapped within an intronic region of the gene coding for the transcription factor BCL11A (Menzel et al. 2007; Uda et al. 2008), which represses gamma globin gene expression (Sankaran et al. 2008). An important advance was the identification of an erythroid-specific enhancer within this region with the aid of ENCODE data (Bauer et al. 2013). As BCL11A is expressed in erythroid and immune stem cells as well as the brain, finding an erythroid-specific enhancer was key to understanding how its expression is regulated specifically in erythroid cells and immediately was suggested as a target for gene therapy to increase fetal hemoglobin expression (Bauer et al. 2013; Hardison and Blobel 2013; Vinjamur et al. 2018). Of particular note is that this is the first instance, to my knowledge, of a gene therapy target residing in a non-coding region. Preliminary results as of 2021 are promising, so far showing that HbF expression is increased and there is a reduction in disease symptoms with recent encouraging news that these benefits are persisting and will likely be sustained (Dunbar 2021; Eisenstein 2021; Esrick et al. 2021; Frangoul et al. 2021; Stein 2021). Gene therapy targeting this erythroid-specific enhancer is one of several gene therapy approaches for hemoglobinopathies (Orkin and Bauer 2019).
This tremendous advance only underscores the power of integrating multiple types of data and the value of a detailed, tissue-specific view of gene regulation. This example of ENCODE data playing a key role in enabling clinical trials to treat these hemoglobinopathies fulfills an early motivation of ENCODE to stimulate the development of new therapies and, at the same time, is a fitting testament to my scientific career coming full circle.
2.3. Completing the Catalog
It was recognized from the start that a significant challenge in identifying all functional elements is that transcripts and other functional elements often are expressed or are biologically active in specific cell types, times in development, or under specific environmental conditions. This is in sharp contrast to the human genome, which is nearly invariant from cell to cell (except for lymphocytes and somatic variation). As the Marker Paper described, a truly comprehensive compendium of all functional elements would require the study of every cell type under all stages of development. The strategy for addressing this challenge was unclear when ENCODE started and developing this path was a major motivation for the pilot project. Even knowing these challenges from the start, the longevity of ENCODE surprised me, and, to my knowledge, it is one of the longest-running genomics consortia.
While early studies focused on immortalized cell lines, advances in technologies have enabled the study of primary cells, tissues and explants, single cells, and induced Pluripotent Stem cells differentiated into a range of cell types and organoids. Single-cell methods to identify transcripts and open chromatin were employed in ENCODE 4 to help dissect out tissue heterogeneity. Determining the most informative biosamples, however, remained a major research activity even in ENCODE 4.
As the project progressed and science and technologies advanced over the years, defining the end goal for creating a truly “complete” catalog became more of a practical exercise rather than a completion end point because of the seemingly unbounded set of biological conditions under which the genome functions, not to mention the impact of genetic variation on function. As such, ENCODE focused on making its resources as useful for the research community as possible. ENCODE data and analyses served as a catalyst for hypothesis generation that directs further studies, often focused in specific biological contexts. Beyond the data, resources such as data standards, computational tools, and data processing pipelines have been used by biomedical community users for their own studies.
One of the inherent tensions for ENCODE was to create a resource that serves a wide range of needs, from “power” users who are highly sophisticated computationally and want to use the basic data for their own analyses to the investigators who want ENCODE to provide a list of all possible functional elements to use to help in further studies of their own biological question and/or genomic regions of interest. As it has been and continues to be an active area of research to devise the best ways to call functional elements (see e.g., Encyclopedia discussion above), ENCODE has endeavored to provide as much information as possible about what information was used to make the calls and what are the limitations.
2.4. Strengths and Limitations of Consortium Research
Managing ENCODE as a highly collaborative research consortium had afforded many advantages, including coordinated data generation at high throughput to take advantage of economies of scale, uniform data processing, standard data formats, metadata, and centralized data access. In addition, having a diversity of expertise and experiences to bear on the enormous scientific challenge enhanced the quality of the product. Nonetheless, running this large and complex project with hundreds of investigators carried with it challenges, most notably balancing the needs of individual researchers with those of the Consortium. Dr. Ewan Birney, PI of ENCODE 2’s Data Analysis Center and lead of the integrated analyses and 2012 paper package, then joint director of the European Bioinformatics Institute, offered his observations about lessons learned from that experience (Birney 2012). Some key observations he made included the importance of participants working together on a common goal of creating a useful community resource, the need for a clear organizational structure and code of conduct, the value of having a diversity of individuals bringing in different ideas, and ensuring that the resource is of high value.
NHGRI has strived to learn from experience to create and evolve a framework by which ENCODE succeeds, for example, the use of Cooperative Agreements and their associated requirements, data sharing policies, evaluating progress toward milestones, outside advisors, and even peer pressure. By design and necessity, ENCODE has been a directed, top-down, managed project, but there are limitations to how much we can (or should) control this large and complex enterprise. Data production groups, the DCC, and the DAC were more heavily managed than technology development and computational analysis research projects. While following NIH rules, successful project management often required approaches and skills imparted to us by our wise leaders Mark Guyer and Francis Collins based on their years of experience managing the HGP. Maintaining long-term focus on a common set of goals at times can be as arduous as herding cats (an apt but overused cliché). Coordination of groups working with disparate technologies and data types at a time when technologies were rapidly changing required a collaborative approach and trust between awardees and NHGRI staff, for example, to decide when technological or strategic changes were worthwhile with respect to the scientific gain vs. the cost (in terms of time, disruptions, funds). Further, relationships between individuals that form over years (both within and outside ENCODE) as well as different communication styles, can influence—for better or worse—group dynamics.
2.4.1. Looking Forward
While the long-term impact of ENCODE won’t be known for years, I anticipate that it will continue to be a valuable resource for the biomedical community for the foreseeable future. Additional data, especially functional characterization, as well as new and improved analyses from ENCODE 4 will enhance the breadth and quality of genome annotations. The community will continue to use the data in novel ways to bring new insights into the genome including understanding regulation of gene expression, human health, and disease. Further, tremendous opportunities exist for understanding the influence of genetic variation on phenotype.
Notes
Many thanks go to my NHGRI colleagues who have been involved in the management of ENCODE over the years. I thank Francis Collins and Mark Guyer for their tremendous leadership and guidance, especially in the early years of the ENCODE Project and Eric Green for his important contributions to the early planning and execution of the Pilot Project as well as his sustained support of ENCODE as NHGRI director. I have had the great fortune to have worked with outstanding fellow program directors in ENCODE. I am especially grateful to have had the opportunity to work closely with Peter Good, my co-pilot from the early planning days of ENCODE through ENCODE 2, an invaluable contributor to the success of the project who had the knack of homing in on areas that needed attention. I also thank Mike Pazin for his phenomenal scientific knowledge along with keen project management skills that emphasized fairness across all parts of the Consortium. Additional thanks go to Daniel Gilchrist and Stephanie Morris for their contributions to the management of ENCODE, providing subject-matter expertise and steadying influences. Additional thanks go to Kris Wetterstrand, who made invaluable contributions to ENCODE in the early days of the Consortium, aided by her experience as a part of the Human Genome Project management team. Many of NHGRI’s Program Analysts were also indispensable members of the ENCODE Project Management team over the years, contributing outstanding organizational, technical, and interpersonal skills along with scientific knowledge. These include Leslie Adams, Omar Al Jammal, Eileen Cahill, Stephanie Calluori, Shaila Chhibba, Julie Coursen, Laura Liefer Dillon, Belinda Jackson, Caroline Kelly, Rebecca Lowdon, Jessica Melone, Samuel Moore, Preetha Nandi, Hannah Naughton, Brianna Nunez, Sandra Kamholz Oza, Michael Pagan, Ella Samer, Yetkaterina Vaydylevich, Judith Wexler, Julia Zhang, and Sherrie Zhou. I am grateful to Richard Myers, Mike Pazin, Michael Snyder, and John Stamatoyannopoulos for their careful reading of the manuscript and their helpful comments. Finally, I am indebted to the hundreds of ENCODE researchers whose hard work and dedication to ENCODE led to its many successes.
1. “deCODE” is the name of an Icelandic-based genetic project. See www.deCODE.com.
2. For example, PsychENCODE, https://www.nimhgenetics.org/resources/psychencode; ZENCODE, http://zebrafish.igib.in/zebrafish-encyclopedia-of-dna-elements-encode; and even fruitENCODE, http://www.epigenome.cuhk.edu.hk/encode.html.
3. RoadMap Epigenomics Project. See https://commonfund.nih.gov/epigenomics.
4. Visualization tool “SCREEN.” See https://screen.encodeproject.org/.
5. I recall it was Richard Lifton MD, PhD, then professor and chair of Human Genetics at Yale University, now Carson Family Professor and president of The Rockefeller University.
6. American Recovery and Reinvestment Act. See https://recovery.nih.gov/.
8. Genotype-Tissue Expression, see https://commonfund.nih.gov/gtex.
9. 4D Nucleome Programs, see https://commonfund.nih.gov/4dnucleome.
10. International Human Epigenome Consortium, see http://ihec-epigenomes.org/welcome/.
11. ENCODE Portal, see https://www.encodeproject.org.
12. PacBio, https://www.pacb.com/.
13. UCSC Genome Browser, http://genome.ucsc.edu/ENCODE.
15. ENCODE portal, see https://encodeproject.org.
References
- Allen, M. A., L. W. Hillier, R. H. Waterston, and T. Blumenthal. 2011. “A Global Analysis of C. elegans Trans-Splicing.” Genome Research 21 (2): 255–64.
- Annual Reviews. 2013. “CRISPR: Editing the Genome.” Retrieved January 7, 2020. https://www.annualreviews.org/page/crispr.
- Araya, C. L., T. Kawli, A. Kundaje, et al. 2014. “Regulatory Analysis of the C. elegans Genome with Spatiotemporal Resolution.” Nature 512 (7515): 400–405.
- Bauer, D. E., S. C. Kamran, S. Lessard, et al. 2013. “An Erythroid Enhancer of BCL11A Subject to Genetic Variation Determines Fetal Hemoglobin Level.” Science 342 (6155): 253–57.
- Berezikov, E., N. Robine, A. Samsonova, et al. 2011. “Deep Annotation of Drosophila melanogaster microRNAs Yields Insights into Their Processing, Modification, and Emergence.” Genome Research 21 (2): 203–15.
- Birney, E. 2012. “The Making of ENCODE: Lessons for Big-Data Projects.” Nature 489 (7414): 49–51.
- Birney, E., J. A. Stamatoyannopoulos, A. Dutta, et al. 2007. “Identification and Analysis of Functional Elements in 1% of the Human Genome by the ENCODE Pilot Project.” Nature 447 (7146): 799–816.
- Boyle, A. P., C. L. Araya, C. Brdlik, et al. 2014. “Comparative Analysis of Regulatory Information and Circuits Across Distant Species.” Nature 512 (7515): 453–56.
- Brooks, A. N., L. Yang, M. O. Duff, et al. 2011. “Conservation of an RNA Regulatory Map Between Drosophila and Mammals.” Genome Research 21 (2): 193–202.
- Brown, J. B., N. Boley, R. Eisman, et al. 2014. “Diversity and Dynamics of the Drosophila Transcriptome.” Nature 512 (7515): 393–99.
- Brunet, T. D., and W. F. Doolittle. 2014. “Getting ‘Function’ Right.” Proceedings of the National Academy of Sciences USA 111 (33): E3365.
- Catarino, R. R., and A. Stark. 2018. “Assessing Sufficiency and Necessity of Enhancer Activities for Gene Expression and the Mechanisms of Transcription Activation.” Genes & Development 32 (3–4): 202–23.
- Celniker, S. E., L. A. Dillon, M. B. Gerstein, et al. 2009. “Unlocking the Secrets of the Genome.” Nature 459 (7249): 927–30.
- Cherbas, L., A. Willingham, D. Zhang, et al. 2011. “The Transcriptional Diversity of 25 Drosophila Cell Lines.” Genome Research 21 (2): 301–14.
- Chung, W. J., P. Agius, J. O. Westholm, et al. 2011. “Computational and Experimental Identification of Mirtrons in Drosophila melanogaster and Caenorhabditis elegans.” Genome Research 21 (2): 286–300.
- Collins, F. S., E. D. Green, A. E. Guttmacher, and M. S. Guyer. 2003. “A Vision for the Future of Genomics Research.” Nature 422 (6934): 835–47.
- Doolittle, W. F. 2013. “Is Junk DNA Bunk? A Critique of ENCODE.” Proceedings of the National Academy of Sciences USA 110 (14): 5294–300.
- Dunbar, C. E. 2021. “A Plethora of Gene Therapies for Hemoglobinopathies.” Nature Medicine 27 (2): 202–4.
- Eaton, M. L., J. A. Prinz, H. K. MacAlpine, G. Tretyakov, P. V. Kharchenko, and D. M. MacAlpine. 2011. “Chromatin Signatures of the Drosophila Replication Program.” Genome Research 21 (2): 164–74.
- Eddy, S. R. 2013. “The ENCODE Project: Missteps Overshadowing a Success.” Current Biology 23 (7): R259–61.
- Eisenstein, M. (2021). “Gene Therapies Close in on a Cure for Sickle-Cell Disease.” Nature 596 (7873): 2–4.
- ENCODE Consortium. 2000. “ENCODE Portal.” Retrieved January 6, 2020. https://www.encodeproject.org/matrix/?type=Experiment&status=released.
- ENCODE Consortium. 2004. “The ENCODE (ENCyclopedia Of DNA Elements) Project.” Science 306 (5696): 636–40.
- ENCODE Consortium. 2007a. “ENCODE Project Data Release Policy (2003–2007).” Retrieved December 29, 2019. https://www.genome.gov/12513440/encode-project-data-release-policy-20032007.
- ENCODE Consortium. 2007b. “Genome Research ENCODE Special Issue.” Genome Research 17 (6): 667–964.
- ENCODE Consortium. 2009. “ENCODE Consortia Data Release, Data Use, and Publication Policies.” Retrieved December 29, 2019. https://www.genome.gov/Pages/Research/ENCODE/Mod-ENCODE_Consortia_Data_Release_Policy_revised_11-22-09.pdf.
- ENCODE Consortium. 2012a. “ENCODE Collection.” Retrieved January 10, 2020. https://www.nature.com/collections/aghcdefffg/.
- ENCODE Consortium. 2012b. “An Integrated Encyclopedia of DNA Elements in the Human Genome.” Nature 489 (7414): 57–74.
- ENCODE Consortium. 2014. “ENCODE Data Use Policy for External Users.” Retrieved December 29, 2019. https://www.genome.gov/Pages/Research/ENCODE/ENCODE_Data_Use_Policy_for_External_Users_03-07-14.pdf.
- ENCODE Consortium. 2019a. “ENCODE Data Organization.” https://www.encodeproject.org/help/data-organization/.
- ENCODE Consortium. 2019b. “ENCODE Data Processing Pipelines.” https://www.encodeproject.org/pipelines/.
- ENCODE Consortium. 2019c. “ENCODE Data Standards.” https://www.encodeproject.org/data-standards/.
- ENCODE Consortium. 2019d. “ENCODE Data Use, Software, and Analysis Release Policies.” Retrieved December 29, 2019. https://www.encodeproject.org/about/data-use-policy/.
- Ercan, S., Y. Lubling, E. Segal, and J. D. Lieb. 2011. “High Nucleosome Occupancy Is Encoded at X-linked Gene Promoters in C. elegans.” Genome Research 21 (2): 237–44.
- Esrick, E. B., L. E. Lehmann, A. Biffi, et al. 2021. “Post-Transcriptional Genetic Silencing of BCL11A to Treat Sickle Cell Disease.” New England Journal of Medicine 384 (3): 205–15.
- Feingold, E. A., and B. G. Forget. 1989. “The Breakpoint of a Large Deletion Causing Hereditary Persistence of Fetal Hemoglobin Occurs Within an Erythroid DNA Domain Remote from the Beta-Globin Gene Cluster.” Blood 74 (6): 2178–86.
- Forget, B. G., D. Tuan, M. V. Newman, et al. 1983. “Molecular Studies of Mutations That Increase Hb F Production in Man.” Progress in Clinical and Biological Research 134: 65–76.
- Frangoul, H., D. Altshuler, M. D. Cappellini, et al. 2021. “CRISPR-Cas9 Gene Editing for Sickle Cell Disease and β-Thalassemia.” New England Journal of Medicine 384 (3): 252–60.
- Ganguly, P. 2019. “NHGRI to Require Explicit Consent for Data Sharing in Genomics Research.” Retrieved January 6, 2020. https://www.genome.gov/news/news-release/NHGRI-to-request-explicit-consent-for-genomic-research-data-sharing.
- Gerhard, D. S., L. Wagner, E. A. Feingold, et al. 2004. “The Status, Quality, and Expansion of the NIH Full-Length cDNA Project: The Mammalian Gene Collection (MGC).” Genome Research 14 (10b): 2121–27.
- Gerstein, M. B., Z. J. Lu, E. L. Van Nostrand, et al. 2010. “Integrative Analysis of the Caenorhabditis elegans Genome by the modENCODE Project.” Science 330 (6012): 1775–87.
- Gerstein, M. B., J. Rozowsky, K.-K. Yan, et al. 2014. “Comparative Analysis of the Transcriptome Across Distant Species.” Nature 512 (7515): 445–48.
- Graur, D., Y. Zheng, N. Price, R. B. Azevedo, R. A. Zufall, and E. Elhaik. 2013. “On the Immortality of Television Sets: ‘Function’ in the Human Genome According to the Evolution-Free Gospel of ENCODE.” Genome Biology and Evolution 5 (3): 578–90.
- Graveley, B. R., A. N. Brooks, J. W. Carlson, et al. 2011. “The Developmental Transcriptome of Drosophila melanogaster.” Nature 471 (7339): 473–79.
- Green, E. D., C. Gunter, L. G. Biesecker, et al. 2020. “Strategic Vision for Improving Human Health at the Forefront of Genomics.” Nature 586 (7831): 683–92.
- Hardison, R. C., and G. A. Blobel. 2013. “Genetics. GWAS to Therapy by Genome Edits?” Science 342 (6155): 206–7.
- Ho, J. W. K., Y. L. Jung, T. Liu, et al. 2014. “Comparative Analysis of Metazoan Chromatin Organization.” Nature 512 (7515): 449–52.
- Hoskins, R. A., J. M. Landolin, J. B. Brown, et al. 2011. “Genome-wide Analysis of Promoter Architecture in Drosophila melanogaster.” Genome Research 21 (2): 182–92.
- Ikegami, K., T. A. Egelhofer, S. Strome, and J. D. Lieb. 2010. “Caenorhabditis elegans Chromosome Arms Are Anchored to the Nuclear Membrane via Discontinuous Association with LEM-2.” Genome Biology 11 (12): R120–R120, 1–20.
- International HapMap Consortium. 2003. “The International HapMap Project.” Nature 426 (6968): 789–96.
- Johnson, S. 2011. “modENCODE: Revealing the Inner Workings of the Genome.” Retrieved January 1, 2020. https://recovery.nih.gov/Stories/ViewStory.aspx?id=455.
- Kellis, M., B. Wold, M. P. Snyder, et al. 2014a. “Defining Functional DNA Elements in the Human Genome.” Proceedings of the National Academy of Sciences USA 111 (17): 6131–38.
- Kellis, M., B. Wold, M. P. Snyder, et al. 2014b. “Reply to Brunet and Doolittle: Both Selected Effect and Causal Role Elements Can Influence Human Biology and Disease.” Proceedings of the National Academy of Sciences USA 111 (33): E3366.
- Kharchenko, P. V., A. A. Alekseyenko, Y. B. Schwartz, et al. 2011. “Comprehensive Analysis of the Chromatin Landscape in Drosophila melanogaster.” Nature 471 (7339): 480–85.
- Lander, E. S., L. M. Linton, B. Birren, et al. 2001. “Initial Sequencing and Analysis of the Human Genome.” Nature 409 (6822): 860–921.
- Leder, P., B. F. Clark, W. S. Sly, S. Pestka, and M. W. Nirenberg. 1963. “Cell-Free Peptide Synthesis Dependent upon Synthetic Oligodeoxynucleotide.” Proceedings of the National Academy of Sciences USA 50: 1135–43.
- Ledford, H. 2010. “Genome ‘Census’ Reveals Hidden Riches.” Nature.
- Liang, F., I. Holt, G. Pertea, S. Karamycheva, S. L. Salzberg, and J. Quackenbush. 2000. “Gene Index Analysis of the Human Genome Estimates Approximately 120,000 Genes.” Nature Genetics 25 (2): 239–40.
- Liu, T., A. Rechtsteiner, T. A. Egelhofer, et al. 2011. “Broad Chromosomal Domains of Histone Modification Patterns in C. elegans.” Genome Research 21 (2): 227–36.
- Lu, Z. J., K. Y. Yip, G. Wang, et al. 2011. “Prediction and Characterization of Noncoding RNAs in C. elegans by Integrating Conservation, Secondary Structure, and High-Throughput Sequencing and Array Data.” Genome Research 21 (2): 276–85.
- Menzel, S., C. Garner, I. Gut, et al. 2007. “A QTL Influencing F Cell Production Maps to a Gene Encoding a Zinc-Finger Protein on Chromosome 2p15.” Nature Genetics 39 (10): 1197–99.
- modENCODE Consortium, S. Roy, J. Ernst, et al. 2010. “Identification of Functional Elements and Regulatory Circuits by Drosophila modENCODE.” Science 330 (6012): 1787–97.
- modENCODE Consortium. 2011. “Genome Research February 2011.” https://genome.cshlp.org/content/21/2.toc.
- Moore, J. E., M. J. Purcaro, H. E. Pratt, et al. 2020. “Expanded Encyclopaedias of DNA Elements in the Human and Mouse Genomes.” Nature 583 (7818): 699–710.
- mouseENCODE, J. A. Stamatoyannopoulos, M. Snyder, et al. 2012. “An Encyclopedia of Mouse DNA Elements (Mouse ENCODE).” Genome Biology 13 (8): 418, 1–5.
- mouseENCODE Consortium. 2014. “Mouse ENCODE Special Issue.” Nature. Retrieved January 16, 2020. https://www.nature.com/collections/yhwmkwknsh.
- National Disease Research Interchange. 2014. “Authorization for Tissue Donation in National Institutes of Health Research Project.” Retrieved December 29, 2019. https://www.genome.gov/Pages/Research/ENCODE/GTEx_Consent_ENCODE_addendum_10-9-14.pdf.
- National Human Genome Research Institute. 2002a. “The Comprehensive Extraction of Biological Information from Genomic Sequence.” https://www.genome.gov/10005568/genomic-sequence-workshop-summary.
- National Human Genome Research Institute. 2002b. “Summary of National Advisory Council for Human Genome Research September 2002 Meeting.” https://www.genome.gov/11007802/september-2002-nachgr-meeting-summary.
- National Human Genome Research Institute. 2003a. “Determination of All Functional Elements in Human DNA Request for Applications (HG-03–003).” https://grants.nih.gov/grants/guide/rfa-files/rfa-hg-03-003.html.
- National Human Genome Research Institute. 2003b. “ENCODE Pilot Project Launch Meeting.” Retrieved January 3, 2020. https://www.genome.gov/10506205/encode-project-launch-meeting.
- National Human Genome Research Institute. 2003c. “Technologies to Find Functional Elements in Genome DNA Request for Applications (HG-03–004).” Retrieved January 3, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-03-004.html.
- National Human Genome Research Institute. 2006a. “Creating the Encyclopedia of DNA Elements (ENCODE) in the Human Genome Request for Applications (HG-07–030).” Retrieved January 2, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-07-030.html.
- National Human Genome Research Institute. 2006b. “A Data Coordination Center for the Encyclopedia of DNA Elements (ENCODE) Project Request for Applications (RFA-HG-07–031).” Retrieved January 2, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-07-031.html.
- National Human Genome Research Institute. 2006c. “A Data Coordination Center for the Model Organism ENCODE Project (modENCODE) Request for Applications (RFA-HG-06–007).” Retrieved December 26, 2019. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-06-007.html.
- National Human Genome Research Institute. 2006d. “ENCODE Phase 2 Concept Clearance.” Retrieved December 26, 2019. https://www.genome.gov/sites/default/files/genome-old/pages/about/nachgr/may2006nachgragenda/scalingrfaclearance.pdf.
- National Human Genome Research Institute. 2006e. “Identification of All Functional Elements in Selected Model Organism Genomes Request for Applications (RFA-HG-06–006).” Retrieved December 26, 2019. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-06-006.html.
- National Human Genome Research Institute. 2006f. “Summary of National Advisory Council for Human Genome Research May 2006 Meeting.” Retrieved December 26, 2019. https://www.genome.gov/19518513/may-2006-nachgr-meeting-summary.
- National Human Genome Research Institute. 2007a. “National Advisory Council for Human Genome Research February 12, 2007 Concept Clearance. A Data Analysis Center for the Encyclopedia of DNA Elements (ENCODE) Project.” Retrieved December 26, 2019. https://www.genome.gov/sites/default/files/genome-old/pages/About/NACHGR/February2007NACHGRAgenda/ConceptClearanceENCODEDataAnalysisCenter.doc.
- National Human Genome Research Institute. 2007b. “Summary of National Advisory Council for Human Genome Research February 2007 Meeting.” Retrieved December 26, 2019. https://www.genome.gov/Pages/About/NACHGR/NACHGRMeetingSummaries/NACHGRFeb07CouncilMinutes.pdf.
- National Human Genome Research Institute. 2011a. “Computational Analysis of the Encyclopedia of DNA Elements (ENCODE) Data Request for Applications (RFA-HG-11–025).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-11-025.html.
- National Human Genome Research Institute. 2011b. “Data Analysis and Coordination Center for the Encyclopedia of DNA Elements (ENCODE) Request for Applications (RFA-HG-11–026).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-11-026.html.
- National Human Genome Research Institute. 2011c. “Expanding the Encyclopedia of DNA Elements (ENCODE) in the Human and Model Organisms Request for Applications (RFA-HG-11–024).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-11-024.html.
- National Human Genome Research Institute. 2011d. “Future of ENCODE.” https://www.genome.gov/sites/default/files/genome-old/pages/about/nachgr/may2011nachgragenda/encodeconceptclearancemay_2011_council.pdf.
- National Human Genome Research Institute. 2011e. “Technology Development for High-Throughput Functional Genomics (R01) (RFA-HG-11–013).” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-11-013.html.
- National Human Genome Research Institute. 2011f. “Technology Development for High-Throughput Functional Genomics (R21) (RFA-HG-11–014).” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-11-014.html.
- National Human Genome Research Institute. 2011g. “Technology Development for High-Throughput Functional Genomics (R43/44) (RFA-HG-11–015).” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-11-015.html.
- National Human Genome Research Institute. 2012a. “ENCODE Pilot Project: Description.” Retrieved December 31, 2019. https://www.genome.gov/Funded-Programs-Projects/ENCODE-Project-ENCyclopedia-Of-DNA-Elements/pilot.
- National Human Genome Research Institute. 2012b. “ENCODE Pilot Project: Participants and Projects.” Retrieved December 31, 2019. https://www.genome.gov/Pages/Research/ENCODE/Pilot_Participants_Projects.pdf.
- National Human Genome Research Institute. 2012c. “Long-Range Planning: Reports and Publications.” https://www.genome.gov/12010624/longrange-planning-reports-and-publications.
- National Human Genome Research Institute. 2014. “The modENCODE Project: Model Organism ENCyclopedia Of DNA Elements (modENCODE).” Retrieved December 26, 2019. https://www.genome.gov/26524507/the-modencode-project-model-organism-encyclopedia-of-dna-elements-modencode.
- National Human Genome Research Institute. 2015a. “Concept Clearances for Functional Genomics.” https://www.genome.gov/Pages/About/NACHGR/May2015AgendaDocuments/Concept_Clearance_Functional_Genomics_Initiatives_5-8-15.pdf.
- National Human Genome Research Institute. 2015b. “From Genome Function to Biomedical Insight: ENCODE and Beyond Workshop.” Retrieved January 6, 2020. https://www.genome.gov/27560819/from-genome-function-to-biomedical-insight--encode-and-beyond.
- National Human Genome Research Institute. 2016a. “Characterizing the Functional Elements in the Encyclopedia of DNA Elements (ENCODE) Catalog Request for Applications (RFA-HG-16–003).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-16-003.html.
- National Human Genome Research Institute. 2016b. “Computational Analysis of the Encyclopedia of DNA Elements (ENCODE) Data Request for Applications (RFA-HG-16–004).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-16-004.html.
- National Human Genome Research Institute. 2016c. “ENCODE Data Analysis Center Request for Applications (RFA-HG-16–006).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-16-006.html.
- National Human Genome Research Institute. 2016d. “ENCODE Data Coordinating Center Request for Applications (RFA-HG-16–005).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-16-005.html.
- National Human Genome Research Institute. 2016e. “Expanding the Encyclopedia of DNA Elements (ENCODE) in the Human and Mouse Request for Applications (RFA-HG-16–002).” Retrieved January 6, 2020. https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-16-002.html.
- National Human Genome Research Institute. 2018a. “ENCODE Consortium Criteria for Participation.” Retrieved December 27, 2019. https://www.genome.gov/Pages/Research/ENCODE/Consortium_Criteria_for_Participation.pdf.
- National Human Genome Research Institute. 2018b. “ENCODE External Consultants Panel.” Retrieved December 31, 2019. https://www.genome.gov/Pages/Research/ENCODE/External_Consultants_Panel.pdf.
- National Human Genome Research Institute. 2018c. “ENCODE Phase 2: Participants and Projects.” https://www.genome.gov/Pages/Research/ENCODE/Phase2_Participants_Projects.pdf.
- National Human Genome Research Institute. 2018d. “ENCODE Phase 3: Participants and Projects.” Retrieved January 6, 2020. https://www.genome.gov/Pages/Research/ENCODE/Phase3_Participants_Projects.pdf.
- National Human Genome Research Institute. 2018e. “Novel Genomic Technology Development (PAR-18–777).” Retrieved December 29, 2019. https://grants.nih.gov/grants/guide/pa-files/PAR-18-777.html.
- National Human Genome Research Institute. 2020. Retrieved January 1, 2020. https://www.genome.gov/Funded-Programs-Projects/ENCODE-Project-ENCyclopedia-Of-DNA-Elements.
- National Human Genome Research Institute. 2020. “ENCODE Funded Programs and Projects.” Retrieved January 1, 2020. https://www.genome.gov/Funded-Programs-Projects/ENCODE-Project-ENCyclopedia-Of-DNA-Elements.
- National Human Genome Research Institute. 2022a. “ENCODE Publications.” Retrieved February 13, 2022. https://www.encodeproject.org/publications/.
- National Human Genome Research Institute. 2022b. “Impact of Genomic Variation on Function (IGVF) Consortium.” https://www.genome.gov/Funded-Programs-Projects/Impact-of-Genomic-Variation-on-Function-Consortium.
- National Human Genome Research Institute. 2022c. “Molecular Phenotypes of Null Alleles in Cells (MorPhiC).” https://www.genome.gov/research-funding/Funded-Programs-Projects/Molecular-Phenotypes-of-Null-Alleles-in-Cells.
- National Institutes of Health. 2009. “National Institutes of Health American Recovery and Reinvestment Act of 2009 Challenge Grant Applications Omnibus of Broad Challenge Areas and Specific Topics.” Retrieved January 1, 2020. https://grants.nih.gov/grants/funding/challenge_award/Omnibus.pdf.
- National Institutes of Health. 2020. “NIH Cooperative Agreement Definition.” Retrieved January 10, 2020. https://grants.nih.gov/grants/glossary.htm#CooperativeAgreement.
- Nature. 2020. “Nature Portfolio: ENCODE 3.” https://www.nature.com/collections/dggcchgghg.
- Nienhuis, A. W., and G. Stamatoyannopoulos. 1978. “Hemoglobin Switching.” Cell 15 (1): 307–15.
- Nirenberg, M. W., and J. H. Matthaei. 1961. “The Dependence of Cell-Free Protein Synthesis in E. coli upon Naturally Occurring or Synthetic Polyribonucleotides.” Proceedings of the National Academy of Sciences USA 47:1588–602.
- Niu, W., Z. J. Lu, M. Zhong, et al. 2011. “Diverse Transcription Factor Binding Features Revealed by Genome-Wide ChIP-seq in C. elegans.” Genome Research 21 (2): 245–54.
- Nordman, J., S. Li, T. Eng, D. Macalpine, and T. L. Orr-Weaver. 2011. “Developmental Control of the DNA Replication and Transcription Programs.” Genome Research 21 (2): 175–81.
- Orkin, S. H., and D. E. Bauer. 2019. “Emerging Genetic Therapy for Sickle Cell Disease.” Annual Review of Medicine 70:257–71.
- Riddle, N. C., A. Minoda, P. V. Kharchenko, et al. 2011. “Plasticity in Patterns of Histone Modifications and Chromosomal Proteins in Drosophila Heterochromatin.” Genome Research 21 (2): 147–63.
- Sankaran, V. G., T. F. Menne, J. Xu, et al. 2008. “Human Fetal Hemoglobin Expression Is Regulated by the Developmental Stage-Specific Repressor BCL11A.” Science 322 (5909): 1839–42.
- Snyder, M. P., T. R. Gingeras, J. E. Moore, et al. 2020. “Perspectives on ENCODE.” Nature 583 (7818): 693–98.
- Spencer, W. C., G. Zeller, J. D. Watson, et al. 2011. “A Spatial and Temporal Map of C. elegans Gene Expression.” Genome Research 21 (2): 325–41.
- Stein, R. 2021. “First Sickle Cell Patient Treated with CRISPR Gene-Editing Still Thriving.” Retrieved January 7, 2026. https://www.npr.org/sections/health-shots/2021/12/31/1067400512/first-sickle-cell-patient-treated-with-crispr-gene-editing-still-thriving.
- Strausberg, R. L., E. A. Feingold, R. D. Klausner, and F. S. Collins. 1999. “The Mammalian Gene Collection.” Science 286 (5439): 455–57.
- Uda, M., R. Galanello, S. Sanna, et al. 2008. “Genome-Wide Association Study Shows BCL11A Associated with Persistent Fetal Hemoglobin and Amelioration of the Phenotype of Beta-Thalassemia.” Proceedings of the National Academy of Sciences USA 105 (5): 1620–25.
- Venter, J. C., M. D. Adams, E. W. Myers, et al. 2001. “The Sequence of the Human Genome.” Science 291 (5507): 1304–51.
- Vinjamur, D. S., D. E. Bauer, and S. H. Orkin. 2018. “Recent Progress in Understanding and Manipulating Haemoglobin Switching for the Haemoglobinopathies.” British Journal of Haematology 180 (5): 630–43.
- Wellcome Trust. 2003. “Sharing Data from Large-Scale Biological Research Projects: A System of Tripartite Responsibility.” Retrieved December 29, 2019. https://www.genome.gov/Pages/Research/WellcomeReport0303.pdf.
- Wright, F. A., W. J. Lemon, W. D. Zhao, et al. 2001. “A Draft Annotation and Overview of the Human Genome.” Genome Biology 2 (7): Research0025.
- Yue, F., Y. Cheng, A. Breschi, et al. 2014. “A Comparative Encyclopedia of DNA Elements in the Mouse Genome.” Nature 515 (7527): 355–64.