Skip to main content

Perspectives on the Human Genome Project and Genomics: 3 The NHGRI Genome Sequencing Cost Curve

Perspectives on the Human Genome Project and Genomics
3 The NHGRI Genome Sequencing Cost Curve
  • Show the following:

    Annotations
    Resources
  • Adjust appearance:

    Font
    Font style
    Color Scheme
    Light
    Dark
    Annotation contrast
    Low
    High
    Margins
  • Search within:
    • My Notes + Comments
    • Notifications
    • Privacy
  • Project HomePerspectives on the Human Genome Project and Genomics
  • Projects
  • Learn more about Manifold

Notes

table of contents
  1. Cover
  2. Half Title Page
  3. Series List
  4. Title Page
  5. Copyright Page
  6. Contents
  7. Preface
  8. List of Abbreviations
  9. Introduction: Complexity, Contingency, and Controversy in Genomics
  10. Part 1. Producing the Genome
    1. 1. Challenges in the Early Years of the Human Genome Project at the National Institutes of Health: A Personal Retrospective
    2. 2. Unsung Contributors to the Human Genome Project: NIH Staff and Advisors
    3. 3. The NHGRI Genome Sequencing Cost Curve: An Indicator of Scientific Progress
    4. 4. History of the Encyclopedia of DNA Elements (ENCODE) Project
    5. 5. NHGRI Genetic Variation Program
    6. 6. Genome Technology Development Grants for the Human Genome Project and Beyond
  11. Part 2. Contextualizing the Genome
    1. 7. The Nature of Genomic Publishing
    2. 8. Europe and the Genome: An Overlooked Strategy for a Translational Genomics
    3. 9. Technological Change Driving Scientific Questions: Genomic Sequencing as a Case Study
    4. 10. Addressing Ethical, Legal, and Social Implications (ELSI): Navigating Ongoing Productive Tensions
    5. 11. “Variations on a Theme”: A History of Errors and Polymorphisms in the Human Genome Project and Beyond
    6. 12. Transforming the Genome into a Clinical Resource: DNA, Data, and Algorithms in Medicine
  12. Part 3. Interpreting the Genome
    1. 13. The Difference Genomics Makes: Characterizing Human Differences After the Human Genome Project
    2. 14. The Trouble with Being “Socially Responsible”: Science, GWAS, and Sexual Orientation
    3. 15. Epigenetics in Public Health: Comments on the “From Cells to Society” Approach
    4. 16. When Eugenic Enhancement Meets the Myth of Genetic Reductionism
    5. 17. modENCODE and the Elaboration of Functional Genomic Methodology
    6. 18. The Cancer Genome Atlas Project: Data-Driven, Hypothesis-Driven, or Something In-Between?
    7. 19. Large-Scale Biology: Philosophical, Historical, and Computational Perspectives
  13. Contributors
  14. Index

3 The NHGRI Genome Sequencing Cost Curve

An Indicator of Scientific Progress

Kris A. Wetterstrand and Jonathan E. LoTempio Jr.

The National Human Genome Research Institute (NHGRI) genome sequencing cost curve is the graph that shows the precipitous drop in the cost of genome sequencing over the last two decades. It has become iconic and is used in lectures, talks, and analyses the world over to convey seismic shifts in genomic technology, research, and development. Less attention, however, is given the programmatic and managerial priorities of collecting those metrics and how those metrics changed over time. Indeed, the genome sequence cost curve reduces much of the complexity of DNA sequencing technology development and scientific progress into a representative picture. Furthermore, the graph itself was not a goal, but rather a tool used to understand costs in real time to manage genome sequencing programs. Here, we outline cost projection and tracking leading up to the International Human Genome Project (HGP) and beyond. We will then outline the methodology and management factors, which allowed for evolution of cost-tracking metrics based on technological and methodological developments and changing program priorities. Finally, we will discuss research projects as they relate to cost and how genome sequencing cost has come to represent the exceptional scientific advancements in the field of genomics over the last few decades.

1. Introduction

Since its establishment as a leader of the HGP, the NHGRI has been responsible for the design of some of the largest and most risky projects at the forefront of genomics (see chapters 2, 6, and 9). The scope of the NHGRI network of collaborators and the reputation for rigorous work allowed for the formation of international consortia to probe aspects of the human heredity (see chapter 5), chromatin architecture and regulatory elements (see chapter 4), and the human microbiome (Proctor 2011). In each of these instances, the scientific design of genome projects was subject to the cost of the research, and thus, its feasibility at that moment in time. The fundamental facet of program management, which enabled the design and execution of these projects, was accurate, regular, and frequent, tracking the real-time cost of genome sequencing.

The first cost estimate for a genome project was born of a University of California philanthropic/development endeavor in 1984, which evolved into a 1985 workshop (Sinsheimer 1990). At this point, the expectation was that a human genome project would cost approximately $100,000,000 and take ten to twenty years. The yearly cost estimate increased from tens of millions of dollars to hundreds of millions of dollars, while the estimated time to completion increased from hundreds of person-years to tens of thousands of person-years over a series of workshops convened by the Department of Energy, University of California, Cold Spring Harbor Laboratory, and the Office of Technology Assessment between 1985 and 1986 (US Congress, Office of Technology Assessment 1986; 1988). The ideas of these meetings and their participants were largely organized and synthesized by the National Academy of Science (NAS) between in 1987 and 1988, wherein a proposal for a $200,000,000 per year, fifteen-year project to sequence the human genome was outlined (National Research Council 1988). At the same time, the need for new, stable funding for a human genome project appeared in National Institutes of Health (NIH) Director James Wyngaarden’s statement to the House and Senate appropriation committees (Committee on Appropriations Subcommittee on Departments of Labor, Health and Human Services, Education, and Related Agencies 1986).

The costs presented in the NAS report were the equivalent to 3 percent of the total US federal budget dedicated to biological science and more than double of the total US federal investment in genetics of any species as of 1986 (US Congress, Office of Technology Assessment 1986). These costs were not unexpected, as the first genome projects were technically laborious endeavors involving thousands of specialists across molecular biology, genetics, and computer science.

Early managers of the HGP carefully considered estimates of genome cost when assessing and then justifying the feasibility of such an immense and important project. This was critical to establish the size and scope of the project, as well as to inform international public and private funders of a realistic number for the financial commitment needed to finish the project. That number, US$3–6 billion, from the 1988 NAS report, served as the basis for initial HGP design based on a cost between US$1 and $2 per base pair. Ultimately, the cost of the HGP was approved at US$3 billion (National Research Council 1988), approximately US$1 per finished base pair (National Human Genome Research Institute 2020).

While completing the project at a lower cost than the budget would have been a laudable outcome, the predominant project driver for the HGP wasn’t cost saving. The quality of human genome sequence was also a key metric; therefore, the project management team was also concerned with tracking which base pairs were usable, contiguous, and placed. The methodology and technology of the day (see chapters 6 and 9), required work across many international centers, each which made capital investment in sequencing technology.

This project was undertaken and funded by numerous governments and nonprofits as the HGP. The US contributions were coordinated by the NIH in 1988 as the NIH Office of Human Genome Research and then given an expanded authority as the National Center for Human Genome Research (NCHGR) in 1989. The programmatic leadership became one of the strongest voices in the management of the project. By this time, the product of the HGP was defined as the full sequence of a human genome with one error in ten thousand base pairs (99.99 percent accurate) (National Human Genome Research Institute 2020).

Nearly ten years later, in the closing weeks of 1998, the NCHGR was newly elevated to status as a full-fledged institute, NHGRI. With the end of the HGP in sight, a group of program staff at NHGRI convened a cost workshop. The meeting was specifically attended by individuals from the NHGRI-funded genome centers who worked most closely with cost reporting, allowing for a high-granularity discussion (of the components) of sequence production costs. Key agenda items were the discussion of cost components, and the need for regular meetings of the individuals responsible for their reporting.

It wasn’t until 2001, however, that NHGRI, in conjunction with representatives from the NHGRI-funded genome centers, refined and implemented processes to collect and assess genome sequencing cost data. The decisions of this group of scientists, administrators, and program managers constructed the framework for genome sequence cost reporting at NHGRI to the present day and produce the NHGRI genome sequence cost curve (see Figure 3.1, Wetterstrand 2022). NHGRI program staff and grantees have also developed detailed methods to track DNA sequencing production metrics. These metrics are an essential tool for the discussion and evaluation of the progress of sequencing programs working to meet scientific goals.

Y-axis runs from 100 to 100 million USD on a log base 10 scale. X-axis runs from August 2001 to May 2022. Graph shows decrease in cost.

Figure 3.1. NHGRI genome sequencing cost curve.

Figure Description

NHGRI genome sequencing cost curve showing the cost of sequencing a human genome. Data is presented on a logarithmic scale. The first data point begins at a cost of roughly one hundred million US dollars per genome in August 2001. This part of the graph contains consistent decreasing cost data representing Sanger-based (dideoxy chain termination) capillary sequencing until January 2008 where second generation or “next-generation” sequencing by synthesis platforms come online at NHGRI sequencing centers. There is an exponential drop from this point until 2012 or so, where cost decreases are smaller each reporting period. There is one further inflection point where high throughput next generation sequencing (branded as Illumina X Ten) platforms come online and drive costs to roughly one thousand US dollars per genome.

Date

x-axis

Cost per Genome

y-axis

September 2001

$95,263,072

March 2002

$70,175,437

September 2002

$61,448,422

March 2003

$53,751,684

October 2003

$40,157,554

January 2004

$28,780,376

April 2004

$20,442,576

July 2004

$19,934,346

October 2004

$18,519,312

January 2005

$17,534,970

April 2005

$16,159,699

July 2005

$16,180,224

October 2005

$13,801,124

January 2006

$12,585,659

April 2006

$11,732,535

July 2006

$11,455,315

October 2006

$10,474,556

January 2007

$9,408,739

April 2007

$9,047,003

July 2007

$8,927,342

October 2007

$7,147,571

January 2008

$3,063,820

April 2008

$1,352,982

July 2008

$752,080

October 2008

$342,502

January 2009

$232,735

April 2009

$154,714

July 2009

$108,065

October 2009

$70,333

January 2010

$46,774

April 2010

$31,512

July 2010

$31,125

October 2010

$29,092

January 2011

$20,963

April 2011

$16,712

July 2011

$10,497

October 2011

$7,743

January 2012

$7,666

April 2012

$5,901

July 2012

$5,985

October 2012

$6,618

January 2013

$5,671

April 2013

$6,618

July 2013

$5,550

October 2013

$5,096

January 2014

$4,008

April 2014

$4,920

July 2014

$4,905

October 2014

$5,731

January 2015

$3,970

April 2015

$4,211

July 2015

$1,363

October 2015

$1,245

May 2016

$4,046

August 2016

$1,363

November 2016

$1,245

February 2017

$1,176

May 2017

$1,508

August 2017

$1,356

November 2017

$1,015

February 2018

$1,333

May 2018

$1,134

August 2018

$1,844

November 2018

$1,232

February 2019

$1,463

May 2019

$1,467

August 2019

$1,392

November 2019

$989

February 2020

$606

May 2020

$942

August 2020

$695

November 2020

$645

February 2021

$702

May 2021

$689

August 2021

$512

November 2021

$851

February 2022

$454

May 2022

$562

It was clear that this sequencing capacity would persist after the completion of the HGP and that the field would benefit from “experience curve effects” where a second iteration tends to be cheaper or easier than a first (Majd and Pindyck 1987). This meant that NHGRI could undertake new scientific projects at a scale not previously imagined when completing the first genome(s). Accordingly, NHGRI program staff worked with key stakeholders at each of the genome sequencing centers to gain insight into sequencing capacity and costs at each center, and how production could be reported and used for program design and management. This work served as the foundation for cost tracking and indeed influences the genome sequence cost curve to this day.

Managing programs in this quantitative mode would require greater cooperation than a traditional research grant. The Federal Grants and Cooperative Agreements Act of 1977 outlines three funding mechanisms, which remain largely intact today: contracts, grants, cooperative agreements. Traditional NIH research grants have a minimal degree of oversight. Formal progress reports are submitted yearly, and program staff have minimal input to influence the work products of that project. Conversely, contracts are designed to procure a very specific product set out at the beginning of the award.

Neither of these mechanisms allow for the coordination of activities across multiple sites, as was honed during the HGP, nor do they allow for the fact that rapid development of technology makes a priori procurement decisions impossible. Given the speed at which DNA sequencing technologies changed (Heather and Chain 2016), the design of large-scale genome sequencing projects at NHGRI was considerate of the need for flexibility, collaboration, and rapid iteration. The cooperative agreement mechanism (National Institutes of Health 2020) allows for program coordination (see chapter 2). Therefore, large-scale genome sequencing projects, which produced data underlying the genome sequence cost graph, were all structured as cooperative agreements, opening the door for the collection of data that allows for the granular tracking of cost and production.

Through the cooperative agreement mechanism and with buy-in from grantee staff responsible for project management, program officers and administrators could help guide multi-site programs in a dynamic and informed way. The cooperative agreement mechanism allowed for the development of tracking products in addition to the obvious metrics of success including papers, patents, presentations, and tool development.

In this chapter, we outline the fundamentals of production metrics and cost tracking through grant recipient progress reports and aim to explore the methodologies that underpinned cost tracking from the end of the human genome project through subsequent genome sequencing projects funded by NHGRI. We will examine how the evolution of progress report metrics reflects the changing priorities of managing the twin needs of technological improvements and research. Additionally, we explore how the NHGRI genome sequencing cost curve compares to reported costs for specific genome projects and from commercial entities with sequencing platforms. Finally, we reflect on how the cost curve has come to represent the remarkable scientific advancements made in the field of genomics that has revolutionized biomedical research and give some thoughts on the future of the genome sequence cost curve.

2. Methodology

2.1. Overview of Genome Sequence Cost Graph Components as a Function of Changing Production Priorities and Sequencing Technology Platforms

While negotiations and planning for progress reporting began in 1998, ramp up and report implementation commenced in 2001. Since then, NHGRI has tracked detailed production and cost metrics associated with DNA sequencing performed at the sequencing centers funded by the institute. These data served as important tools for assessing improvements in DNA sequencing approaches and for coordinating the DNA sequencing capacity of the NHGRI Genome Sequencing Program (GSP). Technology platforms, research methods, and sequencing projects have evolved throughout the program, leading to changes in emphasis on specific progress reporting metrics.

From 2001 to 2020, costs were largely reported at the same frequency. Upon issuance of the Large-Scale Sequencing Center Grants in 2004 (National Human Genome Research Institute 2003a) and subsequent renewal in 2006 (National Human Genome Research Institute 2005a) and 2010 (National Human Genome Research Institute 2010), timing of quarterly reports stabilized through 2015 and the funding of the Centers for Common Disease Genomics (CCDGs) (National Human Genome Research Institute 2015). The overall GSP budget and number of grantees has varied since the beginning of the program with peaks in 2002 and 2001, respectively (see Figure 3.2).

For example, early progress report metrics focused on the success rate, read length, and cost of a sequencing read generated by Sanger-based sequencing, as well as the cost of a finished base (National Human Genome Research Institute 2015). After the advent of so-called next-generation sequencing (NGS) platforms (Metzker 2010), progress reports focused on the quality and cost of bases generated from flow cells and comparative costs between exomes (the part of the genome that consists of exons, segments of a DNA or RNA molecule containing information coding for a protein), and genomes (the entire set of genetic instructions found in a cell).

This resulted in a landscape with two specific metrics of progress: the cost of sequencing per base pair, and the cost of sequencing per human genome. The cost per base pair was a real and reported metric based on the amount of sequence generated by each center, whereas the cost per genome was calculated with certain, changing assumptions about the size of the human genome and the coverage needed to achieve a “useful product” for study. This “useful product” changed as sequencing technology evolved.

The NGHRI genome sequencing budget increases from 1996 to 2003, peaking at 199 million, then drops off each following year to just 60 million by 2020.

Figure 3.2. NHGRI sequencing budget from 1996 to 2020.

Figure Description

The overall NHGRI genome sequencing budget is presented for 1996 through 2020 and broken down by grantee sequencing center in millions of US dollars. Shades of grey represent those different grantee sequencing centers and have not been labeled to preserve anonymity. At the highest overall number of centers in 2000, there were a total of twelve NHGRI-funded sequencing centers. The lowest number, three NHGRI-funded sequencing centers, spanned from 2008-2015. Sequencing centers were budgeted beginning in 1996 with 23.3 million US dollars. The total sequencing budget increased each year until 2002 with a high of 199.3 million US dollars. In 2020, the last budget year reported, the total funds for genome sequencing centers were 60.0 million US dollars. None of these totals were adjusted for inflation and are reported in real dollars at the time of funding.

Center

1996

1997

1998

1999

2000

2001

2002

2003

2004

2005

2006

2007

2008

2009

2010

2011

2012

2013

2014

2015

2016

2017

2018

2019

2020

1

4.1

8.2

8.8

45.2

61.9

69.9

69.0

71.0

58.5

52.7

47.7

52.0

52.0

46.8

46.6

46.6

35.9

34.1

32.4

30.7

20.0

20.0

20.0

20.0

20.0

2

6.7

9.7

27.4

43.3

46.0

55.6

66.5

60.3

49.0

44.1

40.0

40.4

41.0

37.1

36.8

36.8

28.4

26.9

25.6

24.3

15.0

15.0

15.0

15.0

15.0

3

1.1

4.0

9.8

20.0

22.8

29.4

32.9

27.0

35.4

31.9

28.9

29.5

30.0

27.0

27.0

27.0

21.3

20.3

19.2

18.3

15.0

15.0

15.0

15.0

15.0

4

1.0

2.4

8.9

7.2

2.6

7.6

5.3

1.9

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

5

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

10.0

9.0

8.2

1.2

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

6

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

10.0

9.0

8.2

0.7

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

7

No dataNo dataNo data

6.5

4.5

9.3

7.9

0.3

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

8

3.3

5.3

5.1

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

9

3.6

7.3

1.4

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

10

1.0

1.4

2.0

No data

2.5

3.6

3.9

0.5

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

11

2.5

4.2

3.6

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

12

No dataNo dataNo data

3.9

1.1

3.2

2.2

1.0

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

13

No dataNo data

3.1

3.9

3.2

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

14

No dataNo data

1.0

No data

1.4

2.1

2.2

2.4

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

15

No dataNo dataNo dataNo data

1.3

2.6

1.7

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

16

No dataNo dataNo dataNo data

1.2

1.8

2.0

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

17

No dataNo dataNo dataNo data

1.7

2.6

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

18

No dataNo dataNo dataNo dataNo data

2.6

0.0

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

19

No dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo dataNo data

10.0

10.0

10.0

10.0

10.0

In the early days of the HGP, costs were estimated per base pair. Over the HGP, production increased to the point that kilobase pairs, or one thousand base pairs, became the standard unit of production that was used by sequencing centers. During the large-scale sequencing program between the 2000s and early 2010s, the unit was increased to megabase, one million base pairs. In 2015, centers began reporting gigabase pairs and by 2017, terabase pairs (one thousand gigabase pairs) were achieved in any given quarter. The costs associated with each of these units were then used to infer the cost per human genome (see Figure 3.3).

For the purpose of cost tracking, “one human genome” was calculated as a three billion base pair genome at 6x coverage from 2001 to 2007, at 10x coverage in January 2008, and at 30x coverage from April 2008 to the present (see Table 3.1). For most of the time captured in this curve (see Figure 3.1), the cost of generating this theoretical 30x human genome was not actually undertaken, rather, the cost per megabase across organismal, medical, exome, and low-coverage sequences was extrapolated to generate this figure.

Y-axis runs from 0 to 10 thousand USD on a log base 10 scale. X-axis runs from August 2001 to February 2022. Graph shows decrease in cost.

Figure 3.3. Cost of raw megabase pair of DNA sequence from which the cost of sequencing a human genome is calculated.

Figure Description

Cost of raw megabase pair of DNA sequence from which the cost of sequencing a human genome is calculated. Data is presented on a logarithmic scale. These data were captured across organisms and projects—the unit, whether it was an early-days cost per basepair or the unit megabase pair as presented in this figure, is agnostic to scientific intent. From this number, an estimated cost per human genome was calculated even though few whole human genomes were produced during this time. The first data point begins at a cost of roughly 5 thousand three hundred US dollars per megabase pair in August 2001. This part of the graph contains consistent decreasing cost data representing Sanger-based (dideoxy chain termination) capillary sequencing until January 2008 where second generation or “next-generation” sequencing by synthesis platforms come online at NHGRI sequencing centers. There is an exponential drop from this point until 2012 or so, where cost decreases are smaller each reporting period. There is one further inflection point where high throughput next generation sequencing (branded as Illumina X Ten) platforms come online and drive costs to roughly 6 cents per megabase pair.

Date

x-axis

Cost per Megabasepair

y-axis

September 2001

$5,292.39

March 2002

$3,898.64

September 2002

$3,413.80

March 2003

$2,986.20

October 2003

$2,230.98

January 2004

$1,598.91

April 2004

$1,135.70

July 2004

$1,107.46

October 2004

$1,028.85

January 2005

$974.16

April 2005

$897.76

July 2005

$898.90

October 2005

$766.73

January 2006

$699.20

April 2006

$651.81

July 2006

$636.41

October 2006

$581.92

January 2007

$522.71

April 2007

$502.61

July 2007

$495.96

October 2007

$397.09

January 2008

$102.13

April 2008

$15.03

July 2008

$8.36

October 2008

$3.81

January 2009

$2.59

April 2009

$1.72

July 2009

$1.20

October 2009

$0.78

January 2010

$0.52

April 2010

$0.35

July 2010

$0.35

October 2010

$0.32

January 2011

$0.23

April 2011

$0.19

July 2011

$0.12

October 2011

$0.09

January 2012

$0.09

April 2012

$0.07

July 2012

$0.07

October 2012

$0.07

January 2013

$0.06

April 2013

$0.06

July 2013

$0.06

October 2013

$0.06

January 2014

$0.04

April 2014

$0.05

July 2014

$0.05

October 2014

$0.06

January 2015

$0.04

April 2015

$0.05

July 2015

$0.015

October 2015

$0.014

May 2016

$0.013

August 2016

$0.017

November 2016

$0.015

February 2017

$0.011

May 2017

$0.015

August 2017

$0.013

November 2017

$0.020

February 2018

$0.014

May 2018

$0.016

August 2018

$0.016

November 2018

$0.015

February 2019

$0.011

May 2019

$0.007

August 2019

$0.010

November 2019

$0.008

February 2020

$0.007

May 2020

$0.008

August 2020

$0.008

November 2020

$0.006

February 2021

$0.009

May 2021

$0.005

August 2021

$0.006

November 2021

$0.006

February 2022

$0.006

May 2022

$0.006

Table 3.1. Considerations for calculating a theoretical cost per three-billion base pair human genome based on a megabase of sequence produced an aggregate cost of data generated on various sequencing platforms.

Platform

(Method)

First year present in cost graph

Last year present in cost graph

Read length

(base pairs)

Coverage

(x-times)

ABI capillary

(Sanger sequencing)

2001

2007

500–600

6

454

(pyrosequencing)

2008

2008*

300–400

10

Illumina/SOLiD

(sequencing by synthesis)

2008

2021

75–150

30

* Only one report in 2008 used 454 cost data (three-month reporting period ending January 31, 2008).

While whole human genomes were not a high priority for sequence generation between 2001 and 2011, critical work was still undertaken. Some of the important nonhuman genomes included chicken, chimpanzee, honeybee, and sea urchin (National Human Genome Research Institute 2002) as well as the pathogens and vectors project (National Human Genome Research Institute 2006), and the human gut microbiome initiative (Gordon et al. 2006). These projects all leveraged either the technology or sequence developed during the HGP in their completion. The genome sequence generated during these projects informed the theoretical cost of sequencing a whole human genome. It was not until 2009 that human sequencing projects represented the majority of DNA sequence generated by NHGRI centers.

For data from January 2008 (representing data generated using “second-generation” sequencing platforms), Figure 3.1, the “Cost per Genome” graph, reflected projects that involved the “re-sequencing” of the human genome, where an available reference human genome sequence was available to serve as a backbone for downstream data analyses. The required “sequence coverage” may have been greater for sequencing genomes for which no reference genome sequence was available.

A key nuance that both graphs obscure is the technological shift, which occurred 2007–08. In all reporting, the data from 2001 through October 2007 represented the costs of generating DNA sequence using Sanger-based chemistries on capillary-based instruments (“first generation” sequencing platforms). From January 2008, the cost data within progress reports represented the costs of generating DNA sequence using “second-generation” (or NGS) platforms.

For the Sanger-based sequence data, the cost accounting reflected the generation of bases with a minimum quality score of Phred20 (or Q20) (Green and Ewing 1998), which represented an error probability of 1 percent and is an accepted community standard for a high-quality base. For sequence data generated with NGS platforms, there is not yet a single accepted measure of accuracy; each manufacturer provides quality scores that are accepted by the NHGRI sequencing centers as equivalent to or greater than Q20 in the case of Illumina and Pacific Biosciences (“third generation” platforms) (Illumina 2011; Land et al. 2014), or much lower, in the case of Oxford Nanopore Technologies (ONT, also a “third generation platform”) (Laver et al. 2015). This divergence in usefulness of quality scores in current platforms available in 2024 underscores the differences in technologies and how accuracy of a base and length of a read can interplay together for genome analysis. A very accurate short read (150bp, Illumina) may be difficult to place in the genome, while a less accurate long read (150kbp, ONT) may have enough content for placement. In the case of original HGP reads, length (500bp) and accuracy (99 percent) considerations meant that, at 6x coverage, only 1/10,000 bases would be incorrect.

In Figure 3.3, the “Cost per Megabase of DNA Sequence” graph, the data reflect the cost of generating raw, unassembled sequence data; no adjustment was made for data generated using different instruments despite significant differences in the sequence read lengths. In contrast, Figure 3.1, the “Cost per Genome” graph does account for these differences since sequence read length influences the ability to generate an assembled genome sequence.

Regardless of the year, the science, or the grantee, the sequencing cost reports had to include the cost of producing the sequences of interest, associated compute and analysis costs, and the labor and equipment leveraged in the project. While some of the particular definitions were modified over the years, the following so-called “production” categories informed the reported and calculated costs:

  1. Labor, administration, management, utilities, reagents, and consumables
  2. Sequencing instruments and other large equipment (amortized over three years)
  3. Informatics activities directly related to sequence production (e.g., laboratory information management systems and initial data processing)
  4. Submission of data to a public database
  5. Indirect costs, or “overhead” (National Institutes of Health Office of Management 2017) as they relate to the above items
  6. In the case of costs covered by significant subsidies to a sequencing center (e.g., a grantee institution providing funds for purchasing large equipment), NHGRI has attempted to appropriately account for such costs in these analyses

Importantly, “non-production” activities were excluded from reporting of costs. These non-production activities included:

  1. Quality assessment/control for sequencing projects
  2. Technology development to improve sequencing pipelines
  3. 3Development of bioinformatics/computational tools to improve sequencing pipelines or to improve downstream sequence analysis
  4. Management of individual sequencing projects
  5. Informatics equipment
  6. Data analysis downstream of initial data processing (e.g., sequence assembly, sequence alignments, identifying variants, and interpretation of results)

Apart from some changes experienced in the earliest phase of progress reporting, these categories remained stable over time, thereby enabling the collection of data providing the most complete picture of sequencing costs over time.

2.2. Changes in Reporting over Three Periods of Large-Scale Sequencing

NHGRI sequencing progress reports emphasized the capture of data regarding different types of projects in the sequencing pipeline. The genome sequence cost curve then emerged from reports based on these activities, regardless of organism or project. In the section below, three different periods of time organized around the types of genome projects underway at the NHGRI-funded genome sequencing centers will be described.

These periods are the Human Genome Project (HGP) period from 2001 to 2004, the Large-Scale Sequencing and Analysis Centers (LSAC) period from 2004 to 2015, and the GSP period from 2015 to 2020. These time periods will serve as a framework for illustrating the evolution of the NHGRI-funded sequencing pipeline and the progress reports associated with it.

A fact obscured by the stability of twenty years of progress reports is that those reports had to be consistent in the data captured so that NHGRI program staff could design future Request for Applications (RFAs) and programs, while being flexible enough to account for technology development and the roll out of new platforms, flow cells, and chemistries.

In terms of what was included, and what was excluded from the quarterly reports, three eras are important:

  1. The HGP period from 1998 to 2004, where the programmatic driver was success and quality of the first genome, not cost per base pair, and as such, not all data could be included in the sequencing cost graph.
  2. The LSAC period from 2004 to 2015, where the programmatic driver was high-quality, high-impact organismal sequencing from 2004 to 2011 and human mendelian exomes, clinical genomes, informatics development, and cancer and general medical sequencing from 2011 through 2015.
  3. The GSP period from 2015 to 2020, where the programmatic driver was the high production of whole human genomes to find some variation responsible for common human diseases with a heritable component.

This shift in focus from sequencing the tree of life, to exploring the possibilities of lower cost sequencing, to diving deep into common, human genomics reflects the possibilities of genome science at different costs per base pair. (See chapters 2, 6, and 9).

2.3. Period I: International Human Genome Project (HGP) Period (1998–2004)

During the HGP period, the main drivers of genome sequencing projects were data production and quality. Sequencing cost was a factor, in that funds were limited, but the ability to sequence the genome at a level of quality that would allow for genome assembly and comprehensively represent the full human genome was the priority. The first formal sequencing progress reports were collected from the NHGRI sequencing centers in October 2001 and subsequently collected every six months. During this time, NHGRI and the sequencing centers laid the foundations for capturing the cost of activities to be included in the cost metric.

To achieve the goals of this period of genome sequencing, the emphasis of monthly reports was placed on:

  1. The length and quality of reads
  2. The ratio of reads attempted to reads passed quality control
  3. The capability of producing sequence from different library preparation methods
  4. Which reference organism received sequence production in a reporting period
  5. The trajectory of production increases and cost decreases

Of critical importance to this period was the distinction between a working draft genome, defined as 90 percent of bases accurate and placed, and a finished genome, defined as 99.9 percent of bases accurate and placed. While the technology of this era could produce reads approaching one thousand base pairs (National Human Genome Research Institute 2020) using Sanger-based, capillary DNA sequencing, reads produced in this program were optimized for cost and scientific utility at five hundred base pairs. Following the completion of HGP in 2003 (National Human Genome Research Institute 2003a), focus was placed on the ability to produce shotgun sequencing whether whole genome or from clones mapped across a genome. At this time, the methods for library construction, template preparation, sequencing reaction, and data analysis were diverse. As such, items needed to be precisely described and accounted for to help program staff understand the immediate state of the art.

The reported costs were considered “fully loaded,” in that they included technology development and software development, though they did not include all costs associated with the research grants. They included labor, materials and reagents, subclone library construction, template preparation, sequencing reaction, gel electrophoresis, data analysis, and indirect costs. Equipment costs were also included as amortized costs over a three-year time span. Excluded activities spanned bacterial artificial chromosome (BAC) fingerprint mapping, BAC-end sequencing, and light BAC coverage, and other activities not part of data production, as described by the centers.

This period was marked by work to finish the human genome and to sequence model organisms, such as mouse, rat, fruit fly, Brewer’s yeast, and worm. By the end of this period, sequencing had started for organisms such as tetraodon, ciona, chimp, macaque, sea urchin, chicken, dog, and many others. These projects, for the most part, were established through design of specific RFAs, wherein potential parties would submit their plan to contribute to a given organism. More on this process can be found in Felsenfeld and Wetterstrand (see chapter 9).

2.4. Period II: Large-Scale Sequencing and Analysis Centers (LSAC) Period: From 2004 to 2015

The activities of the LSACs represent the bulk of sequencing activity at NHGRI-funded centers over time. They also produced the reports that underpinned the majority of data in the genome sequence cost curve.

During the LSAC period, cost was a significant driver of scientific choices and outcomes, in that at decreased costs the number of organisms which could be sequenced expanded. Furthermore, decreasing costs allowed for targeted resequencing of parts of interest in the human genome to probe variation as it related to disease (Stratton 2008). Cost metrics during this time experienced two major revisions: first, where details regarding subcomponents of the cost were examined separately, and second, where those subcomponents were combined. NHGRI moved away from examining “fully loaded” costs to the concept of “production only” cost elements, which stayed consistent through 2020. The data captured in these reports represent the core activities that are common to all types of sequencing projects, regardless of the organism or target, which produce essential elements needed for complete products of a quality that is useful for any research question. These activities came to define “production only” costs. Costs associated with activities that fell outside this definition were still assessed but are not included in the genome sequencing cost curve.

Progress report frequency increased from six-month reporting periods to three-month reporting periods. This was largely a function of increased capacity and decreased cost. Since more projects were being undertaken and completed quickly, there was a need for more information. In these reports, NHGRI deemphasized cost per read, instead choosing to focus on the cost per kilobase pair of sequence generated. The progress reports captured information about the types of sequencing projects underway (i.e., human microbiome, cancer, medical, organismal, 1000 Genomes Project) and information about where in the sequencing and analysis pipeline individual projects sat.1 NGS platforms (Roche/454, Solexa/Illumina and ABI/SOLiD) were introduced during this time as well.

The NHGRI progress reporting working group defined more precisely what was included in reported cost metrics and started using “production only” terminology. Key to this terminology was the contrast between different kinds of development. Radical developments were contrasted from incremental developments by their completeness. For example, the design and roll out of version 1.0 of a new software tool is radical, while updating to version 1.1 is usually incremental; testing a new platform is radical, while testing new chemistry on an existing platform is usually incremental.

It was then decided that “production only” costs would include personnel, management, coordination with other centers, supplies, reagents, library construction, production-related bioinformatics (LIMS, informatics support, software licensing fees), administrative support, and indirect costs. Also included in production only costs, but presented separately, were equipment (amortized over three years); incremental bioinformatics development (bioinformatics platforms that have been newly incorporated into the production floor and improvements to the existing system); incremental technology development (technology platforms that have been newly incorporated into the production floor and improvements to the existing system); routine, automated BAC assembly; and data submission. In 2009, all subcategories were combined into one metric. Genome finishing costs, pertinent to the completion of genome projects, and the resolution of complex regions or bioinformatic annotation of genes had a similar set of categories.

Activities excluded from reported cost metrics included whole genome assembly, assembly assessment and validation, automated annotation, heterozygosity testing, manual annotation, radical bioinformatics development, radical technology development, development of assembly methods, development of annotation methods, finishing, and mapping.

During the early years of the LSAC period (2004 to 2010), progress report production metrics similar to that collected in the HGP period increased in level of detail by collecting production and cost numbers for different types of libraries (whole genome shotgun, BAC-based shotgun, and fosmid-based shotgun) and their subclones (small insert plasmid, large insert plasmid, fosmid, and BAC end). Metrics were also collected for different finishing activities and data deposition.

In 2005, the NHGRI-funded centers started reporting the number and types of sequencing machines that were online at their production sites. NHGRI often collected metrics on “new” platforms before they were in production.

The cost of genome sequencing curve (see Figure 3.1) did not include so-called next generation, or later second generation, sequence generation platforms (NGS) until 2008. By 2009, the emphasis was almost completely on the NGS platforms (Roche/454, Solexa/Illumina and ABI/SOLiD), although data related sequence improvement/finishing and ABI/Sanger sequencing production were still collected, though this ceased by the end of the period. Also, during this time, NHGRI collected data regarding PCR-directed, 16S/18S metagenomic, cDNA, as well as whole genome sequencing.

NGS metrics include items such as pairing (pair or fragment), runs, pass filter MB, pass filter MB/run, cycles (35, 45, 50, 75, 100) that correspond to read length, pass filter clusters/run, % pass filter align/run, % error rate, % duplicate reads, and cost per run. Definitions Solexa/Illumina, ABI/SOLiD, ABI/Sanger, and Roche/454 production metrics are described in Table 3.2. Like a “read,” a “run” represents a unit of production for NGS platforms. Cost per Mb was calculated from cost per run. The Illumina (originally Solexa) platform quickly dominated in terms of production. At the beginning of 2013, metrics for Roche/454 and ABI/SOLiD were dropped.

Table 3.2. ABI/Sanger, Roche/454, Solexa/Illumina and ABI/SOLiD production metrics

Technology platform

Solexa/Illumina

ABI/SOLiD

ABI/Sanger

Roche/454

Sample type

Sample type: Paired ends or fragments.

No data

Small-insert plasmids should have insert sizes of 7kb or less. Large-insert plasmids should have insert sizes of 7–10kb.

No data

Quality

% pass filter error rate/lane: The percentage of called bases in aligned reads that do not match the reference. A low number here is desirable and would indicate high-quality data that matches against the reference sequence.

Avg % matching beads/run: The percent of beads/reads matching the reference per run.

Total Q20 bases are calculated from the number of successful reads and the average read length.

Avg QV/run: Determined by the vendor software and uses all signals contained in the reads that passed filtering and trimming. Quality scores for individually called bases are determined by comparing the signal intensity measured during the nucleotide flow with models of ideal signals that were generated empirically at 454 Life Sciences. The quality scores are then reported as Phred-equivalents.

Read and runs

Runs: Each Solexa run produces eight individual lanes of data.

Runs: Number of SOLiD runs generated.

Successful read: >=100greater than or equal to one hundred bases of Q20 of organism sequence.

Runs: Number of 454 runs generated.

Read length

Read length: Targeted read length.

Read length: Targeted read length determined by instrument/reagent cycle.

Average read length: avg #number of Q20 bases of organism sequence (total #number of Q20 bases in successful reads/#number of successful reads). Please trim vector sequence before reporting average read length

Avg pass filter reads/run: Total pass filter reads per run.

Total pass filter

Total pass filter MB: Total number of bases collected from the run from passed clusters.

Total pass filter MB: The total MB generated is determined by the number of matching beads/reads to the reference sequence.

Pass rate is calculated as #number successful reads/#number attempted reads.

Total pass filter MB/run: Average MB/run.

Pass filter

Pass filter MB/run: Average number of bases collected per run from passed clusters.

Pass filter MB/run: Average MB generated per run (slide or spot) and is based on the matching beads/run.

No data

Total pass filter reads: Total pass filter reads for each project and Platform 454.

Filtering includes Keypass, dots, mixed, trimmed too short quality (TSQ), and trimmed too short primer (TSP).

Keypass filter verifies that the well contains a valid key sequence, usually >greater than80% of total reads.

Dots are instances of three nucleotide flows that record no incorporation and can include wells with poor chemistry or mixed sequences resulting in poor signal intensity.

Trimmed to short quality (TSQ) reads are instances where the trimmed read has fewer than 3% borderline flows or the trimmed read due to quality is less than 50bp, whichever occurs first.

Trimmed too short primer (TSP) reads are instances where the adaptor sequence is scanned and trimmed resulting in a read shorter than 50bp.

Other, platform specific considerations

% pass filter clusters/lane: The percentage of clusters passing filtering. The higher the number is the better. Ideally >greater than 50–70%, but very dependent on cluster density.

% pass filter align/lane: The percentage of called bases in aligned reads that do not match the reference. The percentage of filtered reads that uniquely align to reference.

Avg #number of beads/run: The average number of beads detected per slide or spot.

Avg matching beads/run: The average number of beads/reads that match the reference sequence per run.

Total number of sub-clones yielding two successful reads (plasmid sub-clones only). Please provide the total number of clones attempted, including those for which neither end was sequenced.

Whole genome finishing is finishing based on whole genome assemblies that are being taken to high (Bermuda) standards for finishing.

Grades of “pre-finishing” will be addressed by a working group. Definitions for these categories will be developed based on working group discussions.

Please report the number of base pairs targeted by pre-finishing efforts.

Finished sequence should meet the NHGRI definition.

Avg % pass filter reads/run: Total pass filter reads to those of keypass.

Total MB: Total pass filter mega bases generated.

Additional emphasis was placed on type of project. In 2009, NHGRI started tracking data production from the collaboratively funded Human Microbiome Project, The Cancer Genome Atlas, The 1000 Genomes Project, as well as internally funded organismal, medical, metagenomics sequencing programs. By 2013, the progress reports emphasized production for individual projects, though cost per kilobase pair per organism was not a part of this. By the end of this period, NHGRI collected detailed information in terms of the number samples and gigabase pairs generated for whole genome; whole exome; targeted sequencing (non-exome); transcriptome, RNA, & cDNA; 16S; and BACs, clones, linked-pair libraries projects that belonged to the following bins: one thousand genomes, cancer, HMP, medical non-cancer, medical nonhuman model, metagenomic, human microbiome, microbe, non-microbe organism, and other (e.g., see Figure 3.4). Subsequently, NHGRI added tracking projects by sampling status, sequencing status, analysis status, and completion status. Projects within these bins were proposed by the community and the sequencing centers based on high research priority needs and the availability of samples. More about this process can be found in Felsenfeld and Wetterstrand (see chapter 9).

Sequencing priorities changed from 2009 to 2016, with a decrease in organismal and cancer studies and an increase in medical sequencing.

Figure 3.4. Sequencing production from 2009 to 2016 broken down by type of project.

Figure Description

Sequencing production from 2009 to 2016 broken down by type of project. Graph shows the change in the types of sequencing projects conducted by the grantee sequencing centers funded by NHGRI. The types of project are: whole genome organismal sequencing (organism), human microbiome project (HMP, https://hmpdacc.org/), Gabriella Miller Kids First (GMKF, https://commonfund.nih.gov/kidsfirst), 1000 Genomes Project (1000 Gs, https://www.internationalgenome.org/), cancer genomics (cancer) and human disease (medical). Data is based on the number of basepairs produced and is given in percentages of the overall pipeline. Graph shows change in priorities over time, with a decrease in organismal sequencing, a decrease in sequencing to study cancer, and an increase in general medical sequencing.

Year

Quarter

Sum of Medical

Sum of Cancer

Sum of 1000 Gs

Sum of GMKF

Sum of HMP

Sum of Organism

Sum of Other

2009

1

0%

40%

7%

0%

5%

31%

17%

2009

2

0%

42%

24%

0%

1%

19%

13%

2009

3

1%

82%

6%

0%

0%

3%

7%

2009

4

3%

75%

9%

0%

0%

3%

10%

2010

1

3%

71%

16%

0%

0%

3%

8%

2010

2

6%

77%

4%

0%

1%

4%

8%

2010

3

9%

51%

7%

0%

22%

5%

7%

2010

4

26%

46%

6%

0%

1%

11%

9%

2011

1

19%

53%

9%

0%

1%

4%

14%

2011

2

19%

57%

3%

0%

0%

7%

13%

2011

3

27%

49%

4%

0%

1%

8%

9%

2011

4

32%

49%

5%

0%

3%

5%

6%

2012

1

36%

47%

5%

0%

2%

5%

5%

2012

2

43%

39%

1%

0%

3%

9%

6%

2012

3

50%

32%

0%

0%

5%

5%

8%

2012

4

54%

30%

0%

0%

0%

11%

5%

2013

1

44%

38%

7%

0%

0%

7%

4%

2013

2

45%

33%

10%

0%

1%

8%

3%

2013

3

57%

35%

0%

0%

1%

5%

2%

2013

4

57%

36%

0%

0%

1%

3%

4%

2014

1

47%

47%

2%

0%

0%

3%

2%

2014

2

69%

23%

0%

0%

1%

4%

3%

2014

3

84%

6%

0%

0%

2%

1%

7%

2014

4

48%

39%

0%

0%

0%

2%

11%

2015

1

56%

25%

0%

0%

0%

0%

18%

2015

2

40%

12%

0%

0%

0%

1%

47%

2015

3

92%

5%

0%

0%

0%

0%

2%

2015

4

91%

5%

0%

0%

0%

0%

5%

2016

1

36%

5%

0%

46%

0%

0%

13%

2.5. Period III: Genome Sequencing Project (GSP) Period (2015–2020)—CCDGs

This period of sequencing projects saw the greatest increase in capacity, which played a key role in the programmatic design of the CCDG RFA (National Human Genome Research Institute 2015a) A new center was added, and project scope changed fundamentally. This required a rethinking of progress reporting.

A key addition to the reporting scheme in this period of the program was monthly sample pipeline reports. Due to the large number of samples in different projects spread over multiple centers, there was programmatic need to have detail regarding each sample, including sex, ethnicity, and disease status, as well as patient or participant consent level.

The necessity for centers to take on this new facet of reporting, the cost, and production reports were streamlined to be less burdensome and more reflective of the work that was to be performed in the completion of the project. These new reports included accounting of production in megabases, as well as the literal number of genomes and exomes. Reported expenditures at each center included costs associated with project management, informatics, and technical developments, as well as analysis and annotation of genomes and exomes.

Since this program included a conceptual leap, progress reports changed a bit in each reporting quarter. In November 2016, the delta between samples sequenced and the goal for each center were reported. In February 2017, reporting fields for the amount of data generated on a given platform were reported. Minor changes were made over the next few years, with the final reporting structure finalized in 2018.

A significant reduction in the cost of sequencing occurred in 2015 with release of the Illumina HiSeq X Ten (Illumina 2020) platform. While chemistry remained largely unchanged, massive gains in throughput were achieved through parallelization across many instruments. This appeared in the genome sequence cost curve before the start of CCDG reporting because the LSACs were still funded into no cost extension. The HiSeq X series was quickly superseded by the NovaSeq series.

These large-scale human genomics projects undertaken by the CCDGs were largely outlined in the grant application responses to the RFA with the specific goal of understanding common disease through high-powered human sequencing studies. After funding, the centers were organized around four project working groups: cardiovascular disease, neurodegenerative disease, metabolic disease, bone disease. Each center’s proposal was analyzed for points of collaboration across the funded sites, and those projects that could achieve sufficient power were prioritized. Knowing this, working groups were chosen based on the availability of large numbers of samples (tens and hundreds of thousands). Each project team included representatives from each center who met frequently to set goals and plan analysis in each common disease domain. This level of coordination was a return to multicenter goals in the spirit of the HGP and were enabled through frequent scientific reports and meetings.

A complete set of NHGRI sequencing progress reports from October 2001 to January 2023 are available through the NHGRI HGP Archive.2

3. Discussion

3.1. How Do Costs in Projects of the Cost Curve Compare to Actual Genome Projects?

The original projected duration and cost of the HGP was fifteen years and $3 billion. In actuality, the project was finished early, in thirteen years and under budget, at $2.7 billion, fiscal year 1991 dollars (National Human Genome Research Institute 2020). This figure includes early technology development, genetic and physical mapping, and sequencing of model organisms done as proof of concept prior to sequencing the human genome. The total cost for the HGP “working draft” genome sequence, achieved in 2000, is approximately $300 million worldwide, with roughly half ($150 million) being funded by the NIH. This amount is comparable to the costs presented in the genome sequencing cost curve (Figure 3.1). The finished sequence, completed in 2003, cost $450M or an added $150M to the working draft (National Human Genome Research Institute 2015).

It cost an estimated $13.8 million in October 2005 (see Figure 3.2) to sequence a human-sized genome of three thousand megabases to a quality comparable to the working draft produced by the HGP in 2001. This assumes that all the sequencing was done at a cost of $767 per megabase of raw sequence and incorporates assumptions regarding the technology that was in use at that time (Wetterstrand 2022) and described previously in this chapter. It is worth noting that $13.8 million is an estimate that represents a snapshot in time based on the three-month reporting period that ended in October 2005. At that time, the LSACs were working on nonhuman genomes, rather than additional human genomes. Actual genome projects occurred over months or years and experienced changes in the cost of sequencing over the length of the projects. For example, the dog genome, published in 2005 and similar in size to the human genome, cost approximately $30 million to sequence over the course of the project (National Human Genome Research Institute 2005a).

In January 2015, NHGRI estimated the cost of sequencing a human-sized genome at just under $4,000. As above, the cost estimates are based on the average cost of sequencing a raw megabase of DNA at the LSACs over a three-month period, during which time the centers were running multiple sequencing platforms to address a variety of genomic projects. As of November 2020, this genome cost is calculated to be approximately $700 (see Figure 3.1). In short, these estimates were not designed to track the number of genomes produced, but rather progress toward that theoretical ambition, the $1,000 genome.

3.2. Costs in an R&D Setting Are Different Than Costs in a Commercial Setting

There are numerous chapters of this book dedicated to NHGRI-funded projects and initiatives. While these chapters engage with work completed across the HGP, HapMap/1000 Genomes, LSAC, and CCDG projects, this chapter describes an emergent property of multiple projects. At the same time that large-scale sequencing was focused on genome production and analysis, the $1000 genome project (see chapter 6) was a technology development project. The two projects, one focused on high production and sequencing platform optimization and the other on lowering costs and breakthrough technology development, created an environment which drove costs down.

The cost curve, while rooted in the data produced (see Figure 3.5) and cost reported by a certain set of NHGRI grantees, evolved to be a project-agnostic marker of technological advancement. This served as a benchmark that would influence projects that aimed to produce genome sequence to answer large-scale questions, pointing to each threshold where new science would become economically feasible. This is seen notably in the creation of the Centers for Mendelian Genomics (National Human Genome Research Institute 2012a) and the Clinical-Sequence Exploratory-Research programs (National Human Genome Research Institute 2012b) in 2011, which required inexpensive exomes to advance genome science and medicine.

However, these numbers are specific to work completed in an NIH-funded research and development setting. These costs were not necessarily achievable in a competitive contract, as they include the specific line items important to the programmatic interests of NHGRI, described in this chapter. They also represent cost at an economy of scale that, until recently was only achievable in specific, large-scale settings. However, companies who only were contracted to provide sequence data, not “fully loaded” genomes, may produce data at a lower cost than the cost curve indicated.

The data produced for research also do not meet clinical standards, which are considerably higher. Indeed, some low coverage sequencing experiments would be unthinkable in oncology, where tumor-normal pairs can exceed 90x coverage of a genome. This is also a circular story, from a cost perspective: The low cost of these later genomes was enabled by the high cost of the first human reference genome. Decades of production and maintenance of one high-quality reference allowed for the reference-guided experiments to be undertaken with the burden sharing of that first reference genome.

Cumulative output of sequencing centers rapidly increased after 2015 from 2 billion megabase pairs and leveled in 2020 at 20 billion megabase pairs.

Figure 3.5. Cumulative sequence production in megabase pairs of DNA.

Figure Description

Figure 3.5. Cumulative sequence production in megabase pairs of DNA. Graph presents the total cumulative output of the NHGRI-funded grantee sequencing centers in megabases pairs (millions of base pairs). Graph shows a sigmoid-like function with a rapid increase in production after 2015 (two billion megabase pairs) and a leveling in 2020 at twenty billion megabase pairs. Illustrative how the majority of data was produced in just a few years after the highest-throughput sequencing platforms were released in 2007 and 2008.

Date

x-axis

Cumulative Megabasepair Production

y-axis

September 2001

16,200

March 2002

32,312

September 2002

51,045

March 2003

74,091

October 2003

112,435

January 2004

130,297

April 2004

151,542

July 2004

173,774

October 2004

199,170

January 2005

224,069

April 2005

249,892

July 2005

279,677

October 2005

310,508

January 2006

340,472

April 2006

373,492

July 2006

406,358

October 2006

440,104

January 2007

472,818

April 2007

502,953

July 2007

532,879

October 2007

563,243

January 2008

656,230

April 2008

1,003,802

July 2008

1,686,609

October 2008

3,265,767

January 2009

5,235,679

April 2009

8,930,534

July 2009

18,575,583

October 2009

29,068,541

January 2010

46,853,123

April 2010

74,396,129

July 2010

113,832,536

October 2010

164,749,992

January 2011

226,504,436

April 2011

299,424,486

July 2011

385,974,524

October 2011

490,761,545

January 2012

575,500,600

April 2012

639,074,158

July 2012

696,741,226

October 2012

790,202,481

January 2013

887,601,392

April 2013

987,356,702

July 2013

1,077,135,829

October 2013

1,222,909,263

January 2014

1,407,011,631

April 2014

1,525,699,638

July 2014

1,637,958,679

October 2014

1,693,702,193

January 2015

1,797,500,940

April 2015

1,923,444,291

July 2015

2,603,795,479

October 2015

3,018,307,365

May 2016

3,934,031,798

August 2016

4,254,972,399

November 2016

5,103,142,914

February 2017

6,057,654,783

May 2017

6,896,660,152

August 2017

7,640,062,863

November 2017

8,194,691,430

February 2018

8,602,458,886

May 2018

9,127,055,872

August 2018

9,729,348,639

November 2018

10,752,910,251

February 2019

11,249,791,541

May 2019

12,161,707,783

August 2019

13,339,931,709

November 2019

14,972,527,311

February 2020

16,528,535,170

May 2020

17,346,372,404

August 2020

18,068,930,638

November 2020

19,251,317,674

February 2021

19,582,271,189

May 2021

19,701,430,086

August 2021

19,762,346,439

November 2021

19,918,879,164

February 2022

19,920,932,632

May 2022

19,922,986,100

3.3. How Cost Represents the Success of the Field and Organized Cooperation

It is an impossible task to catalogue the number of times that the genome sequence cost curve on NHGRI’s website has been screenshotted for an awardee’s acceptance speech (Shendure 2012), a National Academies report (National Research Council Committee 2011), or by a news outlet (Economist 2020). On one hand, it is conjecture to state that the genome sequence cost curve has become a meme; on the other hand, one need only to attend an American Society for Human Genetics (ASHG) annual meeting to see it presented at least once by someone outside of the NHGRI progress reporting team. Alternatively, a reverse image search on Google will yield enough results to satisfy a skeptic.

Broadly, the speed that cost diminished is an accomplishment in itself, which can be seen simply in a graph. In another sense, what the decrease of costs represents was a success in that it allowed NHGRI program staff to organize consortia of disparate players to collectively tackle large-scale research questions in biology. The consistent drumbeat of progress reporting was a powerful tool in managing the changing landscape of technologies, grantee institutions, and researchers involved in NHGRI-funded genome sequencing activities over the last twenty years. The reporting of production numbers showcased not only the success of new sequencing technologies, but also the extraordinary improvements to the sequencing pipeline made by the sequencing centers implementing those technologies. With each quarter’s report, colleagues could appreciate the incremental progress of the sequencing field abstracted from the discrete work products that they had achieved in the regular course of their work.

3.4. What Is the Future of Cost Tracking?

As explored in this paper, cost tracking was a distinctive feature of programmatic management from the HGP through large-scale sequencing at NHGRI. The size of these projects in terms of financial expenditure, persons involved, and data generated required a high level of oversight for coordination. The current nature of projects at NHGRI, which are now more decentralized across many research groups, as well as the vast differences in goals now means that data collected in progress reports are very different from those collected at large-scale production centers. The fact the NHGRI can fund projects in so many realms of genomics is also a function of the relative cheapness of genome data production. Essentially, the barriers to entry are down.

While the next order of magnitude drop from the $1000 genome to the $100 genome will be important for science, there are many disparate players demanding genomes, and so many technology companies working to advance the science, that the ecosystem can no longer be fit to a single genome sequence cost curve. One could imagine a genome sequence cost curve for each platform, maintained by the company who produces that platform. However, even something so simple as a cost curve obscures differences in genomic approaches to research questions. The $1000 genome is really the $1000 alignment, where for around $1000 in 2020, one can generate enough short-read data on an Illumina platform to accurately call bases of a human genome of interest. However, that genome must be aligned to the human reference genome sequence (Genome Analysis Toolkit 2020). Data generated on other platforms that produce longer reads cost more than $1000, but the value of the data is fundamentally different. The possibilities for de novo assembly of a new human genome increase with sequence read length and higher accuracy.

It is also likely that the genome sequence ecosystem will thrive with data generated from multiple platforms, for example, with short-read Illumina data adding accuracy or markers for phasing to already accurate Pacific Biosciences-generated HiFi sequences, which can then be assembled (Nurk et al. 2020) and compared to longer reads or maps from ONT or Bionano Genomics. The complexity of this work product, and the costs of generating each data type for one experimental question, requires a paradigm shift in thinking about the cost of a genome, versus the cost to understand the full context of the sequence and positional data for a given genome. Unless there is a programmatic need to track these developments in a large-scale way, it is unlikely that a single cost-curve equivalent will emerge for this genomic ecosystem.

4. Conclusions

Genomics has evolved before and will evolve again. The precipitous drop in genome sequencing costs in the mid-2000s captured one key development: short-read sequencing by synthesis approach of NGS could enable exciting new experiments at a lower cost and higher throughput than capillary-based Sanger sequencing. At the same time that the $1000 genome was achieved, long-read sequence generation platforms matured to the point where their higher cost data products were adding new value. The cost curve does not account for this, a fact that hints at a future with many different cost curves, which show the diminishing cost of a multi-platform reference-quality genome, or a medical genome, or a personal genome.

The NHGRI sequence cost curve is likely one of the most quoted figures produced by the institute, where it is used to show the power and pace of technological development. This was an unintended consequence but underscores the need for communication of science. Whether a program manager or public consumer, one does not need to know that exomes and organismal genomes were produced to know that the cost of sequence was dropping fast. How, and by whom, these facets of technological progress are measured in the future is presently undecided. However, the ubiquity of the sequence cost curve shows our need to visualize progress, and the positive external benefits that can come from careful progress reporting in program management.

Notes

  1. 1. For 1000 Genomes Project, see https://www.internationalgenome.org/.

  2. 2. See https://archive.genome.gov/.

References

  • Committee on Appropriations Subcommittee on Departments of Labor, Health and Human Services, Education, and Related Agencies. 1986. Departments of Labor, Health and Human Services, Education, and Related Agencies Appropriations for 1987. US Government Printing Office.
  • Economist. 2020. “Genomics Took a Long Time to Fulfil Its Promise.” The Economist. March 12.
  • Genome Analysis Toolkit. 2020. “Reference Genome.” https://gatk.broadinstitute.org/hc/en-us/articles/360035891071-Reference-genome.
  • Gordon, J. I., R. E. Ley, R. Wilson, et al. 2006. “Extending Our View of Self: The Human Gut Microbiome Initiative (HGMI).” https://www.genome.gov/Pages/Research/Sequencing/SeqProposals/HGMISeq.pdf.
  • Green, P., and B. Ewing. 1998. “Phred—Quality Base Calling.” Accessed January 29, 2026. https://www.phrap.com/phred/.
  • Heather, J. M., and B. Chain. 2016. “The Sequence of Sequencers: The History of Sequencing DNA.” Genomics 107 (1): 1–8.
  • Illumina. 2011. “Quality Scores for Next-Generation Sequencing.” https://www.illumina.com/documents/products/technotes/technote_Q-Scores.pdf.
  • Illumina. 2020. “Whole-Genome Sequencing Power: Discover How HiSeq X Ten Breaks the $1000 Genome Barrier for Human Whole-Genome Sequencing.” https://www.illumina.com/systems/sequencing-platforms/hiseq-x.html.
  • Land, M. L., D. Hyatt, S.-R. Jun, et al. 2014. “Quality Scores for 32,000 Genomes.” Standards in Genomic Sciences 9 (1): 20.
  • Laver, T., J. Harrison, P. A. O’Neill, et al. 2015. “Assessing the Performance of the Oxford Nanopore Technologies MinION.” Biomolecular Detection and Quantification 3:1–8.
  • Majd, S., and R. S. Pindyck. “The Learning Curve and Optimal Production Under Uncertainty.” Working Paper No. 2423. National Bureau of Economic Research, October 1987.
  • Metzker, M. L. 2010. “Sequencing Technologies—The Next Generation.” Nature Reviews Genetics 11 (1): 31–46.
  • National Human Genome Research Institute. 2002. “NHGRI Prioritizes Next Organisms to Sequence.” https://www.genome.gov/10002851/2002-release-new-organism-sequencing-priorities.
  • National Human Genome Research Institute. 2003a. “International Consortium Completes Human Genome Project.” https://www.genome.gov/11006929/2003-release-international-consortium-completes-hgp.
  • National Human Genome Research Institute. 2003b. “Large Scale Sequencing capacity (RFA-HG-03–002).” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-03-002.html.
  • National Human Genome Research Institute. 2005a. “Genome Sequencing Centers (RFA-HG-06–001).” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-06-001.html.
  • National Human Genome Research Institute. 2005b. “Researchers Publish Dog Genome Sequence.” https://www.genome.gov/17515860/2005-release-researchers-publish-dog-genome-sequence.
  • National Human Genome Research Institute. 2006. “Pathogens and Vectors.” https://www.genome.gov/26525388/pathogens-and-vectors.
  • National Human Genome Research Institute. 2010. “Genome Sequencing and Analysis Centers (RFA-HG-10–015).” https://grants.nih.gov/grants/guide/rfa-files/rfa-hg-10-015.html.
  • National Human Genome Research Institute. 2012a. “Centers for Mendelian Genomics.” https://www.genome.gov/Funded-Programs-Projects/NHGRI-Genome-Sequencing-Program/Centers-for-Mendelian-Genomics-CMG.
  • National Human Genome Research Institute. 2012b. “Clinical Sequencing Evidence-Generating Research (CSER).” https://www.genome.gov/Funded-Programs-Projects/Clinical-Sequencing-Evidence-Generating-Research-CSER2.
  • National Human Genome Research Institute. 2015a. “Centers for Common Disease Genomics (RFA-HG-15–001).” https://grants.nih.gov/grants/guide/rfa-files/RFA-HG-15-001.html.
  • National Human Genome Research Institute. 2015b. “The Cost of Sequencing a Human Genome.” Retrieved November 23, 2020. https://www.genome.gov/about-genomics/fact-sheets/Sequencing-Human-Genome-cost.
  • National Human Genome Research Institute. 2020. “Human Genome Project FAQ.” Retrieved November 23, 2020. https://www.genome.gov/human-genome-project/Completion-FAQ.
  • National Institutes of Health. 2020. “NIH Cooperative Agreement Definition.” Retrieved January 10, 2020. https://grants.nih.gov/grants/glossary.htm#CooperativeAgreement.
  • National Institutes of Health Office of Management. 2017. “Indirect Cost: Definition and Example.” https://oamp.od.nih.gov/dfas/indirect-cost-branch/indirect-cost-submission/indirect-cost-definition-and-example#:~:text=Indirect%20Costs%20(definition%20extracted%20from,treatment%20as%20a%20direct%20cost.
  • National Research Council. 1988. Mapping and Sequencing the Human Genome. The National Academies Press.
  • National Research Council. 2011. Toward Precision Medicine: Building a Knowledge Network for Biomedical Research and a New Taxonomy of Disease. The National Academies Press.
  • Nurk, S., B. P. Walenz, A. Rhie, et al. 2020. “HiCanu: Accurate Assembly of Segmental Duplications, Satellites, and Allelic Variants from High-Fidelity Long Reads.” Genome Research 30 (9): 1291–305.
  • Proctor, L. M. 2011. “The Human Microbiome Project in 2011 and Beyond.” Cell Host & Microbe 10 (4): 287–91.
  • Shendure, J. 2012. “2012 Curt Stern Award Address.” https://slideplayer.com/slide/15281967/.
  • Sinsheimer, R. 1990. Oral history interview with Dr. Robert L. Sinsheimer, by Charlotte E. (Shelley) Erwin, Caltech Archives Oral History Project, May 30, 1990, March 26, 1991, http://resolver.caltech.edu/CaltechOH:OH_Sinsheimer_R.
  • Stratton, M. 2008. “Genome Resequencing and Genetic Variation.” Nature Biotechnology 26 (1): 65–66.
  • US Congress, Office of Technology Assessment. 1986. Technologies for Detecting Heritable Mutations in Human Beings OTA-H-29. US Government Printing Office.
  • US Congress, Office of Technology Assessment. 1988. Mapping Our Genes—Genome Projects: How Big? How Fast? US Government Printing Office.
  • Wetterstrand, K. A. 2020. “DNA Sequencing Costs: Data from the NHGRI Genome Sequencing Program (GSP).” Retrieved June 19, 2022. https://www.genome.gov/about-genomics/fact-sheets/DNA-Sequencing-Costs-Data.

Annotate

Next Chapter
4 History of the Encyclopedia of DNA Elements (ENCODE) Project
PreviousNext
Copyright 2026 by the Regents of the University of Minnesota

All rights reserved.
Powered by Manifold Scholarship. Learn more at
Opens in new tab or windowmanifoldapp.org