Notes
7 The Nature of Genomic Publishing
Chris Gunter and Magdalena Skipper
1. A Virtuous Cycle
We have each been fortunate to serve as the senior editor at the journal Nature responsible for handling research papers in genetics and genomics: CG from 2002 to 2008, and MS from 2008 to 2015. This position is unlike most editing roles because the relationship between the journal and the field has been incredibly dynamic and synergistic, with a cycle of demands on one side driving community action on the other. To the best of our knowledge, no other community has shown such a willingness en masse to select or shun a journal for its publication policies. In return, in part to keep this community and its position at one of the most highly cited multidisciplinary scientific journals in the world, Nature in particular has modified and created new publication policies based on the demands for openness and for new ways of publishing and displaying vast amounts of genomic data and analysis. In this chapter, we will review the interactions between the genomics community and editors starting after the publication of the 2001 draft human genome paper (IHGS Consortium 2001), and discuss how these interactions changed scientific publishing.
2. The Role of an Editor
Nature’s mission was established in 1869; by the time the 2001 draft human genome paper was being considered, it had minimally changed to the following: “First, to serve scientists through prompt publication of significant advances in any branch of science, and to provide a forum for the reporting and discussion of news and issues concerning science. Second, to ensure that the results of science are rapidly disseminated to the public throughout the world, in a fashion that conveys their significance for knowledge, culture and daily life.”1
One of the most common questions we hear is how, and to what extent, editors are involved in shaping publications. In our experience, it can be anything from minor to instrumental. First, there are our interactions with the authors around the papers themselves.
For many papers, there is little to no contact between the authors and the editors before the paper is submitted to a journal, and the main way in which editors influence these papers is by judiciously selecting peer reviewers. For some papers, the editors may hear about the work through a conference or a lab visit and ask the authors to submit the work for consideration. And then there are genome/“big data” papers, which are almost always an output of large collaborative efforts; in these cases, editors may be working with the authors for years in advance.
There is often a discussion between the journal and the authors on what features might accompany the main paper. These may be developed on the authors’ side: companion papers, commentaries, additional pages/figures/tables/references that go beyond the established format requirements, online material, posters, papers submitted to other journals that should be coordinated to come out at the same time if possible. Some features may be arranged by the journal, including covers, special sections, fold-outs, editorials, news features, press conferences, new online material/formats, or coverage within other journals published by the same publisher. Our experiences with genome papers have included all of the above and have served as a wonderful testing ground for policies and ideas that have been rolled out across journal groups, imprints, publishers, or even across publishing as a result.
One small but important example was the allowance for multiple equal contributing authors (MGS Consortium 2002). The large group of authors on this paper wanted to indicate, using daggers following some of the authors’ names, that a subset of the group had contributed to project leadership. Such an indication was not in the style guide used by Nature’s subeditors at the time. There were heated transatlantic discussions between the primarily US-based authors and UK-based Nature subediting and production team about the daggers. In the end, there were no daggers added; instead, the end of the paper included a note of “Authors’ contributions” indicating “The following authors contributed to project leadership. . . .”
While a form of individual author contribution statement predates the publication of Human Genome Project (Policy on Papers’ Contributors 1999), the rise of multi-author, collaborative papers has prompted the journal to formalize author contributions statements further and make them mandatory across all papers regardless of discipline in 2009 (Authorship policies 2009). This step of giving due credit to every contributing author regardless of where in the author list their name appears can be thought of as a legacy of these early genomics papers.
Since 2002, consortia-authored and “big data” papers have become much more common (previously, they would be almost exclusively encountered in high-energy physics and astronomy), and there has been a movement to indicate credit for authors through daggers, asterisks, ORCiD badges, and other means. Nature Research journals encourage transparency in detailing relative author contributions, and current policy allows for up to three corresponding authors and up to six “equally contributing authors” or who “jointly supervised the work.”2
3. Data Availability
3.1. Microarrays as an Example of Community-Led Changes
Another way in which editors influence scientific publications is by setting standards, in concert with the community. Given the rise of experimental microarrays (glass slides with spots of DNA where millions of genomic variants could be tested at once) by 2002, it became clear that we needed a policy on how to handle papers that reported and used microarray data. As we ultimately stated in an editorial, “Variables in every step of the experiment often make cross-paper comparison virtually impossible” (Microarray Standards at Last 2002). We wanted to ensure that the proper data were included for both reviewers to judge the original paper and for other groups to replicate the analysis or use the data in meta-analyses to ask new questions. Although an international community scientific group called the Microarray Gene Expression Data group had published an open letter outlining recommended minimal guidelines in late 2001 (Brazma et al. 2001), it took months of discussions with community leaders and editorial colleagues for CG to be able to get this policy approved by almost all parties. Nature published a 2002 editorial spelling out standards for both reporting and submitting papers with microarray data:
Harried editors can rejoice that, at last, the community is taming the unruly beast that is microarray information. Therefore, all submissions to Nature and the Nature family of journals received on or after 1 December containing new microarray experiments must include the mailing of five compact disks to the editor. These disks should include the necessary information compliant with the MIAME standard. The information must be supplied in a format that could be read by widely available software packages. Data integral to the paper’s conclusions should be submitted to ArrayExpress or GEO databases, with accession numbers where available, supplied at or before acceptance for publication.3
Read in the light of 2019, the mailing of CDs seems terribly quaint. But, as the editorial stated, editors at Nature and elsewhere had been relying on authors to make data available on private websites during review, until we all realized that the authors could be and were tracking the IP addresses of visitors and therefore getting a good idea of whom the referees were. [In 2019, and beyond, there is increasing interest in open peer review without reviewer anonymity, but at the time this was not very common.] For the 2008 paper reporting the complete genome sequence of one individual, James Watson (Wheeler et al. 2008), we had to move up to hard drives, requiring the authors to send in five of them containing the genomic data.
Since then, reporting standards for data, code, materials used, etc., have increased considerably.4 Data availability statements, separate from the methods section, have become mandatory on all papers and formal citation of data (via separate DOIs) is encouraged both for transparency and to aid recognition of data generators’ contributions.5 For reference genome papers, assembly information must be made available in addition to the sequence reads. Large genomic datasets are expected to be made available in public repositories, for example, while data types for which community-accepted standards do not exist are expected to be deposited in unstructured repositories such as figshare, Zenodo, or Dryad. In the world in which discs containing data are no longer shipped around the globe, authors are encouraged to make use of the option to host their data before peer-review in private, password-protected records offered by many repositories (e.g., GEO, figshare, or Dryad offer this service), access to which is offered confidentially to the reviewers.
As genomic information moved into genotype-phenotype association era and as some studies have begun to focus on vulnerable human populations, access to data has—with good reason—become more sensitive and complex. Accordingly, Nature and other journals have developed policies for allowing data to be deposited in controlled access repositories, such as UKDA.6 See also Respectful Re-Use (2012).
3.2. Data Release
In 2003, the Wellcome Trust sponsored an “International Data Release Meeting” in Fort Lauderdale, which will be covered at length in other chapters in this book. Decisions taken at this meeting obviously had large ramifications for how we handled genomics and other papers, so Nature was one of four journals invited to participate in the otherwise closed meeting.
Discussion in advance focused on “publication etiquette,” including a straw man proposal put forward by Ewan Birney. He suggested that funders could ask sequencing project teams to publish smaller papers that provided contact information about large sequencing projects, along with a schedule for data generation and final publication [CG, personal communication]. These “marker papers” would provide a reference in the literature, which could be cited by others who used the data before the data producers published their own paper. CG and Phillip Campbell, the editor in chief of Nature at the time, agreed that Nature would not be the proper venue for such papers, but we felt that other journals in the Nature Publishing Group (NPG) might be. We were also concerned that this proposal would put public projects at a disadvantage and make it more likely that they could be scooped by outside efforts.
Nature had stated its policy in a February 8, 2001 editorial titled “Handling (mis?)appropriated data.” Specifically, Nature stated that it had “decided to adopt the following practice”:
Appropriation of uncredited data will not prevent us from sending a paper out for prompt review. But we will require written assurance that authors are not violating any originators’ data-licensing agreement. We will encourage our referees to be alert to the use of appropriated unpublished data from databases. Where there are concerns over credit, we will usually seek advice from an originator of the data in addition to the usual refereeing process. We would not be giving originators a veto: where disagreements arise, we will use our judgement, having consulted referees over technical considerations if necessary, and will usually insist on an acknowledgment as a condition of publication.7
Chris Gunter’s opening remarks for the panel addressed many of the tensions we saw in publishing at the time:
We of course insist on release of data at publication, and in genomics we ask for it at the time of review. I often tell authors literally that their papers will sit on my desk until I receive the necessary data for referees to review them. However, the data are still only for the referees at that point. I have had referees refuse to review the paper because the data were not deposited yet. I also hear frequently from sequencers that they are afraid their primary data will be scooped if it’s deposited, and I see an increasing number of genomics papers fail at Nature. This means anyone who did want to scoop the primary sequencers would have more time to do so. In general, we prefer not to be gorillas that enforce unwanted restrictions on the community, but instead prefer to listen to the community and act accordingly.8
Ultimately, the agreement published from Fort Lauderdale called for free and open data access, required by the funding agencies, for genomic projects.9 This did require that Nature and other journals change their policies, to stop requiring letters from the data generators confirming that it was okay to use their data in a paper. In an effort to protect scientists who produced large datasets, the report includes the much-debated sentence, “There should be no restrictions on the use of the data, but the best interests of the community are served when all act responsibly to promote the highest standards of respect for the scientific contribution of others.”
This “respect” would be tested fairly quickly: After the publication of the draft human genome, each publicly funded group that had focused on sequencing and finishing one human chromosome wanted to publish their work, including more biological insight than could fit into any overall human genome paper. Nature felt that these papers were sufficient markers to commit to considering each of them separately, as always, with no guarantee of publication unless the papers passed technical review. In the case of chromosome 7, the publicly funded group was in the process of completing their analyses and draft paper (Hillier et al. 2003) when Science published a chromosome 7 sequence and analysis paper from another group (Scherer et al. 2003). The data in the Science paper were ~85 percent from the Celera (private) human genome project, and ~15 percent from the public human genome project.10 It appeared that this paper was timed to be published before the announcement by the international consortium that they had “completed” the Human Genome Project.11
Of course, under the Fort Lauderdale rules, the public data were free and openly available through NIH and other databases. But vocal members of the genomics community felt outraged at this lack of respect for the original data producers, and angrily recommitted to send their papers to Nature in the future. This felt like an aftershock of the original earthquake that was the Science publication of the Celera human genome paper (Venter et al. 2001). It also served as another example (to us, the editors) of why listening to the community and paying attention to their desired standards was crucial for maintaining their goodwill and for offering them the service and support they could trust.
3.3. Publication Release
For many years, genome papers at Nature and other NPG titles have enjoyed a special status when it comes to publication release. Since 2001 and the draft human genome publication, the editorial and publishing team took the decision that genome papers were to be published not behind a paywall but freely available (sometimes with website registration) at Nature, in recognition of “a consistent character of ‘genome’ papers: they represent the completion of a key and fundamental research resource, describing and reflecting on what has been revealed but not usually providing insights into mechanism” (Shared Genomes 2007). In 2007, NPG decided to not only make these papers freely available but also publish them under a Creative Commons license, or a public copyrighted license that allows creators to share their work under certain conditions, stating:12 “In 1996, as human genome sequencing was getting under way, leading players stated: ‘It was agreed that all human genomic sequence information, generated by centres funded for large-scale human sequencing, should be freely available and in the public domain in order to encourage research and development and to maximise its benefit to society.’13 These principles have continued to guide the field, and NPG has consistently made genome papers freely available in keeping with them. This new licence allows us to formalize the arrangement.”14
More recently, the use of CC-BY license15 for papers in Nature and other journals with Nature in their title has been extended to include papers that describe reporting and experimental standards, consensus statements and white papers, and community experiments to compare the performance of software tools, as well as papers addressing important public health needs.16
4. Changing the Definition of a “Paper”
Genomics is one of those disciplines that can be said to have outgrown the standard format of a scientific paper. While there is already a wealth of information in an individual genome, these days single genome analyses are rare. At a minimum, studies tend to include transcriptome, proteome, metabolome, and increasingly other -omic data; more frequently, multiple genomes and other -omes—from individuals or cells or tissues or cancers—are generated and analyzed in tandem.
A growing tendency to seamlessly integrate data and computational code17 into such analyses is blurring the boundaries between a paper and the so-called research objects.18 But even within the limits of individual papers, genomics papers have pushed the limits of journal format. These papers frequently are longer than papers in many other disciplines, a trend that absolutely extends to the length of their supplementary information (SI). A good example—and a result of a close working relationship between the authors and the editor—is provided by the main paper from the 1000 Genomes Project (1000 Genomes Project Consortium et al. 2015). The SI document runs to 124 pages, has its own table of contents, and reads like a book that tells a story of so much more than an ordinary SI file would.19 For many years, MS used this SI document as an example for other editors at Nature and beyond, who in turn shared it with their authors to illustrate how to make SI informative, well-organized, and engaging. That said, SI documents are frequently unwieldy, raising the issue of how much journals can reasonably ask reviewers to read in the short review period (particularly when SI is not well organized) and how many methodological details must be included in the main paper versus in the SI.
In response to many of these concerns, in 2013, Nature introduced a policy of placing all supplementary data within Extended Data (Nature Papers Enhanced 2013), providing “the online reader with immediate access to many display items (figures and tables) previously buried in the Supplementary Information PDF.” However, the complexity and richness of large, consortium-led efforts has meant that following numerous internal policy discussions led by MS the journal continues to exempt them from this rule. See Lek et al. (2016), specifically the SI.20 They far exceed the guidelines stating, “Extended Data will not normally contain more than ten individual display items (figures and tables) in addition to the limits set for the printed version of the paper (typically four and five display items for Letters and Articles, respectively)” (Nature Papers Enhanced 2013).
Other large consortium-led projects were even more extensive and can be seen as examples that truly broke through the limits of individual papers. Examples include the second phase of the ENCODE project or ENCODE-2 (published in 2012 in Nature, other journals published by NPG, and journals from other publishers—all at once), as well as the 2013 TCGA pan-cancer analysis published mainly in Nature Genetics (The Cancer Genome Atlas Research Network et al. 2013).
For the ENCODE-2 papers (one example being Dunham et al. 2012), MS and editors from other journals worked together with the large consortium of authors to devise a new way to navigate the vast array of data and analyses generated by the project. Nature’s ENCODE Explorer, an interactive and specially designed website published simultaneously with the papers, was a first of its kind. Its introduction read:
You can discover the project’s results in the conventional way, by visiting individual papers, or you can engage with the material by following “Threads,” each one dedicated to a theme discussed in more than one paper. Threads lie at the heart of the Nature ENCODE explorer. They are a new way in which to explore the wealth of information collectively described by the 30 papers published across three different journals: Nature, Genome Research, and Genome Biology. They complement the papers by highlighting and bringing together topics that are otherwise covered only in subsections of individual papers. Each Thread consists of relevant paragraphs, figures and tables from across the papers, united around a specific theme.21
Examples of ENCODE threads can be seen in Nature collections.22 TCGA can also be seen in Nature collections.23
MS recalls numerous calls with representatives of the consortium, followed by her visits to the European Bioinformatics Institute at Hinxton and reciprocal visits from lead authors Ewan Birney and Ian Dunham to the Nature’s offices in Crinan Street, to flesh out how exactly this concept could be put into practice. Once the concept was agreed upon in principle, there remained the non-trivial task of convincing Nature’s publisher and production teams that they should devise a way to implement the “threads” online, in an interactive way. This had never been done before, neither by Nature nor by anyone else.
One particularly lengthy and heated session around a table in a meeting room overlooking Regent’s Canal in the London office involved an in-depth presentation and a Q&A with Ewan Birney, Ian Dunham, and the whole production team of Nature. Much to MS’s and the visitors’ delight, the efforts paid off and the way forward toward timely online implementation was agreed. The final design of what came to be known as the ENCODE Explorer was a result of a collaboration between Nature’s art department and Max Gadney, a graphic designer from Made By Pi hired for this purpose.
As the new threads were composed of display items of sections from individual papers, and these were distributed across journals in three publishing houses, their preparation for publication required MS (at Nature), Clare Garvey (chief editor of Genome Biology, BMC), and Hillary Sussman (chief editor of Genome Research, CSHL Press) to share accepted manuscripts pre-publication, with the authors’ permission. This was unprecedented and made possible by the professionalism and the collaborative spirit with which the genomics editors worked.
The growing volume of genomic information and its publishing was coincident with advances in animation and online visual representation, making it possible to reach broad non-specialist audiences with information that was ultimately about them and for them, not least in the form of rapidly growing precision medicine. The 2012 ENCODE papers were accompanied by an animation “The Story of You,” which looked at the history of genetics and genomics.24 Another example of this type of outreach was the video accompanying the 2015 Epigenome Roadmap (Roadmap Epigenomics Consortium et al. 2015), a suite of papers funded through the NIH Roadmap Epigenomics Program.25 The story board for both of these was devised through close collaboration between MS, the handling editor, and the Nature multimedia team, led by Charlotte Stoddard.
In the last few years, we have seen an increasing trend for accompanying materials being produced by the scientists themselves, for example, by consortia creating detailed online FAQ documents to go along with their papers [one paper was created for Karlsson Linnér et al. (2018)]26 or individual scientists creating “tweetorials” by threading together a number of tweets explaining new papers as they are published.
5. The Future
From our current positions at the National Human Genome Research Institute (CG) and back at Nature (MS), we look forward to continuing to work with the genomics and other communities as they push the boundaries of scientific publishing. Genomic data are no longer a domain of just one community; by now, almost every aspect of biomedical sciences has been touched by the legacy of this field, either in the form of very large data and their analysis or by the collaborative, consortium-led style of working. New data visualizations are the most likely exciting development, as is the continuing desire for open access of data and publications.
Notes
1. Nature’s mission, see https://www.nature.com/nature/about/.
2. Nature Research journal editorial policies, https://www.nature.com/nature-research/editorial-policies/authorship.
3. See Nature 419:323, September 26, 2002.
4. See https://www.nature.com/nature-research/editorial-policies/reporting-standards.
5. See https://www.nature.com/news/announcement-where-are-the-data-1.20541.
7. See Nature 409:649, February 8, 2001.
8. See Personal notes from teleconference call by C. Gunter, January 2003.
9. See https://www.sanger.ac.uk/legal/assets/fortlauderdalereport.pdf.
10. See http://www.genomenewsnetwork.org/articles/07_03/chrom7.shtml.
11. See https://www.genome.gov/11006929/2003-release-international-consortium-completes-hgp.
12. See http://en.wikipedia.org/wiki/Creative_Commons_license.
13. See http://www.ornl.gov/sci/techresources/Human_Genome/research/bermuda.shtml.
14. See Nature 450 (7171):762, 2007.
15. CC-BY license, https://creativecommons.org/share-your-work/licensing-examples/
16. See https://www.nature.com/nature-research/editorial-policies/self-archiving-and-license-to-publish#creative-commons-licences.
17. See http://blogs.nature.com/ofschemesandmemes/2018/08/01/nature-research-journals-trial-new-tools-to-enhance-code-peer-review-and-publication.
19. Supplementary information, https://static-content.springer.com/esm/art%3A10.1038%2Fnature15393/MediaObjects/41586_2015_BFnature15393_MOESM86_ESM.pdf.
21. See https://www.nature.com/encode/about/nature-encode-explorer.
22. Nature collections, https://www.nature.com/collections/aghcdefffg/.
23. Nature collections, https://www.nature.com/collections/ffebghgeba/.
24. “The Story of You,” https://youtu.be/TwXXgEz9o4w.
26. See https://8a649e3c-e96e-4fce-86bb-b117a3d1270f.filesusr.com/ugd/2f9665_7da965f733dd4bb8bfca009743c9737e.pdf.
References
- Authorship policies. 2009. Nature 458:1078.
- Brazma, A., P. Hingamp, J. Quackenbush, et al. 2001. “Minimum Information About a Microarray Experiment (MIAME)—Toward Standards for Microarray Data.” Nature Genetics 29 (4): 365–71.
- The Cancer Genome Atlas Research Network, J. N. Weinstein, E. A. Collisson, et al. 2013. “The Cancer Genome Atlas Pan-Cancer Analysis Project.” Nature Genetics 45 (10): 1113–20.
- Consortium, IHGS. 2001. “Initial Sequencing and Analysis of the Human Genome.” Nature 409:860–921.
- Consortium, MGS. 2002. “Initial Sequencing and Comparative Analysis of the Mouse Genome.” Nature 420:520–62.
- Dunham, I., A. Kundaje, S. F. Aldred, et al. 2012. “An Integrated Encyclopedia of DNA Elements in the Human Genome.” Nature 489:57–74.
- Hillier, L. D. W., R. S. Fulton, L. A. Fulton, et al. 2003. “The DNA Sequence of Human Chromosome 7.” Nature 424:157–64.
- Karlsson Linnér, R., P. Biroli, E. Kong, et al. 2019. “Genome-Wide Association Analyses of Risk Tolerance and Risky Behaviors in over 1 Million Individuals Identify Hundreds of Loci and Shared Genetic Influences.” Nature Genetics 51:245–57.
- Lek, M., K. J. Karczewski, E. V. Minikel, et al. 2016. “Analysis of Protein-Coding Genetic Variation in 60,706 Humans.” Nature 536:285–91.
- Microarray Standards at Last. 2002. Nature 419:323.
- Nature Papers Enhanced. 2013. Nature 498:6.
- 1000 Genomes Project Consortium, A. Auton, L. D. Brooks, R. M. Durbin, E. P. Garrison, H. M. Kang, et al. 2015. “A Global Reference for Human Genetic Variation.” Nature 526:68–74.
- Policy on Papers’ Contributors. 1999. Nature 399:393.
- Respectful Re-Use. 2012. Nature Genetics 44 (10): 1073.
- Roadmap Epigenomics Consortium, A. Kundaje, W. Meuleman, et al. 2015. “Integrative Analysis of 111 Reference Human Epigenomes.” Nature 518:317–29.
- Scherer, S. W., J. Cheung, J. R. MacDonald, et al. 2003. “Human Chromosome 7: DNA Sequence and Biology.” Science 300:767–72.
- Shared Genomes. 2007. Nature 450:762.
- Venter, J. C., A. D. Adams, E. W. Myers, et al. 2001. “The Sequence of the Human Genome.” Science 291:1304–51.
- Wheeler, D. A., M. Srinivasan, M. Egholm, et al. 2008. “The Complete Genome of an Individual by Massively Parallel DNA Sequencing.” Nature 452:872–76.