What is DNA Sequencing?

DNA sequencing is the laboratory technique that determines the exact order of the four nucleotide bases—adenine (A), guanine (G), cytosine (C), and thymine (T)—in a DNA molecule. This sequence acts as the genetic code that dictates cellular functions and encodes an organism’s inherited traits. By deciphering this code, researchers can compare genetic material across species, trace evolutionary relationships, and identify unique genetic markers that define a species.

In taxonomy, DNA sequencing offers a robust, objective method for species delimitation. Unlike physical characteristics, which can vary with environment or developmental stage, DNA sequences are stable and heritable. This makes sequencing a powerful tool for resolving classification ambiguities, especially for organisms with limited morphological features, such as microorganisms or deep-sea species. The ability to sequence DNA from minute samples—a single hair, a drop of blood, or a soil particle—has opened up new frontiers in biodiversity research.

Types of DNA Sequencing Technologies

Over the past several decades, a suite of sequencing technologies has emerged, each with distinct strengths and applications. The choice of technology depends on project scale, required accuracy, and the nature of the genetic material under study.

Sanger Sequencing

Developed by Frederick Sanger in the 1970s, Sanger sequencing remains the gold standard for accuracy, routinely achieving per‑base accuracy exceeding 99.9%. It relies on chain‑termination using dideoxynucleotides (ddNTPs) labeled with fluorescent dyes. While ideal for sequencing short DNA fragments—such as individual genes or small genomic regions—Sanger is relatively low‑throughput and costly for large‑scale projects. Nevertheless, it is indispensable for validating results from other platforms and for targeted species‑discovery applications, such as barcoding the mitochondrial COI (cytochrome c oxidase I) gene in animals. The International Barcode of Life (iBOL) initiative has built extensive COI reference libraries using Sanger sequencing, forming the backbone of many identification systems.

Next‑Generation Sequencing (NGS)

Next‑generation sequencing (also called high‑throughput sequencing) revolutionized genomics by enabling massively parallel sequencing of millions of DNA fragments simultaneously. Major platforms include Illumina (Solexa), Ion Torrent, and SOLiD. NGS generates short reads (typically 50–300 base pairs) but at a dramatically lower cost per base compared to Sanger. This makes it ideal for whole‑genome sequencing, transcriptomics, and metagenomics. For species discovery, NGS allows researchers to sequence entire mitochondrial genomes, capture target genes from environmental samples, or perform shotgun sequencing of complex microbial communities. Projects like the Earth BioGenome Project leverage NGS to sequence all eukaryotic species on Earth, a goal that requires terabytes of data but is now within reach thanks to these technologies.

Third‑Generation Sequencing (Long‑Read)

Third‑generation sequencing (TGS) technologies, primarily Pacific Biosciences (PacBio) single‑molecule real‑time (SMRT) sequencing and Oxford Nanopore Technologies (ONT) nanopore sequencing, produce read lengths often exceeding 10,000 base pairs and sometimes reaching megabases. Long reads are invaluable for assembling complex genomes with repetitive regions, resolving structural variants, and sequencing full‑length genes without computational assembly errors. For species discovery, TGS enables direct sequencing of entire mitochondrial or chloroplast genomes from field samples, bypassing the need for PCR amplification that can introduce bias. Although TGS historically suffered from higher error rates (~10–15%), recent improvements—such as PacBio HiFi reads (99.9% accuracy) and nanopore R10 chemistry—have made it competitive with short‑read platforms. Portable nanopore sequencers like the MinION allow real‑time sequencing in remote locations, from rainforest canopies to deep‑sea hydrothermal vents, accelerating the discovery of new species in their native habitats.

The Role of DNA Sequencing in Discovering New Species

Traditional taxonomy relied on morphological traits—size, shape, color—to classify organisms. Yet many species are morphologically similar, especially among insects, fungi, marine plankton, and soil microbes. DNA sequencing overcomes these limitations by detecting subtle genetic differences invisible to the naked eye. This molecular approach has unveiled numerous cryptic species—distinct species that were once lumped under a single name due to external resemblance.

By analyzing DNA barcodes (short standardized gene regions) or whole genomes, researchers can:

  • Identify cryptic species: For example, sequencing revealed that what was once considered a single species of African elephant (Loxodonta africana) is actually two distinct species: the forest elephant (L. cyclotis) and the savanna elephant (L. africana). Their genomic divergence is greater than that between lions and tigers.
  • Discover new species in unexplored habitats: Deep‑sea vents, cave systems, and the microbiomes of other organisms harbor countless unknown species detectable through environmental DNA (eDNA) sequencing. A single litre of seawater can yield DNA from hundreds of microbial species, many of which cannot be cultured in the lab.
  • Resolve evolutionary relationships: Phylogenetic analyses based on DNA sequences provide insights into evolutionary history and divergence times, improving classification systems and revealing how species adapted to different niches.
  • Accelerate discovery pace: In a single metagenomics study, scientists can identify hundreds of species from a sample, dramatically speeding up biodiversity assessments compared to manual morphological sorting.

Environmental DNA (eDNA) and Metagenomics

A revolutionary extension of DNA sequencing is environmental DNA (eDNA) analysis. Organisms shed genetic material into their surroundings through skin cells, saliva, faeces, or gametes. By collecting and sequencing eDNA from water, soil, or air samples, researchers can detect the presence of multiple species without ever seeing or capturing them. This approach has transformed biodiversity monitoring, especially for rare, elusive, or invasive species. For example, eDNA metabarcoding of river water can reveal the complete fish community upstream, while sediment samples from the deep sea can uncover previously unknown benthic foraminifera and nematodes.

Metagenomics goes a step further by sequencing all DNA in an environmental sample, capturing entire genomes of unculturable organisms. This has been particularly powerful for microbial species discovery. The Tara Oceans expedition used metagenomic sequencing across global oceans, discovering thousands of new viral, bacterial, and eukaryotic species. Many of these play critical roles in carbon cycling and marine food webs, and their discovery has forced a recalibration of global diversity estimates.

Case Studies and Notable Discoveries

Amphibians in Remote Rainforests

In the tropical forests of South America, herpetologists have deployed DNA barcoding to uncover numerous new amphibian species. A landmark study on glass frogs (Centrolenidae) combined mitochondrial and nuclear gene sequences to recognize several cryptic species with distinct vocalizations and distributions. These discoveries carry immediate conservation implications: hidden species often have smaller geographic ranges and face higher extinction risk than previously assumed. For instance, the newly described Hyalinobatrachium yaku from Ecuadorian Amazonia was identified solely through genetic divergence from its sister species, prompting protection measures for its restricted habitat.

Marine Microorganisms

Marine microbiology has been revolutionized by DNA sequencing. The Tara Oceans project used NGS to sample plankton across all major ocean basins, leading to the identification of thousands of new eukaryotic and bacterial species. Sequencing revealed that microbial diversity is far greater than estimated by morphology alone, with many lineages representing entirely new phyla. One unexpected find was the discovery of a new group of marine viruses, called “mirusviruses,” that bridge the gap between giant viruses and conventional nucleocytoplasmic large DNA viruses, reshaping our understanding of viral evolution.

Fungal Diversity in Soils

Fungi are notoriously difficult to identify by morphology due to their variable forms and cryptic life stages. DNA sequencing of ribosomal RNA genes (e.g., ITS regions) from soil samples has uncovered staggering fungal diversity. A 2022 study in Nature Microbiology used high‑throughput amplicon sequencing to estimate that over 90% of fungal species remain undescribed, many belonging to lineages previously unknown to science. These fungal “dark matter” species likely play vital roles in nutrient cycling, plant symbiosis, and carbon sequestration. Reference databases like UNITE are expanding rapidly to catalog this hidden diversity.

Cryptic Fish in Coral Reefs

Coral reefs host spectacular fish diversity, but many species are morphologically similar. Barcoding of the COI gene has revealed multiple cryptic species within what was once considered the clownfish (Amphiprioninae) complex. For example, the widely distributed Amphiprion ocellaris was found to comprise at least three genetically distinct lineages, each with different host anemone preferences and thermal tolerances. These findings affect conservation strategies, as distinct cryptic species may respond differently to bleaching events and habitat degradation.

Challenges and Limitations

Despite its power, DNA sequencing is not without pitfalls. Sequence data alone may be insufficient for formal species description; integration with morphological, ecological, and behavioral data is often required by taxonomic rules. Incomplete reference databases can lead to misidentification or false positives, particularly for poorly studied groups. Costs, though declining, remain significant for large‑scale studies in biodiversity‑rich developing nations. Moreover, high‑quality DNA is essential; degraded samples from museum specimens or old environmental collections often yield incomplete or erroneous sequences. Sequencing errors, especially in long‑read platforms, require rigorous quality control and validation with multiple markers.

Another challenge is the “taxonomic impediment”: there are too few trained taxonomists to describe the flood of new species detected by sequencing. Bioinformatics pipelines must handle massive datasets, and decisions about species boundaries from genetic data require careful statistical frameworks (e.g., species delimitation models). Nevertheless, collaborative efforts like the International Barcode of Life (iBOL) are building comprehensive reference libraries and training the next generation of molecular taxonomists.

Future Directions

The future of DNA sequencing in species discovery is bright, propelled by rapid technological advances. Portable sequencers such as Oxford Nanopore’s MinION enable real‑time sequencing in the field—from rainforests to deep‑sea trenches—reducing sample degradation and accelerating identification. For instance, the MinION has been used to sequence the genome of a weevil pest on location in Papua New Guinea, demonstrating its potential for rapid biodiversity surveys. Handheld sequencers combined with lightweight bioinformatics laptops can soon make on‑site species discovery routine.

Another emerging trend is the integration of machine learning with sequencing data to predict species boundaries from genomic patterns. Automated pipelines that combine eDNA sampling with portable sequencing could enable continuous biodiversity monitoring, aiding conservation efforts in real‑time. Long‑read sequencing and de novo genome assembly will become standard for non‑model organisms, allowing researchers to uncover entire gene families and adaptations unique to newly discovered species. The combination of long reads with chromatin conformation capture (Hi‑C) can produce chromosome‑level genome assemblies that serve as robust references for comparative genomics.

As costs continue to plummet, the goal of sequencing all known eukaryotic species on Earth—a “tree of life for eukaryotes”—becomes increasingly feasible. Such a comprehensive genetic catalog would not only fill gaps in taxonomic knowledge but also provide a resource for bioprospecting, conservation planning, and understanding ecosystem function. The Earth BioGenome Project aims to sequence all 1.5 million described eukaryotic species within a decade, leveraging both short‑ and long‑read technologies.

Conclusion

DNA sequencing technologies have fundamentally reshaped how we discover and classify new species, shifting from morphology‑based approaches to a genetic foundation. From Sanger’s early methods to today’s high‑throughput and portable platforms, these tools have unveiled a hidden world of biodiversity—cryptic amphibians in rainforests, unknown microbes in the deep ocean, and countless fungi beneath our feet. The rapid pace of technological innovation promises to accelerate species discovery even further, providing critical insights for conservation and our understanding of life on Earth. As we continue to explore the planet’s genetic diversity, DNA sequencing will remain an indispensable tool in the quest to document and protect the natural world for generations to come.