S. O. Shuster

August 2026

Pathogen Agnostic DNA Sequencing: Notes from the BlueDot Biosecurity Course

Currently, a number of modern sequencing techniques allow for the pathogen-agnostic detection needed for metagenomic sequencing (MGS) to be an effective early detection tool. However, several improvements need to be made for MGS to become as inexpensive, accessible, and prolific as PCR and Sanger sequencing.

Modern sequencing techniques generally fall into two groups: next-generation sequencing (NGS), pioneered by Illumina, and third-generation sequencing (TGS). NGS works by first extracting DNA (or RNA) from a sample, then fragmenting it into short sequences and forming a "library" through the attachment of adapters to those fragments. The fragments are then loaded onto a microfluidics chip, where they attach to DNA oligomers and are prepared for sequencing (i.e., amplified into clusters and made single-stranded). The DNA is then sequenced through a repetitive process of flowing fluorescently labeled nucleotides over the sample and using a camera to read which colors bind. Finally, the resulting images are analyzed, cleaning up unclear signal and removing overlapping clusters. This is a very high-throughput option and is sequence-agnostic—any DNA, even completely novel viral DNA, can be analyzed. However, sequencing is limited to short fragments, and aligning these fragments into complete stretches of DNA can require a reference genome and significant bioinformatics pipelines.

There are several types of TGS, but I'll focus on the most commercialized: nanopore sequencing by Oxford Nanopore. Nanopore sequencing works by using a specialized motor protein to feed a length of single-stranded DNA through a protein nanopore anchored in a membrane. As each base passes through the nanopore, the resulting change in electric current can be detected and correlated with the type of nucleotide. Nanopore sequencing can handle extremely long sequences in a single run.

Currently, both NGS and TGS are widely used in research and personalized medicine but lack widespread adoption in point-of-care diagnostics and surveillance screening, for a few key reasons. First, both techniques currently cost 3–25x as much as PCR, with nanopore sequencing on the higher end. Second, sample preparation for these techniques is time-intensive and technically difficult. Third, time-to-results varies but is generally much slower than PCR, though nanopore sequencing can already be faster; some of this delay comes from computational bottlenecks, and current bioinformatics pipelines still need better reference sequences. Finally, and perhaps most importantly, NGS and TGS both sequence all DNA in a sample, making it a signal-detection problem to distinguish pathogen DNA from host or background DNA. PCR, by contrast, is targeted and thus blind to background DNA.

Making MGS as ubiquitous as PCR—and a useful tool for identifying new pathogens in patient samples and wastewater—will require faster, cheaper, easier-to-use NGS and TGS. It is likely possible to perform library preparation steps inside a machine using a cartridge-like setup. Techniques can become less expensive through improved microfluidics manufacturing. A promising direction for faster sequencing is reoptimizing away from the extremely high sequence fidelity needed for research and precision medicine but unnecessary for pathogen recognition. To make MGS effective at screening for unknown pathogens, host DNA depletion will also be needed (options include preparation steps, adaptive nanopore sequencing, and computational tools). Since the current customers of MGS are primarily research and precision medicine, explicit government funding and direction, similar to the $1,000 Genome Project, will be needed to drive this innovation, some of which is already underway, including the work of SecureBio.

← Back to all posts