Showing posts with label biotechnology. Show all posts

Wednesday, June 24, 2026

thumbnail

Sanger Sequencing Explained: How One Missing Oxygen Changed DNA Sequencing Forever



In 1977 Frederick Sanger described a method of DNA sequencing using chain-terminating nucleotides called dideoxynucleotides. 

The aim was to determine the sequence of nucleotides in a piece of DNA using these artificial or synthetic nucleotide analogues. Unlike natural DNA nucleotides (deoxyribonucleoside triphosphates, dNTPs), these lab-synthesised forms lack the 3'-hydroxyl group required for DNA strand extension.

This method became known as Sanger sequencing.

These chain-terminating nucleotides are called dideoxyribonucleoside triphosphates (ddNTPs).

DNA is made up of a chain of four different nucleotides called dNTPs. To copy DNA and extend a DNA strand, DNA polymerase adds a complementary nucleotide.

A closer look at its structure shows that a dNTP consists of a deoxyribose sugar, a nitrogenous base, and a triphosphate group. A nucleoside consists of a sugar and a base. In DNA the sugar is deoxyribose, while in RNA it is ribose. The base is one of four bases: adenine, thymine, guanine, or cytosine.

A ddNTP lacks both the 2'-OH and 3'-OH groups found in ribose. Compared with deoxyribose, it is missing the 3'-OH group.

The role of DNA polymerase is to add new nucleotides to a growing DNA strand. During DNA synthesis, the 3'-hydroxyl group (3'-OH) of the growing DNA strand reacts with the α-phosphate of the incoming dNTP, forming a phosphodiester bond and releasing pyrophosphate.

If a ddNTP is incorporated into the DNA strand, synthesis stops because the ddNTP lacks the 3'-OH group required to add the next nucleotide. This absence of a 3'-OH group terminates DNA chain elongation.

It is also useful to understand the naming convention of 5' (five-prime) and 3' (three-prime). The carbons in the deoxyribose sugar are numbered 1' through 5'. The nitrogenous base is attached to the 1' carbon, while the phosphate group is attached to the 5' carbon.

The 3'-OH group attached to the 3' carbon is the chemical group required for DNA strand extension. Because DNA polymerase adds new nucleotides to the existing chain utilising the phosphate group of the new dNTP. Hence, the dogma that DNA synthesis proceeds in the 5'→3' direction. And when DNA sequences are written, they are conventionally written from 5' to 3'.

DNA polymerase adds nucleotides complementary to the template strand, so C pairs with G and A pairs with T.

So how does Sanger sequencing work?

The original Sanger sequencing method was different from the one used today. The original method was completely manual and used radioactive labels.

Let's take a look at the original Sanger sequencing method.

We need a primer, DNA polymerase, dNTPs, a DNA template, and ddNTPs.

One of the dNTPs, usually dATP, is labeled with a radioactive isotope.

A total of four tubes are used, one for each ddNTP.

To begin, the DNA template is heated to denature the double-stranded DNA into single strands. Remember, this was before PCR existed. Because the DNA polymerases available at the time were not thermostable, the enzyme was added after the denaturation step.

The mixture is then cooled to allow the sequencing primer to anneal to the template.

DNA polymerase, all four dNTPs, and one of the four ddNTPs are then added to each tube.

DNA polymerase extends the primer along the DNA template. Occasionally, a ddNTP is incorporated instead of a dNTP, terminating the DNA fragment.

Because the ddNTP is present at a much lower concentration than the corresponding dNTP, incorporation occurs randomly.

The result is a collection of DNA fragments that terminate at every occurrence of that particular base, generating fragments of different lengths.

All fragments in a tube begin with the same primer sequence and end with the same terminating nucleotide.

Low incorporation of the ddNTP allows longer stretches of DNA to be sequenced.

In the original Sanger method, read lengths of approximately 200 nucleotides were achievable.

Next, the sequencing reactions are mixed with loading dye and loaded into separate lanes of a polyacrylamide gel.

The fragments migrate through the gel according to size, with smaller fragments moving faster than larger fragments.

Polyacrylamide gels have sufficient resolution to distinguish DNA fragments that differ by a single nucleotide in length.

At this stage the fragments cannot be seen.

The loading dye indicates when the fragments have migrated through the gel.

To visualise the DNA fragments, the gel is dried onto a support and exposed to X-ray film.

The radioactive labels incorporated into the DNA fragments expose the film, producing a pattern of bands.

The process of determining the DNA sequence from these bands is called base calling.

The gel is read from the bottom upward, starting with the shortest fragment. This reveals the sequence of the newly synthesized DNA strand in the 5'→3' direction.

For example, if the shortest fragment appears in the ddTTP lane, the first base called is T. If the next shortest fragment appears in the ddGTP lane, the next base is G.

Continuing upward through the gel reveals the complete sequence.

The original Sanger sequencing method was very labor-intensive. It could take several days to generate approximately 200 nucleotides of sequence from only a small number of samples.

There was a strong need to streamline and automate the process.

Applied Biosystems created the first commercial automated DNA sequencing instrument in 1987, the AB370A.

Researchers had already demonstrated that fluorescent dyes could replace radioactive labels. These fluorescent dyes were safer and eliminated the need for time-consuming X-ray film detection.

In this instrument, each of the four sequencing reactions was labeled with a different fluorescent dye.

After the sequencing reactions were completed, all four reactions could be combined and loaded into a single lane of a gel.

The AB370A used a laser to detect fluorescent DNA fragments as they migrated through the gel.

The instrument automatically transferred the data to a computer, which performed automated base calling.

Up to 16 samples could be run simultaneously, with read lengths approaching 450 nucleotides.

The AB370A demonstrated that DNA sequencing could be faster and more automated.

Scientists began to think that sequencing the entire human genome might be achievable.

In 1990 the U.S. government launched the Human Genome Project, an international effort to map and sequence the entire human genome.

Sequencing the human genome promised major advances in biology and medicine, including identifying genes associated with inherited diseases and improving our understanding of human biology.

Kary Mullis had invented PCR in 1983, but it was not until 1989 that Vincent Murray applied thermostable Taq polymerase to Sanger sequencing.

In traditional Sanger sequencing, most labeled primers remain unused because the primer is present in excess relative to the DNA template.

The use of Taq polymerase allowed repeated cycles of denaturation, primer annealing, and extension, similar to PCR.

Because only a single sequencing primer is present, newly synthesised strands do not serve as templates for exponential amplification.

As a result, the amount of sequencing product increases approximately linearly rather than exponentially. This process became known as cycle sequencing.

The increased signal generated by cycle sequencing also reduced the amount of input DNA required.

Another important advance was capillary electrophoresis.

In capillary electrophoresis, DNA fragments migrate through a thin capillary filled with a polymer matrix under an electric field.

The narrow capillary efficiently dissipates heat, allowing higher voltages to be used without overheating.

Higher voltages result in faster separations and improved resolution.

Beckman Coulter launched the first commercial capillary electrophoresis instrument in 1989.

This technology paved the way for capillary-based Sanger sequencing systems.

Applied Biosystems launched the ABI Prism 310 in 1995, marking the beginning of modern Sanger sequencing.

The ABI Prism 310 replaced slab gels with a single capillary.

A sequencing run could be completed in under three hours, with read lengths approaching 600 base pairs.

The system also automated sample loading and reduced DNA input requirements through electrokinetic injection.

DNA fragments were separated by size, detected by a laser, and analyzed automatically by software that performed base calling.

Although fluorescently labeled ddNTPs were available, fluorescent primer labeling was initially preferred because it produced more uniform signal intensities.

This changed with the introduction of BigDye Terminator chemistry in 1997.

BigDye Terminators improved the balance of fluorescent signal intensity among dye-labeled ddNTPs, allowing all four termination reactions to be performed in a single tube.

This greatly simplified sequencing workflows.

The Human Genome Project continued to drive demand for faster and more automated sequencing technologies.

In 1998 Applied Biosystems launched the ABI Prism 3700, which contained 96 capillaries.

The ABI Prism 3700 played a major role in sequencing the human genome.

Each run processed 96 samples simultaneously, generated read lengths approaching 800 base pairs, and required minimal hands-on time.

Celera Genomics, led by Craig Venter, purchased hundreds of ABI Prism 3700 instruments and used them to compete directly with the publicly funded Human Genome Project.

Celera produced a draft human genome sequence in 2001, and the Human Genome Project published its own draft sequence the same year.

As an aside, the Human Genome Project did not sequence the DNA of a single individual. The  reference genome was assembled from DNA obtained from multiple anonymous donors. Much of this DNA came from blood samples (while  red blood cells lack nuclei and therefore contain no genomic DNA,  white blood cells contain nuclei and provide a rich source of genomic DNA). The resulting reference genome was therefore a composite sequence rather than the genome of any one person.


Modern Sanger sequencing remains widely used today.

Sanger sequencing typically achieves greater than 99.9% accuracy in high-quality reads, which is why it is often considered the benchmark against which other sequencing technologies are compared.

For small projects involving individual genes, plasmids, or a limited number of samples, Sanger sequencing is often faster and more cost-effective than next-generation sequencing (NGS).

However, Sanger sequencing has lower throughput and lower sensitivity for detecting rare variants, typically requiring variants to be present at roughly 15–20% frequency before they can be reliably detected.

In contrast, NGS can detect variants present at much lower frequencies and can generate billions of reads and terabases of sequence data in a single run.

This allows many whole human genomes to be sequenced simultaneously.

For validating variants, sequencing plasmids, or analysing a small number of genes or samples, Sanger sequencing remains one of the most practical and widely used sequencing methods available.

Related post:

What Is Next-Generation Sequencing? NGS vs Sanger Explained

https://adwoabiotech.blogspot.com/2026/06/what-is-next-generation-sequencing-ngs.html

Oxford Nanopore Sequencing: A Different Philosophy of Reading DNA

https://adwoabiotech.blogspot.com/2026/06/what-is-oxford-nanopore-sequencing-how.html


Friday, June 19, 2026

thumbnail

Complete RPMI-1640 Preparation for Malaria Culture

A Step-by-Step Guide to Plasmodium Culture Media

If you’ve spent any time in a malaria lab, you know that Plasmodium falciparum is a picky eater. 

The goal here is to create a stable, nutrient-rich environment that mimics human physiological conditions while accounting for the common pitfalls of lab life; like the frustrating tendency of buffers to outgas or essential amino acids to degrade.

The media used (RPMI-1640) was developed at Roswell Park Memorial Institute in the 1960s. The "1640" refers to the formulation number assigned during its development, reflecting the extensive empirical (trial-and-error, experiment-based) optimisation that produced the medium.

RPMI-1640 has a fascinating history because it emerged during the period when mammalian cell culture was transitioning from empirical media recipes to more rationally designed formulations. For those like me, wondering what empirical means, it’s approaches that were based on observation, experimentation, and trial-and-error rather than a complete theoretical understanding. 

Origin of RPMI-1640

RPMI stands for: Roswell Park Memorial Institute

The medium was developed at the Roswell Park Comprehensive Cancer Center in Buffalo, New York.

The principal developers were:

  • George E. Moore

  • Robert E. Gerner

  • Harold A. Franklin

during the 1960s.


The general components and amounts in 1L RPMI are: 

RPMI 1640: 10.44 g (The nutritional backbone, containing l-glutamine at a final conc. of 2mM). This is 5.22g if only making 500 mL

HEPES: 5.96 g (Your primary buffering agent). If making 500 mL, add 12.5 mL of 1M HEPES. If you buy the powdered RPMI, it likely comes with this already.


NaHCO3​: 58 mL of 3.6%  (final is 2g/L ; For pH stability and gas exchange). Sodium bicarbonate is notorious for outgassing (releasing CO2​), which can cause your pH to drift upward over time. To combat this, we add it just before use.


Hypoxanthine: 50 mg (200uM, Essential for parasite purine salvage)


Gentamicin: 20 ug/mL stock (Your antibiotic shield). For 1L you can add 20 mg of gentamicin


Ultrapure/sterile H2​O: 960 mL

1M NaOH: For dissolving the hypoxanthine


Concentrated HCL and NaOH: For final pH adjustment





Related Content

Ready to put this media to use? Our step-by-step Plasmodium falciparum culture protocol shows exactly how Complete RPMI-1640 fits into culture initiation and maintenance.


References:

1.     Lopez-Perez, M., & Seidu, Z. (2022). Establishing and Maintaining In Vitro Cultures of Asexual Blood Stages of Plasmodium falciparum. Methods in Molecular Biology. 

2.     Maier, A. G., & Rug, M. (2013). In vitro culturing Plasmodium falciparum erythrocytic stages. Methods in Molecular Biology. 

3.     Trager, W., & Jensen, J. B. (1976). Human malaria parasites in continuous culture. Science, 193(4254), 673–675. 

4.     World Health Organization. (2023). World Malaria Report 2023. Geneva: WHO.


Wednesday, May 13, 2026

thumbnail

Antisense RNA Explained: Why Strand-Specific RNA-Seq Matters

 The Two Strands of DNA Do Not Carry the Same Information: Antisense RNA and Why Strand-Specific RNA-Seq Changed Everything



For the longest time, I subconsciously assumed the two strands of DNA carried the same biological information.

Not identical sequences, obviously. We all learn early on that DNA strands are complementary:

A pairs with T.
G pairs with C.

But conceptually, I still thought of one strand as essentially a mirrored backup copy of the other. Same information, just reversed.

Then I properly encountered antisense RNA and strand-specific RNA sequencing.

And suddenly I realised something that completely changed how I visualise genomes:

The opposite strand of DNA can encode entirely different biological information.

Not just regulatory signals. Entirely different RNAs. Sometimes entirely different proteins.


The Textbook Version vs Reality

Most of us are taught transcription something like this:

DNA → RNA → Protein

A gene sits on DNA. RNA polymerase reads it. mRNA is produced. Protein gets made. 

But real genomes are much messier than that.

Genes can overlap. Transcription can occur in both directions. RNAs can regulate other RNAs. Some RNAs are never translated at all. And sometimes the "opposite" strand of DNA contains completely different instructions.


What Is Antisense RNA?



To understand antisense RNA, we first need to separate two ideas:

Sense strand
The DNA strand whose sequence matches the RNA transcript (except T is replaced with U).

Antisense strand (template strand)
The DNA strand actually used by RNA polymerase as the template during transcription.

Already, the naming is confusing enough to give undergraduate students mental strain.

But here’s the important part:

An RNA molecule can also be produced from the opposite DNA strand in the reverse direction.

That RNA is called an antisense RNA.

This means you can have something like this:

  • One strand producing a normal protein-coding mRNA

  • The opposite strand producing a completely different RNA transcript

And because the sequences are complementary, the RNAs can physically interact with one another.

That interaction can:

  • block translation,

  • alter chromatin structure,

  • regulate transcription,

  • or affect RNA stability.

In other words, antisense RNAs are not necessarily transcriptional "noise". They can have real biological functions.


The moment the light bulb came on for me was realising this:

The opposite DNA strand is not just a passive complementary copy.

It can contain:

  • different promoters,

  • different transcription start sites,

  • different open reading frames,

  • and entirely different regulatory information.

That means the genome is not one-dimensional.

It is layered, so strand direction becomes critically important.


One of the Most Famous Examples: XIST and TSIX

A classic example comes from X chromosome inactivation in mammals.

Female mammals have two X chromosomes, but one must be largely silenced to prevent double dosage of X-linked genes.

This process is controlled by a long non-coding RNA called XIST.

XIST coats one X chromosome and helps silence it.

Interestingly, another RNA called TSIX is transcribed from the opposite strand across the XIST locus.

TSIX is an antisense transcript.

And rather than being meaningless background transcription, TSIX helps regulate XIST expression itself.

So you end up with a regulatory system where:

  • XIST promotes X chromosome silencing

  • TSIX regulates XIST

  • both arising from opposite strands of the same genomic region

A classic paper by Navarro et al. (2005) showed that TSIX transcription alters chromatin conformation at the XIST locus, highlighting that antisense transcription itself can have regulatory consequences.

At this point, the idea that DNA is simply a static storage medium starts to feel very incomplete.


Scientists Initially Thought Antisense Transcription Was Mostly Noise

And honestly, this assumption made sense at the time.

Early transcriptomics methods often struggled to determine which DNA strand an RNA came from. Researchers could detect transcriptional signal, but not always its orientation.

So when overlapping or opposite-direction transcripts appeared, many scientists assumed they were:

  • transcriptional errors,

  • random polymerase activity,

  • or biological noise.

Then strand-specific RNA sequencing methods became more widely adopted.

And suddenly researchers realised antisense transcription was everywhere.

Not just in humans.
Not just in weird edge cases.

Entire layers of genome regulation had been hiding in plain sight simply because earlier methods collapsed strand information together.


Why Ordinary RNA-Seq Can Cause Problems

This is where strand-specific RNA-seq becomes incredibly important.

In conventional (non-stranded) RNA-seq, you can sequence RNA transcripts perfectly well, but you may lose information about which DNA strand they originally came from.

That becomes a major problem when:

  • genes overlap,

  • antisense RNAs exist,

  • or neighbouring genes are transcribed in opposite directions.

Imagine two genes sitting on opposite strands of DNA:

One goes left → right
The other goes right → left

If your RNA-seq data is not strand-specific, all the sequencing reads can appear merged together.

You may detect transcription in that region, but you cannot confidently determine:

  • which gene produced the RNA,

  • whether both genes are active,

  • or whether antisense transcription is occurring.

That ambiguity can completely alter biological interpretation.


Strand-Specific RNA-Seq Preserves Directionality

Strand-specific RNA-seq solves this problem by preserving transcript orientation during library preparation.

In simple terms, it tells you:

"This RNA came from THIS DNA strand."

That sounds like a tiny technical detail but having that information changes how we interpret:

  • gene boundaries,

  • overlapping loci,

  • antisense transcription,

  • non-coding RNAs,

  • and transcript abundance.

This is especially important in compact genomes where genes are tightly packed together.

Without strandedness, transcriptional landscapes can become blurred.


The Bioinformatics Consequences Are Huge

If you accidentally analyse stranded RNA-seq data using the wrong strand settings during alignment or counting, you can:

  • invert expression signals,

  • assign reads to the wrong genes,

  • underestimate transcript abundance,

  • or completely miss antisense transcription.

In other words:
your computational interpretation becomes biologically wrong.

This is why bioinformaticians care so much about strandedness metadata in RNA-seq experiments.

It is not computational nitpicking: it fundamentally affects what the data means.


For me, the most fascinating part of all this is philosophical as much as technical.

At school and university, DNA is often presented as though genes sit neatly along a chromosome like words written left-to-right in a book.

But real genomes are far stranger than that.

They are:

  • bidirectional,

  • overlapping,

  • dynamic,

  • and deeply layered.

The two strands of DNA are not redundant copies carrying the same biological information.

They can encode entirely different transcripts with entirely different functions.


Further Reading and References

Ali, T., Grote, P., & Rosenstiel, P. (2023). Natural antisense transcripts in disease and therapy. Non-Coding RNA, 9(6), 76. https://pmc.ncbi.nlm.nih.gov/articles/PMC10761088/

Goodman, A. J., Chung, D. W. D., McGrath, P. T., et al. (2013). Pervasive antisense transcription is evolutionarily conserved in budding yeast. Molecular Biology and Evolution, 30(2), 409–421. https://academic.oup.com/mbe/article/30/2/409/1017569

Jensen, T. H., Jacquier, A., & Libri, D. (2013). Dealing with pervasive transcription. Molecular Cell, 52(4), 473–484. https://www.sciencedirect.com/science/article/pii/S1097276513007983

Navarro, P., Page, D. R., Avner, P., & Rougeulle, C. (2005). Tsix transcription across the Xist gene alters chromatin conformation without affecting Xist transcription: Implications for X-chromosome inactivation. Genes & Development, 19(12), 1474–1484. https://genesdev.cshlp.org/content/19/12/1474.full

Schurch, N. J., Schofield, P., Gierliński, M., et al. (2014). Improved annotation of alternatively spliced and long intergenic non-coding RNAs by combining strand-specific RNA sequencing, paired-end sequencing and multiple knockout mutants. Nucleic Acids Research, 42(13), e104. https://arxiv.org/abs/1311.2494

Senner, C. E., & Brockdorff, N. (2009). Xist gene regulation at the onset of X inactivation. Current Opinion in Genetics & Development, 19(2), 122–126. https://pubmed.ncbi.nlm.nih.gov/19345091/


Related post: 


Stranded vs. Unstranded RNA-Seq: Why Strand Information Matters in Gene Expression

https://adwoabiotech.blogspot.com/2025/05/stranded-vs-unstranded-rna-seq-why.html


About

Search This Blog

Powered by Blogger.

About Me

My photo
Adwoa Biotech Tools and Techniques Hub offers clear, practical explanations of essential molecular biology and biotechnology methods. Learn PCR primer design, cDNA synthesis, cloning strategies, nucleic acid purification, CRISPR delivery innovations, data analysis concepts, and everyday lab skills. Enjoyed the tutorial, connect with me on YouTube for video content on these topics: @adwoabiotech