DNA and RNA: The Building Blocks of Life
Explore the structure of DNA and RNA, how the double helix stores genetic information, and how the central dogma drives protein synthesis.
Explore the structure of DNA and RNA, how the double helix stores genetic information, and how the central dogma drives protein synthesis.
Every living thing on Earth — from the tiniest bacterium clinging to a deep-sea vent to the blue whale cruising the Pacific Ocean — carries the same type of molecular instruction manual. That manual is written in a language with only four letters, coiled into a shape so elegant that the scientists who first described it could barely believe what they were seeing. That molecule is DNA, and its close cousin RNA is the system that reads the instructions and puts them to work.
Understanding DNA and RNA is crucial to modern biology. These molecules sit at the very heart of what life is. They explain why you have your grandmother’s eyes, why bacteria can become resistant to antibiotics, and why it might one day be possible to cure inherited diseases with a precise molecular edit.
Now, let’s take a deep dive on DNA - how it stores information, different RNA types and what do they do. We will also discuss the central dogma - the flow of genetic information, Human Genome Project and CRISPER gene editing.
For centuries, biologists knew that traits passed from parents to offspring, but they had no idea how. By the late nineteenth century, work by Gregor Mendel on pea plants had revealed that inheritance followed mathematical rules, and chemists had identified a mysterious substance called “nuclein” in the nuclei of white blood cells. But it wasn’t until 1944 that Oswald Avery and colleagues at the Rockefeller Institute provided strong evidence that DNA — not protein, as many had assumed — was the molecule of heredity.
The final, definitive picture of how DNA is structured came in 1953, when James Watson and Francis Crick published a landmark paper in Nature describing the double helix. Their model drew heavily on X-ray diffraction images produced by Rosalind Franklin and Raymond Gosling. It also relied on chemical work by Erwin Chargaff. Chargaff noticed that in any sample of DNA, the amount of adenine always equaled thymine, and guanine always equaled cytosine. That observation — Chargaff’s rules — was the key that unlocked the structure.

DNA: Deoxyiribonucleic Acid.
The molecule is a polymer — a long chain of repeating units called nucleotides. Each nucleotide has three parts: a sugar molecule (deoxyribose), a phosphate group, and one of four nitrogen-containing bases — adenine (A), thymine (T), guanine (G), or cytosine (C).
Two of these nucleotide chains run antiparallel to each other — one running in the 5′ to 3′ direction, the other in the 3′ to 5′ direction — and are wound around each other like a twisted ladder. The sugar-phosphate groups form the backbone (the uprights of the ladder), while the bases project inward and pair with bases on the opposite strand (the rungs). Adenine always pairs with thymine (two hydrogen bonds), and guanine always pairs with cytosine (three hydrogen bonds). These are called Watson-Crick base pairs, and this complementary pairing is the key to how genetic information is copied.

The human genome contains roughly 3.2 billion base pairs. If you stretched all the DNA from a single human cell end to end, it would measure about two metres. Yet it fits inside a cell nucleus that is roughly five to ten micrometres across — smaller than a red blood cell. Nature accomplishes this by wrapping DNA around spool-like protein complexes called histones, forming structures called nucleosomes. These are further compacted into chromatin fibres, and during cell division, the chromatin condenses further still into the dense, compact structures we recognise as chromosomes.
Human cells contain 46 chromosomes arranged in 23 pairs — one of each pair inherited from each parent. Our cells, described in detail in our article on important facts about cells, are remarkably sophisticated machines, and DNA organisation is central to how they work.

The genetic information in DNA is encoded in the sequence of bases along one strand — the “sense” strand. A gene is a specific region of DNA whose base sequence encodes instructions for building a particular protein or functional RNA molecule. The human genome contains roughly 20,000–25,000 protein-coding genes, which account for only about 1.5% of the total DNA. The rest — once dismissed as “junk DNA” — is now known to have important regulatory and structural roles.
Think of DNA as a recipe book written in four-letter code. The book is stored in the nucleus, but it never leaves — it’s too precious and too large. Instead, when a recipe is needed, a working copy is made and sent out to the kitchen. That working copy is RNA.
RNA stands for ribonucleic acid. Like DNA, it is a nucleotide polymer, but with two important differences: the sugar is ribose (not deoxyribose), and instead of thymine, RNA uses uracil (U). RNA is also typically single-stranded, though it can fold back on itself to form complex three-dimensional shapes.
There are several types of RNA, each with a distinct role in translating the genetic code into proteins.
Messenger RNA is the working copy of a gene. When a gene needs to be expressed, an enzyme called RNA polymerase binds to a region of DNA called the promoter and unwinds the double helix. It then reads the template strand of the DNA and synthesises a complementary mRNA molecule — a process called transcription.
The mRNA carries the genetic message from the nucleus to the ribosomes in the cytoplasm, where proteins are built. It is read in triplets of bases called codons. Each codon specifies a particular amino acid (the building blocks of proteins), or a start/stop signal. The genetic code — the mapping of 64 possible codons to 20 amino acids — is nearly universal across all life on Earth, which is remarkable evidence of our shared ancestry.

Transfer RNA is the adapter molecule that bridges the language of nucleotides and the language of amino acids. Each tRNA has a specific anticodon region that base-pairs with a complementary mRNA codon, and at its other end, it carries the corresponding amino acid. Think of tRNA as a molecular taxi: it picks up an amino acid at one location and delivers it to the ribosome at exactly the right moment.

Ribosomal RNA is a structural and catalytic component of ribosomes — the cellular machines that actually build proteins. In humans, ribosomes consist of two subunits, each made of rRNA molecules and dozens of proteins. The rRNA is not merely structural; it is the ribosome’s rRNA that catalyses the formation of peptide bonds between amino acids. Ribosomes are among the most ancient molecular machines in biology, and their rRNA sequences are so conserved across species that biologists use them to reconstruct evolutionary trees.
Beyond these three classical types, cells produce a rich variety of other RNA molecules. MicroRNAs (miRNAs) are short molecules that regulate gene expression by targeting mRNAs for degradation or blocking their translation. Small interfering RNAs (siRNAs) operate similarly and have become powerful tools in research. Long non-coding RNAs (lncRNAs) influence chromatin structure and gene regulation in ways still being actively investigated.
When scientists research deeper about RNAs, they found that it has many different roles which seems too much for such a molecule. They concluded that RNA may be evolving on its own and coined the term, RNA evolution.
[!NOTE] The RNA world hypothesis proposes that RNA, not DNA, was the original molecule of life. RNA can both store information and catalyse chemical reactions — a combination that DNA alone cannot achieve. Explore this idea further in our article on the origin of life on Earth.
In 1958, Francis Crick articulated what he called the “central dogma” of molecular biology: genetic information flows from DNA to RNA to protein. While we now know the flow is more complicated in some cases (RNA viruses can convert RNA back to DNA using an enzyme called reverse transcriptase), the central dogma remains a powerful organising principle.
Transcription occurs in the nucleus. RNA polymerase binds to the promoter sequence upstream of a gene, unwinds the DNA, and reads the template strand in the 3′ to 5′ direction, building a complementary mRNA strand in the 5′ to 3′ direction. In eukaryotes (organisms with nuclei, including all plants, animals, and fungi), the initial RNA transcript is called pre-mRNA. This transcript is processed before leaving the nucleus. A protective cap is added at the 5′ end, and a poly-A tail at the 3′ end. Non-coding regions called introns are spliced out, leaving only the protein-coding exons.
The mature mRNA travels to the ribosome, where translation begins. A ribosome clamps around the mRNA and reads it codon by codon from the start codon (AUG, which codes for the amino acid methionine) to a stop codon. At each step, a tRNA with the matching anticodon delivers its amino acid, the ribosome links it to the growing chain with a peptide bond, and the tRNA is released. When the ribosome reaches a stop codon, the protein chain is released and typically folds into a three-dimensional shape that determines its function.
The entire process can happen with remarkable speed. A bacterial ribosome can add about 20 amino acids per second. Some proteins consist of just a few dozen amino acids; others, like titin (a protein in muscle), contain more than 30,000.

By the early 1990s, molecular biologists had developed the tools to read DNA sequences — determining the exact order of base pairs along a strand. The question arose: could we read the entire human genome? The Human Genome Project (HGP) was launched in 1990 as an international collaboration among research centres in the United States, United Kingdom, France, Germany, Japan, and China. Its goal was audacious: to map every one of the roughly 3 billion base pairs in the human genome.
The first working draft was announced in June 2000 — simultaneously by the public consortium and the private company Celera Genomics, which had mounted a parallel effort using a different computational strategy. The complete sequence was declared finished in April 2003, marking the 50th anniversary of Watson and Crick’s double helix paper.

The HGP revealed surprising findings. The human genome contains far fewer protein-coding genes than expected — about 20,000–25,000, rather than the 100,000 that many had predicted. It also revealed the vast extent of non-coding DNA and the large amount of sequence shared with other species. We share about 98.7% of our DNA with chimpanzees, and even about 85% with mice.
The HGP’s legacy has been enormous. It created reference sequences for medical genetics, accelerated drug development, and laid the groundwork for personalised medicine. This is the idea that treatments can be tailored to an individual’s unique genetic profile, an approach increasingly driven by AI as detailed in how AI is revolutionizing healthcare. The technology that emerged from the project has also driven the cost of genome sequencing from billions of dollars to a few hundred. This opened up applications in forensics, agriculture, ecology, and beyond. Understanding how these discoveries relate to what tissues are made of and how cells organise themselves has deepened our picture of biology at every scale.
If the Human Genome Project was about reading the genetic code, CRISPR-Cas9 is about editing it. Discovered as a bacterial immune system in the early 2010s and adapted into a gene-editing tool by Jennifer Doudna, Emmanuelle Charpentier, and colleagues — work that earned the 2020 Nobel Prize in Chemistry — CRISPR allows scientists to target a precise location in a genome and cut the DNA there.
The system uses a guide RNA (gRNA) to lead the Cas9 protein to the target sequence. Once Cas9 cuts both strands of the DNA, the cell’s own repair machinery takes over. Scientists can exploit this repair process to disable a gene or to insert a new sequence.

The potential applications are staggering: correcting the mutations that cause sickle cell disease or cystic fibrosis, engineering crops that are resistant to drought or disease, potentially eliminating hereditary diseases. The first CRISPR based therapy — for sickle cell disease — received regulatory approval in the United States and United Kingdom in late 2023. At the same time, CRISPR raises profound ethical questions about germline editing (changes that would be inherited by future generations) and equity of access. The technology is evolving faster than the ethical governance around it.
DNA is double-stranded, uses the sugar deoxyribose, and contains thymine as one of its four bases. It is the long-term storage molecule for genetic information, housed mainly in the cell nucleus. RNA is single-stranded, uses ribose, and substitutes uracil for thymine. RNA acts as a working copy of genetic instructions and is involved in carrying, decoding, and executing those instructions to build proteins.
DNA stores the complete genetic instructions for an organism — everything from the proteins that make up its cells to the signals that regulate when and where those proteins are made. During reproduction, DNA is copied so that each new cell receives a complete set of instructions. Sections of DNA called genes are transcribed into RNA and then translated into proteins, which do most of the actual work in cells.
The central dogma, articulated by Francis Crick in 1958, states that genetic information flows in one direction: from DNA to RNA (via transcription) and from RNA to protein (via translation). While some exceptions exist — such as retroviruses that convert RNA back into DNA — this principle describes the fundamental direction of information flow in virtually all living organisms.
The Human Genome Project found that humans have approximately 20,000 to 25,000 protein-coding genes — far fewer than many scientists had predicted. These genes make up only about 1.5% of the total human genome. The rest of the DNA includes regulatory sequences, non-coding RNAs, and repetitive elements whose functions are still being explored.
CRISPR-Cas9 is a gene-editing technology adapted from a bacterial immune system. It uses a guide RNA to direct the Cas9 enzyme to a specific location in a genome, where it makes a precise cut. Scientists can then disable or modify genes with high accuracy. It is important because it offers potential cures for genetic diseases, tools to develop better crops, and new approaches to cancer therapy — though it also raises serious ethical questions.