# From One Letter to One Disease: What a Gene Is, How a Mutation Breaks It, and Why Precise Editing Is a Big Deal

How sickle cell disease arises from one swapped DNA letter, and why moving specificity from engineered proteins into a guide RNA made rewriting genes practical

> genetics · sickle cell disease · gene editing history · About 10 min · Oct 6

## Key points

1. A gene is a stretch of DNA whose letter order specifies the amino acid sequence of one protein; the ribosome reads the letters three at a time as codons, and because 64 codons map to only 20 amino acids plus stop, the code is redundant.
2. Mutations differ in kind: silent changes do nothing, missense swaps one amino acid, nonsense truncates the protein, and frameshift shifts the reading frame and usually destroys the protein — so a mutation's effect depends on where it lands and what the protein does.
3. Sickle cell disease comes from a single A-to-T swap in codon 6 of HBB (GAG to GTG), replacing surface glutamic acid with valine; the oily valine fits a hydrophobic pocket exposed on neighbouring deoxygenated hemoglobin molecules, creating a new sticky contact that builds rigid fibres and sickles the red blood cell.
4. Hemoglobin S still carries oxygen; the disease comes from one new interaction plus quantity, because fibre formation depends on roughly the 34th power of HbS concentration — which is why raising fetal hemoglobin, which dilutes HbS and does not join the fibres, prevents sickling.
5. Earlier approaches could add a working gene copy (useless for dominant mutations, random insertion, wrong regulation) or cut DNA using ZFNs and TALENs, where specificity lived in a protein and every new target meant engineering a new protein.
6. CRISPR-Cas9 moved specificity into a roughly 20-letter guide RNA that base-pairs with the target, so retargeting means changing RNA rather than protein; the only extra requirement is a short PAM motif (NGG for the standard Cas9), present on average every 8–12 bases in the human genome.
7. A cut is not an edit: the cell's NHEJ repair is fast, template-free, works in any cell cycle phase and usually breaks the gene, while HDR can write an exact sequence from a donor template but works mainly in dividing cells, takes many hours, and typically succeeds in only 0.5% to 20% of chromosomes.
8. Casgevy, approved in December 2023, avoids the hard step by targeting the BCL11A erythroid enhancer rather than the HBB mutation: NHEJ disruption of the GATA1 binding site keeps fetal hemoglobin switched on, and 93.5% of 44 patients in a single-arm pivotal study went at least 12 months without a severe vaso-occlusive crisis.
9. The main known risk of CRISPR therapeutics is off-target editing, and because only somatic cells are edited, the change is not inherited — the heritable-editing and regulation debates turn on that distinction.

---

In December 2023, the US Food and Drug Administration approved the first medicine that treats a disease by editing a patient's own DNA with CRISPR. It is called Casgevy, and it is approved for sickle cell disease — a condition caused by a change of exactly one letter in a person's genome. That single fact contains the whole story of this article: one letter can cause a serious disease, and it took decades of work to be able to change DNA on purpose. To see why that is impressive, and why it is still hard, you need three things: what a gene is, how one letter change becomes a disease, and what earlier tools could and could not do.

## A gene is a stretch of DNA that spells out a protein

DNA is a long molecule made of two strands wound around each other. Each strand is a chain of four building blocks — usually written A, T, G and C — and the two strands pair up in a fixed way: A with T, G with C. The order of those letters is the information. Your genome is about 3 billion letters long, split across 23 pairs of chromosomes.

A **gene** is a section of that sequence that carries the instructions for one protein. Cells use a gene in two steps. First they copy it into a working molecule called messenger RNA. Then a structure called the ribosome reads that RNA three letters at a time. Each three-letter group is a **codon**, and each codon means one **amino acid** — the small molecules that get strung together to build a protein. Four possible letters in groups of three give 64 codons, but only 20 amino acids plus a "stop" signal are used. That mismatch means the code is **redundant**: several different codons can specify the same amino acid.

Keep one specific gene in mind for the rest of this article. **HBB** sits on chromosome 11 and encodes β-globin, one of the four protein chains that make up hemoglobin — the molecule in red blood cells that carries oxygen. HBB is read as 146 amino acids in a row. Amino acid number 6 is where our story happens.

## How one letter changes a protein

A **mutation** is any change in the DNA sequence. Because the ribosome reads in groups of three, what a mutation does depends heavily on its type:

- **Silent**: the letter changes but the codon still means the same amino acid. Redundancy protects the protein, and nothing changes.
- **Missense**: the codon now means a different amino acid. One brick in the protein is swapped for another.
- **Nonsense**: the codon becomes a "stop", so the protein is cut short.
- **Frameshift**: one or two letters are inserted or deleted. Since three letters make a codon, the reading frame shifts, and every codon after that point is misread. This usually destroys the protein completely.

Notice that "mutation" and "disease" are not the same thing. Plenty of mutations do nothing noticeable. What matters is where the change lands and what the protein has to do.

## One amino acid, in the wrong place, with the wrong chemistry

Sickle cell disease is caused by a missense mutation in HBB. At codon 6, the letters GAG become GTG — a single A-to-T swap. GAG codes for glutamic acid; GTG codes for valine. So amino acid 6 of β-globin changes from glutamic acid (Glu) to valine (Val). The mutation is written as p.Glu6Val.

The interesting question is why that swap causes disease. The answer is chemistry plus location.

Glutamic acid carries a negative charge and dissolves comfortably in the watery environment inside a cell. Valine is uncharged and oily — it dislikes water. Amino acid 6 sits on the **surface** of the folded β-globin chain, in contact with the surrounding fluid. This matters for a second reason: when hemoglobin releases its oxygen, it changes shape and exposes a greasy pocket on the surface of the β chain. In normal hemoglobin nothing fits there. In hemoglobin S, the oily valine at position 6 does fit — it slides into the pocket of a neighbouring deoxygenated hemoglobin molecule.

That creates a new sticky contact between molecules that simply does not exist in normal hemoglobin. The molecules link up into long, rigid fibres inside the red blood cell. The fibres force the cell out of its usual flexible disc shape into a stiff crescent, or sickle. Sickled cells get stuck in small blood vessels, which causes the episodes of severe pain and organ damage that define the disease, and they rupture early, which causes anemia.

Three things are worth pulling out of this:

1. **The mutation does not break the protein's main job.** Hemoglobin S still binds oxygen reasonably well. The damage comes from one new sticky patch on the outside that makes molecules clump — a new interaction, not a lost function.
2. **Quantity dominates.** The rate at which hemoglobin S fibres form depends on roughly the 34th power of its concentration inside the cell. That extreme sensitivity is why small changes in how much hemoglobin S a red cell contains translate into large changes in disease severity.
3. **One good copy is mostly enough.** A person with one HbS copy and one normal copy has sickle cell trait, which is usually mild. Disease generally requires two HbS copies, or HbS plus another faulty β-globin variant. (The trait persists in some regions partly because carrying it gives some protection against malaria.)

## Why "rewriting genes precisely" is a different kind of power

Before genome editing, medicine had one main genetic strategy: **gene addition**. Package a working copy of the gene into a modified virus, and let cells make the protein from that extra copy. This works for diseases where the problem is a missing or weak protein. It has real limits. It does not remove the faulty gene, so it cannot help **dominant** mutations, where the damaged protein actively causes harm. The added copy lands at semi-random places in the genome, which risks disrupting other genes. And the extra copy is not controlled by the gene's own on/off switches, so the amount and timing of protein may be wrong.

The alternative is to edit the existing gene in place. That requires cutting DNA at a chosen address — which is where the technology history matters.

**Zinc finger nucleases (ZFNs)**, developed in the 2000s, were the first practical targeted cutters. They fuse a DNA-binding domain (zinc fingers, each recognising roughly three bases) to a cutting domain from the enzyme FokI, which only cuts when two copies pair up — so two ZFN proteins are aimed at sites flanking the target. The idea is elegant, but zinc fingers do not behave independently: how one finger binds depends on its neighbours. Assembling a working, specific pair can require selection and real protein-engineering expertise; groups have reported that non-specialists may need months, and the first open-source reagents could only be aimed at a site roughly every 200 bases of DNA.

**TALENs** came next. Their DNA-binding repeats are 33–35 amino acids long, and one repeat reads one base, so designing them is closer to spelling out the target sequence; a pair can be designed in as little as a couple of days. The catch is size: each monomer is around 3,000 base pairs of DNA, which makes delivery into cells harder.

The shared bottleneck is that in both systems, **the specificity lives in a protein**. A new target means designing and building a new protein.

CRISPR changed that. In the CRISPR-Cas9 system, the cutting protein — Cas9 — is the same every time. Specificity comes from a short **guide RNA** of about 20 letters that base-pairs with the target DNA, following A-T and G-C rules you already know. Cas9 is a nuclease with two cutting sites, one for each DNA strand, and the guide RNA tells it where to cut. To aim at a new location, you change the sequence of the RNA, not the protein.

Two details make this practical. First, Cas9 needs a short motif called a **PAM** (for the commonly used Cas9, the letters NGG, where N is any base) immediately next to the target. In the human genome such sites occur on average every 8–12 bases, so this requirement rarely blocks you. Second, because only the RNA changes, it is easy to deliver several guides at once and edit several places in the same cell.

So the shift is this: targeting DNA went from a protein-engineering problem — a lab, months of work, specialist skill — to a base-pairing problem you can solve by ordering a short piece of RNA. That is what made editing cheap and widespread.

## But a cut is not yet an edit

Here is the part that is often glossed over. Cas9 makes a **double-strand break** — it cuts both strands of the DNA. What happens next is decided by the cell, not by the tool. Cells have two main repair routes.

The first is **non-homologous end joining (NHEJ)**. It needs no template, it works throughout the cell cycle, it is fast (around half an hour), and it is the dominant route. It glues the broken ends back together, but imperfectly, leaving small insertions or deletions. If that happens inside a gene, the gene is usually rendered non-functional. NHEJ is therefore wonderful for **knocking a gene out** — and useless for writing a specific sequence back in.

The second is **homology-directed repair (HDR)**. If you supply a DNA template carrying the sequence you want, the cell can copy from it and repair the break exactly as specified. This is the true "rewrite". It is also the hard one: it works mainly in cells that are actively dividing (the S and G2 phases of the cell cycle), it takes seven hours or more, and reported efficiencies are typically in the range of 0.5% to 20% of chromosomes. In non-dividing cells such as neurons, it is much worse.

```mermaid
flowchart LR
  A["Cas9 + guide RNA"] --> B["Cut at the target site"]
  B --> C{"Cell repairs the cut"}
  C -->|"No template: NHEJ"| D["Small indels: gene usually broken"]
  C -->|"Donor template: HDR"| E["Exact sequence written in: a true rewrite"]
  D --> F["Easy: used to disable the BCL11A enhancer"]
  E --> G["Hard: low efficiency, dividing cells only"]
```

You also have to worry about **off-target editing**: the guide RNA can tolerate some mismatches, especially away from the PAM end, so Cas9 occasionally cuts somewhere unintended.

## The first approved CRISPR therapy picked a target where breaking is the goal

Now the sickle cell therapy makes sense. Casgevy is a one-time treatment; the label initially covered patients 12 and older with recurrent vaso-occlusive crises, and the approved age range has since been lowered to 2 and up.

The striking thing is what it edits. It does **not** correct the HBB mutation. Instead, the guide RNA directs Cas9 to cut inside an enhancer of a different gene, **BCL11A** — specifically the binding site for a transcription factor called GATA1 in the enhancer that is active in red blood cell precursors. BCL11A is the switch that turns off fetal hemoglobin (HbF) after birth. Breaking that switch region reduces BCL11A in red cell precursors, so the genes for fetal hemoglobin stay switched on.

Raising fetal hemoglobin works because it dilutes hemoglobin S inside the cell and does not join the fibres the way HbS does. Given how steeply polymerisation depends on HbS concentration, that change is enough to stop the cells from sickling quickly.

Why that target? Because correcting the HBB point mutation exactly would require HDR in blood stem cells — inefficient, and difficult to get into enough cells. Disabling an enhancer requires only NHEJ, which the cell performs readily. The therapy deliberately sidesteps the hardest step in the technology.

The procedure is still heavy. Blood stem cells are collected from the patient, edited outside the body, and infused back after chemotherapy that clears space in the bone marrow. In the pivotal single-arm study of 44 patients, 93.5% went at least 12 months without a severe vaso-occlusive crisis. There was no placebo group, the treatment is expensive, and the main identified risk is unintended off-target editing, which regulators required to be followed up over the long term. Because only the patient's own body cells are edited, the change is not inherited.

## A framework for judging the next big claim

When you read about a new gene-editing treatment, ask three questions. Is the goal to **break** something or to **rewrite** something? Breaking needs only NHEJ and is comparatively easy; rewriting needs HDR plus a template, and its low efficiency in the right cells is usually the real obstacle. Is the edit done in cells that stay in the patient — or in embryos or sperm and egg cells, where it would be inherited by their children? And what happens elsewhere in the genome when the guide RNA lands somewhere unintended? Ethical and regulatory arguments about gene editing are mostly arguments about those last two questions, and they only make sense once you know what a gene is and how a single letter can turn into a disease.

## Sources

1. [FDA: FDA Approves First Gene Therapies to Treat Patients with Sickle Cell Disease — announcement of the first CRISPR/Cas9-based therapy approval (December 2023)](https://www.fda.gov/news-events/press-announcements/fda-approves-first-gene-therapies-treat-patients-sickle-cell-disease)
2. [FDA: December 8, 2023 Summary Basis for Regulatory Action — CASGEVY — BCL11A enhancer target, single-arm trial of 44 patients, 93.5% VF12 response, off-target editing risk](https://www.fda.gov/media/175179/download)
3. [FDA: Package Insert — CASGEVY — current age indication, mechanism, and off-target genome editing warning](https://www.fda.gov/media/174615/download)
4. [Pathophysiology of Sickle Cell Disease (Annual Review of Pathology) — p.Glu6Val, HbS polymerisation, and its steep concentration dependence](https://www.annualreviews.org/content/journals/10.1146/annurev-pathmechdis-012418-012838)
5. [GeneReviews: Sickle Cell Disease — genetic basis and range of genotypes](https://www.ncbi.nlm.nih.gov/books/NBK1377/)
6. [Genome Engineering with Targetable Nucleases (Annual Review of Biochemistry) — ZFN, TALEN and CRISPR/Cas mechanisms and the advantages of RNA-guided targeting](https://www.annualreviews.org/content/journals/10.1146/annurev-biochem-060713-035418)
7. [JCI: Expanding the genetic editing tool kit: ZFNs, TALENs, and CRISPR-Cas9 — retargeting effort, site-selection limits and delivery constraints](https://jci.org/articles/view/72992)
8. [Molecular Therapy Nucleic Acids: CRISPR-Cas9-mediated homology-directed repair for precise gene editing — NHEJ versus HDR, cell-cycle restriction and HDR efficiency ranges](https://www.cell.com/molecular-therapy-family/nucleic-acids/fulltext/S2162-2531(24)00231-2)

---

Original article: https://eulore.ai/articles/single-letter-mutation-to-gene-editing-cfb5d9d8

> **Eulore** · Learn a little. Understand a lot.
>
> Eulore is an AI learning tool that turns what you want to learn into a continuing series. Share a topic, and it gets to know your starting point before creating articles you can read in 5–10 minutes. Ask as you read, and shape what comes next.This article was created in the same way.
>
> Start your own series → https://eulore.ai
