The previous article ended on a striking idea: what changed gene editing from a bespoke protein-engineering project into a routine lab task was moving the specificity out of the protein and into an RNA molecule. A 20-letter guide RNA now decides where the cut happens. But that sentence skips over the interesting part. A 20-letter RNA is a very small address in a genome of about 3 billion letters, and Cas9 does not float around with the DNA unwound and open for business. So how does it find one site out of billions, and where exactly does it cut?

The answer to both questions is the same three-letter sequence: the PAM.

Two parts, one of which you swap out

Cas9 is a single protein of roughly 1,400 amino acids, and it carries two cutting sites in the same molecule. In its native form it holds two RNA molecules: a short one whose 20 letters match the target DNA, and a longer one that acts as a handle the protein grips. In 2012, Jinek and colleagues fused those two into one molecule, the single guide RNA (sgRNA), which is what labs use today. The part that matters for targeting is the 20-letter spacer; the rest is scaffolding that Cas9 needs in order to fold correctly.

The DNA sequence that the spacer matches is called the protospacer. The name comes from bacteria: the same sequence sits inside the bacterial CRISPR array as a stored "spacer" from a past viral infection, and the copy sitting in the virus is the "proto-spacer." Keep the two ideas separate in your head — they are the same letters, in two different places, doing two different jobs.

The PAM is a letter label the protein reads, not a sequence the guide pairs with

For the Cas9 from Streptococcus pyogenes (SpCas9), the PAM is 5′-NGG-3′, where N means any of the four letters. The site is written out as one continuous 23-letter string, and this is the convention every guide-design tool uses:

15′—[ 20 letters that the guide also spells ]—NGG—3′   non-target strand (displaced)
23′—[ the 20 complementary letters         ]—CCN—5′   target strand (pairs with guide)

Those 23 letters all sit on the same physical strand — and it is the strand the guide does not pair with. The guide RNA pairs with the opposite strand, over the 20 letters only. The PAM is not copied into the guide and is not base-paired with anything the guide carries.

The protein reads it directly. A crystal structure of SpCas9 bound to guide RNA and target DNA showed that two conserved arginines, Arg1333 and Arg1335, reach into the major groove of the DNA helix and hydrogen-bond to the two guanines of the PAM 1. Two details follow from that, and both surprise people:

  • Only the two G's are read with real specificity. The third position, N, is untouched, so NGG covers 4 of the 64 possible triplets — on average one PAM every 16 base pairs. In a human genome that is roughly 190 million of them.
  • Only the strand carrying the PAM is read. The partner nucleotides on the target strand are not contacted at all. A mismatch there does not stop cutting. So the PAM is not an "anti-codon" for anything; it is a three-letter label that the protein feels with its own amino acids.

A related motif, NAG, is weakly accepted, which the same structure explains: a lysine at position 1107 contacts a base just outside the PAM and forces a pyrimidine there, and A satisfies that while other letters do not 1.

Job one: the PAM is the entry point for the search

If Cas9 had to unwind the double helix at every position to test 20 base pairs, searching a eukaryotic genome would be hopeless. It doesn't. Single-molecule and biochemical work showed that Cas9 collides with DNA at random and lets go again almost immediately — unless it lands on a PAM 2. A DNA molecule containing more PAMs holds Cas9 longer, in proportion to how many PAMs it has. The decisive control experiment: a sequence perfectly complementary to the guide RNA but with a mutated PAM next to it is ignored, behaving no differently from random DNA. Nothing is read until a PAM is found.

So the search is not "read every letter until the right one appears." It is "sample the cheap labels, then read the fine print only where a label is present." The PAM acts as a filter that keeps the expensive step — unwinding and testing base pairing — rare.

On how Cas9 travels between the labels, the evidence is mixed. Some studies found no sliding along DNA, just repeated collisions; later single-molecule work has found short local sliding, particularly near PAMs, with the contribution of each mode depending on salt conditions and the assay 3. What the sources agree on is that PAM recognition is the obligatory first step.

Job two: the PAM licenses the unzipping, and sets its direction

Finding a PAM is not the same as finding a target. The protein then has to check whether the letters next to that PAM match the guide. Double-stranded DNA resists being pulled apart, and the cell provides no energy source for the job, so the checking step has to pay for its own unwinding as it goes.

The structure suggests how. Reading the PAM positions the duplex so that the phosphate at the position immediately upstream of the PAM — the "+1 phosphate," on the target strand — is gripped and twisted by a short loop of the protein, nicknamed the phosphate lock loop 1. That distortion melts the first base pairs of the adjacent sequence, which lets the guide's PAM-proximal letters pair up. Each new base pair formed is energetically favourable, and it destabilises the base pair just beyond it, so the RNA–DNA hybrid extends step by step away from the PAM. The same work showed the process is directional: it nucleates at the PAM and runs outward toward the far end of the 20-letter site 2. This directional "zip-up" is why three complementary base pairs in a test tube are not enough — the hybrid has to be able to start at the PAM and keep going.

The PAM also switches the cutting machinery on. Early experiments found that PAM recognition is required not only for binding but for cleavage, even when a guide matches its site perfectly: presenting the target as a single DNA strand without a PAM gives binding but no cutting. PAM recognition acts as an allosteric trigger, remodelling the protein into a catalytically competent state [ref2, ref3].

There is a bacterial reason for this arrangement too, and it is a neat piece of design. Inside a bacterium's own CRISPR array, the stored spacer is flanked by repeating sequences, not by a PAM. Because every step — binding, melting, cutting — is gated on a PAM, the system cannot turn on its own stored archive. Self versus non-self is settled by three letters rather than by any special protection mechanism 2.

Why the letters nearest the PAM decide almost everything

Since the hybrid zips outward from the PAM, a mismatch near the PAM stops the process early, while a mismatch near the far end may never be reached in a way that matters. Experiments confirm exactly that asymmetry. The PAM-proximal stretch of roughly 8 to 12 base pairs is called the seed region; two or three mismatches inside it can abolish cutting, whereas mismatches at the far end are often tolerated [ref4, ref5]. Different assays report different seed lengths — as short as 5 letters in some binding experiments — and the number depends on protein concentration and how long the assay allows the complex to sit 4. What holds across methods is the gradient: PAM-proximal mismatches matter more.

Because R-loop formation is reversible, a partially matching site usually ends in the complex falling off rather than a wrong cut. That is the mechanism behind off-target effects: a site that shares the seed and a PAM can be cut even when the far half of the guide does not match.

Where the cut lands

Now the two nuclease domains. The HNH domain cuts the strand that is base-paired with the guide (the target strand); the RuvC domain cuts the displaced strand (the non-target strand, the one carrying the PAM). Both domains cut between the third and fourth nucleotide upstream of the PAM — that is, three base pairs from the PAM. Because the two breaks line up, SpCas9 leaves blunt ends, with no overhang, rather than staggered ones [ref4, ref5].

Notice what determines the cut position: the PAM, not the far end of the guide. Trimming the guide from 20 to 19 letters barely moves the break. Both domains belong to the same protein, so a single binding event produces a double-strand break — a property that made Cas9 so much easier to use than the earlier tools, where two separate proteins had to be aimed at the same spot.

The break itself is not the edit. The cell repairs it, by non-homologous end joining (which usually leaves a few letters inserted or deleted — useful for shutting a gene off) or, if a template is supplied, by homology-directed repair (which can write in a chosen sequence, but is much less efficient). Repair is the next layer of the story; what matters here is that the machinery hands the cell a break at a predictable place.

What this design buys, and what it costs

The gain is enormous: specificity now lives in 20 letters of RNA you can order from a supplier, and it is re-checked at several points — PAM first, then seed pairing, then full hybrid.

The cost is that the PAM restricts where you can cut at all. With SpCas9, an NGG lands about once every 16 base pairs on average, so targets are plentiful but not universal; places with no PAM nearby simply cannot be edited. This is why labs use Cas9 enzymes from other bacteria, and engineered variants with relaxed or altered PAMs, when the obvious site is blocked. The protein family is not uniform: Cas12a uses a T-rich PAM and makes a staggered cut instead of a blunt one 3.

That constraint shaped real medicine. In the sickle cell therapy described in the previous article, the guide was aimed at a control region of the gene BCL11A, not at HBB itself, with the goal of switching fetal hemoglobin back on. Whichever site a team prefers, the 20 letters had to be chosen at a spot where an NGG already sat next to them.

Splitting the machine into three separable events — PAM recognition, directional R-loop formation, and cleavage — is what makes the rest of the field's questions tractable. Off-target effects, guide design rules, why some guides work better than others, and why engineering the protein changes which sites are editable all come back to how well these three steps stay coupled, and to the three letters sitting next to the target.

Quick check

A guide is perfectly complementary to a site, but the PAM next to it reads TCG instead of TGG. What happens? Almost nothing. PAM recognition comes first and is not based on base pairing, so an unusable PAM means the adjacent sequence is never interrogated. Binding and cleavage both stay off — the guide's perfection is irrelevant.

Why can a single mismatch three letters from the PAM kill the edit, while three mismatches at the far end may not? The RNA–DNA hybrid nucleates at the PAM and extends outward. A mismatch at the start aborts the zip before it gets going; a mismatch at the far end is reached only after a long stable hybrid has already formed, and is more often tolerated.

If the guide must spell 20 letters, why doesn't the cut move when you shorten it? Because cleavage is positioned relative to the PAM: both domains cut three base pairs upstream of it. The guide's length affects how well the site is recognised, not where the scissors close.