The Code and Toolkit of the Double Helix
Every test point that looks scattered is really an extension of the same chemical fact: extension can only proceed from a 3'-OH.
Full text
In the small hours, a graduate student watches the indicator light on the PCR machine blink red every thirty seconds, like a heartbeat. In the 0.2 mL tube in front of him, the target fragment starts as only a few hundred copies; after thirty cycles it will be amplified a billionfold. A thought strikes him — this machine can run without stopping only because, decades ago, someone fished a bacterium out of a hot spring in Yellowstone, and that bacterium's DNA polymerase can survive at 95°C.
The story of DNA is, at its core, a chain of reasoning about "why it is built this way." Behind the Chargaff numbers lies the chemistry of base pairing; the seemingly mundane chemical fact of the 3'-OH end turns out to explain the primer, AZT, ddNTPs, and the direction of proofreading, all at once; the division of labor among the four major repair systems is not something to memorize by system name, but something to understand by recognizing "which kind of damage has occurred." This chapter strings these codes together with the molecular biology toolkit that follows — you will discover that why PCR absolutely requires Taq, why a cDNA library must use reverse transcriptase, and why a YAC can hold the largest insert are all consequences of the very same 3'-OH logic, extended one step further.
Chargaff's Rules and the Three Conformations: Do Not Call B-DNA Left-Handed
Chargaff's rules are not an equation to memorize by force; they are a byproduct of chemical base pairing. Step one: A embraces T with two hydrogen bonds, and G embraces C with three. Step two: therefore, in double-stranded DNA, A must equal T, and G must equal C. Step three: the sum of purines (A+G) must equal the sum of pyrimidines (T+C) = 50%. Step four: given the percentage of one base, you can back-calculate the other three the way you divide candy — T=31% → A follows at 31% → together 62% → the remaining 38% is split evenly between G and C → C=G=19%. From chemical pairing to exam arithmetic, the entire chain of causation is only one step long.
- Double-stranded DNA: A=T, G=C; A+G (purines) = T+C (pyrimidines) = 50%.
- Calculation formula: given T=31% → A=31%, the remaining 38% is split between G and C → C=G=19% (not 31%).
- B-DNA = right-handed, 10 bp/turn, base-pair rise of 3.4 Å (the predominant physiological conformation).
- A-DNA = right-handed, 11 bp/turn, 2.6 Å (dehydrated conditions, RNA-DNA hybrids).
- Z-DNA = left-handed, 12 bp/turn, 3.7 Å (alternating GC sequences, transcriptionally active regions).
- Traps: ① B-DNA listed as 3.6 Å (wrong — it is 3.4); ② B-DNA called left-handed (wrong — left-handed is Z); ③ T=31% leads you to write C as 31% too (wrong — it is 19%).
Full text
Conformation questions are a different kind of trap. B-DNA is the leading actor under physiological conditions — right-handed, 10 bases per turn, base-pair rise of 3.4 Å — and these three numbers function like an ID number: write 3.6 Å, or call it left-handed, and it is no longer B-DNA. Z-DNA is the left-handed oddball, fond of haunting alternating GC sequences and common in transcriptionally active regions; A-DNA is the conformation seen under dehydrating conditions or in RNA-DNA hybrids.
The 3'-OH Rule: One Chemical Fact Underpins All of Replication
DNA polymerase can only extend a tail, never start one — it can add a new nucleotide only onto an existing 3'-OH, always in the 5'→3' direction. This one rule generates five consequences in a row. Step one: without a 3'-OH, there can be no extension → primase must first synthesize a short RNA primer to provide a starting point, which is why the primer used in replication is RNA, not DNA. Step two: at the 3' position, AZT and ddNTPs carry an azido group or a bare hydrogen instead of an OH → once incorporated, no further nucleotide can be added → they become chain terminators. Step three: Sanger sequencing exploits ddNTPs to terminate the reaction randomly and read out fragment length. Step four: the two strands run in opposite directions, but the polymerase only works 5'→3' → the leading strand is synthesized continuously, while the lagging strand can only be made as a series of Okazaki fragments. Step five: proofreading and primer removal run in opposite directions → proofreading is a 3'→5' exonuclease, while primer removal uses a 5'→3' exonuclease.
- The primer in DNA replication is RNA (synthesized by primase, not DNA).
- AZT mechanism = chain termination from the missing 3'-OH; target = HIV reverse transcriptase.
- Proofreading activity = Pol III's 3'→5' exonuclease (Taq lacks this activity → low fidelity).
- Primer removal = Pol I's 5'→3' exonuclease; sealing the nick = DNA ligase (not a polymerase).
- The lagging strand is made of Okazaki fragments; once the primer is excised, the gap is filled in and sealed.
- Traps: ① listing the AZT target as protease/RNase H/host polymerase (wrong — it is reverse transcriptase); ② listing nick-sealing as polymerase/helicase (wrong — it is ligase); ③ listing the proofreading direction as 5'→3' (wrong — it is 3'→5').
Full text
Following this rule, the processing of the lagging strand becomes easy to understand. The polymerase produces the lagging strand as a series of Okazaki fragments, each one preceded by a stretch of RNA primer. Next, DNA pol I removes the primer with its 5'→3' exonuclease and fills in DNA, leaving behind a nick, which DNA ligase then seals. So the role of "sealing the nick" belongs to ligase, not polymerase; the discontinuity of the lagging strand is not fundamentally about two different enzymes, but a compromise forced by directionality.
The target of AZT (zidovudine) is HIV reverse transcriptase, not the host's DNA pol α, protease, or RNase H — this direction must be firmly memorized. Although the drug acts on DNA synthesis, its selectivity comes from reverse transcriptase's affinity for AZT-TP being far higher than that of host enzymes. The ddNTPs used in Sanger sequencing, by contrast, target any DNA polymerase indiscriminately, since their job is simply to terminate the reaction randomly and read out the sequence.
The Four Repair Systems and the SOS Response: Which Kind of Damage Is It
The logic that sorts repair systems has only one rule: look at what the damage looks like. Step one: ask how large the lesion is. Step two: match it to the right tool. When a single base is damaged (deamination, oxidation, appearance of uracil), BER (base excision repair) takes over: DNA glycosylase excises the damaged base → AP endonuclease processes the resulting gap → a polymerase fills it in → ligase seals it. A UV-induced pyrimidine dimer is a "large distorting lesion" that BER cannot handle, so NER (nucleotide excision repair) steps in instead: in prokaryotes, UvrABC, and in eukaryotes, the XP protein family, together excise a stretch of nucleotides containing the lesion, which is then resynthesized. When a newly replicated strand incorporates the wrong base, that is a "mismatch," and MMR (mismatch repair) takes over: MutS recognizes the mismatch → MutL mediates → MutH nicks the "unmethylated" new strand (the old strand is already methylated and serves as the reference) → the new strand is resynthesized. Each of the three systems handles one kind of damage, and they must not be mixed up.
- BER: DNA glycosylase excises the abnormal base (deamination, oxidation, uracil) → AP endonuclease.
- NER: handles large distorting lesions such as UV pyrimidine dimers; deficiency = XP (xeroderma pigmentosum).
- MMR: post-replication mismatches; MutS recognizes, MutH nicks the unmethylated new strand; deficiency = Lynch syndrome / HNPCC.
- SOS: RecA activation → LexA autocleavage (the one being cleaved) → repair genes are derepressed.
- Traps: ① listing DNA glycosylase under MMR (wrong — it belongs to BER alone); ② assigning UV dimers to BER (wrong — they need NER); ③ naming UvrA or RecA as the one broken down in SOS (wrong — it is LexA).
Full text
A fair-skinned little boy breaks out in erythema and freckles after the briefest sun exposure and is diagnosed with skin cancer before the age of ten. He does not simply "burn easily" — he has xeroderma pigmentosum (XP): the enzyme system in his body that is specifically meant to recognize UV-induced pyrimidine dimers has failed.
As for the SOS response, it is the "emergency measure" triggered by massive damage, and its logic resembles a coup d'état. Step one: under normal conditions, the LexA repressor keeps the repair genes (uvrA/B, recA, and others) suppressed. Step two: damage exposes large stretches of single-stranded DNA (ssDNA). Step three: RecA is activated and becomes a co-protease. Step four: RecA promotes the autocleavage of LexA. Step five: once the repressor collapses, all the repair genes are derepressed and transcription begins. It is the repressor LexA that gets cleaved, not any repair-gene product — this is the pitfall the exam loves most to dig.
Why PCR Requires Taq: Heat Resistance Is the Key
The three steps of PCR (polymerase chain reaction) follow an inescapable logic. Step one: denaturation at 95°C separates the two strands. Step two: annealing at 50–65°C lets primers pair complementarily with the template. Step three: extension at 72°C has the polymerase add nucleotides. Step four: every cycle must return to that 95°C station → *E. coli* Pol I/Pol III are irreversibly inactivated at that temperature and would have to be replenished every single cycle, making the reaction unworkable. Step five: Taq polymerase, isolated from *Thermus aquaticus* found in a hot spring, is heat-resistant and therefore irreplaceable. Taq lacks 3'→5' exonuclease activity and indeed has relatively low fidelity, but that is a side effect, not the main reason *E. coli* Pol cannot be used.
- Three steps: 95 / 50–65 / 72°C (denaturation / annealing / extension).
- Main reason for using Taq = heat resistance (*E. coli* Pol is inactivated at 95°C); Taq's lack of proofreading is a side effect.
- One primer pair → one specific segment; multiple sites require multiplex PCR.
- Traps: ① listing the main reason as "Taq has high fidelity" (wrong — it is actually low); ② claiming one primer pair can amplify multiple regions (wrong — only one segment); ③ listing the polymerase used in PCR as *E. coli* Pol (wrong — it would be heat-inactivated).
Full text
Primer specificity also deserves careful thought. A single primer pair recognizes only one uniquely complementary sequence on the template, so one pair amplifies only one segment; to test multiple sites at once, you need multiplex PCR (multiple primer sets).
Libraries, Vectors, and the Three Blots: Reverse Transcriptase Is the Dividing Line
- Genomic library = restriction enzyme + ligase (no reverse transcriptase needed; contains introns).
- cDNA library = reverse transcriptase + ligase (no introns; allows eukaryotic protein expression in prokaryotes).
- RFLP is used for paternity testing, linkage analysis, and DNA fingerprinting (not for building a cDNA library).
- Largest vector = YAC (contains an origin of replication, telomere, and centromere).
- Type II restriction enzymes recognize palindromic sequences; transformation = CaCl₂ + 42°C heat shock; site-directed mutagenesis needs no reverse transcriptase.
- The three blots: Southern = DNA, Northern = RNA, Western = protein (using antibodies).
- Traps: ① adding reverse transcriptase to a genomic library (wrong — not needed); ② using RFLP to build a cDNA library (wrong — unrelated); ③ describing transformation as "low-voltage electrophoresis" (wrong — it is heat shock or electroporation).
Full text · 1 table
To clone a gene, you must first decide what starting material to use and whether introns should be retained. A genomic library is made by fragmenting the entire genome and inserting the pieces into vectors — it contains introns and requires no reverse transcriptase. A cDNA library is made by reverse-transcribing mature mRNA into cDNA and then cloning it — it contains no introns and absolutely requires reverse transcriptase. The latter's greatest use is expressing eukaryotic genes in prokaryotic cells: because prokaryotes lack splicing machinery, expressing a eukaryotic protein requires reverse transcriptase to remove the introns beforehand. RFLP (restriction fragment length polymorphism) is an entirely different matter: it compares fragment lengths after restriction-enzyme digestion and is used in paternity testing, linkage analysis, and DNA fingerprinting — it has nothing to do with building a cDNA library.
| Vector | Capacity | Features |
|---|---|---|
| Plasmid | ~10 kb | Smallest, simplest |
| Phage λ | ~15–20 kb | — |
| Cosmid | ~45 kb | — |
| BAC | ~300 kb | Bacterial artificial chromosome |
| YAC | 100 kb – several Mb | Largest; contains a eukaryotic origin of replication + telomere + centromere |
Swipe or scroll sideways to compare every column; keyboard: focus the table and use arrow keys.
The advantage of a YAC is that it holds the largest insert, not that it has the best transformation or expression efficiency — this is a frequently tested direction. A Type II restriction enzyme recognizes a 4–8 bp palindrome: the 5'→3' reading of the top strand matches the 5'→3' reading of the complementary strand, as in GAATTC↔CTTAAG; if the sequence read across the complementary strand is asymmetric, it is not a Type II target. The standard method for plasmid transformation is preparing competent cells with CaCl₂ plus a 42°C heat shock; another commonly used route is electroporation (a high-voltage pulse, not "low-voltage electrophoresis"). Site-directed mutagenesis is accomplished using a mutation-carrying primer together with a polymerase and requires no reverse transcriptase — this direction is also a favorite on the exam.
Distinguishing the three blots is a gimme question — just remember "what molecule is being detected": Southern blot detects DNA, Northern blot detects RNA (both use nucleic-acid probe hybridization), and Western blot detects protein (using antibodies). The mnemonic SNoW DRoP: S-D, N-R, W-P; only Western uses antibodies.
The polymerase can only extend a tail, never start one, so everything revolves around the 3'-OH — without it there is no primer, no AZT, no Okazaki fragments.
Read-aloud version (copy the whole thing into any TTS)
In the small hours the lab's red indicator light blinks every thirty seconds; the graduate student watches the PCR machine as the target fragment in that 0.2 mL tube is amplified a billionfold over thirty cycles. A thought strikes him: this machine can run without stopping only because someone once fished a bacterium out of a hot spring in Yellowstone, and that bacterium's DNA polymerase can survive at ninety-five degrees Celsius. Every test point in this chapter that looks scattered actually grows out of the very same chemical fact — the rule that extension can only proceed from a 3'-OH.
Start by getting Chargaff straight. In double-stranded DNA, A equals T and G equals C simply because A embraces T with two hydrogen bonds and G embraces C with three — that is merely a byproduct of pairing, not a separate rule. So when a question gives you T at thirty-one percent and asks for C and G, the method is like dividing candy: pair A with T first to use up sixty-two, and split the remaining thirty-eight evenly between G and C, so C and G are both nineteen — never impulsively write C as thirty-one. The conformation questions are also just an ID-number game: B-DNA is right-handed, ten bases per turn, with a base-pair rise of three point four angstroms; writing it as three point six angstroms or as left-handed is always wrong. The left-handed one is Z-DNA, which likes to haunt alternating GC sequences in transcriptionally active regions, while A-DNA is the conformation seen under dehydration or in RNA-DNA hybrids.
Next comes the true star of the show: the 3'-OH rule. DNA polymerase can only add a new nucleotide onto an existing 3'-OH, always in the 5' to 3' direction. That sounds like a plain chemical fact, but it explains several things at once. First, why replication needs a primer: the polymerase cannot initiate from scratch, so something must first provide a short tail carrying a 3'-OH for it to extend, and that is why primase synthesizes a short RNA primer as a starting scaffold — the primer in DNA replication is therefore RNA, not DNA. Second, why AZT and dideoxynucleotide triphosphates are chain terminators: their 3' position carries an azido group or a bare hydrogen instead of an OH, so once incorporated, nothing further can be attached. Third, why Sanger sequencing works at all: it exploits dideoxynucleotides to stop the reaction randomly and then reads out fragment length. Fourth, why the leading strand is continuous while the lagging strand is not: the two strands run in opposite directions, but the polymerase only works 5' to 3', so the lagging strand can only be built in the reverse direction as a series of Okazaki fragments. Fifth, proofreading activity and primer-removal activity must be kept separate: proofreading is a 3' to 5' exonuclease, while primer removal is a 5' to 3' exonuclease, each running in the opposite direction to do its own job. Taq lacks proofreading activity and therefore has low fidelity, but that is a side effect, not the main reason PCR cannot use E. coli polymerase — that main reason is always heat resistance.
The gap left behind once Okazaki fragments are finished is sealed by DNA ligase; the exam often swaps this role for polymerase as a trap, but sealing the gap is always ligase's job. Do not misremember the target of AZT either: it locks onto HIV reverse transcriptase, not the host's DNA polymerase alpha, not protease, and not RNase H, because reverse transcriptase's affinity for AZT triphosphate far exceeds that of host enzymes, allowing it to selectively lock down the virus while sparing the host.
Repair systems are not something to memorize by name; what matters is recognizing what the damage looks like. When a single base is damaged — deamination, oxidation, the appearance of uracil — the job calls for base excision repair, in which DNA glycosylase excises the damaged base and hands it off for further processing. A UV-induced pyrimidine dimer is a large distorting lesion that base excision repair cannot handle, so nucleotide excision repair takes over instead, which is why patients with xeroderma pigmentosum have a nucleotide excision repair defect. When a newly replicated strand incorporates the wrong base, that is called a mismatch, and it calls for mismatch repair, in which MutS recognizes the mismatch, MutL mediates, and MutH nicks the unmethylated new strand, since the old strand is already methylated and can serve as the reference; Lynch syndrome is a mismatch repair defect. Remember that DNA glycosylase belongs to base excision repair alone — the exam loves to plant it inside mismatch repair as a trap. The SOS response is the emergency measure triggered by massive damage, and its logic resembles a coup d'état. Under normal conditions the LexA repressor suppresses the repair genes, including uvrA, uvrB, recA, and others; damage exposes large stretches of single-stranded DNA, RecA is activated into a co-protease, and it promotes the autocleavage of LexA — once the repressor collapses, all of these repair genes are derepressed and begin transcription. It is the repressor LexA that gets cleaved, not any repair-gene product, and not RecA itself; RecA is the catalyst, not the one broken down. Remember one line: LexA falls, repair rises.
Why PCR absolutely requires Taq is, in fact, simple. The three-step cycle denatures at ninety-five degrees, anneals at fifty to sixty-five degrees, and extends at seventy-two degrees, and every round must return to that ninety-five-degree station; E. coli polymerase is irreversibly inactivated at that temperature and would have to be replenished every cycle, making the reaction unworkable — which is exactly why the heat-tolerant Taq, fished out of a hot spring, is needed. Taq lacks 3' to 5' exonuclease activity and so has relatively low fidelity, but that is a side effect, not the reason other enzymes cannot be used; the main reason is, in two words, heat resistance. Primer specificity also deserves careful thought: a single primer pair recognizes only one uniquely complementary sequence on the template and therefore amplifies only one segment; testing multiple sites at once requires multiplex PCR.
Distinguishing the two types of library is a frequently tested comparison. A genomic library is made by fragmenting the entire genome and inserting the pieces into vectors, so it contains introns and requires no reverse transcriptase; a complementary DNA library is made by reverse-transcribing mature mRNA into cDNA and then cloning it, so it contains no introns and absolutely requires reverse transcriptase — which is also why a cDNA library must be used to express a eukaryotic gene in prokaryotic cells, since prokaryotes lack splicing machinery and the introns must be removed beforehand. RFLP is an entirely different matter: it compares fragment lengths after restriction-enzyme digestion for paternity testing, linkage analysis, and DNA fingerprinting, and has nothing to do with building a cDNA library — the exam loves to slip it in as a decoy.
Up the ladder of vectors, capacity keeps growing: a plasmid holds about ten kilobases, phage λ about fifteen to twenty, a cosmid about forty-five, a BAC about three hundred, and a YAC can reach anywhere from a hundred kilobases up to several megabases — the YAC is the one that holds the largest insert, because it carries a eukaryotic origin of replication, a telomere, and a centromere, letting it persist stably the way a eukaryotic chromosome does. Remember that its advantage is sheer capacity, not the best transformation efficiency. A Type II restriction enzyme recognizes a palindromic sequence, where the 5' to 3' reading of the top strand matches the 5' to 3' reading of the complementary strand; if the sequence read across the complementary strand is asymmetric, it is not that enzyme's target. The standard method for plasmid transformation is preparing competent cells with calcium chloride plus a forty-two-degree heat shock; another route is electroporation, not low-voltage electrophoresis. Site-directed mutagenesis can be accomplished simply with a mutation-carrying primer plus a polymerase, requiring no reverse transcriptase — this direction is also a favorite on the exam. Finally, for the three blotting techniques, just remember what molecule is being detected: Southern detects DNA, Northern detects RNA, both by nucleic-acid probe hybridization, while Western detects protein using antibodies; the mnemonic is "snow," and only Western uses antibodies. One rule about the 3'-OH, combined with one rule that the type of damage determines the repair system, brings the whole chapter's test points to life.