- Home
- Module 6: Genetics, Evolution and Ecosystems
- Manipulating Genomes
Manipulating Genomes¶
Part of Module 6: Genetics, evolution and ecosystems.
The ability to read, copy and edit DNA sequences has transformed both medicine and biology. This topic covers the key techniques: sequencing (reading the order of bases), amplification (copying a specific sequence), electrophoresis (separating by size), and genetic engineering (inserting sequences from one organism into another). Each technique has direct medical and commercial applications, but also raises ethical questions that are part of the specification.
What You Need to Learn¶
Further detail: AS Biology A (H020) and A Level Biology A (H420).
How DNA is sequenced and what sequencing is used for, how DNA profiling uses PCR and gel electrophoresis, how genes are transferred into other organisms by genetic engineering and how transformed cells are identified, the ethical issues involved, and how gene therapy aims to treat genetic disease.
DNA Sequencing¶
The Sanger Chain-Termination Method¶
The chain-termination (dideoxy) method, developed by Frederick Sanger, was the first widely used DNA sequencing technique and remains the basis of modern sequencing.
Principle: Modified nucleotides (dideoxynucleotides, ddNTPs) lack the 3'-OH group needed to add the next nucleotide. When a ddNTP is incorporated into a growing chain, extension stops. By using a mixture of normal nucleotides and a small proportion of fluorescently labelled ddNTPs, a population of fragments of every possible length is generated.
Procedure:
- The DNA sample to be sequenced is denatured into single strands
- A reaction mixture is prepared containing:
- Single-stranded template DNA
- Primer (complementary to the template sequence)
- Four standard dNTPs
- DNA polymerase
- A small proportion of each of the four fluorescently labelled ddNTPs (each base labelled with a different colour)
- DNA polymerase extends the primer; when a ddNTP is incorporated, extension stops, generating a fragment of a specific length
- Many reactions occur simultaneously, producing fragments of every possible length (terminating at every possible base)
- Fragments are separated by high-resolution capillary gel electrophoresis — shorter fragments migrate faster, separating fragments that differ by a single base
- A laser detects the fluorescent label on the terminal ddNTP of each fragment
- The sequence is read from the electropherogram (the colour and order of fluorescent peaks)
High-Throughput (Next-Generation) Sequencing¶
Modern sequencing has evolved from Sanger's single-reaction method to massively parallel systems that sequence millions of fragments simultaneously. Key features:
- DNA is fragmented into short pieces (~100–300 bp)
- All fragments are sequenced at once (parallelisation)
- Computational assembly reconstructs the whole genome from overlapping fragments
- Whole human genomes can now be sequenced in hours rather than years
Applications of DNA Sequencing¶
| Application | Description |
|---|---|
| Evolutionary comparisons | Comparing gene or genome sequences between species reveals evolutionary relationships; molecular phylogenetics; the more similar the sequence, the more closely related the species |
| Medical diagnosis | Sequencing to identify disease-causing mutations in individual patients |
| Personalised medicine | Genome-wide sequencing identifies variants that influence drug metabolism or disease susceptibility; therapies can be tailored to an individual's genome |
| Prediction of protein sequence | The amino acid sequence of a protein can be inferred from its coding DNA sequence |
| Synthetic biology | Designed DNA sequences can be synthesised chemically and inserted into organisms to produce novel functions |
Genomics, Bioinformatics and Proteomics¶
Advances in sequencing have produced enormous datasets requiring computational analysis:
- Genomics: the study of whole genomes using DNA sequencing and computational biology. The Human Genome Project (completed 2003) mapped the entire human genome and made the data publicly available.
- Bioinformatics: the development of software, computing tools and mathematical models to store, retrieve and analyse biological data (nucleotide sequences, protein sequences, gene expression data).
- Computational biology: uses bioinformatics tools and biological data to model biological systems.
- Proteomics: the study of the complete set of proteins expressed by a genome (the proteome). The proteome is more complex than the genome because of alternative splicing and post-translational modification: one gene can give rise to multiple protein variants.
- DNA barcoding: comparing a short standardised DNA sequence from an unknown organism to a reference database to identify the species or establish evolutionary relationships.
Sequencing pathogen genomes is also valuable for: identifying sources and transmission routes of disease outbreaks; detecting antibiotic-resistant strains; developing new vaccines and drug targets.
DNA Profiling¶
DNA profiling (also called DNA fingerprinting or genetic fingerprinting) is used to identify individuals or determine genetic relationships. It exploits regions of non-coding, repetitive DNA that vary widely between individuals:
- Variable number tandem repeats (VNTRs) (also called minisatellites): repeated DNA sequences where the number of repeats varies between individuals. They are heritable, located across the genome, and have no protein-coding function. With the exception of identical twins, no two individuals share identical VNTR patterns.
- Short tandem repeats (STRs) (also called microsatellites): shorter repeated sequences (2–6 bp) that are used in modern profiling because they can be amplified accurately by PCR from small or degraded samples.
A high similarity in VNTR or STR patterns between two individuals indicates they are likely closely related. The probability of two unrelated individuals sharing the same profile across 10–13 STR loci is approximately 1 in 10¹³.
Polymerase Chain Reaction (PCR)¶
PCR amplifies a specific target DNA sequence exponentially. It is essential before profiling because biological samples (blood, hair, saliva) often contain only tiny amounts of DNA.
Components of a PCR reaction:
- Template DNA (the sample to be amplified)
- Primers (short, single-stranded oligonucleotides complementary to either end of the target sequence)
- Free deoxynucleotide triphosphates (dNTPs)
- Heat-stable DNA polymerase (e.g. Taq polymerase, from Thermus aquaticus)
- Buffer solution
PCR cycle (repeated ~30 times):
| Step | Temperature | What happens |
|---|---|---|
| Denaturation | ~95 °C | Hydrogen bonds between base pairs break; double-stranded DNA → two single strands |
| Annealing | 50–65 °C (primer-dependent) | Primers bind (anneal) to their complementary sequences at each end of the target sequence |
| Extension | ~72 °C | Taq polymerase adds nucleotides from the 3' end of each primer, synthesising new complementary strands |
After 30 cycles: 2³⁰ ≈ 10⁹ copies of the original sequence.
Worked example: how much DNA does PCR make?
Each cycle doubles the number of copies of the target sequence.
Copies after n cycles = starting copies × 2ⁿ
One starting molecule after 30 cycles: 2³⁰ ≈ 1.1 × 10⁹ copies.
This is why a single cell's DNA from a crime scene can give enough material for profiling. The doubling needs a heat-stable polymerase, such as Taq polymerase, because each cycle heats the mixture to about 95 °C to separate the strands.
Gel Electrophoresis¶
Gel electrophoresis separates DNA (or protein) fragments by size using an electric current.
Procedure:
- Agarose gel is prepared (a porous matrix that acts as a molecular sieve)
- DNA samples are loaded into wells at one end of the gel
- An electric current is applied; DNA fragments are negatively charged (due to phosphate groups) and migrate towards the positive electrode
- Smaller fragments move faster and further; larger fragments move slower and not as far
- Fragments are visualised using ethidium bromide (intercalates into DNA; fluoresces orange under UV light) or SYBR Green, or radioactive labels
- A DNA ladder (marker with known fragment sizes) is run in a separate lane for size comparison
DNA profiling using STRs compares the pattern of bands produced after PCR amplification of multiple STR loci. The probability of two unrelated individuals having the same profile across 10–13 STR loci is vanishingly small (approximately 1 in 10¹³).
Uses of DNA profiling:
- Forensic identification (match crime-scene DNA to a suspect or eliminate suspects)
- Paternity testing
- Identification of remains
- Relationship testing in wildlife conservation
- Assessing genetic diversity within populations (e.g. for conservation breeding programmes)
Limitations of DNA profiling:
- Environmental contamination or sample degradation may compromise results
- Close genetic relatives (siblings, parents) can have similar profiles
- Match probability calculations assume population independence; this may not hold in small communities
- A matching profile does not by itself prove presence at a crime scene; other evidence is required
Worked example: comparing DNA profiles
DNA from a crime scene and from three suspects is cut, amplified at three STR sites and run on a gel. The numbers show the repeat counts at each site.
| Sample | STR 1 | STR 2 | STR 3 |
|---|---|---|---|
| Crime scene | 12, 14 | 9, 11 | 15, 15 |
| Suspect A | 12, 14 | 9, 11 | 15, 15 |
| Suspect B | 12, 13 | 9, 11 | 15, 16 |
| Suspect C | 11, 14 | 10, 11 | 15, 15 |
Only Suspect A matches at all three sites. B and C each differ at least once, so they are excluded. Matching at more STR sites makes a chance match less likely, which is why several sites are used. A match shows that the suspect could be the source, and does not prove it.
Genetic Engineering¶
Genetic engineering is the direct manipulation of an organism's genome by inserting, deleting or modifying specific DNA sequences. The key application is the production of transgenic organisms — organisms carrying a gene from a different species.
Key Tools¶
Restriction endonucleases (restriction enzymes):
- Cut DNA at specific recognition sequences (palindromic sequences of 4–8 base pairs)
- Different enzymes recognise different sequences (e.g. EcoRI recognises GAATTC)
- Most produce sticky ends — short, single-stranded overhangs — which are complementary to overhangs produced by the same enzyme in another piece of DNA
- Some cut bluntly
DNA ligase:
- Seals the phosphodiester bonds between fragments of DNA
- Used to join the insert (gene of interest) to the vector
Vectors:
- Vehicles that carry the DNA insert into the host cell
- Most common vector: plasmid (small circular DNA molecule found naturally in bacteria)
- Other vectors: bacteriophages (viruses that infect bacteria), yeast artificial chromosomes (YACs), liposomes (lipid vesicles, used for gene therapy in humans)
Procedure for Producing a Recombinant Plasmid¶
- Isolate the gene of interest from donor DNA using restriction enzymes (or synthesise it chemically using the mRNA sequence and reverse transcriptase to produce cDNA)
- Cut the vector (plasmid) with the same restriction enzyme — both gene and plasmid now have complementary sticky ends
- Mix the gene and plasmid under conditions that allow complementary base pairing between sticky ends
- DNA ligase seals the phosphodiester bonds, creating a recombinant plasmid
- Introduce the recombinant plasmid into a host bacterium (e.g. E. coli) using electroporation — brief high-voltage pulses that temporarily increase membrane permeability — or calcium ion treatment and heat shock
- Bacteria that have successfully taken up the recombinant plasmid are identified using marker genes
Explore Recombinant Plasmid Formation¶
Use the interactive below to walk the sequence from donor gene to transformed bacterium. The route toggle is especially useful here because you need to keep the direct restriction-enzyme route separate from the mRNA -> cDNA -> double-stranded DNA -> sticky-ended DNA route. Open full interactive.
How to read the model
This is a process map, not a full laboratory protocol. It simplifies recognition sequences, later marker-gene screening and colony selection so the core order of isolation, reverse transcription or direct cutting, ligation and transformation stays clear.
Exam technique
For genetic engineering, name each enzyme and its job in sequence: reverse transcriptase or restriction enzymes to obtain the gene, restriction enzymes to cut the plasmid, ligase to join them, and a vector to carry the gene into the host. Explain why the same restriction enzyme is used (complementary sticky ends). For PCR, name the components, the temperature of each step and why each is used.
Identifying Transformed Bacteria Using Marker Genes¶
Antibiotic resistance markers:
- If the plasmid contains a gene for antibiotic resistance (e.g. ampicillin resistance)
- Bacteria are grown on media containing that antibiotic
- Only bacteria that took up the plasmid survive (they are resistant)
- However, this only identifies bacteria that took up any plasmid, not necessarily a recombinant one (with the insert)
Insertional inactivation:
- The gene of interest is inserted into the middle of a second marker gene (e.g. a lacZ gene producing β-galactosidase, which turns a substrate blue)
- Bacteria with non-recombinant plasmid (no insert): marker intact → colonies turn blue
- Bacteria with recombinant plasmid (insert disrupts marker): marker non-functional → colonies remain white
- White colonies contain the recombinant plasmid with the gene of interest
Applications¶
- Human insulin production: insulin gene inserted into E. coli plasmid; bacteria produce human insulin for diabetes treatment
- Human growth hormone: produced in engineered bacteria
- Chymosin (used in cheese production): produced from engineered fungi; replaces animal rennet
- Herbicide-resistant crops (e.g. glyphosate-resistant soya): herbicide-resistance gene inserted into crop genome; allows use of herbicide on fields containing only the crop
- Insect-resistant crops (Bt crops): Cry toxin genes from Bacillus thuringiensis inserted into crop genome; crops produce their own insecticide
Ethical Considerations¶
Arguments for genetic engineering of crops and microorganisms:
- Increased crop yields; reduced pesticide use
- Production of medicines (insulin, vaccines) at low cost
- Potential to alleviate nutritional deficiencies (e.g. Golden Rice with β-carotene)
Arguments against:
- Potential environmental effects (gene flow to wild relatives; effects on non-target organisms)
- Reduced genetic diversity in crops (monocultures at risk)
- Commercial control of seed supply; impact on small-scale farmers in developing countries
- Concerns about long-term safety of consuming GM foods (contested; no confirmed harmful effects yet)
- Ethical objection to crossing species boundaries
Gene Therapy¶
Gene therapy is the insertion of a normal (functional) allele into cells of a person with a genetic disorder caused by a faulty allele. There are two approaches:
| Type | Target cells | Effect passed to offspring? | Notes |
|---|---|---|---|
| Somatic gene therapy | Body cells (e.g. lung epithelial cells in cystic fibrosis) | No | Temporary; cells divide and replace treated cells, so repeated treatment needed |
| Germline gene therapy | Fertilised egg or early embryo | Yes (heritable change) | Permanent cure; affects all cells; raises major ethical concerns; not currently permitted in most countries |
Vectors for gene therapy:
- Retroviruses: integrate the therapeutic gene into the host cell chromosome (permanent expression); risk of insertional mutagenesis (integration near an oncogene could cause cancer)
- Adenoviruses: infect many cell types; do not integrate (transient expression); may trigger immune response
- Liposomes: lipid vesicles that fuse with the cell membrane; low efficiency; no immune response; do not integrate
Example — cystic fibrosis:
- CFTR gene (coding for the cystic fibrosis transmembrane conductance regulator) is inserted into a liposome or adenovirus vector
- Vector is inhaled as an aerosol
- The gene is expressed in lung epithelial cells, producing functional CFTR protein
- Current treatments still experimental; clinical success has been limited
Common Confusions¶
- PCR and DNA replication: PCR copies a chosen section of DNA in vitro using a heat-stable polymerase and primers, whereas replication copies whole chromosomes in the cell.
- Restriction enzymes and ligase: restriction enzymes cut DNA, and ligase joins it.
- Vector and host: the plasmid is the vector, and the bacterium that takes it up is the host.
- Profiling and sequencing: profiling compares repeating patterns between people. Sequencing reads the actual order of bases.
- Somatic and germ-line therapy: somatic therapy treats the patient and the change is not inherited. Germ-line therapy alters gametes or embryos, so the change is passed on.
Check Yourself¶
- Describe the three stages of a PCR cycle and the temperature of each.
- Calculate the number of copies of a DNA sequence after 10 cycles of PCR from one starting molecule.
- Explain why a heat-stable DNA polymerase is used in PCR.
- Explain how gel electrophoresis separates fragments of DNA.
- Describe how a gene is inserted into a plasmid, and how bacteria that have taken up the recombinant plasmid are identified.
- Give one benefit and one ethical concern of gene therapy.
Answers
- Denaturation at about 95 °C separates the strands. Annealing at about 55 °C lets primers bind to the target sequence. Extension at about 72 °C lets Taq polymerase build the new strands from the primers.
- 2¹⁰ = 1,024 copies.
- Each cycle heats the mixture to about 95 °C to separate the strands. A normal polymerase would be denatured, but Taq polymerase from a thermophilic bacterium stays active.
- DNA is negatively charged because of its phosphate groups, so in an electric field it moves towards the positive electrode through the gel. Smaller fragments move faster and further, so the fragments separate by size.
- The gene and the plasmid are cut with the same restriction enzyme to give complementary sticky ends, and ligase joins them. Bacteria that have taken up the plasmid are identified by a marker gene, such as one for antibiotic resistance or fluorescence, so only transformed bacteria grow or glow.
- Benefit: it may cure or treat a genetic disease, such as cystic fibrosis, by replacing a faulty allele. Concern: germ-line changes would be inherited, and there are worries about safety, long-term effects and the use of the technology to select characteristics.
Key Terms¶
- DNA sequencing: determination of the order of bases in a DNA molecule.
- Dideoxynucleotide (ddNTP): chain-terminating nucleotide used in Sanger sequencing.
- Polymerase chain reaction (PCR): technique that amplifies a specific DNA sequence by repeated heating and cooling cycles.
- Gel electrophoresis: method that separates DNA fragments by size as they move through a gel in an electric field.
- DNA profiling: identification of individuals by comparing patterns in their DNA.
- Recombinant DNA: DNA molecule formed by joining DNA from different sources.
- Restriction endonuclease: enzyme that cuts DNA at a specific recognition sequence.
- DNA ligase: enzyme that joins together DNA fragments by forming phosphodiester bonds.
- Plasmid vector: small circular DNA molecule used to carry foreign DNA into a host cell.
- Marker gene: gene used to identify cells that have taken up recombinant DNA.
- Genetic engineering: deliberate modification of an organism’s DNA using biotechnology.
- VNTR (variable number tandem repeat): a non-coding repetitive DNA sequence with variable repeat number.
- Genomics: the study of whole genomes using DNA sequencing and computational tools.
- Bioinformatics: the use of computational tools and databases to analyse biological sequence data.
- Proteomics: the study of the complete set of proteins expressed by a genome.
- cDNA (complementary DNA): a DNA copy made from an mRNA template using reverse transcriptase.
- Reverse transcriptase: enzyme that synthesises DNA from an RNA template.
- Gene therapy: treatment that introduces a functional allele into cells to correct a genetic disorder.
Connected Pages¶
- 6.1.1 Cellular control (gene mutations; the genes being engineered)
- 6.1.2 Patterns of inheritance
- 6.2.1 Cloning and biotechnology (use of microorganisms in biotechnology)
- 5.1.4 Hormonal communication (insulin production by GM bacteria)
- 2.1.3 Nucleotides and nucleic acids (DNA structure and replication)
- Recombinant Plasmid
- Module 6: Genetics, evolution and ecosystems