Posted in

THE GENOME – Self Learning Series # 1, P# 80, Ch # 4

THE GENOME - Self Learning Series # 1, P# 80, Ch# 4
  • Over the past 50 years, major advances have improved our understanding of the genetic basis of both inherited and acquired human diseases.
  • This progress came from a better understanding of:
    • The structure of the human genome.
    • The factors that control gene expression.
  • Sequencing of the human genome at the beginning of the 21st century was a major achievement in biomedical science.
  • Since then:
    • The cost of genome sequencing has rapidly decreased.
    • The ability to analyze huge amounts of genetic data has greatly increased.
    • Together, these advances have transformed our understanding of health and disease.
  • Genome research has also shown that genetic regulation is far more complex than simply reading the linear DNA sequence.
  • These powerful tools can:
    • Improve understanding of disease pathogenesis.
    • Help drive new therapeutic innovations.
  • Early research mainly focused on finding genes that encode proteins.
  • More recent research has shown that noncoding DNA also has an important role in regulating gene expression.

KEY CONCEPT

Human genome sequencing + cheaper sequencing + powerful data analysis → better understanding of genes and gene regulation → better understanding of disease and development of new therapies.

CONCEPTUAL EXAMPLES

  • Protein-coding DNA → provides information for making proteins.
  • Noncoding DNA → can help regulate how genes are expressed.
  • Genome sequencing → helps scientists study the genetic basis of inherited and acquired diseases.

Protein-Coding and Noncoding DNA

  • The human genome contains about 3.3 billion DNA base pairs.
  • It has only slightly more than 19,000 protein-coding genes, which make up only about 1.5% of the genome.
  • These genes produce proteins that act as:
    • enzymes
    • structural components
    • signaling molecules
  • The actual number of proteins is greater than 19,000 because one gene can produce multiple RNA transcripts → different protein isoforms.
  • Surprisingly, worms with fewer than 1000 cells and much smaller genomes also have about 20,000 protein-coding genes, and many of their proteins are similar to human proteins.
  • Therefore, much of the difference between humans and simpler organisms appears to depend on the 98.5% of DNA that does not encode proteins.
  • More than 85% of the human genome is transcribed, and almost 80% is involved in regulation of gene expression.
  • Conceptually:
    • Protein-coding DNA → provides the building blocks and machinery
    • Noncoding DNA → provides the regulatory “architectural plan”
  • Major functional classes of non–protein-coding DNA include (Fig. 4.1):
    • Promoters and enhancers → bind transcription factors and regulate transcription.
    • Sites that bind proteins controlling higher-order chromatin structure.
    • Noncoding regulatory RNAs → especially microRNAs and long noncoding RNAs → regulate gene expression without being translated into proteins.
    • Mobile genetic elements (transposons) → “jumping genes” that can move within the genome and influence gene regulation and chromatin organization.
    • Structural DNA regions:
      • telomeres → chromosome ends
      • centromeres → chromosome attachment/tethering regions
  • More than one-third of the human genome consists of mobile genetic elements.
  • Many disease-associated genetic variations occur in noncoding regions.
  • Therefore, altered gene regulation may sometimes be more important in disease than structural changes in proteins.
  • Any two humans are usually more than 99.5% genetically identical.
  • Thus, much individual variation and disease susceptibility is contained within less than 0.5% of DNA.
  • The two common forms of DNA variation are:
    • single-nucleotide polymorphisms (SNPs)
    • copy number variations (CNVs)
  • SNPs = variation at a single nucleotide position.
  • They are usually biallelic, meaning two common alternatives occur at that position.
  • More than 6 million human SNPs have been identified.
  • SNPs occur in both:
    • coding regions
    • noncoding regions
  • Only about 1% of SNPs occur in coding regions.
  • SNPs in noncoding regulatory regions may alter gene expression → change disease susceptibility.
  • Some SNPs are neutral, producing no direct change in gene function or phenotype.
  • Even a neutral SNP can be useful if it is inherited together with a nearby disease-causing variant.
  • This association is called linkage disequilibrium.
  • Linkage disequilibrium = two genetic markers occur together more often than expected by chance.
  • It may result from:
    • natural selection
    • genetic drift
  • Individual SNPs usually have only a small effect on complex diseases such as diabetes, heart disease, or cancer because many genetic factors contribute.
  • CNVs = differences in the number of copies of large stretches of DNA.
  • These regions may range from thousands to millions of base pairs.
  • CNVs may involve:
    • DNA duplication
    • DNA deletion
    • more complex rearrangements such as inversions
  • CNVs account for millions of base-pair differences between individuals.
  • About 50% of CNVs involve protein-coding genes, so they may contribute importantly to human phenotypic diversity.
  • DNA sequence alone does not determine phenotype.
  • Phenotype results from interaction between:
    genes + environment + chance.
  • Even genetically identical monozygotic twins can develop different phenotypes.
  • Changes such as DNA methylation and histone-related modifications can strongly influence gene expression without changing the underlying DNA sequence.

KEY CONCEPT

  • Only about 1.5% of the genome codes for proteins.
  • Much of the remaining genome regulates when, where, and how genes are expressed.
  • SNP = single-base variation.
  • CNV = variation in the number of copies of large DNA segments.
  • Human differences arise from DNA variation + gene regulation + environment + chance.

CONCEPTUAL EXAMPLES

  • Promoter/enhancer variation → gene may be expressed too much or too little → altered disease risk.
  • Neutral SNP near a disease-causing gene → travels with that gene → acts as a genetic marker.
  • DNA segment duplicated → extra copies of genes → CNV.
  • Identical twins have the same DNA sequence but may still differ because gene regulation and environmental influences differ.

FIG. 4.1 — ORGANIZATION OF NUCLEAR DNA

🧠 Easiest idea

DNA is packed inside the nucleus → wrapped around histones → forms chromatin → forms chromosomes.
When a gene is active: DNA → RNA → protein.
1️⃣ LEFT TOP — CELL → NUCLEUS

🩷 Large pink structure = cell
🟣 Purple compartment = nucleus
🟪 Dark round body = nucleolus

Nucleolus

= place where rRNA is made and ribosome assembly begins.

Inside the nucleus, DNA exists mainly as chromatin.

2️⃣ 🟢 EUCHROMATIN = OPEN + ACTIVE

Green arrow points to the lighter, dispersed chromatin.

Euchromatin

DNA is loosely packed.

➡️ Transcription machinery can reach the genes.

Therefore:

Loose DNA → gene accessible → transcription ON

🧠 Memory:

EU = Easy to Use

3️⃣ 🟣 HETEROCHROMATIN = CLOSED + INACTIVE

Purple arrow points to the dark, dense chromatin.

DNA is tightly packed.

➡️ Transcription machinery cannot easily reach it.

Therefore:

Tightly packed DNA → transcription OFF / low

🧠 Memory:

Heterochromatin = Hidden DNA

4️⃣ BOTTOM LEFT — CHROMOSOME

When chromatin becomes extremely condensed during cell division, we can see a chromosome.

🔵 Blue structure = chromosome.

P arm

Pink label = short arm

🧠 P = petite = short

Q arm

Green label = long arm

🧠 Q comes after P → think longer arm

5️⃣ 🔴 CENTROMERE

Red central area joining the chromosome arms.

Centromere

= constricted region where the kinetochore forms.

The kinetochore attaches to spindle microtubules during cell division.

Simple:

Centromere → kinetochore → chromosome movement

6️⃣ 🟡 TELOMERES

Yellow chromosome tips = telomeres.

They contain repetitive DNA and protect chromosome ends.

Function:

Telomeres = protective caps

They prevent important DNA from being lost or chromosome ends from being mistaken for broken DNA.

7️⃣ MIDDLE — How DNA is packed

The blue DNA strand wraps around 🟠 orange protein units.

🟠 Nucleosome

= DNA wrapped around a histone octamer

Think:

Histones = spool
DNA = thread

So:

DNA → nucleosomes → chromatin fiber → chromosome

8️⃣ Dense vs loose nucleosomes

🟣 Heterochromatin

Nucleosomes are very tightly packed.

➡️ Gene OFF.

🟢 Euchromatin

Nucleosomes are more spread apart.

➡️ DNA is accessible
➡️ Gene can be ON.

9️⃣ RIGHT SIDE — Gene transcription

When chromatin is open:

DNA

⬇️ Transcription

Pre-mRNA

The gene contains several important regions.

🔟 🟢 PROMOTER

Green box = promoter

Promoter is a DNA sequence where the transcription machinery assembles to start transcription.

Think:

Promoter = START button

1️⃣1️⃣ 🔵 EXONS

Blue boxes = exons

Exons are sequences that remain in mature RNA.

For protein-coding genes, the coding portions of exons can ultimately specify protein.

1️⃣2️⃣ Gray loops = INTRONS

Introns are intervening sequences present in the initial RNA transcript.

They are removed during:

✂️ Splicing

So:

Pre-mRNA = exons + introns
⬇️ splicing
Mature mRNA = introns removed

1️⃣3️⃣ 🟣 ENHANCER

Purple box = enhancer.

An enhancer increases gene transcription.

It can act from a distance by DNA looping so regulatory proteins can influence the promoter.

Think:

Enhancer = volume-up button for a gene

1️⃣4️⃣ Mature mRNA

After splicing:

5′ UTR

= untranslated region before the coding sequence.

Open-reading frame

= protein-coding region.

3′ UTR

= untranslated region after the coding sequence.

UTRs are not translated into protein, but they help regulate the RNA.

1️⃣5️⃣ TRANSLATION

Mature mRNA leaves the nucleus and is read by ribosomes.

mRNA
⬇️ translation
🟣 Protein

So the fundamental flow is:

DNA → RNA → Protein

🎨 COLOR / STRUCTURE MAP

  • 🟢 Green = active/open regions or regulatory promoter
  • 🟣 Purple = dense heterochromatin / regulatory structures
  • 🟠 Orange = histone-containing nucleosomes
  • 🔵 Blue = DNA, exons, chromosome
  • 🟡 Yellow = telomeres
  • 🔴 Red = centromere
  • Gray loops = introns

⭐ WHOLE FIGURE IN 4 STEPS

1. DNA packing

DNA → nucleosome → chromatin → chromosome

2. Chromatin state

Euchromatin = open/active
Heterochromatin = closed/inactive

3. Gene expression

Promoter + enhancer → transcription → pre-mRNA

4. RNA processing

Splicing removes introns → mature mRNA → translation → protein

🧠 Fastest exam recall

Euchromatin = loose, light, active.
Heterochromatin = dense, dark, inactive.
Nucleosome = DNA wrapped around histone octamer.
P arm = short; Q arm = long.
Centromere = kinetochore site; telomeres protect chromosome ends.
DNA → pre-mRNA → splicing → mature mRNA → protein.✕Compare with Claude Opus 4.8

Epigenetic Changes

  • Almost all body cells contain the same DNA, but different cell types have different structures and functions because they express different sets of genes.
  • These differences are controlled by epigenetic modifications → changes in chromatin that alter gene expression without changing the DNA sequence.
  • One major mechanism is DNA methylation:
    • Cytosine residues in gene promoters become methylated.
    • Heavy promoter methylation → RNA polymerase cannot easily access the gene → transcription is silenced.
  • In many cancers:
    • tumor suppressor gene promoters become methylated → tumor suppressor genes are silenced → uncontrolled cell growth.
  • Another major mechanism involves histones.
  • DNA wraps around histone proteins to form nucleosomes.
  • Histones can undergo reversible changes such as:
    • methylation
    • acetylation
  • These changes alter chromatin structure → change the accessibility of DNA → increase or decrease transcription.
  • Abnormal histone modifications occur in diseases such as cancer → abnormal gene expression.
  • Histone deacetylases and DNA methylation inhibitors are targets used in treatment of certain cancers.
  • Normal epigenetic silencing occurring during development is called imprinting.

Micro-RNA and Long Noncoding RNA

  • Some genes are transcribed into RNA but never translated into protein.
  • These noncoding RNAs can regulate gene expression.
  • Two important types are:
    • microRNAs (miRNAs)
    • long noncoding RNAs (lncRNAs)
  • MicroRNAs (miRNAs):
    • Short RNAs, about 22 nucleotides long.
    • Mainly regulate how target mRNAs are translated into proteins.
    • Therefore, they cause posttranscriptional gene silencing.
  • The human genome contains almost 6000 miRNA genes.
  • One miRNA can regulate multiple protein-coding genes → allowing coordinated control of whole gene-expression programs.
  • miRNA formation occurs as:
    miRNA gene → pri-miRNA → processing → Dicer → mature miRNA → RISC complex (Fig. 4.2).
  • Mature miRNA is about 21–30 nucleotides long.
  • It joins the RNA-induced silencing complex (RISC).
  • miRNA then base-pairs with its target mRNA.
  • RISC causes either:
    • mRNA cleavage, or
    • inhibition of translation
  • Result → less protein is produced from that mRNA.
  • Small interfering RNAs (siRNAs) use a similar pathway.
  • Synthetic siRNAs can be introduced into cells → processed through Dicer/RISC → selectively silence a target mRNA.
  • This is called knockdown technology.
  • Some siRNAs are also used therapeutically.
  • Long noncoding RNAs (lncRNAs) are more than 200 nucleotides long.
  • They regulate gene expression through several mechanisms (Fig. 4.3).
  • Some lncRNAs bind chromatin → prevent RNA polymerase from reaching nearby genes → gene silencing.
  • Important example: XIST
    • produced from the X chromosome
    • coats that X chromosome with a repressive “cloak”
    • causes X-chromosome inactivation and gene silencing
    • XIST itself escapes this inactivation.
  • Other lncRNAs are produced at enhancers and can increase transcription of nearby gene promoters.
  • lncRNAs are being studied in diseases including atherosclerosis and cancer.

KEY CONCEPT

  • Epigenetics = control of gene expression without changing DNA sequence.
  • Promoter methylation → gene OFF.
  • Histone modification → changes chromatin structure → alters transcription.
  • miRNA → RISC → destroys mRNA or blocks translation → ↓ protein.
  • lncRNA → can silence or enhance gene transcription.
  • XIST → X-chromosome inactivation.

CONCEPTUAL EXAMPLES

  • Tumor suppressor promoter becomes heavily methylated → gene cannot be transcribed → tumor suppressor is silenced.
  • miRNA binds a target mRNA → RISC blocks or destroys it → less protein produced.
  • XIST coats one X chromosome → its genes become largely silent → X-chromosome inactivation.

FIG. 4.2 — microRNA (miRNA) → GENE SILENCING

🧠 Simplest idea

miRNA is a tiny regulatory RNA that finds a target mRNA and prevents it from making protein.

Whole figure in one line

miRNA gene → pri-miRNA → pre-miRNA → export from nucleus → Dicer → mature miRNA → RISC → target mRNA → translation blocked OR mRNA destroyed → GENE SILENCING

1️⃣ TOP — miRNA gene

🟢/black double helix = DNA containing the miRNA gene

⬇️ Orange arrow = transcription

DNA produces:

pri-miRNA

= primary miRNA, the first long RNA transcript.

It forms several hairpin loops because parts of the RNA pair with each other.

2️⃣ pri-miRNA → pre-miRNA

Inside the nucleus, pri-miRNA is cut down to a smaller hairpin:

pre-miRNA = precursor miRNA

🧠 Think:

PRI = Primary, long
PRE = Precursor, shorter

This nuclear processing is mainly done by Drosha.

3️⃣ 🟠 Export protein

The orange structure at the nuclear membrane = export protein.

It carries:

pre-miRNA
➡️ out of the nucleus
➡️ into the cytoplasm

So:

Nucleus → export protein → cytoplasm

4️⃣ Dicer cuts pre-miRNA

🔵 Gray-blue enzyme = Dicer

Dicer acts like molecular scissors.

It cuts the hairpin loop and produces a short:

double-stranded miRNA duplex

⬇️

Usually about 21–30 nucleotides long in this figure.

Memory:

Dicer = DICES RNA

5️⃣ miRNA duplex unwinds

The two small RNA strands separate.

⬇️ Orange branching arrows

Usually one strand becomes the guide strand.

The other strand is discarded.

6️⃣ 🟡 RISC complex

The guide miRNA enters:

RISC = RNA-Induced Silencing Complex

Think of RISC as:

miRNA + protein machinery = search-and-stop machine

The miRNA tells RISC:

“Find an mRNA with a matching sequence.”

7️⃣ LEFT SIDE — Target gene makes target mRNA

🔴/black DNA = target gene

⬇️

The target gene is transcribed into:

🔴 strand = target mRNA

Normally:

mRNA → ribosome → protein

But miRNA can stop this.

8️⃣ LEFT PATH — Imperfect match

miRNA binds the target mRNA partly, not perfectly.

Imperfect miRNA–mRNA pairing

⬇️

Translational repression

This means:

mRNA remains present, but the ribosome cannot efficiently make protein from it.

🔵 Large structure = ribosome
🔴 curved stop symbol = translation blocked

Result:

↓ Protein production

9️⃣ RIGHT PATH — Perfect match

If miRNA matches the target mRNA very closely:

Perfect match

⬇️

mRNA cleavage

RISC cuts the target mRNA.

The red strand is shown broken into pieces.

Result:

mRNA destroyed → no protein made

🔟 FINAL — GENE SILENCING

Both pathways end at the same result:

Imperfect match

→ translation blocked

Perfect match

→ mRNA destroyed

⬇️

GENE SILENCING

Meaning:

The gene may still exist in DNA, but its protein output is reduced or stopped.

🎨 COLOR / ARROW GUIDE

  • 🟢/black double helix = DNA / miRNA gene
  • 🟢 hairpin RNA = pri-miRNA / pre-miRNA
  • 🟠 structure = export protein
  • 🔵 gray enzyme = Dicer
  • 🟡 complexes = RISC
  • 🔴 horizontal strand = target mRNA
  • 🔵 large oval = ribosome
  • 🟠 arrows = direction of miRNA processing
  • 🔴 stop mark = translation inhibited

⭐ Fastest conceptual memory

Make it

miRNA gene → pri-miRNA → pre-miRNA

Process it

Export → Dicer → mature miRNA

Use it

miRNA + RISC → finds target mRNA

Silence it

  • Imperfect match → translation OFF
  • Perfect match → mRNA CUT

🧠 One-line exam recall

miRNA + RISC silences genes after transcription by either blocking translation or degrading target mRNA.

FIG. 4.3 — ROLES OF long noncoding RNA (lncRNA)

🧠 Simplest idea

lncRNA does NOT make protein. Instead, it controls whether genes are ON or OFF and helps organize chromatin/protein complexes.

Whole figure in one line

lncRNA → can activate genes, suppress genes, modify chromatin, or organize protein complexes.

A️⃣ GENE ACTIVATION

🟣 Curved strand = lncRNA
🟢 Blob = ribonucleoprotein transcription complex
🔵 Double helix = DNA
🟠 Beads = nucleosomes/histones

Arrow-by-arrow:

lncRNA binds transcription proteins
→ brings/stabilizes them on DNA
→ transcription machinery works better
gene turns ON

🟢 Green arrow = gene activation

Easy idea:

lncRNA acts like a helper/recruiter.

B️⃣ GENE SUPPRESSION

Here the lncRNA behaves as a:

Decoy lncRNA

= a “fake target.”

It binds a transcription factor before the factor can bind DNA.

So:

Transcription factor trapped by lncRNA
→ cannot reach the gene
→ transcription stops

🔴 Red stop symbol = gene suppression

Easy idea:

Decoy lncRNA distracts the transcription factor.

C️⃣ PROMOTE CHROMATIN MODIFICATION

lncRNA can bring enzymes to chromatin.

Black curved arrow points toward the nucleosomes.

These enzymes add/remove chemical marks such as:

  • Methylation
  • Acetylation

🟡/🩷 small colored dots = chemical modifications on histones/chromatin.

These marks can change whether chromatin is:

Open → gene active

or

Closed → gene inactive

Easy idea:

lncRNA tells chromatin-modifying enzymes where to work.

D️⃣ ASSEMBLY OF PROTEIN COMPLEXES

Multiple colored protein shapes = multisubunit protein complex.

lncRNA acts like a:

Scaffold

= a framework holding several proteins together.

So:

lncRNA binds several proteins
→ assembles them into one complex
→ complex acts on chromatin
→ changes gene activity/chromatin structure

Easy idea:

lncRNA = molecular platform for building protein teams.

🎨 COLOR / STRUCTURE GUIDE⭐ 4 MAIN JOBS OF lncRNA

1. Guide/Recruit

→ brings transcription machinery
Gene ON

2. Decoy

→ traps transcription factors
Gene OFF

3. Chromatin modifier guide

→ directs methylation/acetylation enzymes
→ changes gene accessibility

4. Scaffold

→ holds multiple proteins together
→ controls chromatin/gene activity

🧠 Fastest exam memory

lncRNA = G-D-S-C

G = Guide gene activation
D = Decoy → gene suppression
S = Scaffold protein complexes
C = Chromatin modification

⭐ One-line recall

lncRNA regulates gene expression without becoming protein—it can turn genes ON, turn them OFF, modify chromatin, and assemble regulatory protein complexes.

Leave a Reply

Your email address will not be published. Required fields are marked *