Skip to main contentSkip to navigation
ThisIsHowItWorks.in

Complex systems, clearly explained.

An independent visual publication explaining the invisible protocols, networks, infrastructure, and mechanisms that run our world.

Explainers

  • How UPI Works
  • Offline UPI Mechanisms
  • All Explainers (Archive)
  • Topics & Roadmap
  • Search Index

Publication

  • About Publication
  • Editorial Principles
  • Changelog
  • RSS / Atom Feed

Legal & Contact

  • Privacy Policy
  • Terms of Use
  • Editorial & Legal Notice
  • Contact Us

Connect

  • Instagram
  • Discord Community
© 2026 ThisIsHowItWorks.in. All rights reserved.
Durable technical understanding built from first principles.
ThisIsHowItWorks.in
ExploreTopicsAbout
  1. Home
  2. /Topics
  3. /Life & Evolutionary Biology
  4. /Life & Evolutionary Biology
  5. /Life & Evolutionary Biology
  6. /How Genes Actually Build Proteins
Biology · Life & Evolutionary Biology/ Explainer

How Genes Actually Build Proteins

The Central Dogma, mRNA transcription, the 64-codon dictionary, and the nanomechanics of ribosomal translation

Updated for clarity
The Short AnswerFirst-Principles Core

“How does a one-dimensional string of chemical letters in DNA turn into a functioning three-dimensional biological machine?”

DNA does not digest food, pump oxygen, or fight off viruses. DNA is an inert reference library. The actual work of life is performed by proteins: molecular machines folded into precise three-dimensional shapes. To convert genetic blueprints into physical reality, cells execute the Central Dogma of Molecular Biology: Transcription and Translation. Inside the nucleus, RNA Polymerase unzips a gene and copies its sequence into a mobile messenger RNA (mRNA) transcript. That mRNA travels to the cytoplasm, where a 2.5-megadalton molecular factory—the Ribosome—reads the genetic tape in three-letter words called codons. Using transfer RNA (tRNA) adaptors, the ribosome matches each codon to one of twenty standard amino acids, stitching together a polypeptide chain at twenty links per second. Driven by hydrophobic collapse and hydrogen bonding, this linear chain folds into a functional nanotech machine.

Recommended Background

To understand the failure modes and edge cases detailed in this piece, we recommend familiarizing yourself with these foundational mechanisms first:

How Cells Actually Work
Understanding How Cells Actually Work is required before reading How Genes Actually Build Proteins
How DNA Stores and Replicates Information
Understanding How DNA Stores and Replicates Information is required before reading How Genes Actually Build Proteins
In this Explainer7 Sections

The Blueprint vs. The Machine

In How DNA Stores and Replicates Information, we saw that DNA is a magnificent digital library.

Yet if you put a purified vial of DNA on a laboratory table, nothing happens. DNA has no catalytic power. It cannot contract a muscle fiber, transport an oxygen atom, break down a sugar molecule, or detect a photon of light.

DNA is like an architectural blueprint on paper. You cannot live inside a blueprint. You need a construction crew to read the two-dimensional paper drawing and assemble three-dimensional steel beams, bricks, and glass windows.

In living systems, the physical machines are Proteins.

Proteins are the workhorses of life:

  • Enzymes like amylase that break down food starch in milliseconds.
  • Structural beams like collagen holding your skin and bones together.
  • Molecular motors like kinesin walking along microtubule highways.
  • Immune weapons like antibodies that bind to viral surfaces.
  • Transporters like hemoglobin carrying oxygen through your blood vessels.

Every protein in your body begins its existence as an abstract sequence of chemical letters written inside a gene.

In 1958, Francis Crick formulated the grand governing rule of life on Earth: The Central Dogma of Molecular Biology:

$$\textbf{DNA} \quad \xrightarrow{\text{Transcription}} \quad \textbf{RNA} \quad \xrightarrow{\text{Translation}} \quad \textbf{PROTEIN}$$

Information flows in one direction: from nucleic acid storage down into physical protein hardware. Once information has passed into a protein, it can never get back out into the genetic code.

Here is the mechanical reality of how a string of code becomes a physical machine.


Phase 1: Transcription (Copying the Gene)

Your complete genome contains roughly 20,000 genes, stored on 23 pairs of chromosomes safely locked inside the cell nucleus.

The cell does not drag the entire two-meter DNA archive out into the cytoplasm every time it needs to manufacture a protein. That would risk damaging the master master copy.

Instead, the cell makes a cheap, disposable photocopy of a single gene: Messenger RNA (mRNA).

                  THE CHEMICAL DIFFERENCES: DNA vs. RNA

         Feature                 DNA                          RNA
      ──────────────────────────────────────────────────────────────────────────
         Strands                 Double-stranded helix        Single-stranded tape
         Sugar                   Deoxyribose (Stable)         Ribose (Reactive)
         Nitrogenous Bases       A, T, C, G                   A, U, C, G
                                                              (Uracil replaces Thymine)
         Function                Permanent master archive     Disposable working copy

The Transcription Machinery: RNA Polymerase

Transcription is executed by a massive molecular machine called RNA Polymerase.

                     THE TRANSCRIPTION BUBBLE

                                   Non-template Strand
                         5' ─────────────────────────────────────► 3'
                                  /                     \
                             ┌───┴───────────────────────┴───┐
                             │        RNA POLYMERASE         │
                             └───┬───────────────────────┬───┘
                                  \  [ U-A-C-G-G-U ]    /
                                   ════════════════════► 3' (Growing mRNA Transcript)
                         3' ─────────────────────────────────────◄ 5'
                                    Template Strand
  1. Promoter Recognition: RNA Polymerase scans along the double helix until it finds a landing strip: a specific chemical sequence called a Promoter (often containing a repeating sequence like TATA).
  2. Unwinding: Once locked on, RNA Polymerase pulls the two DNA strands apart, creating a 14-base-pair Transcription Bubble.
  3. Synthesis: The polymerase reads the exposed DNA template strand in the $3' \to 5'$ direction and synthesizes a complementary single strand of RNA in the $5' \to 3'$ direction at roughly 50 letters per second:
    • If DNA has A, RNA Polymerase adds Uracil (U).
    • If DNA has T, RNA Polymerase adds Adenine (A).
    • If DNA has C, RNA Polymerase adds Guanine (G).
    • If DNA has G, RNA Polymerase adds Cytosine (C).
  4. Termination: When the polymerase hits a terminator sequence, it releases the finished mRNA transcript and lets the DNA double helix snap back together undamaged.

The single-stranded mRNA transcript is capped, polyadenylated with a protective tail of adenines, and exported through a nuclear pore into the cytoplasm.


Phase 2: The Cipher (The 64-Codon Genetic Code)

Now the cell faces a profound mathematical translation dilemma:

  • The language of mRNA has only 4 letters: ${\text{A}, \text{U}, \text{C}, \text{G}}$.
  • The language of proteins has 20 standard amino acids: Alanine, Cysteine, Glutamate, Lysine, Valine, etc.

How do you translate a 4-letter alphabet into a 20-word vocabulary?

                     THE COMBINATORIC PROOF

   1-Letter Words (4¹):      4 possible words    (Not enough for 20 amino acids!)
   2-Letter Words (4²):     16 possible words    (Still not enough!)
   3-Letter Words (4³):     64 possible words    (PLENTY of room!)

In the early 1960s, Marshall Nirenberg and Heinrich Matthaei cracked the code using synthetic RNA test tubes.

Nature uses three-letter words called Codons.

Because $4^3 = 64$, but there are only 20 amino acids, the genetic code is degenerate (redundant): multiple different codons code for the exact same amino acid.

                  THE UNIVERSAL GENETIC CODE TABLE (SUMMARY)

    First   ───────────── Second Letter ─────────────   Third
    Letter       U            C           A           G         Letter
   ┌───────┬───────────┬───────────┬───────────┬───────────┬───────┐
   │   U   │ Phe, Leu  │ Ser       │ Tyr, STOP │ Cys, Trp  │ U,C,A,G
   │   C   │ Leu       │ Pro       │ His, Gln  │ Arg       │ U,C,A,G
   │   A   │ Ile, Met* │ Thr       │ Asn, Lys  │ Ser, Arg  │ U,C,A,G
   │   G   │ Val       │ Ala       │ Asp, Glu  │ Gly       │ U,C,A,G
   └───────┴───────────┴───────────┴───────────┴───────────┴───────┘
   
   * AUG is the universal START codon (Codes for Methionine)
   * UAA, UAG, UGA are STOP codons (Signal the end of translation)

Notice the engineering brilliance:

  1. The Start Signal: Every protein translation begins with the codon AUG, which codes for the amino acid Methionine.
  2. The Stop Signals: Three codons—UAA, UAG, and UGA—do not code for any amino acid. They are punctuation marks that command the translation machine to halt and release the finished protein.
  3. Error Tolerance: In most codons, the third letter does not matter. GCU, GCC, GCA, and GCG all code for Alanine. If a mutation accidentally changes the third letter, the protein remains completely unaffected (a silent mutation)!

Phase 3: The Bilingual Adaptor (Transfer RNA)

How does a ribosome physically connect an abstract 3-letter mRNA codon to a specific physical amino acid?

Molecules have no eyes. A codon cannot grab an amino acid directly.

Francis Crick predicted the existence of an Adaptor Molecule: a bilingual physical bridge.

That molecule is Transfer RNA (tRNA):

                     THE TRANSFER RNA (tRNA) ADAPTOR

                             Amino Acid Attachment Site (3' CCA)
                                        [ Glycine ]
                                             │
                                          ┌──┴──┐
                                          │     │
                                          │tRNA │  ◄── Cloverleaf-folded
                                          │     │      RNA molecule
                                          └──┬──┘
                                             │
                                         [ C C A ]  ◄── ANTICODON
                                             │
      ═══════════════════════════════════════╪═════════════════════════════
                                         [ G G U ]  ◄── mRNA CODON

A tRNA molecule is a piece of RNA folded into a 3D L-shape:

  • At its bottom tip is a 3-letter sequence called the Anticodon, which forms complementary hydrogen bonds with a matching mRNA codon.
  • At its top tip (the $3'\text{-OH}$ tail) is a covalently attached specific Amino Acid.

An enzyme called Aminoacyl-tRNA Synthetase acts as the quality-control inspector: it verifies that each tRNA is loaded with its exact correct amino acid, consuming ATP to charge the connection.


Phase 4: Inside the Ribosome Factory

The assembly of the protein occurs inside the Ribosome: a colossal, 2.5-megadalton molecular machine made of 60% catalytic ribosomal RNA (rRNA) and 40% structural proteins.

A ribosome consists of two interlocking halves:

  • Small Subunit (30S/40S): Binds the mRNA tape and checks codon-anticodon geometry.
  • Large Subunit (50S/60S): Contains the catalytic active site that forges peptide bonds.

Inside the joined ribosome are three distinct operational bays:

                     THE THREE SLOTS OF THE RIBOSOME

                   ┌─────────────┬─────────────┬─────────────┐
                   │   E SITE    │   P SITE    │   A SITE    │
                   │   (Exit)    │  (Peptidyl) │(Aminoacyl)  │
                   └──────┬──────┴──────┬──────┴──────┬──────┘
                          │             │             │
                    Spent tRNA     Growing Protein   Incoming
                    is ejected      Chain holds      tRNA matches
                    back to cell    here!            new codon!

The Translation Cycle (20 Amino Acids Per Second)

  1. Docking at the A Site: A fresh tRNA carrying its amino acid enters the A (Aminoacyl) Site. If its anticodon matches the mRNA codon, the small subunit clamps down, locking it in place.
  2. Peptide Bond Formation: The catalytic core of the large subunit—a Ribozyme made entirely of pure RNA—catalyzes a chemical attack:
    • It cuts the growing protein chain free from the tRNA in the P site.
    • It welds the chain onto the new amino acid sitting in the A site, forming a covalent Peptide Bond ($\text{-CO-NH-}$).
  3. Translocation: Burning a molecule of GTP fuel, the ribosome lunges exactly three nucleotides forward along the mRNA tape:
    • The empty tRNA in the P site is shunted to the E (Exit) Site and ejected into the cytoplasm.
    • The tRNA in the A site (now holding the enlarged chain) moves into the P site.
    • The A site is now empty, positioned over the next codon, ready for the next incoming tRNA.

The sequence diagram below illustrates this choreographed cycle:

The Transcription and Ribosomal Translation Flow
dataNuclear DNA Gene
processRNA Polymerase
dataMessenger RNA (mRNA)
processRibosome Factory
devicetRNA Adaptors
101 RNA Polymerase binds promoter sequence; unwinds double helix
202 Transcribes complementary 5' to 3' pre-mRNA transcript
303 mRNA exports through nuclear pore; binds small ribosomal subunit
404 Scans for AUG start codon; recruits Initiator Methionine-tRNA
505 Aminoacyl-tRNA docks in A-site matching triplet codon
606 Peptidyl transferase ribozyme welds peptide bond (-CO-NH-)
707 Translocates 3 nucleotides forward; spent tRNA ejected from E-site
808 Encounters Stop Codon (UAA/UAG/UGA); release factor frees protein chain
Sequence diagram tracing the flow of genetic information from nuclear DNA gene transcription by RNA Polymerase into mRNA, nuclear pore export, small and large ribosome subunit docking, and iterative tRNA codon matching.

Phase 5: The Folding Miracle (From String to Machine)

As the ribosome spits out the linear string of amino acids, the polypeptide chain is floppy and useless—like a long strand of wet spaghetti.

For the protein to do work, it must fold into a rigid, precise, three-dimensional geometric structure:

               THE FOUR LEVELS OF PROTEIN ARCHITECTURE

  1. Primary Structure:      [ Met ] ─── [ Ala ] ─── [ Gly ] ─── [ Cys ] (Linear chain)
                                            │
  2. Secondary Structure:            ┌──────┴──────┐
                             Alpha Helix (Spiral)   Beta Sheet (Accordion)
                                            │
  3. Tertiary Structure:     Complete 3D Globular Fold (Active catalytic clefts!)
                                            │
  4. Quaternary Structure:   Multiple folded subunits locked together (e.g. Hemoglobin: 4 units)

In 1973, Christian Anfinsen won the Nobel Prize in Chemistry by proving that a protein's 3D shape is completely determined by its 1D amino acid sequence.

The folding process is driven by basic physics:

  • Hydrophobic Collapse: Some amino acids (like Leucine, Isoleucine, Valine) have oily, non-polar side chains that hate water. In an aqueous cell, these oily side chains violently tuck themselves into the dark, dry interior of the folding ball to escape water molecules.
  • Hydrophilic Exterior: Electrically charged amino acids (like Lysine, Glutamate) face outward, interacting comfortably with surrounding water.
  • Hydrogen Bonds & Disulfide Bridges: Chemical bonds snap into place across loops, locking the folds into rigid $\alpha$-helices and corrugated $\beta$-sheets.

The protein folds in milliseconds.

If it needs help folding properly without getting tangled in the crowded cytoplasm, specialized barrel-shaped protein chambers called Chaperonins (such as GroEL/GroES) isolate the unformed chain like an incubator until it snaps into its correct energetic minimum.

The result is a functional machine: an active enzyme with a catalytic pocket sculpted down to tenths of an angstrom, ready to bind a substrate and catalyze chemical reactions millions of times faster than uncatalyzed chemistry.


The Immortal Relay

Every biological structure you possess—the keratin in your fingernails, the rhodopsin in your retinas that catches light, the insulin balancing your blood sugar, and the actin fibers moving your fingers—was forged by this exact molecular assembly line.

$$\textbf{Nuclear Archive (DNA)} ;\longrightarrow; \textbf{Working Script (mRNA)} ;\longrightarrow; \textbf{Factory (Ribosome)} ;\longrightarrow; \textbf{Molecular Machine (Protein)}$$

It is a manufacturing protocol shared by every bacterium at the bottom of the Mariana Trench, every redwood tree in California, and every human heart.

Yet this manufacturing pipeline is only as reliable as the DNA blueprints that feed it.

What happens when a stray cosmic ray, an environmental toxin, or a microscopic copying slip alters a single nucleotide in the DNA code?

In our next explainer, How Mutations Actually Happen, we examine the physical mechanisms of genetic failure: how chemical damage and replication errors alter the code of life, and why mutations are the raw fuel of both cancer and evolutionary adaptation.

Core Concepts Introduced10 Concepts
The Central Dogma (DNA ──► RNA ──► Protein)RNA Polymerase & Promoter RecognitionMessenger RNA (mRNA) vs Transfer RNA (tRNA)The 64-Codon Genetic Code TableStart Codon (AUG / Methionine) & Stop Codons (UAA, UAG, UGA)Ribosomal Architecture (Large 50S/60S & Small 30S/40S Subunits)Ribosomal Active Sites: A (Aminoacyl), P (Peptidyl), E (Exit)Peptidyl Transferase Ribozyme ReactionProtein Folding: Hydrophobic Collapse & ChaperoninsPost-Translational Modifications (Phosphorylation, Glycosylation)
Knowledge Graph Connections

Where to Go From Here

Explore companion architectures or dive deeper into downstream mechanisms.

Next Question

How Cellular Respiration and ATP Power Living Cells

Why do living cells need oxygen to extract energy from food, and how does the burning of glucose forge sixty kilograms of ATP inside your body every day?

Explore How Cellular Respiration and ATP Power Living Cells
Next Question

How Enzymes Catalyze the Reactions of Life

Why would the chemical reactions that sustain human life take millions of years to happen on their own at body temperature without enzymes?

Explore How Enzymes Catalyze the Reactions of Life
Research Grounding & Primary Sources

Verified Specifications & Architectural References

3 Authoritative References

This explainer is grounded in primary-source engineering specifications, regulatory circulars, and standard documentation.

Primary SourceNature (Francis Crick)• 1970

Central Dogma of Molecular Biology

Crick's definitive clarification of the sequence hypothesis, establishing that sequence information can be transferred from nucleic acid to protein, but never from protein to nucleic acid.

Primary SourceCold Spring Harbor Laboratory (Marshall W. Nirenberg)• 1966

The Genetic Code (Cold Spring Harbor Symposia on Quantitative Biology)

The historical decipherment of the universal 64-codon dictionary through synthetic poly-U cell-free translation assays.

Springer (G. E. Schulz & R. H. Schirmer)• 1979

Principles of Protein Structure

Foundational structural biology text detailing peptide geometry, the Ramachandran plot, alpha helices, beta sheets, and thermodynamic folding pathways.

Previous ExplainerHow DNA Stores and Replicates InformationNext Explainer How Mutations Actually Happen
More from Life & Evolutionary Biology•Topic Hub: Life & Evolutionary BiologyTopic Hub: Life & Evolutionary Biology
Ground Truth Engineering Publication