Genetics & Breeding
Gene Expression Part 1: Reading Genes to Make Proteins
Overview
This lesson describes the steps involved in a cell as DNA sequence information is read to make RNA and RNA is read to make proteins. A gene will only control a trait in an organism when the gene is expressed. This means the gene is read in the cell to make a protein that carries out a specific function. This lesson describes the basic steps in the gene expression process.
Authors
- Don Lee, Department of Agronomy and Horticulture, University of Nebraska-Lincoln
- Patricia Hain, Department of Agronomy and Horticulture, University of Nebraska-Lincoln
Published 2021
Learning Objectives
- Define the roles of DNA and proteins in cell development and metabolism
- Determine the amino acid sequence of a protein given the nucleotide sequence of a gene.
- Describe the roles that the promoter, coding region and termination sequence of a gene play in gene expression.
- Recognize the differences between the structure of proteins, amino acids, genes and nucleotides.
- Draw the process of gene expression and include the following in your drawing. Gene, RNA polymerase, promoter, coding region, termination sequence, intron, cell, nucleus, cytoplasm, mRNA, tRNA, ribosome, anticodon, codon, amino acid, protein, peptide bond.
The Need for Gene Expression
Genes are DNA sequences that control traits in an organism by coding for proteins. Organisms such as plants and animals have tens of thousands of genes. The impact that a single gene’s information can have on an organism, however, is tremendous. Furthermore, organisms have all of their genes in each of their cells but they only need to use the information from a subset of these genes, depending on the type of cell and the cell’s stage of development. Therefore the key to gene function is controlling its expression.



Its importance is evident when you observe the changes plants go through during their lifecycle or during a season. Trees and bushes, for example, have dormant buds through the winter. Environmental signals affiliated with the coming of spring induce genes in the buds to turn on and drive the dramatic changes of leaf development and flowering. The genes were always present in those bud cells but were controlled to turn on at the proper time. Understanding gene expression thus requires an examination of two processes, the activation of gene expression to make an information message and the reading of this message to build a specific protein.


Expression of Acetolactate Synthase Enzyme
The gene encoding the enzyme acetolactate synthase (ALS) is essential for organisms that need to make their own supply of the amino acids leucine, isoleucine and valine. The ALS enzyme is needed in the cell to catalyze a reaction in the chemical pathway used to make these three amino acids.

There are 20 different amino acids making up the subunits of proteins. All proteins are made from differing combinations of the amino acids (Fig. 6, 7) Therefore if you are going to make proteins you need to either make amino acids first or consume amino acids in your diet. Animals consume some of their amino acids (the essential amino acids) but plants and bacteria are able make all their own amino acids.


The ALS gene is not found in animals but is found in plants and bacteria.

The ALS gene consists of several thousand nucleotides of DNA and like all genes has three main parts, the promoter, coding region and termination sequences (Fig. 8, gene sequence abbreviated for demonstration). Each part plays a role in controlling gene expression. The promoter is the on/off switch of the gene, the coding region determines what protein will get made and the termination sequence signals where the gene information ends. Each of these will now be described more fully.

Transcription: control from the promoter and termination sequence
The promoter’s role is to selectively turn on the ALS gene in cells that will need the ALS enzyme. The promoter accomplishes this by its specific nucleotide sequences. We will discuss the regulation of gene expression in more detail later. What cells in a plant will need the ALS enzyme? All cells that are making proteins. Obviously this will be all cells in the plant at some stage of development. Therefore the ALS gene has a promoter sequence that allows the gene to be turned on in all types of cells.
When the ALS gene is turned on, transcription can take place. Like most chemistry in the cell, enzymes will be needed for gene expression. RNA polymerase is the RNA polymerasetranscription enzyme (Fig. 9) which will bind to the DNA sequences in the ALS promoter and interact with one strand of the gene (the coding strand).


RNA polymerase can literally move along the gene promoter until it encounters a series of A and T nucleotides called the TATA box (Fig 10.). The TATA box signals RNA polymerase that it has reached the end of the prom oter. Now the enzyme has the “green light” to read the coding strand DNA and make RNA. The procedure works much like DNA replication. RNA polymerase will read the DNA in a 3’ to 5’ direction and build the RNA 5’ to 3’. Unlike DNA replication though, only one DNA strand is read and the transcription process must end once the RNA polymerase reaches the end of the ALS gene. If RNA polymerase kept going along that strand of the DNA making up the chromosome, other genes that are on the same chromosome might be expressed in the wrong cells or wrong times.

How is the RNA polymerase signaled that it has reached the end of the gene? That is the role of a termination sequence in the gene (Fig. 11). The termination sequences signal the end of the gene and can work in a number of ways. One transcription termination strategy will be described here. The nucleotides that make up the termination sequence of the ALS gene could be ordered to form a palindrome. This means that the nucleotides being placed in the RNA can fold back on themselves and form a hairpin loop (Figs. 12-14)




While no geneticist has actually viewed this hairpin loop formation in action, it is believed that the hairpin formation snaps the RNA polymerase off the DNA coding strand and releases the RNA message. Transcription of an RNA that has the coding region information is now completed. Each time an RNA polymerase goes through the process, one copy of the RNA is made by reading the DNA template.The type of RNA made from transcription of the ALS gene is called a messenger RNA or mRNA for short. Some genes can encode transfer RNA (tRNA) or ribosomal RNA (rRNA). These RNA molecules have special roles in the next part of gene expression called translation. Once an mRNA is transcribed some modifications will occur in the nucleus. We will describe some of the post-transcriptional processes later. For now lets follow the ALS mRNA onto the translation step of gene expression.
*This animation has no audio.*
Translation: reading the code to make proteins
The translation process requires a conversion of information from a chain of nucleotides in RNA to a chain of amino acids in protein. The translation term was coined based on the idea that we are converting to a new molecular language at this stage of gene expression. The conversion of information in translation is a little more complex than transcription and requires a number of molecules to interact and work together (Fig 15).
Ribosomes and tRNAs work together to translate a mRNA. The tRNA is a small RNA (70 to 80 nucleotides) and has a sequence that allows hairpins to form. This gives a secondary structure to the molecule.

Three of the nucleotides near the middle of the tRNA are the anticodon. When the tRNA folds, the anticodon becomes one end of the molecule and the other end will be attached to a specific amino acid. In the figures for this lesson, the tRNA is shown with it’s secondary structure and only the anticodon sequence is shown (Fig. 16). Ribosomes are the molecular workbench of translation. Ribosomes are composed of rRNAs and proteins that will have a complex secondary structure. Active ribosomes are assembled from small and large subunits when the ribosome “workbench” is ready to translate a mRNA . When mRNAs are present in the cytoplasm, the small and large ribosome subunits assemble at the 5’ end of these messages and are ready to read mRNAs like blueprints to build specific proteins.

The ribosome reads the mRNA quickly and accurately but it must know where to start. As the ribosome moves in the 5’ to 3’ direction along the mRNA it “looks” for the sequence AUG (Fig. 16). This sequence is called the start codon and signals the ribosome that this is where to begin building the protein sequence. From this point, the ribosome reads the continuous chain of nucleotides in groups of three and builds the protein according to the instructions in the mRNA sequence. Typically, proteins are several hundred or more amino acids in length. When a complete protein is built, the ribosome will leave the mRNA and look for the 5’ end of another mRNA to begin translation. Each time one ribosome translates an mRNA, one copy of a protein is made.
Why a Triplet Code?
Prior to understanding the details of transcription and translation, geneticists predicted that DNA could encode amino acids only if a code of at least three nucleotides was used. The logic is that the nucleotide code must be able to specify the placement of 20 amino acids. Since there are only four nucleotides, a code of single nucleotides would only represent four amino acids, such that A, C, G and U could be translated to encode amino acids. A doublet code could code for 16 amino acids (4 x 4). A triplet code could make a genetic code for 64 different combinations (4 x 4 x 4) genetic code and provide plenty of information in the DNA molecule to specify the placement of all 20 amino acids. When experiments were performed to crack the genetic code it was found to be a code that was triplet. These three letter codes of nucleotides (AUG, AAA, etc.) are called codons.
The genetic code only needed to be cracked once because it is universal (with some rare exceptions). That means all organisms use the same codons to specify the placement of each of the 20 amino acids in protein formation. A codon table can therefore be constructed and any coding region of nucleotides read to determine the amino acid sequence of the protein encoded. A look at the genetic code in the codon table below reveals that the code is redundant meaning many of the amino acids can be coded by four or six possible codons. The amino acid sequence of proteins from all types of organisms is usually determined by sequencing the gene that encodes the protein and then reading the genetic code from the DNA sequence.
| First Position | Second Position | Third Position | |||||||
|---|---|---|---|---|---|---|---|---|---|
| U | C | A | G | ||||||
| Code | Amino Acid | Code | Amino Acid | Code | Amino Acid | Code | Amino Acid | ||
| U | UUU | Phenylalanine (Phe, F) | UCU | Serine (Ser, S) |
UAU | Tyrosine (Tyr, Y) |
UGU | Cysteine (Cys, C) |
U |
| UUC | UCC | UAC | UGC | C | |||||
| UUA | Leucine (Leu, L) |
UCA | UAA | Stop | UGA | Stop | A | ||
| UUG | UCG | UAG | UGG | Tryptophan (Trp, W) |
G | ||||
| C | CUU | Leucine (Leu, L) |
CCU | Proline (Pro, P) |
CAU | Histidine (His, H) |
CGU | Arginine (Arg, R) |
U |
| CUC | CCC | CAC | CGC | C | |||||
| CUA | CCA | CAA | Glutamine (Gln, Q) |
CGA | A | ||||
| CUG | CCG | CAG | CGG | G | |||||
| A | AUU | Isoleucine (Ile, I) |
ACU | Threonine (Thr, T) |
AAU | Asparagine (Asn, N) |
AGU | Serine (Ser, S) |
U |
| AUC | ACC | AAC | AGC | C | |||||
| AUA | ACA | AAA | Lysine (Lys, K) |
AGA | Arginine (Arg, R) | A | |||
| AUG | Methionine Start (Met, M) |
ACG | AAG | AGG | G | ||||
| G | GUU | Valine (Val, V) |
GCU | Alanine (Ala, A) | GAU | Aspartic acid (Asp, D) |
GGU | Glycine (Gly, G) |
U |
| GUC | GCC | GAC | GGC | C | |||||
| GUA | GCA | GAA | Glutamic acid (Glu, E) |
GGA | A | ||||
| GUG | GCG | GAG | GGG | G | |||||
The Reading Frame, Codons and Anticodons
The mechanics of reading mRNAs to build proteins is an orchestration of several interacting molecules. The mRNAs, tRNAs, ribosomes and amino acids are the key players.
The mRNA has the codons which signal to the ribosome when to start the protein building process and when to end it. The codons in the middle known as the reading frame, determine which amino acids will be placed into the protein. The AUG start codon establishes the beginning of the reading frame on a mRNA. The ribosome must follow this reading frame to build the correct protein. How are the appropriate amino acids placed into the protein as it is built? The amino acids need to be transferred to the ribosomes with the assistance of tRNAs.
The tRNAs are able to perform their transfer and placement function because of their structure (Fig. 17). The tRNA molecule is small, only 70-80 nucleotides in length. Those sequences promote hairpin loops to form, giving tRNA a stable secondary structure. The structure of the tRNA is recognized by special enzymes in the cell that attach the proper amino acid to the tRNAs. The tRNA also has a sequence of three nucleotides called the anticodon. Anticodons on the tRNA will complement and bind to the codon on the mRNA to specify the correct amino acid placement in the growing protein chain. Let’s describe how the key players get protein building started, how they keep it going, and how they end it.

Getting Translation Started with the Start Codon
The ribosome workbench uses the AUG codon as a universal signal to begin translation. The AUG start codon signals the ribosome to place in the amino acid methionine because the tRNA that has methionine attached to it has the anticodon sequence UAC. Therefore the tRNA will temporarily bind to the mRNAs sequence. The ribosome can attach to the mRNA and then allow the tRNA to come in because it has a position called the ‘A’ site. This ribosome site works somewhat like a vice on the workbench, holding stuff in place so it can be worked on. If the codon and anticodon complement, the ribosome will slide over one codon on the mRNA in the 5′ to 3′ direction. This places the start codon part of the mRNA at a new ribosome position called the ‘P’ site. This site acts like a second vice on the workbench. The ‘A’ site will be open and a second tRNA can come in (Fig. 18). If the codon on the mRNA and anticodon on this second tRNA complement well enough to satisfy the ribosome, it will hold both tRNAs in place. Now the ribosome is ready to start hooking together amino acids to form a protein.


Peptide Bond Formation and Protein Building
A ribosome “workbench” has catalytic tools to go along with it’s ‘P’ and ‘A’ “vice sites”. The ribosome has enzymatic functions that allow it to break and form bonds. The ribosome will break the bond that binds the amino acid (met) to the tRNA at the ‘P’ site. Simultaneously the ribosome forms a peptide bond between the two amino acids that were brought in by the tRNAs. The tRNA at the ‘A’ site will now have two amino acids connected to it. The close proximity of these amino acids at the ribosome sites provides the opportunity for peptide bond formation (Fig. 19). Peptide bonds always bind the acid end of one amino acid with the amino end of the next (Fig. 6). Peptide bonds are strong covalent bonds which keep the amino acids connected during and after translation. The tRNA at the ‘P’ site has now accomplished its task by bringing in the first amino acid. This tRNA will then move off the ribosome (Fig. 19). The first two amino acids in the protein have now been put together.
The synthesis of our protein is far from over. With the ‘P’ site now open, the ribosome will again shift one codon in the 5’ to 3’ direction and open up the ‘A’ site. A tRNA can now come into this site if it has the anticodon that is complementary to the next codon on the mRNA (Figs. 20, 21). The ribosome will then form the peptide bond between the third and second amino acids and kick off tRNA number two (Fig. 21). The third tRNA now binds a chain of three amino acids. The ribosome shifts in preparation for tRNA number 4. The process continues as long as the ribosome encounters the proper tRNAs bringing in amino acids.


The tRNAs are Recycled
Proteins are hundreds of amino acids long so the steps described above must be repeated many times in cells that actively make proteins. The tRNA released from a ribosome is recycled by the cell. It can be charged again by binding another amino acid in the cytoplasm and contribute to the synthesis of another protein.
The End of Translation: stop codons looking for something they cannot find
Start codons start translation so it is logical that stop codons stop translation (Fig 22). How do stop codons do this? There are 61 tRNAs with different anticodons (see codon table). That means there are three codons that do not have corresponding tRNAs with complementary anticodons. These three codons serve as stop codons. When a ribosome encounters a stop codon on a mRNA it will wait for a tRNA with the right anticodon to come over. It will not skip the codon or shift over one nucleotide to form a new reading frame. The ribosome waits for the right tRNA, but it does not wait for long. A stalled ribosome will quickly cleave off the bound tRNA with the growing protein chain and then move on to translate another mRNA. No tRNAs in the cell have anticodons that complement any of the three possible stop codons. Therefore, stop codons are able to end the translation process when the completed protein is made.


| First Position | Second Position | Third Position | |||||||
|---|---|---|---|---|---|---|---|---|---|
| U | C | A | G | ||||||
| Code | Amino Acid | Code | Amino Acid | Code | Amino Acid | Code | Amino Acid | ||
| U | UUU | Phenylalanine (Phe, F) | UCU | Serine (Ser, S) |
UAU | Tyrosine (Tyr, Y) |
UGU | Cysteine (Cys, C) |
U |
| UUC | UCC | UAC | UGC | C | |||||
| UUA | Leucine (Leu, L) |
UCA | UAA | Stop | UGA | Stop | A | ||
| UUG | UCG | UAG | UGG | Tryptophan (Trp, W) |
G | ||||
| C | CUU | Leucine (Leu, L) |
CCU | Proline (Pro, P) |
CAU | Histidine (His, H) |
CGU | Arginine (Arg, R) |
U |
| CUC | CCC | CAC | CGC | C | |||||
| CUA | CCA | CAA | Glutamine (Gln, Q) |
CGA | A | ||||
| CUG | CCG | CAG | CGG | G | |||||
| A | AUU | Isoleucine (Ile, I) |
ACU | Threonine (Thr, T) |
AAU | Asparagine (Asn, N) |
AGU | Serine (Ser, S) |
U |
| AUC | ACC | AAC | AGC | C | |||||
| AUA | ACA | AAA | Lysine (Lys, K) |
AGA | Arginine (Arg, R) | A | |||
| AUG | Methionine Start (Met, M) |
ACG | AAG | AGG | G | ||||
| G | GUU | Valine (Val, V) |
GCU | Alanine (Ala, A) | GAU | Aspartic acid (Asp, D) |
GGU | Glycine (Gly, G) |
U |
| GUC | GCC | GAC | GGC | C | |||||
| GUA | GCA | GAA | Glutamic acid (Glu, E) |
GGA | A | ||||
| GUG | GCG | GAG | GGG | G | |||||
Summary
Organisms use the process of transcription to make mRNA as a temporary molecule to transport protein coding information from a gene. The translation process reads the mRNA and makes the intended protein, one amino acid at a time. This process is controlled by the promoter of the gene to insure that proteins are made in the right cells at the right time to drive growth, development, metabolism, stress response and other processes needed to complete the organisms life cycle. The next lesson (Gene expression part 2) will apply the gene expression principles to the observation that some plants can be resistant to ALS inhibitor herbicides.
Acknowledgements
Development of this lesson was supported in part by Cooperative State Research, Education, & Extension Service, U.S. Dept of Agriculture under Agreement Number 98-EATP-1-0403 administered by Cornell University and the American Distance Education Consortium (ADEC). Any opinions, findings, conclusions or recommendations expressed in this publication are those of the author(s) and do not necessarily reflect the view of the U.S. Department of Agriculture.
(deoxyribonucleic acid) The molecule that encodes genetic information. DNA is a double-stranded molecule held together by weak bonds between base pairs of nucleotides. It is the fundamental substance of which genes are composed.
Large molecules composed of one or more chains of amino acids in a specific order. Proteins are necessary for the structure, function, and regulation of the organism's cells, tissues, and organs. Each protein has a unique function determined by its shape.
The fundamental unit of heredity that carries genetic information from one generation to the next. A gene is an ordered sequence of nucleotides located on a particular position on a particular chromosome that encodes a specific functional protein.
process.
A protein that catalyzes, or speeds up, a specific biochemical reaction without changing the nature of the reaction.
The basic building blocks of proteins. The sequence of amino acids in a protein and protein function are determined by the genetic code.
A specific DNA sequence to which RNA polymerase binds and initiates transcription. This region contains information which regulates when and how often the gene is transcribed and ultimately the amount of protein it produces.
An enzyme that catalyzes the synthesis of RNA by copying the nucleotide sequence of the DNA.
The process by which the nucleotide sequence of DNA is copied into a single-stranded molecule of RNA. The nucleotide sequence of the RNA created is complementary to the DNA sequence except all thymine molecules are replaced with uracil molecules.
The association of genes on the same chromosome. The shorter the distance between two genes, the greater the probability they will be inherited together.
The sequence of DNA which signals the transcription to stop.
The process following transcription during which the nucleotide sequence of mRNA is read and 'translated' into a chain of amino acids (protein). The mRNA sequence is read three nucleotides (codon) at a time, and each codon codes for a specific amino acid.
The building blocks of DNA and RNA; adenine, cytosine, guanine, thymine (DNA only), and uracil (RNA only).
(messenger ribonucleic acid) The message made during transcription by reading the DNA sequence to build a particular protein. A single-stranded nucleic acid similar to DNA but having a ribose sugar rather than deoxyribose sugar and a uracil rather than thymine as one of the bases.
Region of the DNA sequence between the promoter and the termination sequence. It contains the instructions about how to make a specific protein.