The Degeneracy of the Genetic Code

Author: Shenzhou
Reviewed by: Shenzhou

  During protein translation, the genetic code establishes the main informational connection between a nucleotide sequence and an amino acid sequence, while tRNA goes on to create the physical connection between them. In other words, the genetic code is the essential bridge linking proteins with nucleotide sequences.
  Early in the study of the genetic code, mathematical analysis led some scientists to conclude that only a system in which three bases specify one amino acid could encode all 20 amino acids. Crick later performed experiments in which different numbers of bases were inserted into or deleted from bacteriophages and the resulting phenotypes compared. His results confirmed the hypothesis that the genetic code consists of nucleotide triplets.
  At the same time, Nirenberg and Matthaei used artificially synthesized mRNA and an in vitro protein-translation system to decipher several simple correspondences between codons and amino acids, including UUU, CCC, and AAA. Soon afterward, Nirenberg continued deciphering the rest of the genetic code by taking advantage of nitrocellulose membranes, which retain complexes formed by ribosomes, mRNA, and tRNA carrying the corresponding amino acid.
  They first prepared 20 in vitro synthesis systems, each containing all 20 amino acids and ribosomes. A different one of the 20 amino acids was labeled with carbon-14 (C14) in each system. Artificially synthesized trinucleotide RNA was then allowed to react separately in the systems, after which each reaction mixture was filtered through a nitrocellulose membrane. In theory, only a complete ribosome–amino acid–tRNA–mRNA complex would remain on the membrane and produce a strong radioactive signal there. This method could therefore identify, one by one, the nucleotide sequence corresponding to any amino acid. Using it, scientists ultimately deciphered all 61 codons other than the stop codons.
  Because stop codons do not encode any amino acid, they could not be deciphered with the method above, and a different approach was needed. In 1965, Garen found one: studying reversions in E. coli nonsense mutants to infer and analyze the sequences of stop codons.
  An amber mutant of the E. coli alkaline phosphatase gene arose when its tryptophan (UGG) site mutated into a stop codon. After obtaining large numbers of revertants, Garen studied the new amino acids produced at this site after reversion and analyzed their corresponding sequences: Ser (UCG), Leu (UUG), Tyr (UAC), Lys (AAG), Gln (CAG), and Glu (GAG). Analysis of these sequences showed that the stop mutation produced from Trp (UGG) was UAG, completing the first proof of a stop codon. Finally, in 1967, Brenner and Crick used this method to prove the final stop codon. The three stop codons were named after the nonsense mutations used in the proof: UAG—the amber codon, UAA—the ochre codon, and UGA—the opal codon.
  Once all the codon sequences that encode amino acids had been determined, it was easy to see that there are far more codons than amino acids. This phenomenon is known as the degeneracy of the genetic code. Leu, for example, is encoded by four codons: CUU, CUC, CUA, and CUG. Examining the codons that encode other amino acids also reveals that for any given amino acid, the first two positions tend to be distinctive and fixed, while the third varies more. This suggests that a codon's third position is not important to its encoding of an amino acid. That is an important mechanism in amino acid coding and contributes to the degeneracy of the genetic code to some extent.
  Because the tRNA anticodon lies in a loop, its three bases follow a curved arrangement. An mRNA codon, by contrast, is linear, so the two do not align with perfectly regular geometry. The first base of the anticodon also has considerable freedom because of the structure of tRNA itself. As a result, the first anticodon base and the third codon base do not need to pair in strict accordance with complementary base-pairing rules. If this anticodon position is modified, still more pairings become possible. This wobble pairing allows 61 codons to be read by only 32 tRNAs, reducing to some extent the number of genes needed to encode tRNA.
  The degeneracy of the genetic code is highly significant. It gives an organism's genes a relatively large “margin for error” and greatly reduces the harm caused by mutation. Under this decoding mechanism, a nucleotide mutation at some sites has a good chance of leaving the organism's expressed traits unaffected, helping to stabilize the propagation of a species. The way amino acids are encoded also contains a further mechanism that reduces the risk of mutation.
  We said earlier that the third codon position is not important to amino acid coding. What roles, then, do the first two positions play? The answer is connected to the physicochemical properties of amino acids.
  Comparing codons with amino acid polarity shows that codons sharing the same second base encode amino acids with similar physical or chemical properties. In general, amino acids that are nonpolar or have uncharged side chains have C in the second position of their codons; more strongly hydrophobic amino acids mostly have U in the second position; and codons with A or G in the second position generally correspond to hydrophilic amino acids. Under this coding mechanism, as long as a codon's second position remains unchanged, even a change in the amino acid ultimately translated can leave it with a function similar to that of the original amino acid, further reducing the risk posed by genetic mutation.