Protein primary structure is the linear sequence of amino acids in a peptide or protein. By convention, the primary structure of a protein is reported starting from the amino-terminal (N) end to the carboxyl-terminal (C) end. Protein biosynthesis is most commonly performed by ribosomes in cells. Peptides can also be synthesized in the laboratory. Protein primary structures can be directly sequenced, or inferred from DNA sequences.
Formation
Biological
Amino acids are polymerised via peptide bonds to form a long backbone, with the different amino acid side chains protruding along it. In biological systems, proteins are produced during translation by a cell's ribosomes. Some organisms can also make short peptides by non-ribosomal peptide synthesis, which often use amino acids other than the encoded 22, and may be cyclised, modified and cross-linked.
Chemical
Peptides can be synthesised chemically via a range of laboratory methods. Chemical methods typically synthesise peptides in the opposite order (starting at the C-terminus) to biological protein synthesis (starting at the N-terminus).
Notation Protein sequence is typically notated as a string of letters, listing the amino acids starting at the amino-terminal end through to the carboxyl-terminal end. Either a three letter code or single letter code can be used to represent the 22 naturally encoded amino acids, as well as mixtures or ambiguous amino acids (similar to nucleic acid notation). Peptides can be directly sequenced, or inferred from DNA sequences. Large sequence databases now exist that collate known protein sequences.
Modification
In general, polypeptides are unbranched polymers, so their primary structure can often be specified by the sequence of amino acids along their backbone. However, proteins can become cross-linked, most commonly by disulfide bonds, and the primary structure also requires specifying the cross-linking atoms, e.g., specifying the cysteines involved in the protein's disulfide bonds. Other crosslinks include desmosine.
Isomerisation The chiral centers of a polypeptide chain can undergo racemization. Although it does not change the sequence, it does affect the chemical properties of the sequence. In particular, the L-amino acids normally found in proteins can spontaneously isomerize at the C α {\displaystyle \mathrm {C^{\alpha }} } atom to form D-amino acids, which cannot be cleaved by most proteases. Additionally, proline can form stable trans-isomers at the peptide bond.
Post-translational modification Additionally, the protein can undergo a variety of post-translational modifications, which are briefly summarized here. The N-terminal amino group of a polypeptide can be modified covalently, e.g.,
acetylation − C ( = O ) − C H 3 {\displaystyle \mathrm {-C(=O)-CH_{3}} }
The positive charge on the N-terminal amino group may be eliminated by changing it to an acetyl group (N-terminal blocking). formylation − C ( = O ) H {\displaystyle \mathrm {-C(=O)H} }
The N-terminal methionine usually found after translation has an N-terminus blocked with a formyl group. This formyl group (and sometimes the methionine residue itself, if followed by Gly or Ser) is removed by the enzyme deformylase. pyroglutamate
An N-terminal glutamine can attack itself, forming a cyclic pyroglutamate group. myristoylation − C ( = O ) − ( C H 2 ) 12 − C H 3 {\displaystyle \mathrm {-C(=O)-\left(CH_{2}\right)_{12}-CH_{3}} }
Similar to acetylation. Instead of a simple methyl group, the myristoyl group has a tail of 14 hydrophobic carbons, which make it ideal for anchoring proteins to cellular membranes. The C-terminal carboxylate group of a polypeptide can also be modified, e.g.,
amination (see Figure) The C-terminus can also be blocked (thus, neutralizing its negative charge) by amination. glycosyl phosphatidylinositol (GPI) attachment Glycosyl phosphatidylinositol(GPI) is a large, hydrophobic phospholipid prosthetic group that anchors proteins to cellular membranes. It is attached to the polypeptide C-terminus through an amide linkage that then connects to ethanolamine, thence to sundry sugars and finally to the phosphatidylinositol lipid moiety. Finally, the peptide side chains can also be modified covalently, e.g.,
… excerpt ends here. Continue reading the full article.





