Using genome projects - AL only (3.8.3)
On this page
DNA sequencing can be used to determine the sequence of base pairs in an organism’s genome.
The Human Genome Project successfully sequenced the first complete human genome. Since then, thousands of human genomes have been sequenced.
Only about of the human genome is made up of coding DNA, which directly codes for proteins. The remaining 98 of the human genome is non-coding DNA, which does not directly code for proteins, but has other useful functions.
DNA is sequenced by chopping up the DNA into fragments using enzymes. These fragments are then sequenced by size, and computer algorithms are used to analyse the fragments and collate the information into the mapped genome.
Over time, this process has become quicker and cheaper. Whole genomes can now be sequenced through a process which is mostly automated.
The proteome consists of all the proteins produced by an organism’s genome.
Pre-mRNA contains introns and exons. Each gene can code for multiple different proteins when the pre-mRNA is spliced in different ways prior to translation.
Proteomics can aid in understanding how genes code for multiple proteins, depending on how the pre-mRNA is altered.

Sequencing human genomes has identified various single nucleotide polymorphisms (SNPs), which are common inherited genetic variants. These are the cases where one base has been replaced by another. SNPs may or may not have observable impacts.

Some SNPs are associated with the development of certain diseases, such as dementia or Alzheimer’s disease.
Other SNPs can be associated with how people respond to drugs.
As DNA sequencing becomes more accessible, genetic factors such as disease risk and drug response are becoming easier to predict.
Bacteria were the first organisms to have their genomes sequenced. These genomes are easier to sequence than those of larger, multicellular organisms; bacterial DNA does not have histones or non-coding DNA.
Sequencing the genomes of microorganisms is useful as it helps identify the genetic code relating to antigens. This information can then be used to develop vaccines.
Determining the bacterial proteome can also help identify the DNA that encodes the useful proteins produced by the microorganism.

Determining the proteome from a multicellular organism’s genome can be challenging.
- A large proportion of the DNA is non-coding, so it does not directly code for proteins.
- Due to introns and exons, the same section of DNA could also code for multiple different proteins.
- Once transcribed, proteins are often altered. This means the final protein can differ from the protein coded for by the DNA.
- There is variation among organisms of the same species, so mapping the proteome will vary between organisms within a species.


