Gene Annotation: How Scientists Identify and Describe Human Genes
Sequencing the human genome was a monumental achievement — but producing a raw sequence of three billion base pairs is only the beginning. Making that sequence biologically meaningful requires annotation: the systematic identification and description of all functional elements encoded in the DNA. Gene annotation is the process of locating gene boundaries, predicting gene structure, inferring function, and connecting sequences to existing biological knowledge. It is an ongoing endeavour, and our picture of the genome continues to be refined as new data and analytical methods emerge.
Computational Gene Prediction
The first phase of gene annotation relies on computational tools that search DNA sequences for signatures of genes. Gene prediction algorithms look for open reading frames — stretches of DNA bounded by start and stop codons that could encode a protein. They also look for donor and acceptor splice sites (the sequences at the junctions between exons and introns), promoter elements, and codon usage patterns consistent with protein-coding genes. Comparative genomics adds another layer: sequences conserved across multiple species are more likely to be functional, and tools like BLAST and synteny analysis help identify functional elements by comparing human sequences to those of other vertebrates.
Experimental Validation: Transcriptomics and Proteomics
Computational predictions are validated and refined using experimental data. RNA sequencing (RNA-seq) captures all the RNA molecules produced by cells under specific conditions, directly revealing which genomic regions are transcribed and the structure of the resulting transcripts. Mass spectrometry-based proteomics identifies actual protein sequences in biological samples, confirming that predicted protein-coding genes are translated. The integration of multiple experimental data types — including CAGE-seq for transcription start sites, CLIP-seq for RNA-binding protein interactions, and ribosome profiling for active translation — allows annotation databases to move beyond predictions toward evidence-based gene models.
Functional Annotation: What Does a Gene Do?
Identifying a gene's location and structure is only half the annotation task — determining its function requires additional evidence. Gene Ontology (GO) terms provide a standardised vocabulary for describing biological processes, molecular functions, and cellular components associated with gene products. GO annotations are inferred from multiple types of evidence: direct experimental data, sequence similarity to characterised proteins in other organisms, and computational prediction. Pathway databases like KEGG and Reactome place genes in the context of biological pathways and metabolic networks, helping researchers understand not just what a gene does in isolation but how it participates in broader cellular processes.
Annotation Challenges and Ongoing Work
Despite decades of effort, human gene annotation remains incomplete and imperfect. The exact number of protein-coding genes is still debated as annotation methods improve. Thousands of predicted genes lack functional characterisation. The non-coding portion of the genome — particularly long non-coding RNAs and regulatory elements — is far less well annotated than protein-coding regions. Major international efforts including GENCODE, ENCODE, and the Functional Annotation of the Mammalian Genome (FANTOM) consortium continue to generate new data and refine existing annotations. The result is a genome annotation that improves continuously, benefiting every researcher who relies on it.
Browse our gene annotation resources and database tools, or contact us for guidance on using gene annotation in your research.