How the Human Gene Database Is Organised and Used
Modern genetics would be unmanageable without databases. The human genome contains roughly 20,000 protein-coding genes and many thousands of non-coding RNA genes and regulatory elements — far too much information for any individual or team to track manually. The public gene databases that emerged from the Human Genome Project and continue to grow today form the backbone of genomics research and clinical genomics, providing standardised, curated information that enables scientists worldwide to collaborate and build on each other's findings.
NCBI Gene: The Primary Reference Database
The National Center for Biotechnology Information (NCBI) Gene database is the most widely used human gene reference resource. Each gene entry has a unique numerical identifier (GeneID) and contains standardised information including the gene's official symbol and full name, its chromosomal location, the protein-coding transcripts it produces, known functional roles, associated genetic variants, links to the scientific literature, and connections to orthologous genes in other species. NCBI Gene serves as the hub linking to many related databases including GenBank (DNA sequences), RefSeq (reference sequences), dbSNP (single nucleotide variants), and ClinVar (clinical variant interpretations).
Ensembl: Genome Browser and Annotation Resource
Ensembl, jointly maintained by the European Bioinformatics Institute (EMBL-EBI) and the Wellcome Sanger Institute, provides a genome browser interface for exploring gene structure and context within the genome. It is particularly valuable for visualising how genes relate to each other and to genomic features like regulatory elements, repetitive sequences, and epigenetic marks. Ensembl uses its own gene naming system (ENSG identifiers) and integrates data from multiple annotation pipelines, making it a powerful tool for researchers working at the level of genome structure rather than individual gene function.
OMIM: Linking Genes to Disease
Online Mendelian Inheritance in Man (OMIM) focuses specifically on the relationship between genes and human genetic disorders. Originally compiled by Victor McKusick in the 1960s, OMIM is now a curated database of human genes and genetic phenotypes, providing detailed clinical summaries, inheritance information, molecular genetic findings, and references to the primary literature. Clinicians and clinical geneticists rely heavily on OMIM when evaluating patients with suspected genetic conditions, and it remains the authoritative source for understanding genotype-phenotype relationships in Mendelian disease.
How Researchers Use These Databases
Database use in genomics research spans the full scientific workflow. Bioinformaticians query databases to retrieve reference sequences for sequence alignment and variant analysis. Molecular biologists consult functional annotations when designing experiments. Clinical geneticists look up variant classifications in ClinVar and OMIM when interpreting diagnostic sequencing results. Computational biologists mine database records to discover patterns across hundreds or thousands of genes simultaneously. The interoperability of these databases — made possible by shared identifiers and standardised data formats — is what makes large-scale genomics research tractable.
Explore our curated gene database resources and research tools, or contact us for guidance on navigating human gene databases for your research or educational needs.