The Public Databases Behind Your Genetic Profile
A credible consumer genomics report is only as trustworthy as the data it draws from. GeneticCoding.com doesn't generate its own variant interpretations from scratch — instead, every annotated SNP in your file is cross-referenced against a set of established public and government genomics databases. These are the same resources clinicians, researchers, and regulators rely on, and citing them directly is what lets a report show the evidence behind each finding rather than presenting an opaque score.
The largest source of clinical classifications is ClinVar, a public archive run by the US National Center for Biotechnology Information. ClinVar records whether a variant is classified as pathogenic, likely pathogenic, of uncertain significance, likely benign, or benign, along with the submitting laboratory and a last-evaluated date. The GWAS Catalog, maintained by the European Bioinformatics Institute, curates genome-wide association studies — the research that links specific variants to traits and common disease risk across large populations. For medication response, PharmGKB holds pharmacogenomic annotations describing how variants change the way your body processes particular drugs.
Two further sources anchor the reference layer. dbSNP is the canonical registry of rsIDs — the reference SNP identifiers that let every database talk about the same position consistently — and is also a US government resource. MedlinePlus Genetics, from the National Library of Medicine, supplies the plain-language, publicly funded explanations of genes and conditions that give a report its readable, non-technical tone. We additionally consult ClinGen for curated gene-disease validity assessments, GenCC for gene-disease validity curation, and Ensembl for variant consequence annotation. We also draw on gnomAD, the Genome Aggregation Database, for population frequency, gene constraint, and rare-variant evidence across hundreds of thousands of sequenced genomes, and on the Human Phenotype Ontology (HPO) for standardized phenotype and disease-feature vocabulary. For broader literature reach, every finding also links to Google Scholar, Google's academic search engine, which indexes the same peer-reviewed research behind each association so you can explore the primary studies yourself. A Google Searchlink is also added to every finding as a broad, real-time cross-check of the variant across the open web — the broadest-reach source and a useful gap-finder for variants the curated databases miss.
Three newer arrivals round out the panel. Open Targets is an open partnership between academia and industry that aggregates genetics, genomics, and literature evidence into scored gene–disease associations — a second independent read on how strongly a gene is linked to a condition. ClinPGx carries forward the PharmGKB and CPIC curation tradition with up-to-date gene–drug dosing annotations for medication response findings. And MyGene.info, a public gene-annotation service, supplies the lightweight gene-level annotation layer — summaries, identifiers, and cross-references — that sits under every catalog entry.
Two practical consequences follow from sourcing this way. First, the reports are educational, not diagnostic: a ClinVar classification or a GWAS association describes a statistical or clinical signal in a population, not a prediction about you personally. Second, the knowledge is constantly moving — ClinVar publishes updates roughly weekly, and the slower sources still change regularly. That is exactly why ongoing monitoring for new findings matters, and why a report should always cite the specific database, entry, and date behind every claim so you can verify it yourself.
We are also a source in our own right: the GeneticCoding SNP Encyclopedia is our public, crawlable library of long-form research profiles — one page per rsID in our catalog, with genotype-by-genotype interpretation, evidence strength, population frequency, and citations to the primary databases above. It is generated and refreshed by the same interpretation pipeline behind our reports, so finding pages and paid reports never disagree.
Each report type draws on the sources most relevant to it. Health Predisposition and the newer Nutrigenomics, Fitness, Sleep, Longevity & Healthspan, Skin & Photo-aging,Stress Response, Reproductive, Allergies, andAppearance sections cite the GWAS Catalog and PubMed for the association studies behind each variant. Pharmacogenomics findings add PharmGKB, CPIC, and ClinPGx dosing guidance for the gene in question. Health Predisposition and Carrier Status entries also draw scored gene–disease association evidence from Open Targets. Carrier Status and Vision findings point to ClinVar and MedlinePlus Genetics, and Ancestral findings reference population-frequency data from Ensembl and 1000 Genomes. Across every category, Google Scholar provides a supplementary academic-literature search so the primary studies behind each finding are always one click away. Every claim links out to its primary source so the evidence is verifiable yourself.
