Literature Corpus Collection and Search Methodology
The Folklore Literature Corpus is a governed, source-traceable database of genetics publications. It is built to support reproducible literature discovery while keeping publication identity, provenance, and known record status visible.
This page describes the public methodology and quality controls. It intentionally focuses on corpus behavior, coverage, and interpretation boundaries rather than proprietary infrastructure or model implementation.
Published and reviewed by Helena Bioinformatics · Last reviewed 17 August 2026
Where the Corpus Is Used
Inside the Folklore app
Folklore retrieves publications for the genes and variants in a case, incorporates available phenotype context, and ranks the candidate set for professional evidence review. The corpus supplies the publication record and evidence trail; it does not change a variant classification by itself.
Public semantic search
Anyone can search the public Literature Corpus without a Folklore account. The interface supports natural-language semantic discovery and exact PMID, DOI, or PMCID lookup without using patient context.
How the Corpus Is Built
The collection workflow separates source observations from canonical scientific works. This makes it possible to update a source record without silently duplicating the publication or losing where the information came from.
Source intake
PubMed is the primary source. The annual baseline and subsequent update files provide new, revised, and deleted citations. Additional approved sources may contribute records under the same governance rules.
Scope and relevance screening
Records are screened for genetics relevance, eligible publication types, and the documented date scope. Deletion and retraction status is carried forward rather than hidden.
Record normalization
Titles, abstracts, authors, journal details, publication dates, source links, and stable identifiers are normalized into a consistent publication record.
Genetics enrichment
Publications are enriched with validated gene, variant, phenotype, HPO, and OMIM concepts. Controlled validation and confidence requirements limit unsupported mentions.
Identity reconciliation
Normalized PMID, DOI, and PMCID identifiers are used to represent one scientific work once, while preserving the observations and provenance supplied by each contributing source.
Quality-controlled release
A candidate corpus is checked for identity conflicts, duplicate or incomplete records, retractions, deletions, and referential integrity before it replaces the current searchable release.
How Public Semantic Search Works
Public search combines semantic discovery with deterministic identifier lookup. It searches the meaning expressed in publication titles and abstracts while preserving exact identifiers as authoritative anchors.
Meaning, not only keywords
Natural-language queries are matched by meaning across publication titles and abstracts, so related wording can be discovered without requiring an exact phrase.
Exact identifiers remain exact
PMID, DOI, and PMCID queries act as deterministic anchors. Identifier lookup is not replaced by semantic similarity.
Genetics concepts are searchable
Gene symbols, variant notation, phenotype names, HPO terms, and OMIM identifiers can be used to reach relevant publications and related work.
Results retain provenance
Search results are returned from the canonical corpus with publication identity and source links. Retracted works are excluded from discovery results.
Broad semantic queries are intended for discovery. Exact identifier searches are the appropriate route when a known publication must be retrieved deterministically. Results can be ordered by relevance or publication date.
Coverage and Updates
The primary PubMed intake follows the official NLM annual baseline and update-file model. The present genetics-focused scope includes eligible records from 1990 onward. Updates may add citations, revise metadata, or remove deleted records.
The corpus indexes bibliographic metadata, titles, abstracts, identifiers, and extracted genetics concepts when available. It is not a full-text archive and does not imply access to paywalled articles, supplements, or every publication relevant to a clinical question. Current aggregate counts and contributing sources are shown on the live Literature Corpus page.
Interpretation Boundary
Corpus inclusion, semantic similarity, ranking, and extracted concepts are discovery and triage aids. They do not establish causality, validate a functional assay, determine applicability to a patient, provide a diagnosis or treatment recommendation, or independently establish or change an ACMG/AMP classification. Review the source publication before using it as clinical evidence.
Related Resources