Folklore

Literature Corpus Collection and Search Methodology

The Folklore Literature Corpus is a governed, source-traceable database of genetics publications. It is built to support reproducible literature discovery while keeping publication identity, provenance, and known record status visible.

This page describes the public methodology and quality controls. It intentionally focuses on corpus behavior, coverage, and interpretation boundaries rather than proprietary infrastructure or model implementation.

Published and reviewed by Helena Bioinformatics · Last reviewed 17 August 2026

Where the Corpus Is Used

Inside the Folklore app

Folklore retrieves publications for the genes and variants in a case, incorporates available phenotype context, and ranks the candidate set for professional evidence review. The corpus supplies the publication record and evidence trail; it does not change a variant classification by itself.

Public semantic search

Anyone can search the public Literature Corpus without a Folklore account. The interface supports natural-language semantic discovery and exact PMID, DOI, or PMCID lookup without using patient context.

How the Corpus Is Built

The collection workflow separates source observations from canonical scientific works. This makes it possible to update a source record without silently duplicating the publication or losing where the information came from.

1

Source intake

PubMed is the primary source. The annual baseline and subsequent update files provide new, revised, and deleted citations. Additional approved sources may contribute records under the same governance rules.

2

Scope and relevance screening

Records are screened for genetics relevance, eligible publication types, and the documented date scope. Deletion and retraction status is carried forward rather than hidden.

3

Record normalization

Titles, abstracts, authors, journal details, publication dates, source links, and stable identifiers are normalized into a consistent publication record.

4

Genetics enrichment

Publications are enriched with validated gene, variant, phenotype, HPO, and OMIM concepts. Controlled validation and confidence requirements limit unsupported mentions.

5

Identity reconciliation

Normalized PMID, DOI, and PMCID identifiers are used to represent one scientific work once, while preserving the observations and provenance supplied by each contributing source.

6

Quality-controlled release

A candidate corpus is checked for identity conflicts, duplicate or incomplete records, retractions, deletions, and referential integrity before it replaces the current searchable release.

How Public Semantic Search Works

Public search combines semantic discovery with deterministic identifier lookup. It searches the meaning expressed in publication titles and abstracts while preserving exact identifiers as authoritative anchors.

Meaning, not only keywords

Natural-language queries are matched by meaning across publication titles and abstracts, so related wording can be discovered without requiring an exact phrase.

Exact identifiers remain exact

PMID, DOI, and PMCID queries act as deterministic anchors. Identifier lookup is not replaced by semantic similarity.

Genetics concepts are searchable

Gene symbols, variant notation, phenotype names, HPO terms, and OMIM identifiers can be used to reach relevant publications and related work.

Results retain provenance

Search results are returned from the canonical corpus with publication identity and source links. Retracted works are excluded from discovery results.

Broad semantic queries are intended for discovery. Exact identifier searches are the appropriate route when a known publication must be retrieved deterministically. Results can be ordered by relevance or publication date.

Coverage and Updates

The primary PubMed intake follows the official NLM annual baseline and update-file model. The present genetics-focused scope includes eligible records from 1990 onward. Updates may add citations, revise metadata, or remove deleted records.

The corpus indexes bibliographic metadata, titles, abstracts, identifiers, and extracted genetics concepts when available. It is not a full-text archive and does not imply access to paywalled articles, supplements, or every publication relevant to a clinical question. Current aggregate counts and contributing sources are shown on the live Literature Corpus page.

Interpretation Boundary

Corpus inclusion, semantic similarity, ranking, and extracted concepts are discovery and triage aids. They do not establish causality, validate a functional assay, determine applicability to a patient, provide a diagnosis or treatment recommendation, or independently establish or change an ACMG/AMP classification. Review the source publication before using it as clinical evidence.

Related Resources