# Folklore — full public content corpus

---



---

> Consolidated machine-readable text for every canonical URL declared in https://folklore.helena.bio/sitemap.xml.

---

> Generated from public server-rendered HTML on 2026-08-08.

---

> Canonical HTML pages and their cited sources remain authoritative.

---



---

Pages included: 108

---

## Page: Folklore | Public Genomic Search

Source: https://folklore.helena.bio
Canonical: https://folklore.helena.bio
Description: Search a germline variant and follow its annotation, classification evidence and source record in Folklore.

Turn genomic questions
## into clear evidence.
Search one germline variant and follow its annotation, classification evidence and source record in one place.
Examples
Coordinates, HGVS, SPDI and rsID
Germline · public search
### Start with one variant
Enter the chromosome, position, reference and alternate allele in one line.
### Resolve the record
Folklore normalises the locus and resolves its current GRCh38 annotation.
### Follow the evidence
Classification, sources and provenance stay together from result to record.

---

## Page: About Folklore | Clinical Variant Interpretation Platform

Source: https://folklore.helena.bio/about
Canonical: https://folklore.helena.bio/about
Description: Explore Folklore, Helena Bioinformatics’ evidence-traceable platform for variant classification, phenotype matching, literature evidence, screening and clinical reporting.

## A variant classification platform for clinical genetics labs.
Clinical genomic interpretation
Folklore classifies and prioritises a whole-genome VCF in under 10 minutes. Nuclear variants, mtDNA, SV/CNV, phenotype, inheritance, screening and literature come together in one case view, with an AI-native assistant available throughout the workflow.
Methodology
Under 10 min
WGS classification
and prioritisation
8 modules
One integrated case record
AI-native
Full case context
Brochure
Workflow
### From VCF to a signed, traceable report.
1
#### Ingest
VCF from WGS or WES enters with its sample and quality context.
2
#### Annotate
45 curated databases attach frequency, conservation, and predictions.
3
#### Classify
Deterministic ACMG/AMP, MMDWG and Riggs rules assign a class.
4
#### Prioritise
Phenotype, screening and cohort signals order the gene list.
5
#### Review
A geneticist inspects fired criteria and evidence, then signs.
6
#### Report
A tiered, traceable report carries every result to its source.
See the full pipeline
Modules
### One story. Eight ways to read it.
Variant Analysis & Classification
Built for speed
### A whole genome, ready for review in under 10 minutes.
Locally indexed reference data, parallel processing and framework-specific classifiers keep annotation, classification and patient-specific prioritisation inside one pipeline.
Local Indices
Reference data is available directly inside the analytical pipeline.
Parallel processing
Independent analytical stages run concurrently.
Case Output
Classification and patient-specific prioritisation arrive in one review workspace.
AI-native interpretation
### Ask the whole case.
The assistant is available throughout Folklore and works from the context already assembled in the case: classifications, phenotype, inheritance, quality data and linked literature.
Could you please explain the findings and how they relate to the patient's phenotype?
The most significant finding is a Pathogenic variant in the AP5Z1 gene.
Pathogenic c.1033C>T (p.Arg345Ter)
Biallelic pathogenic variants in AP5Z1 are associated with hereditary spastic paraplegia type 48 (SPG48), an autosomal-recessive, often complicated form of HSP characterised by progressive lower-limb spasticity and weakness.
The geneticist reviews the interpretation and signs the report. Clinical infrastructure
Built for routine
### laboratory work.
Patient data remains in the EU. Access is role-based, data is encrypted in transit and at rest, and every classification records the evidence and classifier version used.
Data and privacy
EU data residency
Helsinki, Finland. Patient data stays in the EU.
Versioned evidence
Each result records its classifier version and supporting evidence.
Public methodology
Classification criteria, thresholds and decision rules are documented openly.
Qualified sign-off
Every final report remains under specialist review.
EU data residency
Encrypted
Role-based access
Decision support, not a diagnostic device
Not a CE-marked IVD SCALABLE INTERPRETATION
### Make interpretation as scalable as sequencing.
Bring panels, exomes and whole genomes into one fast, integrated and case-aware workflow, from VCF to specialist review.
Read the Methodology

### Links and cited sources

- [Methodology](https://folklore.helena.bio/methodology)
- [Brochure](https://folklore.helena.bio/brochure)
- [See the full pipeline](https://folklore.helena.bio/how-it-works)
- [Variant Analysis & Classification](https://folklore.helena.bio/platform/variant-classification)
- [Data and privacy](https://folklore.helena.bio/docs/data-and-privacy)

---

## Page: Folklore in 2 Minutes | Product Overview

Source: https://folklore.helena.bio/video
Canonical: https://folklore.helena.bio/video
Description: This overview follows one whole-genome case through VCF intake, variant classification, phenotype matching, structural-variant review and report generation.

Product overview
## Folklore in 2 minutes
This overview follows one whole-genome case through VCF intake, variant classification, phenotype matching, structural-variant review and report generation.
Your browser does not support HTML5 video.
2:27 / Product overview
Read the methodology

### Links and cited sources

- [Read the methodology](https://folklore.helena.bio/methodology)

---

## Page: Folklore Brochure | Clinical Genomic Variant Interpretation

Source: https://folklore.helena.bio/brochure
Canonical: https://folklore.helena.bio/brochure
Description: Read the 22-page Folklore brochure covering deterministic ACMG/AMP classification, mitochondrial DNA, structural variants, phenotype matching, clinical screening, cohort and trio analysis, literature evidence, reporting, EU data residency, validation and regulatory status.

Product brochure
## Reading the Genome
A 22-page account of Folklore, its classification methods, clinical workflow, evidence model and validation work.
01 / 22
### Inside the brochure
- 01 The inherited story
- 02 The platform
- 03 Nuclear variant interpretation under ACMG/AMP
- 04 Evidence and variant review
- 05 Mitochondrial DNA under MMDWG
- 06 Structural variants under ClinGen/ACMG Riggs
- 07 Phenotype matching
- 08 Clinical screening
- 09 Cohort analysis
- 10 Family and trio analysis
- 11 Segregation and its evidentiary ceiling
- 12 Literature evidence
- 13 AI-assisted interpretation and reporting
- 14 Genome browser
- 15 Review board and audit trail
- 16 Process from VCF to clinical report
- 17 Technical specification
- 18 Audience and deployment model
- 19 Validation against an external laboratory
- 20 Regulatory status
Read the full text version +
HELENA BIOINFORMATICS RUO · MMXXVI READING THE GENOME The genome is an inherited story. Folklore reads it. SOFIA · BULGARIA A CLINICAL GENOMICS PLATFORM FOR VARIANT INTERPRETATION VOL. I FOLKLORE AN INHERITED STORY Folklore is knowledge passed from generation to generation. Never formally written, yet carrying the identity of a whole nation. The genome is exactly that kind of inherited memory, a story recopied at every generation, with an occasional new addition. Every variant is a line added along the line of descent, a heritage carried through time. The folklorist does not invent the tale. Only reads it. Folklorists read, interpret and systematise them, gathering scattered oral knowledge and drawing out meaning, structure and significance. That is exactly what Folklore does: reading the genome and interpreting the results. The platform takes called variant data, runs it through a deterministic interpreter backed by dozens of curated evidence sources, and returns every variant sorted and classified: by pathogenicity across the five ACMG classes, and by closeness to the patient’s phenotype. We do not invent the story. We read it. FOLKLORE 002 FOLKLORE THE PLATFORM One story. Eight ways to read it. Each module is one reading of the genome, and together they turn a raw VCF into a prioritised, review-ready result. Variant classification V 3.39.1 Millions of variants → five ACMG classes, ranked by priority. Mitochondrial DNA Rule-based mtDNA under the MMDWG 2020 framework. Structural variants · SV / CNV Deletions and duplications, Riggs 2020, five tiers. Phenotype matching HPO terms → clinical priority, five tiers. Clinical screening Review order, not pathogenicity, Tier 1–4. Cohort analysis Groups of samples → burden, enrichment, candidates. Family & trio analysis Child–mother–father → the origin of a variant. Literature evidence Variant → relevant, anchored publications. C LINGEN MT-VCEP R IGGS 2020 H PO · 0–100 7 COMP · 9 BOOSTS F ISHER · SKAT-O T RIO · DE NOVO P UBMED-ANCHORED The modules sit within a navigable genome browser, an on-premise AI assistant, a review board, and a tiered clinical report. Pathogenicity is decided by the deterministic classifier, not by AI. Folklore is a decision-support tool for the qualified geneticist. (Research Use Only.) I · THE PLATFORM 003 FOLKLORE INTERPRETATION · THE ENGINE Nuclear variants, under ACMG/AMP. Every variant is annotated through Ensembl VEP, then classified against the 2015 ACMG/AMP guidelines as deterministic, rule-based logic. Folklore implements the ACMG/AMP 2015 framework (Richards et al.) through the Tavtigian Bayesian point system, calibrated to the ClinGen Sequence Variant Interpretation recommendations. Of the 28 evidence criteria, 19 are computed automatically against a curated reference layer of 45 databases: gnomAD, ClinVar, dbNSFP, SpliceAI, AlphaMissense, HPO, UniProt, and more, billions of records held locally on EU infrastructure. The remaining 9, which need segregation, functional, or de novo evidence, stay with the reviewing geneticist. Each applicable criterion carries a weight and the combined score maps onto a single verdict. Every classification is recorded with its evidence and version, so the result is reproducible, auditable and resolves to one of five classes: PATHOGENIC LIKELY PATHOGENIC VUS LIKELY BENIGN BENIGN Every classification traces back to the rule and the source behind it. II · INTERPRETATION 004 FOLKLORE WHAT YOU SEE Every result carries its evidence. The variants arrange by gene and by ACMG priority, Pathogenic through Benign. Along the top, five class cards hold the live totals for each class and filter the list beneath them. A gene search narrows it to one. 3 7 214 1.1k 18k PATHOGENIC LIKELY PATH. VUS LIKELY BENIGN BENIGN #1 BRCA1 c.5266dupC · p.Gln1756fs #2 MYH7 c.1988G>A · p.Arg663His #3 TTN c.21A>G · p.Ser7= PATH DE NOVO? CURATED LIK. PATH VUS Each variant carries its whole case. The ACMG criteria that fired, shown as colourcoded evidence. ClinVar significance and review stars, with any ClinGen expertpanel assertion. In-silico predictions, gnomAD frequency, conservation and constraint. The sequencing quality behind the call. Every value traces to the record that produced it. The record is always one step beneath the result. III · WHAT YOU SEE 005 FOLKLORE MITOCHONDRIAL DNA Mitochondrial variants, under MMDWG. Mitochondrial variants are classified under MMDWG 2020 (McCormick), an independent module running alongside nuclear ACMG/AMP. Every variant in the output carries an explicit framework-provenance label, so the geneticist always knows which framework produced the class. Of the 28 nuclear criteria, seven are excluded for the mitochondrial genome with verbatim rationale. The rest are respecified with mtDNA-specific thresholds: 11 automated, 10 curated with strength tiers. Combining rules from Richards 2015 are preserved unchanged. A curation panel exposes the ClinGen mtDNA VCEP criteria for the reviewer, alongside heteroplasmy assessment and haplogroup-aware interpretation (APOGEE2, MitoTIP, MITOMAP). The classification follows the genome’s biology: PP1 cannot contribute when a variant is homoplasmic across all maternal members (VCEP §5.2.7). The panel then shows the pipeline class and the curated class side by side: P IPELINE CLASS VUS W ITH CURATION → PM2 · PP3 LIKELY PATH. + PS3 functional · heteroplasmy 78% Two frameworks, each named on every variant. IV · MITOCHONDRIAL DNA 006 FOLKLORE STRUCTURAL VARIANTS Structural variants, under ClinGen/ACMG. Copy-number variants are classified under the ClinGen/ACMG 2020 semiquantitative point system (Riggs 2020): losses on Table 1, gains on Table 2, as two distinct point systems. Dosage sensitivity comes from pHaplo / pTriplo (Collins 2022) and overlap with established haploinsufficient (HI) and triplosensitive (TS) regions from the ClinGen dosage map. Population frequency comes from gnomAD-SV. Points sum through a single five-tier ladder into the same P / LP / VUS / LB / B labels. chr16:29,580,000–30,180,000 del · 600 kb S V SPAN GRCh38 DELETION G ENES C LINGEN · HI A structural variant renders as a span (deletion red, duplication blue) over the genes it crosses. A haploinsufficient region overlap and the count of genes and exons feed the Riggs score. 16p13.11 deletion · 600 kb Pathogenic The CNV classification (Riggs) panel shows the total score and tier with each criterion met or not met: established-HI overlap, gene count, dosage, population frequency, alongside copy state, breakpoint precision and paired/split-read support. Its own tab, its own metric, landing on the same five-class scale. V · STRUCTURAL VARIANTS 007 FOLKLORE PHENOTYPE MATCHING Variants, ranked by phenotype. ACMG classification tells you how pathogenic a variant is. It does not tell you which variant explains your patient. Phenotype matching answers that second question. The patient’s HPO terms are correlated with gene–disease associations by semantic similarity, scored 0–100 for each term, and folded into a clinical priority score. Genes are ranked and placed into five tiers (Tier 1, Tier 2, Incidental Findings, Tier 3, Tier 4). Each variant shows its strongest HPO-term matches. THE CORRECT BEHAVIOUR A VUS with strong phenotypic relevance ranks above a Pathogenic variant for an unrelated condition. Relevance to this patient, not pathogenicity in the abstract. The HPO terms entered at intake become filters over the gene list. A selected phenotype collapses the list to the genes that match it. #1 SCN1A HPO match 0.92 · seizures, DEE TIER 1 #2 KCNQ2 HPO match 0.71 TIER 2 #3 TTN HPO match 0.08 · unrelated TIER 4 The story is read for the patient in front of you, not for a textbook. VI · PHENOTYPE 008 FOLKLORE CLINICAL SCREENING Clinical screening, tier by tier. Screening sets the review order from the patient’s full context: age, sex, ethnicity, family history, sample structure, and the chosen panels. A seven-component score (constraint, deleteriousness, phenotype, dosage, consequence, compound-het, age relevance) is adjusted by up to nine clinical boosts (ACMG strength, ethnicity, family history, de novo, and more), then placed into four tiers. Each gene carries a clinical-actionability signal: immediate, monitoring, or future. Screening modes tune the run for neonatal, paediatric, adult diagnostic, or carrier testing. T1 RYR1 score 0.86 · malignant hyperthermia IMMEDIATE T1 BRCA2 score 0.79 MONITORING T2 MYBPC3 score 0.55 FUTURE Screening sets the order of reading. Pathogenicity stays with ACMG. VII · SCREENING 009 FOLKLORE COHORT ANALYSIS Cohort analysis, across samples. A study gathers classified samples into one matrix and runs population-level statistics, from two samples upward. Gene-burden testing reports Fisher exact p, FDR q and odds ratios (with CMC and SKAT-O), read off a volcano plot. A loss-of-function pass adds pLI and LOEUF. Pathway enrichment and cross-sample compound-heterozygote detection round out the signal. A candidate-gene nomination fuses seven independent lines of evidence (burden, pLoF, disease, constraint, pathway, GWAS, compound-het) into a ranked list and an evidence-matrix heatmap. Where samples disagree on a class, the discordance is surfaced, not hidden. #1 LDLR Fisher p 3.1e-5 · OR 6.4 · 5/7 evidence axes SIGNIFICANT #2 APOB Fisher p 8.0e-4 · OR 3.2 FDR Q<0.05 #3 PCSK9 pathway + GWAS + constraint CANDIDATE The same reading, now across a population, with every disagreement between samples shown. VIII · COHORT 010 FOLKLORE FAMILY & TRIO Inheritance, read from the trio. A trio (proband, mother, father) adds what a single sample can never carry: the parental genotypes. Three inheritance workflows run on them: de novo, compound heterozygous, and segregation. De novo. The proband carries the alternate allele; both parents are reference. A per-member table shows the genotype, depth, quality and class behind each call. A clinical-grade filter scopes the drill. M EMBER GENOTYPE DP · QUAL Proband het (0/1) 42x · 99 Father hom_ref (0/0) 38x · 99 Mother hom_ref (0/0) 45x · 99 Compound het. Two variants in one gene, one inherited from each parent, are phased from the trio genotypes: paternal when father is het and mother reference, maternal for the reverse. Only a trio-phased pair is called in trans. PATERNAL GENE · variant A father het · mother 0/0 IN TRANS ↔ MATERNAL GENE · variant B father 0/0 · mother het Cis-ambiguous pairs are shown as such. The trio phases only what the genotypes support. IX · FAMILY & TRIO 011 FOLKLORE FAMILY & TRIO Segregation, and its ceiling. Segregation testing computes a LOD score for each variant. But a trio (one proband and two parents) has too few informative meioses. The maximum LOD attainable is about 0.30, which sits exactly at the ClinGen SVI 2016 supporting evidence threshold. THE CEILING A trio reaches PP1_Supporting at most, never PP1 Moderate or Strong. The platform states the ceiling and recommends extending the pedigree for stronger evidence, rather than borrowing a strength the data cannot support. Before any of this, cross-sample QC checks that the family is the family: PLINK identity-by-descent relatedness catches sample swaps, unexpected consanguinity, and duplicates. A critical alert holds the line: inheritance findings should not drive clinical interpretation until the pedigree checks out. MAX LOD 0.30 Trio ceiling = ClinGen SVI 2016 supporting threshold. Extend the pedigree for more. CROSS-SAMPLE QC PLINK IBD (PI_HAT): sample-swap, consanguinity and duplicate detection. The limit is named, not worked around. A stated ceiling holds where a borrowed strength would not. X · FAMILY & TRIO 012 FOLKLORE LITERATURE EVIDENCE Literature, anchored to the variant. A local, genetics-filtered PubMed mirror with pre-extracted gene, variant and phenotype mentions turns literature search into a sub-second, variant-anchored query. Publications rank by a six-component model aligned to ACMG evidence categories, each badged by strength: strong, moderate, supporting, weak. An exact-match badge flags a paper that names this variant. A functional badge flags experimental data. The gene ranking fuses clinical priority and literature relevance (60/40), and every hit keeps full PMID / PMC / DOI traceability. EXACT MATCH The paper reports this exact variant, not just the gene. SIX-COMPONENT SCORE Phenotype, publication type, gene focus, functional data, variant, recency. TRACEABLE PMID / PMC / DOI on every citation, evidence you can open. Every citation opens to its source. The folklorist names where the claim came from. XI · LITERATURE 013 FOLKLORE INTERPRETATION & REPORT AI assistant, under the deterministic gate. An on-premise AI clinical assistant, a self-hosted open-weight model on EU infrastructure where no external AI API ever sees patient data, integrates classification, phenotype and literature into a structured clinical narrative. You can ask it questions in plain language: it runs queries against the session database and charts them, searches the literature, and explains which ACMG criteria support a class. The division of labour is the whole point: the AI drafts interpretation text, the deterministic classifier decides pathogenicity. The assistant carries its own caution. The finished output is a single tiered PDF: Tier 1 (actionable) and Tier 2 (potentially actionable) variants, each with a complete evidence chain, ready for the geneticist’s review and sign-off. THE DIVISION OF LABOUR AI writes the sentence; rules assign the class. The geneticist makes the call. Research Use Only. Not for diagnostic use without professional review. CLINICAL REPORT · TIERED · EVIDENCE CHAIN TIER 1 BRCA1 c.5266dupC · Pathogenic · PVS1 PS4 PM2 · ClinVar ★★★ TIER 2 MYH7 c.1988G>A · Likely path. · PM1 PM2 PP3 RUO Reviewed & signed: ________________ · qualified clinical geneticist A draft is a draft until a human signs it. Folklore never forgets which is which. XII · INTERPRETATION & REPORT 014 FOLKLORE GENOME BROWSER The genome browser. A navigable per-session window over the patient’s GRCh38: variants as ACMGcoloured lollipops sized by clinical salience, structural variants as spans, genes as MANE models with strand direction, ClinGen dosage regions, alongside evidence lanes for BayesDel, SpliceAI, pext, conservation, ClinVar and constraint. Zoom, pan, jump to a locus, an HGVS change, or a gene. The browser adds no reading of its own. It shows what was already computed. chr17:43,044,295–43,170,245 MANE Select GRCh38 V ARIANTS S TRUCTURAL DELETION G ENES · MANE C LINGEN · D OSAGE Variants as lollipops by ACMG class, height tracks salience. Structural variants as spans, genes with direction along MANE. A graph mode reads insertions and deletions as divergence. Where phase is not measured, it says “inferred”. BRCA1 · c.5266dupC (p.Gln1756fs) PVS1 PS4 Pathogenic PM2 Click a feature for the full readout: fired ACMG criteria as badges, BayesDel and SpliceAI, the pext level, MANE, and the ClinVar stars. Only what has a value is shown. One coordinate frame holds every finding, each shown where the analysis placed it. XIII · GENOME BROWSER 015 FOLKLORE THE REVIEW BOARD The review board. A case carries the exact classifier version it was run with, so every result stays tied to the rules that produced it. The review board is where a qualified geneticist takes over from the pipeline. A reviewer can reclassify any variant, with a written justification that is attributed and revertable, or leave threaded notes for a colleague. The reviewer’s call always overrides the pipeline’s. A case-level outcome (solved, likely solved, unsolved) is tracked with its full change history. Every reclassification is signed, sourced, and reversible. The audit trail is the product. XIV · THE REVIEW BOARD 016 FOLKLORE PROCESS From VCF to clinical report. 1 Ingest & build detection 2 Quality control 3 Annotation 4 Classification 5 Prioritisation 6 Report G RCH37 / 38 VCF in; automatic GRCh37/38 detection and liftover to GRCh38. C LINVAR-SAFE Quality filters; ClinVar-pathogenic variants protected from being dropped. 4 5 DATABASES Against the reference layer: 45 databases, billions of rows. A CMG · MMDWG · RIGGS ACMG/AMP for nuclear variants; MMDWG for mtDNA; Riggs for CNVs. 3 –20 Phenotype matching and screening; down to 3–20 candidates. R UO A clinical-interpretation draft for review by a geneticist. Every stage traceable, every stage reproducible. XV · PROCESS 017 FOLKLORE SPECIFICATION The specification. Genome builds GRCh38 native; GRCh37 auto-lifted Variant classes SNV, indel, mtDNA, SV / CNV (Riggs 2020), all active Reference layer 45 databases; SpliceAI ~3.4B rows, gnomAD ~909M, dbNSFP ~80.7M, ClinVar ~4.1M Classification ACMG/AMP 2015, deterministic; MMDWG (ClinGen mt-VCEP); Riggs 2020; 5 classes Genome browser Navigable GRCh38 view; variants, SVs, genes, dosage; evidence drawer; graph mode Speed Analysis pipeline ~7–12 min (WGS) EU data residency Helsinki, Finland; data never leaves the EU; GDPR Art. 9 Audit & security SHA-256 append-only chain; encryption in transit and at rest; RBAC Artificial intelligence Self-hosted open-weight model (Qwen) on EU infra; no external AI API Status Research Use Only. A decision-support tool, not a diagnostic device EU DATA RESIDENCY GDPR ART. 9 SHA-256 AUDIT RBAC RUO One reading, traceable down to the last rule. XVI · SPECIFICATION 018 FOLKLORE FOR WHOM The tale. Every capability carries both a clinical proof and a signal of scale. THE READING THE MOAT Clinical-grade, traceable, reproducible. Category. A European clinical-genomics Five classes and an audit trail on every software pure-player, no wet lab, no action. A decision-support tool that leaves capex. Moat. Curation, not code: a the last word to the qualified geneticist. reference layer (45 databases, billions of R are disease N ewborn screening C arrier screening rows) and a discordance-gated classifier, built over years. Defensibility. Live rulebased mtDNA and live Riggs CNV, capabilities others treat as afterthoughts. Sovereignty. Self-hosted EU AI where no external API sees patient data; GDPR by design. E uropean software R eference layer S elf-hosted AI E U data The tool reads the story. The last word is the human’s. XVII · FOR WHOM 019 FOLKLORE VALIDATION Validation, against an external lab. 0 68.4% 3 BENIGN ↔ PATHOGENIC INVERSIONS (OPPOSITE POLES), ACROSS ALL THREE COHORTS EXACT FIVE-CLASS AGREEMENT · COHORT 2, CLASSIFIER V3.28.0, 19 EVALUABLE REAL-WORLD COHORTS · 20 / 20 / 57 CASES, ACROSS CLASSIFIER GENERATIONS Five-class agreement sits in the 60–75% range that the literature reports between laboratories (Amendola 2016 ~66%, Harrison 2017 72–76%). Where Folklore differs, it differs by one step and in one direction: it returns VUS rather than asserting a pathogenicity it cannot fully evidence. No benign becomes pathogenic, nor the reverse. WHAT WE HAVE VALIDATED WHAT WE HAVE NOT Concordance against an external clinical Large-scale prospective outcomes, multi- laboratory across three real-world laboratory benchmarking, SV/CNV cohorts. Zero opposite-pole inversions. A phenotype correlation. The claim does not conservative skew toward VUS. exceed the evidence. External clinical laboratory as a peer comparator, not a reference standard. An ongoing, iterative real-world audit. XVIII · VALIDATION 020 FOLKLORE REGULATORY STATUS Regulatory status. Folklore is a Research Use Only tool. It is a clinical decision-support system, not a diagnostic device, and it is not a CE-marked IVD. Every classification Folklore produces requires review and confirmation by a qualified clinical geneticist. The platform drafts and orders evidence. The clinical decision belongs to the reviewing professional. Folklore does not replace clinical judgement, established laboratory procedure, or the reporting standards of the laboratory that uses it. Results are intended to support qualified interpretation, not to stand alone. Research Use Only Not a diagnostic device Not a CE-marked IVD Geneticist sign-off The tool reads the evidence. The clinical decision is the geneticist’s. XIX · REGULATORY STATUS 021 HELENA BIOINFORMATICS EOOD RUO · MMXXVI The genome is an inherited story. Folklore reads it. A CONTACT@HELENA.BIO HELENA BIOINFORMATICS PLATFORM · SOFIA, BULGARIA HELENA.BIO

### Links and cited sources

- [folklore.helena.bio](https://folklore.helena.bio/brochure/folklore-brochure.pdf)

---

## Page: Variant Interpretation Software for Clinical Geneticists | Folklore

Source: https://folklore.helena.bio/for-geneticists
Canonical: https://folklore.helena.bio/for-geneticists
Description: Evidence gathering for VUS review with Bayesian ACMG classification, HPO phenotype matching, AI-assisted summaries, and report generation for genetics laboratories.

## Prepare the Evidence Before Clinical Review
Your browser does not support the video tag.
Interpreting genetic variants requires years of specialized training, deep clinical knowledge, and the kind of judgment that cannot be automated. What can be automated is the hours spent cross-referencing databases, searching literature, and compiling evidence before your interpretation begins.
Folklore prepares the evidence for review; the geneticist retains the clinical judgment.
### The Real Bottleneck
Manual database searches, literature review, and evidence compilation consume much of the time required for variant interpretation.
Without Folklore
Per case, typical workflow
ClinVar lookup per variant ~2 hours PubMed literature search ~3 hours gnomAD frequency checks ~1 hour ACMG criteria mapping ~2 hours Report compilation ~1 hour Total evidence gathering 5 - 10 days
With Folklore
Same case, same rigor
Automated evidence gathering ~30 minutes Your clinical review 30 - 60 minutes Your interpretation Your expertise Total time to report Under 2 hours
Same evidence. Same standards. Your clinical judgment throughout.
### Clear Division of Responsibility
Folklore's scope ends at evidence preparation and rule application. Clinical interpretation remains with the geneticist.
What Folklore Does
Automated evidence gathering
Cross-references ClinVar, gnomAD, dbNSFP, and ClinGen in seconds
Searches millions of PubMed publications for relevant literature
Maps ACMG/AMP criteria against variant evidence systematically
Matches patient phenotype (HPO) against gene-disease profiles
Formats structured reports with full evidence attribution
Completes evidence preparation in minutes, not days
What You Do
Clinical expertise that cannot be automated
Applies clinical judgment that no algorithm can replicate
Integrates patient history, family context, and clinical presentation
Evaluates edge cases where guidelines require expert interpretation
Makes the final classification decision on every variant
Communicates findings to patients and referring physicians
Determines clinical actionability and management recommendations
### Your Expertise, Amplified
Folklore reduces the time spent retrieving and organizing evidence before clinical review.
#### Deeper Evidence Access
Each variant is annotated with 60+ data points from established databases. Literature search covers millions of publications with pre-extracted gene and variant mentions.
#### Consistent, Reproducible Workflow
The pipeline follows the configured search steps for each case and records the resulting evidence.
#### More Time for Complex Cases
Reducing routine evidence gathering leaves more review time for rare variants, conflicting evidence, novel gene-disease associations, and difficult phenotypes.
#### Audit Trail
The report records the databases queried, criteria applied, and publications cited so the reviewer can trace each evidence source.
### What Always Stays in Your Hands
The following decisions remain with a qualified clinical geneticist.
Folklore will never:
Make a diagnostic decision
Override your clinical judgment
Classify a variant without your confirmation
Replace the context only you have about your patient
Communicate results to patients or clinicians
Determine treatment or management plans
Folklore prepares the evidence. The geneticist makes the clinical decision.
### Review a Case Workflow
A demo shows the evidence and report that reach the geneticist for review.
Contact Us Rare Disease Newborn Screening Carrier Screening How It Works Methodology

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [Rare Disease](https://folklore.helena.bio/use-cases/rare-disease)
- [Newborn Screening](https://folklore.helena.bio/use-cases/newborn-screening)
- [Carrier Screening](https://folklore.helena.bio/use-cases/carrier-screening)
- [How It Works](https://folklore.helena.bio/how-it-works)
- [Methodology](https://folklore.helena.bio/methodology)

---

## Page: How Folklore Processes and Interprets a Genome

Source: https://folklore.helena.bio/how-it-works
Canonical: https://folklore.helena.bio/how-it-works
Description: The Folklore workflow for VCF quality control, annotation, ACMG classification, phenotype matching, literature retrieval, screening, and report preparation.

## How Folklore Processes a Genomic Case
Folklore checks the VCF, annotates each variant, applies the deterministic classification rules, matches the phenotype, retrieves literature, and prepares the evidence for clinical review. Model-generated text remains outside the class calculation.
1 Quality Control
2 Annotation
3 Classification
4 Phenotype
5 Literature
6 Screening
7 Interpretation
Under 15 minutes
Full genome processing and clinical interpretation
### The Analysis Pipeline
The core path covers quality control, annotation, classification, phenotype matching, literature retrieval, screening, and report preparation. Optional modules add family, cohort, and mitochondrial analysis.
1
#### VCF Processing & Quality Control
Standard VCF parsing accepts files from any sequencing platform, whole genome, whole exome, or targeted panels. Quality metrics are assessed per variant, applying configurable filters for read depth, genotype quality, and allelic balance.
Variants with documented clinical significance in ClinVar are preserved regardless of quality score. The quality filter therefore does not remove them from clinical review.
VCF standard format Quality filtering ClinVar protection Maximum sensitivity
Output: Quality-filtered variant set with clinically significant variants preserved
2
#### Variant Annotation
Each variant is annotated through Ensembl VEP (Variant Effect Predictor) for consequence prediction, protein impact, and functional domain mapping. Annotation runs in parallel across the variant set.
Multi-source database enrichment adds population frequencies from gnomAD (global and population-specific allele frequencies), clinical significance from ClinVar, functional impact predictions from 12+ computational tools including SIFT, PolyPhen-2, CADD, REVEL, AlphaMissense, DANN, MetaSVM, GERP++, PhyloP, and PhastCons, gene constraint metrics (pLI, LOEUF, o/e loss-of-function), and gene-disease associations from ClinGen.
Ensembl VEP gnomAD ClinVar dbNSFP ClinGen 12+ predictors 60+ annotations per variant
Output: Fully annotated variants with population, functional, conservation, and clinical data
3
#### ACMG/AMP Classification
Variant classification follows the 2015 ACMG/AMP guidelines (Richards et al., Genetics in Medicine), the international standard for clinical variant interpretation. All 28 evidence criteria are systematically evaluated: PVS1, PS1–4, PM1–6, PP1–5, BA1, BS1–4, and BP1–7.
Classification is strictly rule-based. No AI model determines variant pathogenicity. Each variant receives one of five standard classifications, Pathogenic, Likely Pathogenic, Variant of Uncertain Significance (VUS), Likely Benign, or Benign, with an explicit listing of every criterion applied.
ACMG/AMP 2015 guidelines Rule-based 28 evidence criteria 5-tier classification Explicit criteria listing
Output: Classified variants with the applied ACMG criteria and audit record
4
#### Phenotype-Genotype Correlation
Patient phenotype, described using Human Phenotype Ontology (HPO) terms, is compared with the known phenotypic profiles of genes carrying candidate variants. Semantic similarity uses the HPO hierarchy, term specificity, and information content in addition to exact matches.
Each gene receives a normalized relevance score (0–100) with tiered clinical classification. A Pathogenic BRCA1 variant is not flagged as clinically relevant when the patient was referred for epilepsy. Phenotype matching connects technical classification to clinical relevance for the specific patient.
HPO ontology Semantic similarity Information content Normalized scoring Tiered relevance
Output: Ranked gene list prioritized by phenotype match strength for this patient
5
#### Literature Evidence
A locally maintained, genetics-filtered database supports sub-second queries across millions of PubMed publications. Gene mentions, variant mentions, and phenotype associations are extracted during ingestion and used to retrieve evidence for a case.
Multi-component relevance scoring ranks publications for the case. Each citation includes its PubMed identifier (PMID), DOI, and extracted evidence context so the source publication can be checked.
Local PubMed database Pre-extracted entities Relevance scoring PMID/DOI tracking Sub-second queries
Output: Ranked literature evidence with traceable citations per gene and variant
6
#### Clinical Screening & Prioritization
After classification, annotation, phenotype matching, and literature review, a multi-dimensional prioritization algorithm ranks variants by overall clinical relevance. Scoring adapts to the clinical context, patient age, sex, family history, and indication for genetic testing.
The system supports multiple screening strategies including neonatal intensive care, pediatric genetics, adult diagnostic workup, proactive screening, and carrier testing. The output is a tiered shortlist: Tier 1 (actionable findings requiring immediate clinical attention), Tier 2 (potentially actionable, warranting further review), with incidental findings identified and flagged separately.
Context-aware scoring Age/sex adaptation Neonatal Pediatric Adult diagnostic Carrier screening Tiered output
Output: Focused shortlist of clinically actionable variants from hundreds of candidates
7
#### Family and Trio Analysis (Optional)
When family members are sequenced alongside the proband, Folklore adds inheritance-aware evidence on top of the upstream classification. Three algorithms run sequentially on pre-classified data: de novo detection with confidence tiers, compound heterozygous phasing with parental origin determination, and segregation scoring per the ClinGen SVI 2021 framework. Sample QC via PLINK identity-by-descent runs first to detect sample-swap, consanguinity, and duplicate-sample issues before any inheritance call.
No variant re-calling. The service consumes the existing classified DuckDB files and joins them on chromosome, position, and allele. A typical WGS trio completes the full inheritance analysis in approximately 30 to 90 seconds, with explicit feasibility flags recording which phases were planned as feasible and why.
Trio, duo, sibling support De novo detection Compound het phasing ClinGen SVI 2021 segregation PLINK sample QC Feasibility flags
Output: Inheritance-annotated variants with de novo, compound het, and segregation evidence
8
#### Cohort Analytics (Optional)
For population-level research, Folklore aggregates classified samples into a cohort matrix with a deduplicated variant catalog and sparse genotype storage. Six statistical analyses share the matrix: gene-level burden testing with Fisher, CMC, and SKAT-O methods plus FDR correction; pathway enrichment with background correction; pLoF analysis; cohort versus gnomAD frequency analysis; GWAS signal replication; and polygenic risk scoring with PGS Catalog weight files.
A weighted candidate gene nomination engine integrates evidence across all six analyses and produces a ranked list with per-component breakdowns and human-readable evidence summaries. Power analysis is reported per gene so non-significant results in underpowered genes are interpretable.
Cohort matrix Burden testing Pathway enrichment pLoF analysis GWAS replication Polygenic scores Candidate ranking
Output: Ranked candidate genes with statistical evidence and per-component scoring
9
#### Mitochondrial DNA Analysis
Mitochondrial DNA biology differs fundamentally from the nuclear genome. Maternal inheritance, heteroplasmy with tissue-specific threshold effects, lack of recombination, and haplogroup structure all change how variants should be classified. Folklore routes mitochondrial variants through a dedicated classifier following the McCormick 2020 specifications produced by the ClinGen Mitochondrial Disease Variant Curation Expert Panel.
Twenty ACMG criteria are applied with mtDNA-specific thresholds and tools including APOGEE2 for protein-coding genes, MitoTIP and HmtVAR for tRNA, MITOMAP and HmtDB for population frequency. Seven criteria are explicitly excluded with verbatim biological rationale. Haplogroup-aware BA1 and NUMT pseudogene detection prevent the most common false-positive patterns.
MMDWG 2020 ClinGen Expert Panel Heteroplasmy aware Haplogroup aware NUMT detection APOGEE2 / MitoTIP / HmtVAR
Output: mtDNA variants classified per the ClinGen Mitochondrial Expert Panel framework
10
#### Model-Assisted Clinical Interpretation
An AI model summarizes upstream classifications, phenotype correlations, literature findings, and screening results in a structured narrative. Step 3 remains the rule-based class calculation; the model prepares its output for clinical review.
The report can include the classification alone or combine classification, phenotype, literature, and screening data when available. PDF and DOCX outputs identify the supporting evidence. AI inference runs on dedicated EU infrastructure, and no data is sent to an external AI service.
Evidence synthesis Adaptive depth PDF/DOCX reports On-premise AI EU data residency
Output: Downloadable clinical interpretation report with structured evidence and recommendations
### Clinical Control Principles
The architecture separates rule-based classification, evidence synthesis, and the geneticist's clinical decision.
#### Rule-Based Classification
ACMG/AMP criteria applied through deterministic rules determine the class. AI supports evidence gathering and presentation outside that calculation.
#### Evidence Record
Classifications link to the applied ACMG criteria, literature references to their PMID, and phenotype scores to their HPO terms.
#### Reproducible Results
The pipeline records the applied criteria and evidence sources for review.
#### Geneticist Authority
Folklore gathers evidence, applies classification rules, and presents findings. A qualified geneticist reviews the evidence and makes the clinical decision.
### See the Pipeline in Action
Follow a case from VCF upload to a report prepared for clinical review.
Contact Us For Geneticists Rare Disease Newborn Screening Carrier Screening Methodology

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)
- [Rare Disease](https://folklore.helena.bio/use-cases/rare-disease)
- [Newborn Screening](https://folklore.helena.bio/use-cases/newborn-screening)
- [Carrier Screening](https://folklore.helena.bio/use-cases/carrier-screening)
- [Methodology](https://folklore.helena.bio/methodology)

---

## Page: ACMG Methodology - Bayesian Point-Based Framework | Folklore

Source: https://folklore.helena.bio/methodology
Canonical: https://folklore.helena.bio/methodology
Description: 20 automated ACMG criteria using Tavtigian Bayesian point system, BayesDel_noAF with ClinGen SVI calibration, SpliceAI splice impact, and ClinVar integration.

## ACMG Methodology
Classification Engine v 3.39.1 | Updated July 2026
This page documents the production processing stages, thresholds, database versions, and classification rules for clinical geneticists, laboratory directors, and accreditation auditors.
Variant classification follows the ACMG/AMP 2015 framework ( Richards et al., Genetics in Medicine, 2015 ) implemented through the Bayesian point-based system ( Tavtigian et al., Hum Mutat. 2018;39(11):1485-1492. PMID: 30311386 ) with BayesDel ClinGen SVI calibrated thresholds ( Pejaver et al., Am J Hum Genet. 2022;109(12):2163-2177. PMID: 36413997 ) and SpliceAI integration aligned to ClinGen SVI 2023 recommendations ( Walker et al., Am J Hum Genet. 2023;110(7):1046-1067. PMID: 37352859 ). An optional ClinGen VCEP gene-specific overlay is available for approximately 50-60 genes. Deterministic rules, outside the machine-learning layer, determine pathogenicity. Versions v3.24 through v3.28 added eligibility guards for computational-only evidence and evidence inconsistent with the gene mechanism (AR-only LoF genes, dual-mechanism bypass, gnomAD LoF tolerance, ClinGen negative evidence).
Contents
01 Pipeline Overview 02 Reference Databases 03 ACMG Classification 04 Automated Criteria 05 Computational Predictors 06 SpliceAI Integration 07 Classification Logic 08 VCEP Gene-Specific Specifications 09 ClinVar Override Logic 10 Manual Review Criteria 11 Quality Filtering 12 Limitations 13 Version History 14 References
### Pipeline Overview
Six processing stages annotate and classify a raw VCF file. A whole genome of approximately 4 million variants runs in under 15 minutes on dedicated hardware.
1
VCF Parsing
~60s
Standard VCF file parsed into columnar in-memory database. Multi-allelic handling, genome build detection (GRCh38 required).
2
Quality Filtering
~5s
Configurable quality, depth, and genotype quality thresholds. ClinVar-listed pathogenic variants are protected from filtering.
3
VEP Annotation
~3-4 min
Ensembl Variant Effect Predictor for consequence, impact, transcript, and protein annotations. Parallel processing across chromosomes.
4
Reference DB Annotation
~5-10s
Population frequencies, clinical significance, functional predictions, gene constraint, phenotype associations, and dosage sensitivity loaded from 16 production classification and reference databases in this core stage. Across all Folklore modules, the platform integrates 45 reference databases and curated sources.
5
ACMG Classification
<1s
SQL-based ACMG/AMP 2015 classification using Bayesian point framework (Tavtigian et al. 2018). 20 automated criteria evaluated with calibrated evidence strength, point-based classification thresholds applied, continuous confidence scores assigned. Optional VCEP gene-specific overlay for ~50-60 genes. Homozygous reference genotypes (hom_ref) from multi-allelic sites excluded from classification.
6
Export
~5s
Gene-level summaries exported for streaming. Classified variants persisted to analytical database for downstream services.
Maximum Sensitivity Approach
Folklore classifies all variants that pass quality filtering. No frequency-based or impact-based pre-filter runs before classification. A common variant (for example, gnomAD allele frequency 40%) remains in the set and can receive a Benign classification through BA1. The geneticist then assesses clinical relevance from the recorded classification and annotation data.
### Reference Databases
All reference data is stored locally on EU-based infrastructure. No variant data is sent to external APIs during processing. Database versions are fixed per deployment and documented here.
gnomAD
v4.1.0 ~759M variants
Population allele frequencies (global and population-specific), allele counts, homozygote counts
Used by: BA1 (allele frequency > 5%), BS1 (elevated frequency), PM2 (absent in controls), BS2 (homozygote count)
Source: gnomad.broadinstitute.org
ClinVar
2025-01 ~4.1M variants
Clinical significance assertions, review star levels, disease associations, submitter information
Used by: PS1 (known pathogenic), PM5 (different pathogenic missense substitution at the same gene and codon, v3.39.1), PP5 (reputable source pathogenic), BP6 (reputable source benign), ClinVar override logic, quality filter rescue
Source: ncbi.nlm.nih.gov/clinvar
dbNSFP
4.9c ~80.6M variant sites
Functional impact predictions from multiple algorithms, conservation scores
Used by: PP3/BP4 primary tool: BayesDel_noAF with ClinGen SVI calibrated thresholds (Pejaver et al. 2022). Display predictors: SIFT, AlphaMissense, MetaSVM, DANN, PhyloP, GERP (available for clinical review, not used in classification logic)
Source: sites.google.com/site/jpopgen/dbNSFP
SpliceAI
Ensembl MANE (Release 113) Precomputed for all coding variants
Splice impact predictions (4 delta scores: acceptor gain, acceptor loss, donor gain, donor loss)
Used by: PP3_splice (max score >= 0.2), BP4 guard (max score < 0.1), BP7 (synonymous splice check)
Source: Illumina / Ensembl
gnomAD Constraint
v4.1.1 ~18.2K genes
Gene-level constraint metrics indicating tolerance to loss-of-function and missense variation
Used by: PVS1 (LoF intolerance), PP2 (missense constraint), BP1 (LoF tolerance)
Source: gnomad.broadinstitute.org
MANE Select Transcripts
v1.4 (GRCh38) 19,354 transcripts (19,288 Select + 66 Plus Clinical)
Canonical transcript definitions with CDS start/end coordinates for each protein-coding gene. MANE Select is the single representative transcript per gene agreed upon by NCBI and Ensembl. MANE Plus Clinical covers additional clinically relevant transcripts for 66 genes.
Used by: PVS1 MANE positional guard (v3.22.0, HELIX-CR-2026-055): variants outside MANE Select CDS do not receive PVS1 or PM4. Exon-level pext aggregation (CR-056): pext scores computed per MANE Select CDS exon.
Source: NCBI / EMBL-EBI / MANE Consortium
gnomAD pext (Proportion Expressed)
v4.1 (GTEx v10, GRCh38) 189,856 exon records for 18,923 genes
Per-exon expression proportion across 49 GTEx v10 tissues. For each base in the coding sequence, pext measures what proportion of the gene total expression passes through that position. Aggregated to exon level as mean_pext (arithmetic mean across tissues) and max_tissue_pext (maximum across 49 tissues). Max tissue pext used for classification to avoid false negatives for tissue-specific genes.
Used by: PVS1 expression-aware guard (v3.23.0, HELIX-CR-2026-056): max_tissue_pext >= 0.9 = PVS1 Very Strong, 0.1-0.9 = PVS1 downgraded to Strong, < 0.1 = PVS1 blocked. PM4 binary guard (blocked at < 0.1). Scientific basis: Cummings 2020 PMID:32461655.
Source: gnomAD / Broad Institute
HPO (Multi-Source Enriched)
v3.12.0 (6 sources) ~927K records, 5,688 genes
Gene-to-phenotype associations aggregated from 6 curated sources: HPO Consortium (5,173 genes), Orphanet disease-to-HPO (3,176 genes with frequency data), DECIPHER G2P 7 clinical panels (2,125 genes), Monarch Initiative (4,791 genes), ClinVar-MedGen P/LP chain mapping (5,258 genes), and manual clinical curation. Source priority ordering with confidence-based filtering.
Used by: PP4 (patient phenotype matching), Phenotype Matching Service (semantic similarity scoring), Screening Service (gene clinical breadth proxy)
Source: HPO Consortium + Orphanet + G2P + Monarch + ClinVar-MedGen
ClinGen
Latest release ~1.6K genes
Dosage sensitivity scores (haploinsufficiency, triplosensitivity)
Used by: PVS1 (haploinsufficiency_score = 3 as constraint gate fallback for X-linked genes, v3.11.5), BS1 (inheritance-aware frequency threshold proxy), BP2 (trans with pathogenic in recessive)
Source: clinicalgenome.org
Orphanet / Orphadata
June 2025 ~3,200 genes
Gene-disease-inheritance mode annotations for rare diseases (autosomal dominant, autosomal recessive, X-linked dominant, X-linked recessive)
Used by: BS1 (inheritance-aware frequency threshold: Orphanet AD/XLD -> 0.1%, Orphanet AR -> 5%), PVS1 disease association gate (gene must have Orphanet entry for PVS1 to apply)
Source: orphadata.com (INSERM, France)
ClinGen VCEP Specifications
Latest release ~50-60 genes
Gene-specific ACMG criteria thresholds from ClinGen Variant Curation Expert Panels
Used by: BA1, BS1, PM2 (gene-specific frequency thresholds), PVS1 (applicability gate for gain-of-function genes), PVS1 disease association gate
Source: cspec.genome.network
DECIPHER G2P
2026-02-28 (7 panels + mechanism) ~43K records, 2,125 genes; 2,372 mechanism records
Gene-disease associations from 7 curated clinical panels: DD (developmental disorders), Eye, Cardiac, Skin, Skeletal, Cancer, Ear. Each entry includes allelic requirement, molecular mechanism (gain of function, loss of function, dominant negative), and confidence level (definitive, strong, moderate). Molecular mechanism data used for PVS1 GoF/DN guard (v3.17.0).
Used by: HPO enrichment pipeline (source 3 of 6, v3.12.0). PVS1 GoF/DN guard: gof_genes_unified view provides 348 pure GoF/DN/GoE monoallelic genes where PVS1 is blocked (v3.18.0, HELIX-REF-003). Sources: G2P (199 genes), GoFCards (179 genes), manual curation (27 genes). Dual-mechanism genes excluded from view to preserve PVS1 eligibility for LoF phenotypes.
Source: DECIPHER / Wellcome Sanger Institute
Monarch Initiative
Latest release ~151K records, 4,791 genes
Gene-phenotype associations aggregated from HPO Consortium redistribution via Monarch Knowledge Graph. Broad coverage of gene-disease relationships from multiple biomedical ontologies.
Used by: HPO enrichment pipeline (source 4 of 6). Large-scale gene coverage complements primary HPO Consortium data.
Source: Monarch Initiative (monarchinitiative.org)
ClinVar-MedGen HPO Mapping
Derived from ClinVar 2025-01 ~245K records, 5,258 genes
Gene-phenotype associations derived from ClinVar P/LP variant submissions mapped through MedGen CUI to HPO terms. Chain: ClinVar P/LP variant -> gene -> disease (MedGen CUI) -> HPO terms. Largest single contributor of new genes to the HPO enrichment pipeline.
Used by: HPO enrichment pipeline (source 5 of 6). Covers 5,258 genes including many not present in primary HPO Consortium or Orphanet datasets.
Source: NCBI ClinVar + MedGen
UniProt SwissProt
2026_01 (human reviewed proteome) 20,431 proteins, 113,126 PM1-eligible features
Expert-curated protein sequence and feature annotation for the reviewed human proteome. Per-residue functional annotations parsed from the SwissProt flat file format: DISULFID (disulfide bonds), ACT_SITE (enzyme active sites), BINDING (substrate/cofactor binding sites), MOD_RES (post-translational modifications), and MOTIF (functional motifs). Includes UniProt-to-Ensembl cross-references for direct ENSP-based lookup at classification time.
Used by: PM1 Tier 1 (primary) -- residue-level critical functional residue evidence (v3.30.0, HELIX-CR-2026-082). 113,126 PM1-eligible features across 14,278 proteins replace previous domain-level Pfam-only logic (2,262x evidence density increase over v3.28). Eligibility determined by qualifier presence (any /evidence= or /note=); naked features excluded by default with optional config override for clinical cases.
Source: UniProt Consortium (EMBL-EBI / SIB / PIR)
InterPro / Pfam
March 2026 (27,481 entries) ~27.5K Pfam domain entries, ~50 critical for PM1
Pfam entries loaded with domain accession, type (domain/family/repeat/motif), and functional annotation. Critical functional domains provide PM1 Tier 2 fallback evidence for VCEP-defined regions and domain-level evidence not yet captured at residue resolution by UniProt Tier 1 (v3.30.0).
Used by: PM1 Tier 2 fallback (critical functional domain detection via is_critical flag, v3.15.0). ~50 domains marked as critical across 14 categories: kinase, DNA binding, ion channel, GTPase, protease, signaling, tumor suppressor, calcium binding, structural, receptor, hearing, nuclear hormone receptor, ubiquitin, RNA binding. Migrated to Parquet-first Pattern A architecture in v3.30.0 (HELIX-CR-2026-082 Phase 5).
Source: EMBL-EBI InterPro
ClinGen Gene-Disease Validity
Latest release ~3,374 gene-disease pairs
Gene-disease validity classifications (Definitive, Strong, Moderate, Limited) with mode of inheritance for autosomal dominant, autosomal recessive, and X-linked disorders. Curated by ClinGen Gene Curation Expert Panels.
Used by: PVS1 constraint gate (AD Definitive/Strong genes bypass pLI/LOEUF, v3.16.6). PVS1 AR LoF gene bypass (AR Definitive/Strong with verified LoF mechanism, v3.16.0). Disease association gate for PVS1, PP3_Strong, LP4/LP5/LP6, and ClinVar override (v3.16.0).
Source: clinicalgenome.org
GoFCards
Release 1.0 579 genes, 3,161 curated GoF variants
Curated gain-of-function variant database with experimental evidence (animal model, cell model), pathways, disorder associations, and proprietary evidence score (Pscore). Each variant entry includes PMID, functional description, and experimental validation status.
Used by: PVS1 GoF guard via gof_genes_unified view (v3.18.0, HELIX-REF-003). Gene-level aggregation: genes with 3+ curated GoF variants, average Pscore >= 2.0, and at least one variant with animal or cell model evidence qualify as GoF with confidence "strong". 217 genes passed threshold and are integrated into gene_disease_mechanism table.
Source: National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University
ClinGen Dosage Sensitivity (Full)
Latest release (2026-03-21) ~1,626 genes
Full dosage sensitivity curation including haploinsufficiency score (0-3, 30, 40), triplosensitivity score, up to 6 PMIDs per gene, date last evaluated, and disease IDs. Replaces the legacy 3-column clingen table (v3.18.0). HI scores: 0 = no evidence, 1 = little evidence, 2 = emerging evidence, 3 = sufficient evidence, 30 = AR phenotype, 40 = dosage sensitivity unlikely.
Used by: PVS1 constraint gate (HI=3 as fallback for X-linked genes, v3.11.5). Disease association gate (HI IN 2, 3, 30 per HELIX-CR-2026-025, v3.18.1). BS1 inheritance proxy. BP2 (HI=30 for AR phenotype).
Source: clinicalgenome.org
Ensembl VEP
Release 113 All coding/non-coding consequences
Variant consequence prediction, protein impact, transcript annotation, functional domain mapping
Used by: PVS1 (consequence type), PM1 (Pfam domains), PM4 (in-frame indels), BP3 (non-critical regions), BP7 (synonymous), all impact-based criteria
Source: ensembl.org
### ACMG/AMP Classification
Variant classification follows the 2015 ACMG/AMP guidelines with 28 evidence criteria evaluated systematically. 20 criteria are fully automated; 8 require manual curation by the reviewing geneticist.
Classification Priority Order
Classification logic is applied in strict priority order. Higher-priority rules are evaluated first, and the first matching rule determines the final classification:
1
ClinGen ERepo Expert-Panel Override
A ClinGen Expert Curation (ERepo) Pathogenic or Likely Pathogenic assertion from an FDA-recognized Variant Curation Expert Panel is applied at the top of the staircase, above the mechanical frequency-benign branches. An expert-panel P/LP assertion supersedes BA1/BS1/BS2 because the panel already weighed population frequency when it set the classification (canonical BTD c.1330G>C precedent: 5.56% Finnish frequency, yet Pathogenic). The join is variant-level only. Added v3.34.0 (HEL-CR-2026-CLASSIFIER-EREPO-OVERRIDE-001) and re-ordered to the top of the staircase at v3.35.0 (HELIX-CR-2026-093).
2
Established / Low-Penetrance Risk-Allele Suppress
An established or low-penetrance risk allele carried by a ClinGen expert-panel or ClinVar assertion is suppressed to VUS instead of remaining Benign, including at BA1-band frequency. Because the five-tier ACMG output has no risk-allele tier, VUS is the interim class and [RiskAllele] records the assertion. This path cannot produce P/LP. It uses ClinVar risk-allele significance with a documented Factor V Leiden carve-out and became active in v3.38.0 (HELIX-CR-2026-107) on the v3.33-era substrate (HELIX-CR-2026-102A).
3
BA1 Stand-alone
Allele frequency > 5% is always classified Benign. BA1 is the only stand-alone criterion in the ACMG framework and cannot be overridden by any other evidence, including ClinVar assertions.
4
Conflicting Evidence
If a variant has pathogenic evidence at moderate strength or above (PVS, PS, or PM criteria triggered) AND strong benign evidence (BS criteria triggered), the variant is classified as VUS regardless of the individual evidence strength. This is a conservative approach that prioritizes clinical safety.
5
ClinVar Override
ClinVar classification is applied only when no conflicting computational evidence exists. Requires minimum review star level (default: 1 star). ClinVar VUS does not override computational classification.
6
Bayesian Point System
Each triggered criterion contributes points based on its evidence strength: Very Strong (+8), Strong (+4), Moderate (+2), Supporting (+1) for pathogenic; Strong (-4), Supporting (-1) for benign. Total points determine classification: >= 10 Pathogenic, 6-9 Likely Pathogenic, 0-5 VUS, -1 to -5 Likely Benign, <= -6 Benign. This system is mathematically equivalent to the original 18 ACMG combining rules while filling gaps for evidence combinations not explicitly covered.
7
Default
Variants that do not meet any of the above criteria are classified as Uncertain Significance (VUS).
Classification Output
Each variant receives one of five standard ACMG classifications, a list of all criteria that were triggered with evidence strength levels (e.g., "PVS1,PM2,PP3_Strong"), a Bayesian point total, and a continuous confidence score derived from the distance between the point total and the nearest classification boundary:
Pathogenic
>= 10 pts
confidence: 0.80-0.99
Likely Pathogenic
6-9 pts
confidence: 0.70-0.90
VUS
0-5 pts
confidence: 0.30-0.60
Likely Benign
-1 to -5 pts
confidence: 0.70-0.90
Benign
<= -6 pts
confidence: 0.80-0.99
### Automated Criteria (20 of 28)
These criteria are evaluated automatically for every quality-passing variant. Exact conditions and thresholds are documented below. Each criterion lists the databases it depends on and known limitations.
Pathogenic Evidence
PVS1
Null variant in gene where loss-of-function is a known disease mechanism
Very Strong / Strong
Conditions
- Impact = HIGH
- Consequence: frameshift, stop_gained, splice_acceptor, or splice_donor variant
- Gene constraint: pLI > 0.9 OR LOEUF < 0.45 OR gene is in the curated autosomal recessive LoF gene list (~150 genes with Definitive/Strong evidence for biallelic LoF disease mechanism) OR ClinGen haploinsufficiency_score = 3 OR gene is in the curated autosomal dominant ClinVar LOF gene list (AD genes with ClinVar P/LP LOF variants at 2+ review stars but uninformative gnomAD constraint, v3.13.0). LOEUF threshold raised 0.35 -> 0.45 at v3.35.0 (HELIX-CR-2026-093) per the gnomAD v4.1.1 (2026-03-30) LoF-constraint recommendation.
- Last-exon NMD downgrade (v3.9.0): Truncating variants in the last exon are downgraded from PVS1 Very Strong (+8 points) to PVS1_Strong (+4 points). Last-exon variants escape nonsense-mediated mRNA decay (NMD) because there is no downstream exon-exon junction. The resulting truncated protein may retain partial function or exert a dominant-negative effect. Detection uses the VEP exon_number column (format X/Y): last exon when X == Y. Single-exon genes (1/1) correctly receive PVS1_Strong. NULL or missing exon_number defaults to Very Strong (conservative). Reference: Abou Tayoun et al. 2018, Figure 1, Node 3.
- MANE Select positional guard (v3.22.0, HELIX-CR-2026-055): Variants outside the MANE Select/Plus Clinical CDS region do not receive PVS1 or PM4. VEP may annotate variants against non-canonical transcripts that extend beyond the canonical coding sequence, producing false loss-of-function predictions. 19,354 MANE transcripts loaded with CDS coordinates. NULL MANE data = PVS1 applied (conservative fallback). Reference: Abou Tayoun et al. 2018, PMID:30192042.
- Expression-aware guard (v3.23.0, HELIX-CR-2026-056): Per-exon pext scores from gnomAD v4.1 (GTEx v10, 49 tissues) modulate PVS1. Three tiers based on max_tissue_pext: >= 0.9 = PVS1 Very Strong (unchanged), 0.1-0.9 = PVS1_Strong[pext] (downgraded to Strong), < 0.1 = PVS1 blocked (exon not expressed in any tissue). NULL pext = Very Strong (conservative fallback). PM4 blocked at pext < 0.1. Scientific basis: Cummings 2020 (PMID:32461655) - 22.8% false pLoF filtered, < 4% true pathogenic lost.
Exclusions
- NMD-escaping transcripts (consequence contains NMD_escaping_variant). Note: NMD_transcript_variant is NOT excluded - it indicates the transcript undergoes NMD, which is evidence FOR loss-of-function (v3.11.3).
- Non-canonical splice site variants: splice_donor_5th_base_variant (position +5), splice_donor_region_variant (positions +3 to +8), splice_acceptor_5th_base_variant (position -5). ClinGen PVS1 Decision Tree specifies canonical +/-1,2 only. Extended splice variants retain PP3_splice access (v3.11.4).
- Gain-of-function / dominant-negative / gain-of-expression genes: multi-source guard (v3.18.0, HELIX-REF-003). Primary: gof_genes_unified view (348 pure GoF/DN/GoE monoallelic genes from G2P + GoFCards + manual curation, confidence definitive/strong/moderate). Fallback: GOF_AD_GENES curated list (29 genes including ANKRD26, C1QTNF5, CALM2, MYH3, PRKAG2, SKI, TUBB3 and 22 others with PMID-documented non-LoF mechanism). Dual-mechanism genes (e.g., SCN5A, LMNA, KCNH2, MYH7) excluded from view and retain PVS1 eligibility. ClinGen PVS1 Decision Tree (Abou Tayoun 2018): "Is LoF a known mechanism of disease?" - if NO, PVS1 not applicable.
- Gain-of-function gene guard history: v3.16.8 (HELIX-CR-2026-022) introduced GOF_AD_GENES (36 curated genes). v3.17.0 (HELIX-CR-2026-023) added G2P as primary source, reduced GOF_AD_GENES to 14 fallback genes, and fixed GNAS dual-mechanism error.
- Stop-retained and stop-lost variants
- HLA gene family (HLA-A, HLA-B, HLA-C, HLA-DRA, HLA-DRB1, HLA-DRB5, HLA-DQA1, HLA-DQB1, HLA-DPA1, HLA-DPB1, HLA-E, HLA-F, HLA-G, HLA-DMA, HLA-DMB, HLA-DOA, HLA-DOB)
- Homozygous reference genotypes (hom_ref) from multi-allelic VCF sites
Databases: VEP (consequence, impact), gnomAD Constraint (pLI, LOEUF), curated AR LoF gene list (ClinGen Gene-Disease Validity), curated AD ClinVar LOF gene list (v3.13.0), Orphanet (gene-disease association), ClinVar (gene-level P/LP check), ClinGen (haploinsufficiency_score), VCEP (gene coverage), gof_genes_unified view (348 GoF/DN/GoE genes from G2P + GoFCards + manual, v3.20.1), GOF_AD_GENES fallback (29 genes, v3.20.1), gene_disease_mechanism table (disease association gate, v3.18.2), MANE Select transcripts v1.4 (positional guard, v3.22.0), gnomAD pext v4.1 GTEx v10 (expression-aware guard, v3.23.0)
Known limitations
- Does not evaluate reading frame rescue via downstream in-frame reinitiation
- Tissue-specific expression is assessed via gnomAD pext (proportion expressed across transcripts, GTEx v10) for PVS1 and PM4 modulation (v3.23.0, HELIX-CR-2026-056). Exons with max tissue pext < 0.1 block PVS1; 0.1-0.9 downgrade PVS1 to Strong. Alternative transcript usage addressed via MANE Select positional guard (v3.22.0, HELIX-CR-2026-055): variants outside MANE Select CDS do not receive PVS1 or PM4.
- VCEP gene-specific PVS1 applicability gate available for ~50-60 genes (e.g., PVS1 disabled for gain-of-function genes like MYOC). Generic GoF/DN/GoE guard via gof_genes_unified (347 genes from G2P + GoFCards + manual curation) + GOF_AD_GENES fallback (28 genes) covers all other known non-LoF mechanism genes (v3.18.0, HELIX-REF-003). Generic thresholds used for all other genes.
- Autosomal recessive gene bypass uses a curated list of ~150 genes with Definitive/Strong ClinGen validity or equivalent published evidence for biallelic LoF disease mechanism. VCEP AR genes (e.g., CDH23, GJB2, MYO7A, PAH, SLC26A4, USH2A) are handled via the VCEP pvs1_applicable gate and are not duplicated in the curated list. Categories covered: neurodegeneration, metabolic (amino acid, fatty acid oxidation, glycogen storage, lysosomal, peroxisomal), ciliopathies, sensory (hearing, vision), immune/hematologic, cardiac, neuromuscular, connective tissue, kidney, endocrine, liver, and respiratory.
- Autosomal dominant ClinVar LOF gene bypass (v3.13.0) uses a curated list of AD genes with ClinVar P/LP LOF variants (2+ review stars) but uninformative gnomAD constraint (pLI < 0.9, LOEUF >= 0.35, HI != 3). Targets small genes where gnomAD constraint is statistically underpowered. New genes added only with clinical trigger and verification. Reference: Abou Tayoun 2018 PMID:30192042, HELIX-CR-2026-010.
PS1
Same amino acid change as an established pathogenic variant
Strong
Conditions
- ClinVar clinical significance: Pathogenic, Pathogenic/Likely_pathogenic, or Likely_pathogenic
- ClinVar review stars >= 2
Databases: ClinVar (clinical_significance, review_stars)
Known limitations
- PS1 matches the exact variant position and allele. PM5 separately evaluates a different pathogenic missense substitution at the same gene and codon from v3.39.1 onward (CR-2026-PM5-ENABLEMENT-001).
- ClinVar assertions may lag behind current evidence for recently reclassified variants
PM1
Located in a mutational hot spot or well-established functional domain
Moderate
Conditions
- Consequence: missense_variant (v3.20.1, HELIX-CR-2026-035). PM1 is defined for missense variants in critical functional domains (Richards 2015 Table 3, ClinGen SVI Walker 2023 Section 3.3). Synonymous, frameshift, splice, and other non-missense variants do not receive PM1.
- Three-tier PM1 evidence model (v3.30.0, HELIX-CR-2026-082):
- Tier 1 (PRIMARY) - UniProt residue-level evidence: variant overlaps a curated critical residue from UniProt SwissProt human reviewed proteome. 113,126 PM1-eligible features across 14,278 proteins covering DISULFID (disulfide bonds), ACT_SITE (enzyme active sites), BINDING (substrate/cofactor binding), MOD_RES (post-translational modifications), and MOTIF (functional motifs). Match performed via Ensembl protein (ENSP) cross-reference at the exact protein position parsed from hgvs_protein. Evidence eligibility uses UniProt qualifier presence (any /evidence= or /note=); naked features without qualifiers excluded by default (2,067 features, 1.8%).
- Tier 2 (FALLBACK) - InterPro Pfam critical domains: variant overlaps a critical functional domain from interpro_pfam_domains reference table (~50 domains marked is_critical=true, sourced from InterPro Pfam catalog, v3.15.0). Retained for VCEP-defined regions and domain-level evidence not yet captured at residue resolution.
- Tier 3 (RESERVED) - VCEP-specific PM1 region overrides: future expansion for gene-specific PM1 region definitions from published ClinGen Variant Curation Expert Panel specifications.
- PM1 boolean: triggered if Tier 1 OR Tier 2 matches. Single Moderate strength preserved (no tiered Moderate/Supporting split) to maintain compatibility with existing combining rules.
Databases: VEP (consequence, hgvs_protein), UniProt SwissProt human reviewed (uniprot_features, uniprot_ensembl_mapping reference tables, v3.30.0), InterPro Pfam (interpro_pfam_domains reference table)
Known limitations
- UniProt naked features (no /evidence= or /note= qualifiers, 2,067 features) excluded from PM1 eligibility by default. Config flag pm1_uniprot_include_naked allows clinical override for specific cases where curator inclusion in SwissProt is treated as curation evidence (e.g., INSR Cys234Tyr disulfide bond Cys223-Cys234).
- UniProt evidence is residue-level; some VCEP gene-specific PM1 regions (multi-residue critical regions defined by Variant Curation Expert Panels) are not yet integrated and rely on the Pfam fallback tier.
- Tier 2 Pfam fallback retains the curation scope of v3.15.0: ~50 critical domains across 14 categories (kinase catalytic, DNA-binding, ion channel pores, enzyme active sites, and other VCEP-documented domains). Generic structural domains (Caveolin, coiled-coil, DUF) and repetitive regions remain excluded.
- The pm1_evidence_source audit column records each PM1 source (format: uniprot:DISULFID:223-234:P06213 or pfam:PF00069) for per-variant review and frontend display.
PM2
Absent from controls or at extremely low frequency in population databases
Moderate
Conditions
- gnomAD global allele frequency < 0.0001 (0.01%) OR variant absent from gnomAD (NULL frequency = never observed in ~800K individuals)
- NULL gnomAD allele frequency correctly treated as "absent from controls" per Richards et al. 2015 Table 3
Databases: gnomAD v4.1 (global_af)
Known limitations
- Does not apply population-specific frequency adjustments
- ClinGen SVI PM2_Supporting downgrade implemented for VCEP genes (v3.20.0, HELIX-CR-2026-034). 46 VCEP genes use PM2 at Supporting strength; 11 genes retain Moderate (RASopathy, ENIGMA BRCA1/2, HBOP). Non-VCEP genes use full Moderate strength.
- When VCEP gene-specific specifications are enabled, PM2 threshold may differ from the generic 0.01% (e.g., 0% for RASopathy genes where any population frequency argues against pathogenicity)
PM3
Detected in trans with a pathogenic variant for recessive disorders
Moderate
Conditions
- Variant flagged as compound heterozygote candidate (compound_het_candidate = true)
- Compound-heterozygote splice-partner expansion (final v3.39.1, CR-2026-COMPHET-SPLICE-PARTNER-001): Tier A accepts splice donor, splice acceptor, fifth-base donor, and donor-region consequences directly; Tier B accepts splice-region and polypyrimidine-tract consequences only with SpliceAI >= 0.2 or VEP HIGH/MODERATE impact.
- Carrier-guard release (final v3.39.1, CR-2026-COMPHET-SPLICE-PARTNER-001) is tied to an actual PM3 trigger, not the presence of a bare compound_het_candidate flag.
- ClinVar partner validation guard (v3.25.1, HELIX-CR-2026-058): trans partner must be ClinVar Pathogenic or Likely Pathogenic with review_stars >= 2 and not Conflicting. Aligns with Richards 2015 Table 3: PM3 explicitly requires "detected in trans with a pathogenic variant". ClinVar annotation is external/immutable, avoiding circular dependency where partner classification depends on PM3.
Exclusions
- Gain-of-function / dominant-negative gene exclusion (v3.20.2, HELIX-CR-2026-036): PM3 not applied in pure GoF/DN genes. PM3 presupposes AR mechanism (Richards 2015 Table 3), biologically irrelevant for dominant disease genes.
- Dual-mechanism bypass (v3.20.3, HELIX-CR-2026-037): genes with both GoF/DN monoallelic AND biallelic LoF mechanism (gof_genes_exclusive view, 55 dual-mechanism genes excluded from GoF guard) retain PM3 for the AR LoF pathway.
Databases: Pipeline-internal (compound heterozygote detection), ClinVar (partner clinical_significance, review_stars), gof_genes_exclusive view (G2P + GoFCards + manual, 347 GoF/DN/GoE genes minus 55 dual-mechanism), GOF_AD_GENES fallback
Known limitations
- Compound heterozygote status inferred from genotype data without formal phasing
- Trio data or long-read phasing would provide definitive trans confirmation
- ClinVar partner requires 2+ review stars; partners with 1-star P/LP submissions do not satisfy PM3
- A candidate pair that fails the PM3 ClinVar partner guard does not release the AR carrier or computational-only Likely Pathogenic guards (final v3.39.1).
PM4
Protein length change in a non-repetitive region
Moderate
Conditions
- Consequence: in-frame insertion or in-frame deletion
- Located within a Pfam functional domain
- Not in a repetitive or low-complexity region (domains not containing "tandem", "repeat", "lowcomplexity", or "Seg")
Exclusions
- HLA gene family (same exclusion list as PVS1)
Databases: VEP (consequence, domains)
Known limitations
- Repetitive region detection based on VEP domain annotations only
- Does not evaluate whether the in-frame change disrupts a critical functional residue
PM5
Novel missense change at an amino acid residue where a different pathogenic missense change has been seen
Moderate
Conditions
- Consequence: missense_variant with a parseable gene symbol and protein codon position
- A different alternate amino acid at the same gene and codon is present in the ClinVar-derived pathogenic-codon asset as Pathogenic or Likely Pathogenic with review_stars >= 2
- Enabled unconditionally in the final v3.39.1 classifier when the reference asset is available (CR-2026-PM5-ENABLEMENT-001)
Exclusions
- The same alternate amino acid is excluded because it is PS1 evidence, not PM5
- Non-missense, nonsense, benign, and conflicting ClinVar records are excluded when the reference asset is built; the 2-star minimum is enforced when the classifier consumes the asset
Databases: ClinVar variant_summary.txt.gz-derived refdb.pm5_pathogenic_codons (gene_symbol, codon, ref_aa, alt_aa, review_stars)
Known limitations
- Requires a parseable protein-level HGVS codon and an available pm5_pathogenic_codons reference asset; a missing asset degrades safely to no PM5 evidence
- ClinVar assertions may lag behind current evidence for recently reclassified variants
PM5 was activated in the final v3.39.1 implementation (CR-2026-PM5-ENABLEMENT-001). It contributes Moderate observed evidence and can satisfy the observed-evidence requirement of the computational-only Likely Pathogenic guard.
PP2
Missense variant in a gene with low rate of benign missense variation
Supporting
Conditions
- Consequence: missense variant
- Gene constraint: pLI > 0.5 AND mis_z > 2.0 (missense constraint)
- Disease association gate (v3.18.5): gene must have established disease association via same 7 sources as PVS1 gate. Prevents PP2 in genes without known Mendelian disease mechanism.
- MPC regional constraint guard (v3.18.6, HELIX-CR-2026-032): mpc_score >= 1.0 required when MPC data is available. Regions with mpc_score < 1.0 are unconstrained for missense - PP2 blocked. NULL MPC data falls back to gene-level constraint (bypass). Reference: Samocha et al. 2014.
Databases: VEP (consequence), gnomAD Constraint (pLI, mis_z), MPC regional constraint (mpc_score), disease association (7 sources)
Known limitations
- pLI measures loss-of-function constraint; mis_z measures missense constraint
- Both metrics are required: pLI > 0.5 indicates LoF intolerance, and mis_z > 2.0 indicates missense constraint (Samocha et al. 2014).
- MPC regional constraint data available for ~18K genes. NULL MPC falls back to gene-level constraint only.
PP3
Computational evidence supports a deleterious effect (two independent paths)
Supporting / Moderate / Strong
Conditions
- Path A (Missense - BayesDel): Requires consequence = missense_variant (v3.12.1, HELIX-CR-2026-009). BayesDel_noAF is calibrated exclusively for missense variants (Pejaver et al. 2022); non-missense variants (stop_gained, frameshift) receive spurious scores from dbNSFP positional overlap and are excluded. Three strength levels: PP3_Strong (>= 0.518, +4 points), PP3_Moderate (0.290-0.517, +2 points), PP3_Supporting (0.130-0.289, +1 point). Scores below 0.130 are indeterminate. Requires missense relevance guard v2.0 (v3.19.0, HELIX-CR-2026-033): 5 gates determine if missense is a disease mechanism. Gate 1: mis_z > 2.0 (direct constraint). Gate 2: mis_z > 1.5 AND pLI < 0.5 (AR genes). Gate 3: mis_z > 0.5 AND pLI > 0.9 (AD LoF-intolerant). Gate 4a: gene in refdb.clinvar_missense_genes (2,136 genes with ClinVar P/LP missense evidence). Gate 4b: pLI > 0.7 AND mis_z > 0.5 AND gene has disease association (gray-zone genes). NULL fallback: pLI > 0.5 without mis_z data. PP3_Strong additionally requires established disease association (v3.8.1 gate, same 7 sources as PVS1 gate).
- PM1 + PP3 double-counting guard: When PM1 (functional domain) applies alongside PP3_Strong, PP3 is downgraded to PP3_Moderate. Combined PM1 + PP3 capped at Strong equivalent (4 points) per ClinGen SVI.
- Path B (Splice evidence - 3 strength levels, v3.15.1): SpliceAI max_score evaluated with evidence strength modulation. PP3_splice_Strong (>= 0.8, +4 points), PP3_splice_Moderate (0.5-0.799, +2 points), PP3_splice (0.2-0.499, +1 point Supporting). PP3_splice_Strong requires disease association gate (same as PP3_Strong missense). All levels excluded when PVS1 applies (ClinGen SVI 2023 double-counting guard).
- Paths A and B are independent. A variant can trigger both if it has a BayesDel score and a splice prediction.
Exclusions
- PP3_splice not applied when PVS1 is triggered at any strength level (Very Strong or Strong/last-exon). Prevents double-counting loss-of-function and splice evidence per ClinGen SVI 2023.
- PP3_splice missense-tolerated guard (v3.25.7, HELIX-CR-2026-065): PP3_splice (all levels) blocked for pure missense variants without VEP splice consequence when BayesDel noAF data is available. ClinGen SVI (Walker 2023 Section 3.2): PP3 is a single criterion - when BayesDel indicates tolerated and SpliceAI predicts cryptic splice, evidence is conflicting, not additive. BayesDel IS NULL bypass preserved.
- PP3 + PP3_splice mutual exclusion (v3.25.8, HELIX-CR-2026-066): for dual-consequence variants (missense AND splice_region/donor/acceptor), PP3 BayesDel is blocked when PP3_splice applies. SpliceAI is the more specific tool when VEP confirms splice proximity. Together with v3.25.7, exactly one PP3 path per variant - never both.
- BayesDel indeterminate range (-0.180 to 0.129): no PP3 or BP4 applied
Databases: dbNSFP 4.9c (bayesdel_noaf_score), SpliceAI (max_score)
Known limitations
- BayesDel_noAF used to avoid circular reasoning with PM2/BA1/BS1 frequency criteria (Pejaver et al. 2022)
- PM1 + PP3 cap assumes PM1 = Moderate (2 points). Extensible if future VCEP upgrades PM1 strength.
- BayesDel does not reach PP3_Very_Strong per Pejaver calibration data
PP4
Patient phenotype is highly specific for a disease with a single genetic etiology
Supporting
Conditions
- Requires patient HPO terms to be provided for the analysis session
- Trigger condition A: >= 3 patient HPO terms match the gene HPO profile
- Trigger condition B: >= 2 patient HPO terms match AND gene has <= 5 total HPO associations (highly specific gene-phenotype relationship)
Exclusions
- Not evaluated when no patient HPO terms are provided
Databases: HPO (gene-phenotype associations)
Known limitations
- HPO matching is exact term overlap, not semantic similarity (ontology hierarchy not used at this stage)
- Semantic similarity-based matching is performed by the downstream Phenotype Matching Service
PP5
Reputable source reports variant as pathogenic
Supporting
Conditions
- ClinVar clinical significance: Pathogenic, Pathogenic/Likely_pathogenic, or Likely_pathogenic
- ClinVar review stars >= 1 AND < 2 (lower confidence than PS1)
Exclusions
- Does not apply when PS1 already applies (prevents double-counting ClinVar evidence at different strength levels)
Databases: ClinVar (clinical_significance, review_stars)
Known limitations
- ClinGen SVI has recommended retiring PP5 as a standalone criterion; retained for maximum sensitivity
Benign Evidence
BA1
Allele frequency is above 5% in population databases
Stand-alone
Conditions
- gnomAD Grpmax Filtering AF (af_grpmax) > 0.05 (5%). v3.31.0 (HELIX-CR-2026-090) switched the effective BA1/BS1 frequency from global allele frequency to the gnomAD v4 Grpmax Filtering AF per ClinGen SVI guidance (March 2024), with global allele frequency retained only as a fallback. A bottleneck guard skips BA1/BS1 where af_grpmax is NULL or zero (gnomAD intentionally excludes the bottlenecked AJ/FIN/AMI/MID populations from the Grpmax FAF).
Databases: gnomAD v4.1 (af_grpmax Grpmax Filtering AF primary, global_af fallback)
Known limitations
- Uses the gnomAD Grpmax Filtering AF (af_grpmax) as the effective frequency with global allele frequency as fallback (v3.31.0); per-gene population-specific BA1 thresholds beyond the VCEP overlay are not implemented
BA1 is the only stand-alone ACMG criterion. A variant meeting BA1 is classified as Benign regardless of any other evidence, including ClinVar assertions. When VCEP gene-specific specifications are enabled, the BA1 threshold may be lower than 5% for specific genes (e.g., 0.1% for Cardiomyopathy genes).
BS1
Allele frequency is greater than expected for the disorder
Strong
Conditions
- Effective frequency: gnomAD Grpmax Filtering AF (af_grpmax), with global allele frequency as fallback and a NULL/zero bottleneck guard (v3.31.0, HELIX-CR-2026-090, per ClinGen SVI March 2024). The cascade threshold is then selected inheritance-aware:
- Inheritance-aware frequency threshold determined by a 5-level cascade (v3.7.0):
- Priority 1 - VCEP gene-specific: BS1 threshold from published ClinGen VCEP specification (~100 genes)
- Priority 2 - ClinGen HI score = 3: allele frequency >= 0.001 (0.1%) (~400 genes)
- Priority 3 - Orphanet AD inheritance: allele frequency >= 0.001 (0.1%) for genes with autosomal dominant disease association (~1,694 genes)
- Priority 3b - Orphanet XLD inheritance: allele frequency >= 0.001 (0.1%) for genes with X-linked dominant disease association (~90 genes). XLR genes intentionally excluded (fall to AR default).
- Priority 4 - HPO AD fallback: allele frequency >= 0.001 (0.1%) for genes annotated with HP:0000006 (autosomal dominant inheritance) (~200-400 additional genes)
- Recessive-LoF legitimacy guard (v3.39.0, HELIX-CR-2026-108): genes supported by both a definitive/strong/moderate biallelic LoF mechanism in gene_disease_mechanism and autosomal-recessive inheritance in GenCC bypass the generic pure-AD 0.1% clauses and fall through to the AR threshold unless a higher-priority VCEP or ClinGen-HI rule applies.
- HPO-AD completion (v3.39.1, HELIX-CR-2026-108 follow-up): the same biallelic LoF set also bypasses the adjacent HPO-AD fallback, closing the TSHR-class fall-through found after v3.39.0.
- Default - AR threshold: allele frequency >= 0.05 (5%) for all remaining genes
- Upper bound: allele frequency <= BA1 threshold (BA1 takes precedence)
Exclusions
- Does not apply when BA1 applies (BA1 takes precedence)
- XLR (X-linked recessive) genes do not receive the AD threshold - carrier females are unaffected and higher carrier frequencies are expected
Databases: gnomAD v4.1 (af_grpmax Grpmax Filtering AF primary, global_af fallback), ClinGen (haploinsufficiency_score), Orphanet/Orphadata (orphanet_gene_inheritance: has_ad, has_ar, has_xld), HPO (hpo_ids for HP:0000006), gnomAD Constraint (pLI, oe_lof_upper for constraint-implied AD fallback), AR_LOF_GENES curated list, gene_disease_mechanism plus GenCC autosomal-recessive evidence (v3.39.0-v3.39.1)
Known limitations
- Orphanet covers rare diseases only; some common AD conditions may not be catalogued (covered by HPO AD fallback or AR default)
- Dual AD+AR inheritance genes (1,027 genes, e.g. KCNJ11) receive the AD threshold (0.1%) as conservative default - AR pathogenic variants are typically well below 0.1% AF
- Orphanet data updated twice per year (June + December releases). Coverage expands with each release.
- When VCEP gene-specific specifications are enabled, VCEP thresholds take highest priority and override all other cascade levels
- Criteria string includes inheritance source annotation: BS1[ClinGen-HI], BS1[Orphanet-AD], BS1[Orphanet-XLD], BS1[HPO-AD]
- If either recessive-LoF reference source is absent or their intersection is empty, the v3.39 guard is a no-op and the pre-v3.39 threshold cascade is preserved.
Data source: Orphadata (INSERM, France). License: CC-BY-4.0. Attribution: Orphadata: Free access data from Orphanet. INSERM 1999. Available on https://www.orphadata.com.
BS2
Observed in a healthy adult individual for a fully penetrant early-onset disorder
Strong
Conditions
- Inheritance-aware threshold (v3.26.8, HELIX-CR-2026-074): AR-only genes (Orphanet has_ar=TRUE, has_ad=FALSE OR ClinGen HI=30) with documented early-onset-only diseases use threshold of 10 homozygotes. Generic and AD genes use threshold of 15 homozygotes.
- Onset guard: AR threshold activated only when gene exists in refdb.orphanet_disease_onset (no data falls back to generic 15) AND no associated disease has Adult/Elderly/All ages onset. Aligns with Richards 2015 Table 3: BS2 "with full penetrance expected at an early age".
- Reduced-penetrance safeguard (v3.39.0, HELIX-CR-2026-108): in the same evidence-backed biallelic LoF gene set used by BS1, BS2 is suppressed when the variant carries canonical pathogenic signal - a HIGH frameshift/stop-gained/canonical splice/start-lost consequence or a ClinVar P/LP assertion meeting the configured review-star minimum.
- AR path criteria string marker: BS2[AR] when AR threshold path was used.
Exclusions
- AD BS2 from gnomAD homozygote count not implemented (Harrison 2019 PMID:31159682: gnomAD is general population, not healthy controls; APC VCEP Curia 2023 PMID:37805481 excludes gnomAD for AD BS2).
- X-linked hemizygous BS2 deferred (ac_hemi column not in current annotation pipeline).
Databases: gnomAD v4.1 (global_hom), Orphanet (orphanet_gene_inheritance, orphanet_disease_onset), ClinGen (haploinsufficiency_score for HI=30 AR proxy), gene_disease_mechanism plus GenCC autosomal-recessive evidence, VEP consequence/impact, ClinVar clinical significance and review stars (v3.39.0)
Known limitations
- AR threshold of 10 calibrated for early-onset full-penetrance AR diseases. Genes with adult-onset AR diseases (e.g., Wilson disease ATP7B) fall back to generic 15 threshold via onset guard.
- Compound heterozygote-only AR diseases require homozygote count interpretation in carrier-frequency context; not currently modelled separately.
- The v3.39.0 suppression is not a blanket exemption for recessive genes: it requires both independently supported biallelic LoF legitimacy and a canonical pathogenic signal on the variant.
BP1
Missense variant in a gene for which primarily truncating variants are known to cause disease
Supporting
Conditions
- Consequence: missense variant
- Impact: MODERATE
- Gene constraint: pLI < 0.1 (gene is tolerant to loss-of-function)
- pLI value must be present (non-NULL)
- Missense constraint guard (v3.6.8): mis_z must be < 2.0 or absent. Genes with high mis_z have missense as primary disease mechanism.
- ClinVar pathogenic missense guard (v3.8.2): BP1 does not apply when ClinVar has a session-local Pathogenic/Likely Pathogenic missense variant in the same gene. If ClinVar confirms pathogenic missense in the patient VCF, the premise of BP1 is falsified.
- Reference-based ClinVar missense guard (v3.20.9, HELIX-CR-2026-044): BP1 additionally blocked when gene is in refdb.clinvar_missense_genes (>= 2 ClinVar P/LP missense at 2+ stars). Reference-level check covers genes with established missense disease mechanism not present in the current patient VCF (extends session-local guard).
- GoF / DN gene exclusion (v3.20.6, HELIX-CR-2026-039v2): BP1 not applied for missense variants in gain-of-function or dominant-negative genes. BP1 presupposes LoF mechanism (Richards 2015 Table 3); GoF/DN genes have non-LoF truncating mechanism, so the BP1 premise is inapplicable. Same guard sources as PVS1 GoF guard: refdb.gof_genes_exclusive + GOF_AD_GENES fallback.
Exclusions
- Genes with mis_z >= 2.0 (missense is a known disease mechanism)
- Genes with session-local ClinVar Pathogenic/Likely Pathogenic missense variants
- Genes with reference-level ClinVar missense evidence (refdb.clinvar_missense_genes, 2,136 genes)
- GoF/DN genes (refdb.gof_genes_exclusive view + GOF_AD_GENES fallback)
- Curated AR LoF genes (~150 genes with biallelic LoF disease mechanism)
Databases: VEP (consequence, impact), gnomAD Constraint (pLI, mis_z), ClinVar (clinical_significance, consequence), refdb.clinvar_missense_genes (reference-level missense mechanism), refdb.gof_genes_exclusive (G2P + GoFCards + manual), GOF_AD_GENES, AR_LOF_GENES
Known limitations
- Low pLI used as proxy for "primarily truncating variants cause disease"
- mis_z measures heterozygous missense constraint. AR enzyme genes (e.g. MCCC2) may have low mis_z despite pathogenic homozygous missense - the ClinVar missense guards (v3.8.2 session-local + v3.20.9 reference-based) address this gap.
- GoF/DN gene coverage limited to G2P 2026-02-28 release + GoFCards Release 1.0 + curated GOF_AD_GENES. Genes with novel non-LoF mechanism not yet curated may receive incorrect BP1.
BP2
Observed in trans with a pathogenic variant for a fully penetrant dominant disorder
Supporting
Conditions
- Compound heterozygote candidate (compound_het_candidate = true)
- ClinGen haploinsufficiency score = 30 (dosage sensitivity unlikely)
Databases: Pipeline-internal (compound heterozygote detection), ClinGen (haploinsufficiency_score)
Known limitations
- Trans observation inferred without formal phasing
- Haploinsufficiency score of 30 is a specific ClinGen code for "dosage sensitivity unlikely"
BP3
In-frame insertion or deletion in a repetitive region without a known function
Supporting
Conditions
- Consequence: in-frame insertion or in-frame deletion
- Located in a repetitive/low-complexity region (domains contain "tandem", "repeat", "lowcomplexity", or "Seg"), OR not in any Pfam domain, OR no domain annotation available
Databases: VEP (consequence, domains)
Known limitations
- Complementary to PM4 (PM4 requires Pfam domain; BP3 requires absence of critical domain)
BP4
Computational evidence suggests no impact on gene or gene product
Supporting / Moderate
Conditions
- Path A (Missense - BayesDel): Requires consequence = missense_variant (v3.12.1, HELIX-CR-2026-009). BayesDel_noAF score evaluated with ClinGen SVI calibrated thresholds (Pejaver et al. 2022). Two strength levels: BP4_Moderate (<= -0.361, -2 points), BP4_Supporting (-0.360 to -0.181, -1 point). SpliceAI max_score must be < 0.1 or absent.
- Path B (Non-canonical splice - BP4_splice): v3.25.2 (HELIX-CR-2026-060). Supporting benign for non-canonical splice region variants (splice_donor_region, splice_donor_5th_base, splice_acceptor_5th_base, splice_region, splice_polypyrimidine_tract) when SpliceAI max_score <= 0.1. Excludes synonymous (BP7 handles) and missense (BP4 BayesDel handles). PVS1 guard: not applied when PVS1 condition is satisfied (no double-counting splice/LoF evidence). Threshold consistent with BP7 (SpliceAI <= 0.1). References: Jaganathan 2019 PMID:30661751, Walker 2023 PMID:37352859, de la Hoya 2022 PMID:35202600.
- Criteria string marker: BP4 (BayesDel path) and BP4_splice (non-canonical splice path) reported separately.
Exclusions
- BayesDel indeterminate range (-0.180 to 0.129): no BP4 missense path applied
- BP4_splice not applied when PVS1 is triggered (canonical splice +/-1,2 with constraint pass)
Databases: dbNSFP 4.9c (bayesdel_noaf_score), SpliceAI (max_score), VEP (consequence)
Known limitations
- BayesDel_noAF does not reach BP4_Strong per Pejaver calibration data
- SpliceAI guard prevents BP4 missense path for variants with any predicted splice impact
- BP4_splice covers only non-canonical splice region consequences. Deep intronic variants without VEP splice tags require RNA validation per ClinGen SVI (Walker 2023).
BP6
Reputable source reports variant as benign
Supporting
Conditions
- ClinVar clinical significance: Benign, Benign/Likely_benign, or Likely_benign
- ClinVar review stars >= 1
Databases: ClinVar (clinical_significance, review_stars)
Known limitations
- ClinGen SVI has recommended retiring BP6 as a standalone criterion; retained for maximum sensitivity
BP7
Synonymous variant with no predicted impact on splicing
Supporting
Conditions
- Consequence: synonymous variant
- Not in a splice region (consequence does not contain "splice_region")
- SpliceAI max_score <= 0.1 or absent
Databases: VEP (consequence), SpliceAI (max_score)
Known limitations
- Conservation filter intentionally omitted per Walker et al. 2023 Table S13 recommendation ("no improvement in negative predictive value" with conservation filter)
Aligned with ClinGen SVI 2023 (Walker et al.) Figure 4 decision tree for synonymous variant classification.
### Computational Predictors (PP3 / BP4)
PP3 and BP4 use BayesDel_noAF as the primary classification tool with ClinGen SVI calibrated thresholds (Pejaver et al. 2022). BayesDel is a Bayesian framework that integrates deleteriousness scores from multiple underlying predictors into a single calibrated score. The noAF variant (without allele frequency) is used to avoid circular reasoning with PM2, BA1, and BS1 frequency criteria.
BayesDel_noAF Score Evidence Strength ACMG Code Bayesian Points
>= 0.518 Strong pathogenic PP3_Strong +4
0.290 - 0.517 Moderate pathogenic PP3_Moderate +2
0.130 - 0.289 Supporting pathogenic PP3_Supporting +1
-0.180 to 0.129 Indeterminate None 0
-0.360 to -0.181 Supporting benign BP4_Supporting -1
<= -0.361 Moderate benign BP4_Moderate -2
PM1 + PP3 Double-Counting Guard
When PM1 (functional domain, Moderate = 2 points) applies alongside PP3_Strong (4 points), the combined evidence would exceed the ClinGen SVI recommended cap of Strong equivalent (4 points). In this case, PP3_Strong is downgraded to PP3_Moderate (2 points), yielding a combined 2 + 2 = 4 points. PP3_Moderate and PP3_Supporting are not affected by this cap.
Why BayesDel
The ClinGen SVI Working Group calibrated four computational tools (BayesDel, MutPred2, REVEL, VEST4) and demonstrated that a single calibrated tool with evidence strength modulation provides more accurate classification than a fixed-threshold multi-predictor consensus. BayesDel_noAF was selected because it explicitly excludes allele frequency from its model (avoiding circular reasoning with PM2/BA1/BS1), is precomputed in dbNSFP 4.9c, and reaches both PP3_Strong and BP4_Moderate in the Pejaver calibration - providing the widest evidence strength range among available tools.
Display Predictors (Not Used in Classification)
The following predictors are available alongside every variant for the reviewing geneticist to inspect, but they are not used in the PP3/BP4 classification logic. They were used in the weighted consensus approach prior to v3.4 and are retained for clinical reference:
SIFT AlphaMissense MetaSVM DANN PhyloP GERP
### SpliceAI Integration
Splice impact predictions are integrated following ClinGen Sequence Variant Interpretation (SVI) Working Group recommendations (Walker et al., 2023).
SpliceAI predicts the impact of each variant on mRNA splicing through four delta scores: acceptor gain (DS_AG), acceptor loss (DS_AL), donor gain (DS_DG), and donor loss (DS_DL). The maximum of these four scores is used for classification thresholds.
Scores come from Ensembl precomputed MANE transcript predictions and are not computed at runtime. The recorded prediction source supports reconstruction without a runtime dependency on external services.
Thresholds
- SpliceAI >= 0.8 → PP3_splice_Strong: Strong evidence for spliceogenicity (+4 points). Requires disease association gate (v3.15.1).
- SpliceAI 0.5 - 0.799 → PP3_splice_Moderate: Moderate evidence for spliceogenicity (+2 points).
- SpliceAI 0.2 - 0.499 → PP3_splice_Supporting: Supporting evidence for spliceogenicity (+1 point).
- SpliceAI < 0.1 → BP4 guard : Required for BP4 to apply. Prevents benign classification when splice impact is predicted.
- SpliceAI <= 0.1 → BP7: Required for synonymous variant BP7 classification. Confirms no splice impact.
PVS1 Double-Counting Guard
When PVS1 (loss-of-function) is triggered for a variant, PP3_splice is not applied. This prevents double-counting the same biological mechanism (splice disruption leading to loss of function) as both PVS1 and PP3 evidence, per ClinGen SVI recommendation.
Reference: Walker et al., Am J Hum Genet. 2023;110(7):1046-1067. PMID: 37352859
### Classification Logic
Classification uses the Bayesian point-based framework (Tavtigian et al. 2018, 2020) which is mathematically equivalent to the original 18 ACMG combining rules while providing proper classifications for evidence combinations not explicitly covered by the 2015 guidelines.
Bayesian Point System (Primary)
Each evidence criterion contributes points based on its strength level. The total determines classification:
Pathogenic Evidence Points
- Very Strong (PVS) +8
- Strong (PS) +4
- Moderate (PM) +2
- Supporting (PP) +1
Benign Evidence Points
- Stand-alone (BA1) Override to Benign
- Strong (BS) -4
- Supporting (BP) -1
Classification Thresholds
Pathogenic
>= 10 pts
Likely Path.
6 - 9 pts
VUS
0 - 5 pts
Likely Benign
-1 - -5 pts
Benign
<= -6 pts
High-Confidence Conflict Safety Check
When pathogenic evidence at Strong or Very Strong level conflicts with Strong benign evidence (BS), the variant is flagged for manual review regardless of the point total. This prevents automated resolution of genuinely conflicting high-quality evidence.
ACMG 2015 Combining Rules (Reference)
The original 18 ACMG 2015 combining rules are a special case of the Bayesian point system - every rule produces the same classification under both approaches. They are retained here as a reference.
Pathogenic (8 rules)
- P1 : 1 Very Strong (PVS) + >= 1 Strong (PS)
- P2 : 1 Very Strong (PVS) + >= 2 Moderate (PM)
- P3 : 1 Very Strong (PVS) + 1 Moderate (PM) + 1 Supporting (PP)
- P4 : 1 Very Strong (PVS) + >= 2 Supporting (PP)
- P5 : >= 2 Strong (PS)
- P6 : 1 Strong (PS) + >= 3 Moderate (PM)
- P7 : 1 Strong (PS) + 2 Moderate (PM) + >= 2 Supporting (PP)
- P8 : 1 Strong (PS) + >= 4 Moderate (PM)
Likely Pathogenic (7 rules)
- LP1 : 1 Very Strong (PVS) + 1 Moderate (PM) - gated by ar_biallelic_het_lof_guard (v3.25.5): blocked for het LoF in AR-only LoF genes without compound het partner (carrier state)
- LP1b : 1 Very Strong (PVS) or PVS1_Strong + >= 1 Supporting (PP) - ClinGen SVI PM2 downgrade accommodation (v3.20.0); same het LoF guard as LP1
- LP2 : 1 Strong (PS) + 1-2 Moderate (PM) - gated by computational-only LP guard (v3.25.6 CR-064 / v3.25.9 CR-067): blocked when sole Strong is PP3_Strong or PP3_splice_Strong with PM2 only and no PP4/PP5/PS1/PM3/PM4. v3.27.2 CR-081: bypass for PP3_splice_Strong in AD/XL LoF-intolerant genes with disease association and splice proximity.
- LP3 : 1 Strong (PS) + >= 2 Supporting (PP) - same het LoF guard as LP1
- LP4 : >= 3 Moderate (PM) - requires gene_has_disease_association (v3.16.7 CR-021); blocked by ar_biallelic_missense_guard (v3.25.3 CR-061), computational_only_lp_guard (v3.25.4 CR-062), ar_biallelic_het_lof_guard (v3.25.5 CR-063)
- LP5 : 2 Moderate (PM) + >= 2 Supporting (PP) - same disease association gate and three guards as LP4
- LP6 : 1 Moderate (PM) + >= 4 Supporting (PP) - same disease association gate and three guards as LP4
Benign (2 rules)
- B1 : 1 Stand-alone (BA1) - frequency > 5%
- B3 : >= 1 Very Strong benign (BS1_VeryStrong, -8 pts) - v3.33.0 (HELIX-CR-2026-094): a single Very Strong benign criterion is sufficient alone for Benign under the ClinGen SVI Bayesian framework (Tavtigian 2018, Pejaver 2022). Fires for the MYO15A and OTOF Hearing-Loss genes at the multi-tier BS1 Supporting band (ClinGen CSPEC GN023). Evaluated after B1 and before B2 (CASE order B1 -> B3 -> B2).
- B2 : >= 2 actual Strong benign (BS1, BS2) - v3.26.6 (BB-007): explicit Strong-only count. BP4_Moderate (Moderate benign per Pejaver 2022) is included in bs_count for LB1 participation but excluded from B2. Tavtigian 2020 (PMID:32720330): BS1(-4) + BP4_Moderate(-2) = -6 pts = LB, not B (requires <= -7 pts). v3.33.0 (HELIX-CR-2026-094): Strong count uses has_bs1_strong, excluding BS1_VeryStrong (handled by B3) to prevent double-counting.
Likely Benign (2 rules)
- LB1 : 1 Strong benign (BS or BP4_Moderate mapped to bs_count) + 1 Supporting benign (BP)
- LB1b : 1 actual Strong benign (BS1 or BS2) + BP4_Moderate - v3.26.6 (BB-007) explicit fallthrough: Tavtigian BS1(-4) + BP4_Moderate(-2) = -6 pts = LB. Prevents BS1 + BP4_Moderate (without other BP) from incorrectly falling to VUS.
- LB2 : >= 2 Supporting benign (BP)
Conflicting Evidence Handling
The Bayesian point system sums conflicting evidence. For example, PM2 (+2) and BS1 (-4) yield -2, or Likely Benign; v3.3 would have defaulted this combination to VUS. A direct conflict between Strong or Very Strong pathogenic evidence and Strong benign evidence still triggers manual review regardless of the point total (see High-Confidence Conflict Safety Check above). BA1 remains a higher-priority stand-alone override.
### VCEP Gene-Specific Specifications
ClinGen Variant Curation Expert Panels (VCEPs) adapt generic ACMG/AMP 2015 criteria to specific genes or diseases. Folklore implements approved VCEP specifications as an optional overlay on top of the standard classification.
How It Works
The standard classification pipeline runs first, producing a generic ACMG classification for all variants. For variants in genes with available VCEP specifications, gene-specific thresholds are applied via a lightweight overlay that modifies frequency cutoffs (BA1, BS1, PM2) and criterion applicability (PVS1) based on the published VCEP specification. The Bayesian point total is then recalculated with the modified criteria.
What VCEPs Modify
- BA1 : Gene-specific allele frequency threshold for stand-alone benign (e.g., 0.1% for Cardiomyopathy genes instead of generic 5%)
- BS1 : Gene-specific elevated frequency threshold, already mode-of-inheritance-aware
- PM2 : Gene-specific absent-in-controls threshold (e.g., 0% for RASopathy genes)
- PVS1 : Gene-specific applicability gate. Set to FALSE for gain-of-function genes (e.g., MYOC in glaucoma)
Toggle Behavior
VCEP overlay is enabled by default and can be disabled per case in case settings. When enabled, variants in VCEP-covered genes display an audit trail marker (e.g., "[VCEP:Hearing Loss v1.0]") in the criteria string.
Coverage
Approved VCEP specifications are sourced from the ClinGen Criteria Specification Registry (CSpec). Approximately 50-60 genes have published specifications, including panels for Hearing Loss, Cardiomyopathy (MYH7, MYBPC3), RASopathy (PTPN11, BRAF, SOS1), PTEN, CDH1, TP53, PAH, and BRCA1/BRCA2 (ENIGMA). For all other genes, generic ACMG 2015 thresholds are used.
Source: ClinGen Criteria Specification Registry (cspec.genome.network)
### ClinVar Override Logic
ClinVar clinical significance assertions are used as classification evidence, but only under specific conditions that prevent overriding computational evidence when conflicts exist.
When ClinVar override IS applied
ClinVar has a Pathogenic, Likely Pathogenic, Benign, or Likely Benign assertion with at least 1 review star, AND no conflicting computational evidence exists (no BA1, no conflicting pathogenic+benign criteria at moderate+ strength). For P/LP assertions with only 1 review star (single submitter), the gene must also have an established disease association via at least one of five sources: ClinVar P/LP (gene-level, requires 2+ review stars to prevent circular self-validation), Orphanet, ClinGen haploinsufficiency score, curated AR LoF gene list, or VCEP coverage (v3.10.0 disease association gate, v3.10.1 CTE fix). P/LP assertions with 2+ review stars bypass this gate. Benign/Likely Benign assertions are not gated.
When ClinVar override is NOT applied
BA1 applies (frequency > 5% always overrides ClinVar). OR conflicting evidence exists (pathogenic + benign criteria both triggered). OR ClinVar asserts VUS (VUS does not override computational classification). OR ClinVar review stars are below the minimum threshold. OR ClinVar P/LP with 1 review star in a gene without established disease association (v3.10.0 gate, v3.10.1 circular reference fix - the gene-level ClinVar disease check requires 2+ stars to prevent a 1-star submission from self-validating its own gate).
Default minimum review stars for override: 1 (configurable per deployment)
When ClinVar classification is used, the criteria string includes "ClinVar" as the first element (e.g., "ClinVar,PM2,PP3") to make the evidence source explicit. When a ClinVar P/LP override is blocked by the disease association gate (v3.10.0), the criteria string shows "ClinVar_gated" as the first element to indicate that the ClinVar assertion was present but not applied. ClinVar review star level is available alongside every variant for the reviewing geneticist to assess assertion quality.
### Manual Review Criteria (8 of 28)
These criteria require information that cannot be determined from a single-sample VCF file - family segregation data, functional study results, confirmed de novo status, or case-level clinical context. They must be evaluated by the reviewing geneticist.
PS2
De novo variant (confirmed paternity and maternity)
Requires trio sequencing data and confirmed parental relationships. Cannot be determined from single-sample VCF analysis.
PS3
Well-established in vitro or in vivo functional studies show a deleterious effect
Requires curation of published functional assay data. Automated literature extraction of functional evidence is not yet implemented.
PS4
Prevalence of the variant in affected individuals is significantly increased compared with controls
Requires case-control study data or odds ratios not available in standard annotation databases.
PM6
Assumed de novo without confirmation of paternity and maternity
Requires family structure information not available in single-sample analysis.
PP1
Cosegregation with disease in multiple affected family members
Requires multi-generational pedigree data and segregation analysis.
BS3
Well-established in vitro or in vivo functional studies show no deleterious effect
Requires curation of published functional assay data (benign counterpart of PS3).
BS4
Lack of segregation in affected members of a family
Requires family segregation data not available in single-sample analysis.
BP5
Variant found in a case with an alternate molecular basis for disease
Requires clinical case-level information about alternative diagnoses.
### Quality Filtering
Three configurable quality presets control the stringency of variant filtering. Quality filtering occurs before annotation and classification.
Preset Quality (QUAL) Depth (DP) Genotype Quality (GQ) Recommended Use
Strict >= 30 >= 20 >= 30 High-confidence clinical reporting
Balanced >= 20 >= 15 >= 20 Standard clinical analysis (default)
Permissive >= 10 >= 10 >= 10 Maximum sensitivity / research
ClinVar Rescue Mechanism
ClinVar Pathogenic or Likely Pathogenic variants that fail quality thresholds are flagged as rescued variants and continue through classification, keeping low-coverage findings available for review.
### Limitations and Disclaimers
- Folklore is a clinical decision support tool, not a diagnostic device. All classifications require review and confirmation by a qualified clinical geneticist.
- 8 of 28 ACMG criteria require information not available from single-sample VCF analysis (segregation, functional studies, de novo confirmation). These criteria must be evaluated manually by the reviewing geneticist.
- SpliceAI predictions are computational. RNA splicing studies remain the gold standard for confirming splice-altering effects.
- Population frequency data from gnomAD may underrepresent certain ethnic groups and geographic populations. Allele frequency thresholds should be interpreted in the context of the patient's ancestry.
- ClinVar assertions vary in quality and currency. Review star levels are displayed alongside all ClinVar-derived evidence to enable informed interpretation.
- SV/CNV processing is mandatory for every case. Constitutional copy-number losses and gains that meet the shipped applicability predicates are classified under the Riggs 2020 ClinGen/ACMG framework; typed INS, INV, BND, ambiguous CNV, and other non-applicable SV rows retain no nuclear ACMG, MMDWG, or Riggs verdict. The final v3.39.1 implementation makes the nuclear boundary explicit with sv_type IS NULL and deterministic mandatory SV routing (CR-2026-VA-MANDATORY-DETERMINISTIC-SV-PROCESSING-001). Repeat expansions are not classified by this pipeline.
- Mitochondrial variants are routed to a dedicated classification module implementing the MMDWG 2020 specifications. They are not approximated with nuclear ACMG rules; mitochondrial-specific criteria, evidence sources, and limitations are documented on the mtDNA methodology page.
- VCEP gene-specific specifications are implemented as a threshold overlay for BA1, BS1, PM2, and PVS1 applicability. The overlay does not implement VCEP-specific functional assay interpretation (PS3/BS3) or gene-specific segregation logic (PP1/BS4), which require manual curation. Coverage is limited to approximately 50-60 genes with published ClinGen CSpec specifications; all other genes use generic ACMG 2015 thresholds.
- PM5 uses a normalized ClinVar-derived pathogenic-codon asset and requires a different alternate amino acid at the same gene and codon with a P/LP assertion at 2+ review stars (final v3.39.1, CR-2026-PM5-ENABLEMENT-001). If that asset is unavailable, the classifier degrades safely to no PM5 evidence.
- Compound heterozygote detection is inferred from genotype data without long-read phasing or trio analysis. Final v3.39.1 expands splice-partner eligibility through a two-tier consequence model and releases carrier guards only when PM3 actually fires (CR-2026-COMPHET-SPLICE-PARTNER-001). Formal phasing should still be performed for clinical confirmation.
- The curated autosomal recessive LoF gene list (~150 genes) for PVS1 bypass covers genes with Definitive/Strong evidence for biallelic LoF disease mechanism. AR genes outside this list that lack constraint (low pLI, high LOEUF) will not trigger PVS1 unless they have VCEP gene-specific specifications.
- Homozygous reference genotypes (hom_ref) from multi-allelic VCF sites are excluded from classification. These represent alleles not present in the patient.
- Classification guards (v3.25.3 ar_biallelic_missense_guard, v3.25.4 computational_only_lp_guard, v3.25.5 ar_biallelic_het_lof_guard, v3.25.6 LP2 PP2 bypass fix, v3.25.9 LP2 PP3_splice_Strong extension, v3.26.6 B2 Strong-only count) downgrade Likely Pathogenic to VUS or Benign to Likely Benign when the evidence profile is computational-only or conflicts with the gene mechanism. Final v3.39.1 ties three carrier/computational guard releases to a validated PM3 trigger rather than a bare compound-heterozygote candidate flag. Subsequent functional or segregation evidence may support re-evaluation of a downgraded variant.
- gnomAD LoF tolerance signal (v3.24.0, HELIX-CR-2026-057): for HIGH impact variants outside MANE Select CDS in genes with documented homozygous or hemizygous LoF carriers in healthy gnomAD individuals, a BP_regional Supporting benign criterion is applied (v3.25.0, CR-057 Phase 3). AD genes are excluded - het LoF tolerance in AD is a different clinical argument. Compound het candidates are excluded - they may contribute to biallelic LoF.
- ClinGen negative evidence guard (v3.26.9, HELIX-CR-2026-078): genes with ClinGen Gene-Disease Validity classifications limited to Limited / Disputed / Refuted (no Definitive / Strong / Moderate) are no longer accepted as having disease association via Orphanet entry alone. This blocks ClinVar 1-star override and LP rules in genes with negative ClinGen evaluation, preventing single-submitter ClinVar promotion in genes that ClinGen has examined and found insufficient evidence.
- Results should always be interpreted in the context of the patient's clinical presentation, family history, and other available clinical information.
### Version History
Methodology changes are recorded under the corresponding production classifier version.
Show older versions ( 82 more ) v3.39.1 June-July 2026 Current
- Completed the BS1 recessive-LoF legitimacy guard for the HPO-AD fallback (HELIX-CR-2026-108 follow-up). Evidence-backed biallelic LoF genes now bypass both the pure-AD clauses introduced in v3.39.0 and the adjacent HPO-AD clause, closing the TSHR-class fall-through while retaining higher-priority VCEP and ClinGen-HI rules.
- Activated PM5 Moderate evidence from a normalized ClinVar pathogenic-codon asset (CR-2026-PM5-ENABLEMENT-001). A different P/LP missense alternate at the same gene and codon with 2+ review stars now contributes observed evidence; the final implementation applies PM5 whenever the asset is available and degrades safely when it is absent.
- Expanded compound-heterozygote splice partners and tightened carrier-guard semantics (CR-2026-COMPHET-SPLICE-PARTNER-001). Canonical/fifth-base/donor-region splice partners are accepted directly, broader splice-region/polypyrimidine candidates require SpliceAI or HIGH/MODERATE impact support, and guard release now requires PM3 to fire after ClinVar partner validation.
- Made the nuclear/SV boundary explicit in the final v3.39.1 code (CR-2026-VA-MANDATORY-DETERMINISTIC-SV-PROCESSING-001): nuclear ACMG evaluates only rows with sv_type IS NULL, while applicable constitutional losses and gains route deterministically to the Riggs methodology.
v3.39.0 June 2026
- Added an evidence-backed recessive-LoF legitimacy set built as the intersection of definitive/strong/moderate biallelic LoF entries in gene_disease_mechanism and autosomal-recessive support in GenCC (HELIX-CR-2026-108). Missing or empty reference data preserves the previous behavior.
- Guarded BS1 pure-autosomal-dominant 0.1% clauses for genes in that set so biologically legitimate recessive LoF variants can fall through to the 5% AR threshold instead of receiving overly strong benign frequency evidence (HELIX-CR-2026-108).
- Suppressed BS2 for those genes only when the variant also has canonical pathogenic signal - HIGH frameshift/stop-gained/canonical splice/start-lost consequence or qualifying ClinVar P/LP evidence - preventing healthy-carrier homozygote counts from being treated as blanket strong benign evidence (HELIX-CR-2026-108).
v3.38.0 June 2026
- Established and low-penetrance risk-allele suppression activated (HELIX-CR-2026-107). The classification staircase now assigns an established or low-penetrance risk allele to VUS at STEP -0.5, above BA1, instead of leaving it Benign at BA1-band frequency. The signal comes from ClinVar clinical significance ('risk_allele'), with a documented Factor V Leiden (ClinVar 642) carve-out. This widens the v3.33-era 102A predicate, which was inactive on the ERepo-only substrate. A risk allele cannot produce a P/LP classification; the assertion appears as a [RiskAllele] annotation.
v3.37.0 June 2026
- Inheritance- and penetrance-aware BS1/BS2 (HELIX-CR-2026-100): benign-frequency strength is gated by inheritance mode, including a G6PD X-linked-recessive guard. Conservative-for-benign.
- BP4_splice eligibility narrowed (HELIX-CR-2026-101): BP4_splice is withheld inside the conserved canonical donor/acceptor motif, where a low SpliceAI score is not sufficient evidence of benign splicing.
- Risk-allele SUPPRESS substrate introduced (HELIX-CR-2026-102A): the established / low-penetrance risk-allele suppress branch added to the staircase (inert until its ClinVar substrate is wired at v3.38.0).
- PVS1 AR loss-of-function reference coverage extended and audit re-gated (HELIX-CR-2026-103); PVS1 canonical-donor +5 position exclusion fix (HELIX-CR-2026-104). Cohort-3 directed under-call batch.
v3.36.0 May 2026
- ClinGen-AR completeness set supersedes the GenCC autosomal-recessive fallback (HELIX-CR-2026-096): a clean cutover of the BS1 AR-evidence gene set to the ClinGen-curated source, reclassifying SPG7.
- gnomAD VCF FILTER carried into the gnomAD asset (HELIX-CR-2026-099): BA1/BS1 are refused on a non-PASS gnomAD site, correcting the PRSS1 A16V frequency artifact.
v3.35.0 May 2026
- ERepo expert-panel P/LP override re-ordered to the TOP of the classification staircase (HELIX-CR-2026-093), above BA1 stand-alone and the conflict guard: an FDA-recognized VCEP P/LP assertion supersedes the mechanical frequency-benign branches (predicates unchanged; only the position moved).
- PVS1 LOEUF constraint threshold raised 0.35 -> 0.45 with the move to gnomAD constraint v4.1.1 (chrX/chrY constraint added), per the gnomAD v4.1.1 (2026-03-30) LoF-constraint recommendation.
- GenCC dual-mechanism BS1 inheritance fallback (HELIX-CR-2026-092).
v3.34.0 May 2026
- ClinGen ERepo expert-panel P/LP override inserted into the classification staircase (HEL-CR-2026-CLASSIFIER-EREPO-OVERRIDE-001, Phase B): an ERepo Pathogenic assertion yields P and an ERepo Likely Pathogenic assertion yields LP, placed above the ClinVar override and below the conflict guard. Variant-level join only; surfaced as an [ERepo] annotation. (Re-ordered to the top of the staircase at v3.35.0.)
v3.33.0 May 2026
- B3 Very Strong benign combining rule added (HELIX-CR-2026-094): >= 1 Very Strong benign criterion (BS1_VeryStrong, -8 pts) is sufficient alone for Benign under the ClinGen SVI Bayesian framework (Tavtigian 2018, Pejaver 2022). Evaluated in CASE order B1 -> B3 -> B2; B2 now counts has_bs1_strong to avoid double-counting Very Strong evidence.
- Multi-tier BS1 Supporting band for 7 Hearing-Loss autosomal-recessive genes (CDH23, GJB2, MYO15A, MYO7A, OTOF, SLC26A4, USH2A) per ClinGen CSPEC GN005/GN023; helena_curated BS1 source label.
v3.32.0 April 2026
- BS1 MOI threshold-selection fix for dual-mechanism (AD+AR) and X-linked genes (HELIX-CR-2026-091): inheritance-mode threshold selection stratified to the ClinGen Definitive/Strong autosomal-recessive evidence gene set (~226 genes), with an HPO autosomal-dominant guard. Subtractive for BS1; conservative-for-benign.
v3.31.0 April 2026
- BS1/BA1 effective frequency aligned to the gnomAD v4 Grpmax Filtering AF (af_grpmax) (HELIX-CR-2026-090), per ClinGen guidance to VCEPs on gnomAD v4 (March 2024): Grpmax Filtering AF is the allele frequency for BA1/BS1, and gnomAD intentionally excludes the bottlenecked AJ/FIN/AMI/MID populations from it. global allele frequency is retained as a fallback only, with a NULL/zero bottleneck guard; audited by the [af_grpmax] marker. Benign-only - cannot create new P/LP.
v3.30.0 April 2026
- PM1: UniProt residue-level evidence integration (HELIX-CR-2026-082). Three-tier PM1 evidence model replaces domain-level Pfam-only logic. Tier 1 (primary): UniProt SwissProt human reviewed proteome, 113,126 PM1-eligible features across 14,278 proteins (DISULFID, ACT_SITE, BINDING, MOD_RES, MOTIF). Tier 2 (fallback): InterPro Pfam critical domains (~50 domains, retained for VCEP-defined regions). Tier 3 (reserved): VCEP-specific PM1 region overrides (future). PM1 boolean unchanged (single Moderate strength), preserving existing combining rules.
- PM1 evidence density: 2,262x increase over v3.28 (50 hardcoded Pfam domains -> 113,126 expert-curated UniProt residues). Architecture follows Folklore Pattern A (Parquet as source of truth, DuckDB as disposable cache). hgvs_protein parsing handles NULL, empty, multi-transcript pipe-delimited, version-suffix-stripped ENSP, and missense-only consequence guard.
- Naked feature handling (Strategy D): UniProt features lacking /evidence= and /note= qualifiers (2,067 features, 1.8% of total) excluded from PM1 eligibility by default. Config flag pm1_uniprot_include_naked allows clinical override for specific cases (e.g., INSR Cys234Tyr disulfide bond Cys223-Cys234, naked DISULFID). Default conservative; opt-in for full UniProt coverage when curator inclusion in SwissProt is treated as curation evidence.
- Audit trail: pm1_evidence_source column populated for every PM1 hit (uniprot:DISULFID:223-234:P06213 or pfam:PF00069). Frontend variant detail panel displays human-readable evidence source with UniProt entry link. Replaces generic "PM1: critical functional domain" annotation.
- InterPro Pfam loader Pattern B remediation: load_interpro_pfam.py migrated from direct DROP TABLE + INSERT to Parquet-first architecture, matching all other reference datasets. CRITICAL_DOMAINS dict retained as manual override layer.
- Reference: Richards 2015 PMID:25741868 (PM1 definition), Walker 2023 PMID:37352859 (ClinGen SVI Section 3.3), UniProt Consortium 2023 PMID:36408920.
v3.28.0 April 2026
- Processing notes refactor (HELIX-CR-HEL-VA-2026-001): sentinel-guarded deduplication of classifier-generated processing notes. All 14 ACMG note categories prefixed with [ACMG] sentinel. Symmetric refactor in artifact detection stage with [ARTIFACT] sentinel. Eliminates stale note preservation across reclassification runs (Finding-1: Case 13 IFT140 missing Cat 4 bypass note) and N-fold note duplication on repeated classifier runs (Finding-2: AIRE sessions 11-18x duplication). Zero classification logic changes.
- Note structure invariant: [ACMG] and [ARTIFACT] note text must not contain "; " internally. MNV notes structurally safe (split/rejoin lossless for internal STRING_AGG).
- Scope: acmg_classifier.py (18 hunks), artifact_detection_stage.py (3 hunks), rerun_classifier_pipeline.py (1 hunk).
v3.27.2 April 2026
- LP2 splice bypass for AD/XL LoF genes (HELIX-CR-2026-081). LP2 computational-only guard (CR-042/064/067) extended with bypass for PP3_splice_Strong (SpliceAI >= 0.8) in autosomal dominant or X-linked dominant LoF-intolerant genes with established disease association. SpliceAI >= 0.8 in a LoF gene is qualitatively different from BayesDel missense PP3_Strong + PM2: high-confidence splice disruption prediction in an established LoF-mechanism gene approaches a functional LoF proxy.
- Bypass conditions (six mandatory): has_pp3_splice_strong AND NOT has_pp3_strong AND ad_lof_constraint_satisfied AND gene_has_disease_association AND splice_proximity_satisfied AND NOT outside_mane_region. MANE guard prevents bypass from reactivating LP through PP3_splice_Strong backdoor when PVS1 is blocked for positional reasons.
- ad_lof_constraint_satisfied excludes AR-only paths (carrier state, not haploinsufficiency), GoF/DN genes via gof_genes_exclusive + GOF_AD_GENES (analogous to PVS1 CR-022 v3.16.8), and dual-mechanism biallelic LoF genes that retain biallelic-only mechanism (CR-079 compatible: IFT140 retains AD path).
- splice_proximity_satisfied excludes pure missense (CR-066 handles dual-consequence), pure synonymous without splice VEP tag, UTR variants, and deep intronic without splice VEP (requires RNA per Walker 2023).
- Additive change: VUS to LP. Cannot create new B/LB or remove existing P/LP. Clinical trigger: HG2024_252 TCOF1 c.1279-6C>G VUS vs CG Pathogenic.
- Reference: Jaganathan 2019 PMID:30661751, Walker 2023 PMID:37352859, Abou Tayoun 2018 PMID:30192042, Richards 2015 PMID:25741868.
v3.27.1 April 2026
- Processing notes CASE priority fix (HELIX-CR-2026-080). Single first-match-wins CASE structure replaced with independent CASE WHEN concatenation. Variants matching multiple conditions now receive all applicable notes separated by semicolons.
- Cat 2 pext guard: Added NOT outside_mane_region guard to prevent conflicting notes when variant is both outside MANE and has pext data.
- Outer CASE guard: Checks if at least one condition matches before entering concatenation. ELSE preserves original processing_notes unchanged.
- Zero classification impact: no changes to acmg_class, acmg_criteria, or confidence_score. Additive notes only. Trigger: BBS9 p.Glu501Ter (Case 3, HG2022_042) - pext downgrade note was masking ar_biallelic_het_lof_guard note.
v3.27.0 April 2026
- Dual-mechanism AD/AR LoF guard bypass (HELIX-CR-2026-079). ar_biallelic_het_lof_guard (CR-063) extended with bypass for genes with documented monoallelic LoF mechanism. Guard was designed for AR-only LoF genes (AMACR, CFTR, GBA1) where het LoF is carrier state. Dual-mechanism genes (biallelic LoF + monoallelic LoF) have an AD pathway where het LoF is pathogenic.
- Two bypass paths: gene_disease_mechanism monoallelic LoF (definitive/strong/moderate) OR CLINVAR_LOF_AD_GENES Python constant (IFT140, GCM2). ClinGen AD Definitive/Strong path removed in v2 review: ClinGen AD moi does not distinguish LoF from GoF/DN.
- Additive change: VUS to LP for het LoF in dual-mechanism AD/AR LoF genes. Cannot create new B/LB. Cannot remove existing P/LP. Affected gene: IFT140 (3rd ADPKD gene per Senum 2022 PMID:34890546).
- Clinical trigger: IFT140 c.640del p.Val214CysfsTer55 (session 6692b88f). Phenotype: HP:0000107 (Renal cyst). Precedent: CR-037 (gof_genes_exclusive, v3.20.3) - analogous dual-mechanism bypass for PVS1/PM3 GoF guard.
v3.26.9 April 2026
- ClinGen negative evidence guard (HELIX-CR-2026-078). gene_has_disease_association Orphanet branch guarded by clingen_negative_genes CTE. When ClinGen has evaluated a gene AND all evaluations are non-actionable (Limited/Disputed/Refuted, no Definitive/Strong/Moderate), Orphanet entry alone is insufficient for gene_has_disease_association.
- clingen_negative_genes CTE: SELECT DISTINCT gene_symbol from ClinGen GDV where classification IN (Limited/Disputed/Refuted) AND gene NOT IN disease_associated_genes_clingen. ~480 genes. Materialized by DuckDB.
- Affects ClinVar 1-star override gate, LP4/LP5/LP6 disease gate, PVS1 disease gate, PP3_Strong disease gate, de novo projection P6/P7/P8 and LP4/5/6.
- Subtractive change: LP -> VUS for 1-star ClinVar override in genes with ClinGen negative-only evaluation. Cannot create new P/LP. 205 affected genes (Orphanet + ClinGen negative-only out of 480 total).
- Clinical trigger: TNC p.Arg108His LP [ClinVar,BP4] (PD2025_120). ClinGen Limited (Hearing Loss VCEP). BayesDel = -0.322 (benign). pLI = 0.0.
- Reference: Strande 2017 PMID:28552198, Richards 2015 PMID:25741868.
v3.26.8 April 2026
- BS2 inheritance-aware AR threshold (HELIX-CR-2026-074). AR-only genes (Orphanet has_ar=TRUE, has_ad=FALSE, or HI=30) with early-onset-only diseases use bs2_ar_homozygote_threshold (10) instead of generic (15).
- Onset guard from refdb.orphanet_disease_onset (HELIX-REF-005). EXISTS guard: gene must be in orphanet_disease_onset (no data = generic threshold). NOT EXISTS guard: gene must not have Adult/Elderly/All ages onset. Aligns with Richards 2015 Table 3: BS2 "with full penetrance expected at an early age".
- AD BS2 from gnomAD not implemented. Harrison 2019 PMID:31159682: gnomAD is general population, not healthy controls. APC VCEP (Curia 2023 PMID:37805481) excludes gnomAD for AD BS2. XL BS2 deferred (ac_hemi column not in annotation pipeline).
- Criteria string: BS2[AR-hom:N] for AR path, BS2 for generic path. Subtractive change: BS2 adds benign evidence only. Cannot create new P/LP. 2 P -> VUS via conflicting evidence (TGM5 p.Gly113Cys 14 hom, DUOX2 p.Phe966SerfsTer29 14 hom).
- Clinical trigger: HG2022_073 CBS p.Ala114Val P with 6 hom in gnomAD (CBS excluded by 6 < 10 threshold and Adult onset blocking onset guard).
v3.26.7 April 2026
- PS1 minimum review stars correction (HELIX-CR-2026-073). ps1_min_stars: 1 -> 2 in processing.yaml. YAML override from commit 39f97c7e (2025-12-16) allowed PS1 (Strong) at ClinVar 1-star, bypassing v3.10.0 disease association gate. Code default was already 2.
- PP5 restoration: review_stars = 1 ClinVar P/LP now correctly receives PP5 (Supporting) instead of PS1 (Strong). PP5 range (>= 1 AND < 2) no longer empty set.
- PM3 partner validation: partner.review_stars >= ps1_min_stars (shared config, identical semantics per Richards 2015 Table 3: both PS1 and PM3 require established pathogenic variant). 0 PM3 variants in production have partner with review_stars = 1 (verified 151 sessions).
- Subtractive change: PS1 removed from 1-star variants. Cannot create new P/LP. 3 LP -> VUS across 151 sessions (P4HA1, HLA-DRB1, NDUFAF7). All 3 genes without disease association in Orphanet/ClinGen/gene_disease_mechanism.
- Clinical trigger: HG2022_016 P4HA1 chr10:73043951 LP [PS1,PM2] - deep intronic in gene without disease association.
- Reference: Richards 2015 PMID:25741868 (Table 3, PS1/PM3 definitions), Walker 2023 PMID:37352859 (ClinGen SVI PS1 2-star minimum).
v3.26.6 April 2026
- B2 combining rule BP4_Moderate exclusion (Bug Bounty BB-007). B2 (>= 2 Strong benign) replaced with explicit (has_bs1 + has_bs2) >= 2. BP4_Moderate (Moderate benign, -2 pts Tavtigian) was incorrectly counted in bs_count for B2 firing, causing B2 (>= 2 Strong) to fire with 1 Strong (BS1) + 1 Moderate (BP4_Moderate). Richards 2015 Table 5 rule (ii): B2 requires >= 2 Strong benign. Tavtigian 2020 PMID:32720330: BS1(-4) + BP4_Moderate(-2) = -6 pts = LB, not B (requires <= -7).
- bs_count unchanged: BP4_Moderate remains in bs_count for LB1 participation (bs_count >= 1 AND bp_count >= 1). Only B2 check uses explicit Strong-only count.
- LB1b fallthrough added: 1 actual Strong benign (BS1/BS2) + BP4_Moderate = LB. Prevents BS1 + BP4_Moderate (without other BP) from falling to VUS.
- Subtractive change for B: B -> LB when BS1 + BP4_Moderate was the sole B2 trigger. Cannot create new P/LP. Cannot remove existing LB.
- Reference: Richards 2015 PMID:25741868 (Table 5 rule ii), Tavtigian 2020 PMID:32720330, Pejaver 2022 PMID:36413997.
v3.26.5 April 2026
- PM2 af_grpmax = 0.0 edge case fix (Bug Bounty BB-006). PM2 af_grpmax path: af_grpmax = 0.0 now treated as absent from controls, consistent with global_af path (v3.21.2, CR-050). gnomAD records AC=0 variants with af_grpmax=0.0 (NOT NULL). At pm2_threshold=0.0 (RASopathy "absent from controls"), 0.0 < 0.0 is false, blocking PM2.
- Affects only VCEP genes with pm2_frequency_field=af_grpmax AND pm2_threshold=0.0 (currently RASopathy genes). Neutral/additive change: restores PM2 for variants incorrectly blocked. Cannot remove existing P/LP.
- Reference: HELIX-CR-2026-050 (v3.21.2 global_af fix), HELIX-CR-2026-047 (v3.21.0 af_grpmax path), Richards 2015 PMID:25741868 (PM2 absent from controls).
v3.26.4 April 2026
- De novo projection compound het + hemizygous (Bug Bounty BB-003). compound_het_candidate guard removed from de_novo_projection outer CASE. De novo and compound het are compatible: allele 1 de novo + allele 2 inherited from carrier parent = legitimate compound het. PS2 (origin) and PM3 (biallelic configuration) are orthogonal evidence types.
- Hemizygous genotype support: rv.genotype IN (het, hom_alt) with chromosome guard (X/chrX/Y/chrY). Hemizygous males on X/Y correctly eligible for de novo projection. Autosomal hom_alt excluded.
- Dead LP rules documentation: LP1/LP1b/LP4/LP5/LP6 annotated as unreachable in de novo context (ps_count+1 >= 1 tautological, P-rules pre-empt). Retained for ACMG completeness.
- Additive change: new de_novo_candidate=true for compound het VUS and hemizygous X/Y VUS. No change to acmg_class (de novo projection is prospective only).
- Reference: Richards 2015 PMID:25741868 (PS2, PM3 definitions), Abou Tayoun 2018 PMID:30192042.
v3.26.3 April 2026
- P8 combining rule fix (Bug Bounty BB-001). P8 combining rule: ps_count >= 1 AND pm_count >= 4 replaced with ps_count >= 1 AND pm_count >= 1 AND pp_count >= 4. Richards 2015 Table 5 rule (viii): >=1 Strong AND >=1 Moderate AND >=4 Supporting.
- Previous implementation was dead code (subsumed by P6: ps>=1 AND pm>=3) and encoded a non-existent rule. De novo projection P8: same fix applied to de_novo_projection CTE. Docstring P8 description corrected.
- Subtractive/neutral change: dead code replaced with correct rule. Profile 1S+1M+4PP rare in current 19-criteria implementation (PS1 triggers ClinVar override before ACMG scoring). No classification changes expected in simulation.
- Reference: Richards 2015 PMID:25741868 (Table 5 rule viii), Tavtigian 2018 PMID:29300386 (Bayesian validation: 4+2+4=10 pts = P).
v3.26.0 April 2026
- De novo projection AF ceiling + post-projection conflict guard (HELIX-CR-2026-071). De novo AF ceiling: de_novo_projection CTE filters by AF. global_af > 0.001 OR af_grpmax > 0.001 blocks de novo projection. PS2 (de novo) presupposes variant not inherited; AF > 0.001 implies ~150+ carriers in gnomAD v4.1 (152K genomes). Config parameter de_novo_af_ceiling (default 0.001).
- Post-projection conflict guard: bs_count > 0 now blocks de novo projection unconditionally. Previous guard checked pre-projection conflict (pvs/ps/pm > 0 AND bs > 0), missing cases where PS2 projection itself creates the conflict.
- Subtractive change: removes false de novo candidates. Cannot create new P/LP or new de novo candidates.
- Clinical trigger: PEX2 intron_variant HG2022_079 (AF=0.0039, false de_novo_candidate=true).
v3.25.9 April 2026
- LP2 computational-only guard PP3_splice_Strong extension (HELIX-CR-2026-067). LP2 computational-only guard (CR-042/064) extended: has_pp3_strong replaced with (has_pp3_strong OR has_pp3_splice_strong). PP3_splice_Strong + PM2 without observed evidence now blocked, same as PP3_Strong + PM2.
- Subtractive change: LP -> VUS. Cannot create new P/LP. De novo projection unchanged (PS2 is observed, LP2 legitimate).
- Clinical trigger: IRF1 p.Arg123Gly LP [PP3_splice_Strong,PM2] (CR-065/066 simulation, computational-only profile).
v3.25.8 April 2026
- PP3 + PP3_splice mutual exclusion (HELIX-CR-2026-066). PP3 (BayesDel, all levels: Strong, Moderate, Supporting) blocked for dual-consequence variants where VEP consequence includes splice_region, splice_donor, or splice_acceptor in addition to missense. ClinGen SVI (Walker 2023 Section 3.2): PP3 is single criterion. SpliceAI (PP3_splice) is the more specific tool when VEP confirms splice proximity.
- Symmetric with CR-065: CR-065 blocks PP3_splice for pure missense (no splice VEP). CR-066 blocks PP3 for missense WITH splice VEP. Together: exactly one PP3 path per variant, never both.
- Subtractive change: PP3 removed from dual-consequence variants where PP3_splice also applies. Cannot create new P/LP. Clinical trigger: CR-065 simulation IRF1 p.Arg123Gly VUS to LP (PP3_Supporting + PP3_splice_Strong + PM2, double-counting).
v3.25.7 April 2026
- PP3_splice missense-tolerated guard (HELIX-CR-2026-065). PP3_splice (all levels: Strong, Moderate, Supporting) blocked for pure missense variants (no VEP splice_region/splice_donor/splice_acceptor consequence) when BayesDel noAF data is available. ClinGen SVI (Walker 2023 Section 3.2): when BayesDel indicates tolerated/indeterminate and SpliceAI predicts cryptic splice, evidence is conflicting, not additive.
- Subsumes CR-043 (v3.20.8) double-counting guard. New guard covers both BayesDel >= 0.130 (CR-043 case) and BayesDel < 0.130 (CR-065 case) in single condition. BayesDel IS NULL bypass preserved (SpliceAI as sole in silico evidence). Missense + splice VEP consequence bypass preserved (VEP confirms splice proximity).
- Subtractive change: PP3_splice removed from pure missense with BayesDel. Cannot create new P/LP.
- Clinical trigger: PD2025_101 DICER1 c.5524A>G p.Ile1842Val LP [PP3_splice_Strong,PM2,PP2] vs ClinVar Expert Panel VUS (3 stars).
v3.25.6 April 2026
- LP2 computational-only guard PP2 bypass fix (HELIX-CR-2026-064). LP2 combining rule guard (CR-042) extended: pp_count = 0 condition replaced with (pp_count = 0 OR (pp_count > 0 AND NOT has_pp4 AND NOT has_pp5)). PP2 (gene-level constraint) no longer bypasses the PP3_Strong+PM2 computational-only LP guard. PP4 (HPO match, observed) and PP5 (ClinVar 1-star, observed) still override the guard legitimately.
- Subtractive change: LP -> VUS. Cannot create new P/LP. De novo projection unchanged: PS2 + PP3_Strong = ps_count >= 2 -> P5.
- Clinical trigger: PD2025_002 MEN1 p.Arg115His LP [PP3_Strong,PM2,PP2] vs ClinVar VUS (2 stars, 70 P/LP missense in gene).
v3.25.5 April 2026
- Het LoF LP guard for AR/biallelic-only LoF genes (HELIX-CR-2026-063). All LP combining rules (LP1-LP6) blocked for het LoF (HIGH impact) variants in AR/biallelic-only LoF genes without compound het partner. Het LoF in AR-only gene = carrier state, not pathogenic event. Richards 2015 Table 3: PVS1 presupposes pathogenic context.
- Guard logic: ar_biallelic_het_lof_guard boolean. True when genotype=het, impact=HIGH, compound_het_candidate=false/NULL, AND gene in gene_disease_mechanism (LoF, biallelic, definitive/strong/moderate) OR (Orphanet AR-only AND gene_disease_mechanism LoF). P rules NOT blocked: P1 with ClinVar PS1 is legitimate override.
- De novo projection: blocked by guard. pLI~0 + LOEUF>1 = het LoF tolerated in gnomAD. De novo het LoF not informative for AD mechanism.
- Interaction: orthogonal to CR-061 (missense), CR-062 (computational). Subtractive change: LP -> VUS. Cannot create new P/LP.
- Clinical trigger: PD2025_084 AMACR c.857del LP [PVS1_Strong,PM2] vs ClinVar VUS (2 stars). AMACR: pLI=0.000001, LOEUF=1.108, AR-only Orphanet, biallelic LoF definitive G2P.
v3.25.4 April 2026
- Computational-only LP guard for LP4/LP5/LP6 (HELIX-CR-2026-062). LP4/LP5/LP6 combining rules blocked when ALL pathogenic criteria are computational/annotation-based: PM1 (domain), PM2 (frequency), PP2 (gene constraint), PP3 (BayesDel, any level), PP3_splice (SpliceAI, any level). At least one observed/clinical criterion required: PS1 (ClinVar observed), PM3 (compound het), PM4 (inframe indel), PP4 (phenotype match), or PP5 (ClinVar low confidence).
- De novo projection: LP4/LP5/LP6 NOT gated by this guard. PS2 (de novo) is observed evidence - if confirmed, LP is legitimate even with computational moderates.
- Subtractive change: cannot create new P/LP. LP -> VUS when entire evidence profile is computational. 0 impact on PVS1-based LP, PP3_splice LP, ClinVar override LP, compound het LP, or PM4 LP.
- Clinical trigger: PD2025_106 JAG1 p.Gly309Arg LP [PM1,PM2,PP2,PP3_Supporting] vs ClinVar VUS (2 stars).
v3.25.3 April 2026
- Monoallelic missense LP guard for AR/biallelic LoF genes (HELIX-CR-2026-061). PP2 blocked for het missense in AR/biallelic LoF genes without documented missense disease mechanism. Richards 2015 Table 3: PP2 requires "missense variants are a common mechanism of disease". AR LoF genes without ClinVar P/LP missense (clinvar_missense_genes) do not satisfy this premise.
- LP4/LP5/LP6 combining rules blocked (ar_biallelic_missense_guard) for het missense variants in AR/biallelic LoF genes without compound het partner and without ClinVar P/LP missense evidence.
- De novo projection: LP4/LP5/LP6 blocked by same guard. LP2 (PS2 + PM) intentionally NOT blocked - de novo is direct phenotypic observation, not computational prediction.
- Subtractive change: cannot create new P/LP. ~44 LP -> VUS (het missense in AR/biallelic LoF genes). 0 impact on PVS1-based LP, PP3_splice LP, ClinVar override LP, or compound het candidates.
- Clinical trigger: PD2025_127 ABCA2 p.Arg1064Cys LP [PM1,PM2,PP3_Moderate,PP2] vs ClinVar VUS (2 stars, multiple submitters, no conflicts).
v3.25.2 April 2026
- BP4_splice_benign for non-canonical splice (HELIX-CR-2026-060). Supporting benign for non-canonical splice region variants with SpliceAI <= 0.1. Covers splice_donor_region, splice_donor_5th_base, splice_acceptor_5th_base, splice_region, and splice_polypyrimidine_tract consequences. Excludes synonymous (BP7 handles) and missense (BP4 BayesDel handles).
- SpliceAI threshold <= 0.1: consistent with BP7 synonymous guard. Jaganathan 2019 PMID:30661751, de la Hoya 2022 PMID:35202600.
- PVS1 guard: NOT pvs1_cond (pre-split, covers both full and last-exon). Same guard as PP3_splice paths. No double-counting splice/LoF evidence. Criteria string: BP4_splice (Supporting benign).
- Subtractive change: adds benign evidence only. Cannot create new P/LP. VUS with BS1/BS2 + BP4_splice = LB (LB1 rule). VUS without Strong benign remain VUS (1 BP insufficient for LB2).
- Clinical trigger: PD2025_087 HNF4A c.426+4C>T (splice_donor_region, SpliceAI=0.0, VUS without criteria). Geneticist flag by D. Maneva.
v3.25.1 April 2026
- PM3 ClinVar partner validation guard (HELIX-CR-2026-058). compound_het_candidate = true alone no longer satisfies PM3. Partner must have ClinVar P/LP (clinical_significance LIKE Pathogenic% OR Likely_pathogenic%, NOT LIKE %Conflicting%) with review_stars >= ps1_min_stars (currently 2). Richards 2015 Table 3: detected in trans with a pathogenic variant - partner pathogenicity is explicit.
- ClinVar-only guard (no acmg_class dependency): avoids circular reference where partner classification depends on PM3 which depends on partner classification. ClinVar annotation is external/immutable.
- Subtractive change: PM3 can only be removed, not added. 241 LP -> VUS expected (PM3 was decisive). 5 LP unchanged (validated partner). De novo projection inherits PM3 change via pm3_cond.
- Clinical trigger: HG2022_037 F5 p.Met1874Thr LP [PM2,PM3,PP3_Moderate,BP2] vs ClinVar VUS (2 stars). Partner: Factor V Leiden (p.Arg534Gln, AF=2.14%, ClinVar drug_response 3 stars). PM3 incorrectly fired without partner pathogenicity validation.
- Impact: 188 sessions, 52,929 PM3 variants, 241 LP -> VUS, 0 new P/LP.
v3.25.0 April 2026
- BP_regional Phase 3 - gnomAD LoF population tolerance as Supporting benign (HELIX-CR-2026-057 Phase 3). Supporting benign (bp_count) for HIGH impact variants outside MANE CDS where gnomAD demonstrates regional LoF tolerance in healthy individuals. Builds on v3.24.0 informational signal.
- AD genes excluded - het LoF tolerance in AD is a different clinical argument. Compound het candidates excluded - they may contribute to biallelic LoF.
- Conditions: gls.gene_symbol IS NOT NULL, position outside MANE CDS, impact=HIGH, NOT compound_het_candidate, inheritance_mode IS NULL OR IN (AR, AD/AR, XL). Autosomal: max_nhomalt >= 2. X-linked: max_ac_hemi >= 2.
- Criteria string marker: BP_regional[gnomAD_LoF].
v3.24.0 April 2026
- gnomAD LoF population tolerance signal (HELIX-CR-2026-057 Phase 1). Boolean gnomad_lof_tolerant: gene has HC LoF variants with hom/hemi > 0 outside MANE CDS in gnomAD genomes. Phase 1 informational only, no ACMG criteria change.
- Reference data: gnomad_lof_tolerant_gene_summary table loaded. Aggregates max_nhomalt, max_ac_hemi, lof_variant_count per gene with example_variant audit trail.
- Activated by mt.gene_symbol IS NOT NULL (positional context mandatory) and position outside MANE CDS region.
- Phase 2 (CR-057 Phase 3, v3.25.0): gnomad_lof_tolerant signal is converted into BP_regional Supporting benign criterion.
v3.23.0 April 2026
- PVS1: Expression-aware guard using gnomAD v4.1 pext data (GTEx v10, GRCh38, 49 tissues). Per-exon max_tissue_pext modulates PVS1 application (HELIX-CR-2026-056, Cummings et al. 2020 PMID:32461655). Three tiers: max_tissue_pext >= 0.9 = PVS1 Very Strong (unchanged), 0.1-0.9 = PVS1_Strong[pext] (downgraded from Very Strong to Strong), < 0.1 = PVS1 blocked (exon not expressed). NULL pext = Very Strong (conservative fallback). PM4 receives binary pext guard (blocked at < 0.1).
- PVS1: 189,856 exon records for 18,923 MANE Select genes loaded from gnomAD v4.1 base-level pext TSV. Distribution: 56.3% full (>= 0.9), 42.5% downgrade (0.1-0.9), 1.2% block (< 0.1). Uses max_tissue_pext (maximum across 49 GTEx v10 tissues) to avoid false negatives for tissue-specific genes.
- PVS1: Human-readable annotations in acmg_criteria ("[PVS1 blocked: low exon expression (pext X.XX)]", "[PVS1 downgraded: partial exon expression (pext X.XX)]") and processing_notes (full sentence with gene name, threshold, source).
- Reference databases: gnomAD pext (v4.1, GTEx v10) and MANE Select transcripts (v1.4) added as reference data sources.
- Scientific basis: Cummings 2020 demonstrated that expression-based annotation selectively filters 22.8% of falsely annotated pLoF variants while removing less than 4% of true pathogenic variants.
v3.22.0 April 2026
- PVS1: MANE Select positional guard (HELIX-CR-2026-055, Abou Tayoun 2018 PMID:30192042). Variants outside the MANE Select/Plus Clinical CDS region do not receive PVS1 or PM4. Addresses false PVS1 application on variants annotated by VEP against non-canonical transcripts that extend beyond the canonical coding sequence.
- PVS1: 19,354 MANE transcripts (19,288 Select + 66 Plus Clinical) with CDS coordinates loaded from MANE GFF3 v1.4. LEFT JOIN with GROUP BY prevents row duplication for 66 genes with both Select and Plus Clinical transcripts.
- PVS1: Inline NULL-safe guards on 6 sites (pvs_count, ps_count/pvs1_last_exon, pm_count/pm4, has_pvs1_full, has_pvs1_last_exon, has_pm4). NULL MANE data = PVS1 applied (conservative fallback for genes not yet in MANE).
- Criteria string: [PVS1_BLOCKED:outside_MANE_CDS] marker appended for HIGH impact variants outside MANE CDS. Processing notes: PVS1_BLOCKED:outside_MANE_CDS(gene,mane_gene) audit trail.
- Clinical trigger: HG2022_057 FANCB chrX:14797196 p.Gln864Ter classified LP (PVS1+PM2) against non-canonical transcript ENST00000696353. Position is 46 kb upstream of MANE Select CDS (14,843,567-14,865,510). Reclassified to VUS (PM2 only).
v3.21.4 March 2026
- CLCN1 removed from AR_LOF_GENES (HELIX-CR-2026-053). Myotonia congenita has both AD (Thomsen) and AR (Becker) forms. Het LoF can cause AD Thomsen disease. CLCN1 receives AD/AR via dual HPO check.
v3.21.3 March 2026
- Annotation Phase 7.5: dual HPO inheritance detection (HELIX-CR-2026-051). Genes with both HP:0000006 (AD) and HP:0000007 (AR), confirmed by gene_disease_mechanism biallelic, now receive AD/AR instead of AD. 312 genes annotation improved.
- CYP11A1 added to AR_LOF_GENES. Congenital adrenal insufficiency, AR. OMIM #613743.
v3.21.0 March 2026
- PM2 frequency field alignment for VCEP FAF thresholds (HELIX-CR-2026-047). VCEP genes with pm2_frequency_field=af_grpmax now compare PM2 against af_grpmax instead of global_af. K1 bottlenecked NULL guard: af_grpmax NULL + global_af NOT NULL = PM2 not satisfied.
- PM2 criteria string audit trail: PM2[af_grpmax] and PM2_Supporting[af_grpmax] markers.
v3.20.9 March 2026
- BP1: reference-based missense mechanism guard (HELIX-CR-2026-044). BP1 blocked when gene is in refdb.clinvar_missense_genes (>= 2 ClinVar P/LP missense at 2+ stars). Extends session-local ClinVar guard with reference-level check.
v3.20.8 March 2026
- PP3_splice missense double-counting guard (HELIX-CR-2026-043). PP3_splice blocked when variant already receives PP3 from BayesDel missense path. ClinGen SVI: PP3 is a single criterion.
v3.20.6 March 2026
- BP1: GoF/DN gene guard (HELIX-CR-2026-039v2). BP1 blocked in GoF/DN genes where truncating variants do not cause disease via LoF. PLIN1 added to GOF_AD_GENES (DN mechanism, FPLD4).
v3.20.3 March 2026
- PVS1/PM3: Dual-mechanism gene guard (HELIX-CR-2026-037). gof_genes_exclusive view excludes 55 dual-mechanism genes from GoF guard, restoring PVS1/PM3 for AR LoF pathway.
v3.20.2 March 2026
- PM3: GoF/DN gene exclusion (HELIX-CR-2026-036). Compound het in pure GoF/DN genes no longer satisfies PM3. PM3 presupposes AR mechanism.
v3.18.2 March 2026
- Disease association gate: gene_disease_mechanism table added as 7th source in pvs1_disease_gate and gene_has_disease_association (HELIX-CR-2026-026). 256 genes in gene_disease_mechanism (G2P 225, GoFCards 35, manual 2) were not covered by existing 6 disease gate sources. Confidence filter: definitive/strong/moderate only (predicted excluded per CTO directive). Clinical trigger: SNRPN c.-391+1_-391+2insGA false VUS after v3.18.1 HI fix.
v3.18.1 March 2026
- Disease association gate: haploinsufficiency_score IS NOT NULL replaced with IN (2, 3, 30) in pvs1_disease_gate and gene_has_disease_association (HELIX-CR-2026-025). HI=0 (no evidence), HI=1 (little evidence), HI=40 (dosage sensitivity unlikely) no longer satisfy disease gate. 475 genes lost false disease association, 117 with no other disease source. Clinical trigger: SLC22A18 c.-133+2T>C false LP in gene without Mendelian disease.
v3.18.0 March 2026
- PVS1 GoF/DN guard: gof_genes_g2p view (198 genes) replaced by gof_genes_unified view (347 genes) as primary PVS1 GoF exclusion source (HELIX-REF-003 Phase 2A). gof_genes_unified draws from 3 sources: G2P (199 genes), GoFCards (179 genes), manual curation (27 genes). Dual-mechanism genes correctly excluded. GOF_AD_GENES Python constant retained as fallback.
- Reference database: gene_disease_mechanism table created in reference_db (2,617 records). Unified schema for molecular mechanism data from G2P, GoFCards, and manual curation. Columns: gene_symbol, disease_id, disease_name, mechanism (LoF/GoF/DN/GoE), confidence, source, allelic_requirement, pmid, curation_date, curator, last_reviewed, notes.
- Reference database: GoFCards Release 1.0 integrated (579 genes, 3,161 curated GoF variants). Gene-level aggregation threshold: 3+ variants, Pscore >= 2.0, at least 1 variant with animal or cell model evidence = confidence "strong". 217 genes qualified.
- Reference database: ClinGen Dosage Sensitivity full TSV integrated (1,626 genes, 23 columns). Replaces legacy 3-column clingen table. Backward-compatible clingen table recreated from full data.
- Reference: GoFCards PMID:39578693, ClinGen clinicalgenome.org, HELIX-REF-003.
v3.17.2 March 2026
- PVS1: 12 GoF/DN/GoE AD genes added to GOF_AD_GENES (HELIX-REF-003 Phase 2B). Systemic gap analysis: 28 AD genes depend ONLY on ClinGen AD Definitive path for PVS1, without mechanism check. 13 identified as HIGH risk (established non-LoF mechanism). TRAF7 deferred (insufficient germline functional evidence).
- Genes added: CALM2 (DN calmodulinopathy), CALM3 (DN calmodulinopathy), DNAJB6 (DN myofibrillar myopathy), GJB6 (DN deafness, ClinGen AR=Refuted), GREM1 (regulatory GoE), HSPB8 (toxic GoF CMT2L), INF2 (DN formin FSGS5), NEFL (DN neurofilament CMT), PRKAG2 (GoF AMPK activation), SKI (DN Shprintzen-Goldberg), TNNC1 (GoF Ca2+ sensitivity HCM), TUBB3 (DN microtubule CFEOM3A).
- Cross-referenced with LoGoFunc (variant-level GoF/LoF predictions), GoFCards (curated GoF variants), ClinGen Dosage (HI=0 signal), and G2P mechanism data.
v3.17.1 March 2026
- PVS1: ANKRD26 and C1QTNF5 added to GOF_AD_GENES (HELIX-CR-2026-024). ANKRD26: Thrombocytopenia 2, gain-of-expression mechanism - 5UTR mutations disrupt RUNX1/FLI1 silencing. ClinGen Dosage explicitly states "gain-of-function rather than haploinsufficiency". C1QTNF5: Late-onset retinal degeneration, dominant-negative mechanism - mutant destabilizes wildtype in gC1q domain co-expression.
- Clinical trigger: PD2025_095 ANKRD26 chr10:27044191 classified P [PVS1,PM2,PM3], should be VUS. C1QTNF5 chr11:119339479 received PVS1_Strong for frameshift despite all known pathogenic being missense.
v3.17.0 March 2026
- PVS1: G2P molecular mechanism integration (HELIX-CR-2026-023). DECIPHER Gene2Phenotype (G2P) database now provides primary GoF/DN guard for PVS1. gof_genes_g2p view contains 198 pure GoF/DN monoallelic genes where PVS1 is blocked. G2P is queried at classification time via reference_db ATTACH.
- PVS1: GOF_AD_GENES reduced from 36 to 14 fallback genes (genes not in G2P: 6 neurodegeneration/toxic aggregation, 3 somatic/neomorphic, 5 absent from G2P 2026-02-28). PRSS1 added to fallback (GoF trypsinogen autoactivation, clinical trigger PD2025_090).
- PVS1: GNAS removed from GOF_AD_GENES (dual-mechanism: McCune-Albright = GoF, pseudohypoparathyroidism 1A = LoF). G2P view correctly excludes GNAS. GNAS LoF variants now receive PVS1.
- PVS1: 49 dual-mechanism genes (SCN5A, LMNA, KCNH2, KCNQ1, FGFR1, etc.) excluded from GoF view, preserving PVS1 for legitimate LoF phenotypes.
- PVS1: G2P confidence filter: only definitive, strong, moderate included. Limited confidence excluded to prevent blocking PVS1 on weak GoF evidence.
- PVS1: Guard order: G2P view (primary) -> GOF_AD_GENES (fallback). Identical pattern to refdb.ar_lof_genes_clingen and refdb.disease_associated_genes_clingen.
- DECIPHER G2P: Molecular mechanism data (g2p_gene_mechanism table, 2,372 records) added to reference_db alongside existing HPO enrichment data (v3.12.0).
- Loader script: scripts/load_g2p_mechanism.py with dry-run, force flags, and 9-point validation (total records, mechanism distribution, view count, dual-mechanism exclusion, spot checks).
- Subtractive change: can only remove PVS1 from GoF/DN genes. Cannot create false P/LP.
- Clinical trigger: PD2025_090 PRSS1 p.Gly177Ter (stop_gained). GoF gene, PVS1 scientifically incorrect. BA1 protected in this case (AF 41%), but rare PRSS1 truncating variants without BA1 could reach LP[PVS1,PM2] incorrectly.
- Reference: Abou Tayoun 2018 PMID:30192042, Josephs 2023 PMID:37884986, HELIX-CR-2026-023.
v3.16.8 March 2026
- PVS1: GoF AD gene exclusion (HELIX-CR-2026-022). GOF_AD_GENES curated list (36 genes) blocks PVS1 for autosomal dominant genes where the disease mechanism is gain-of-function, dominant-negative, or toxic aggregation. ClinGen PVS1 Decision Tree Node 1: "Is LoF a known mechanism of disease?" - if NO, PVS1 not applicable.
- PVS1: GOF_AD_GENES categories: RAS/MAPK pathway (12 genes), receptor tyrosine kinases (10), PI3K/AKT/mTOR (2), GTPases (3), metabolic enzymes (2), kinases (2), neurodegeneration (6). Each gene documented with PMID for GoF mechanism.
- PVS1: ClinGen AD Definitive/Strong genes retain PVS1 access even if in GOF_AD_GENES - VCEP and ClinGen disease validity take precedence.
- Clinical trigger: PD2025_082 RAC1 p.Val14SerfsTer4 (frameshift in GoF gene). SOD1 discordance in same session.
- Reference: Abou Tayoun 2018 PMID:30192042, HELIX-CR-2026-022.
v3.16.7 March 2026
- LP classification: Disease association gate for LP rules LP4, LP5, LP6 (HELIX-CR-2026-021). LP rules that rely solely on moderate and supporting criteria now require established disease association via at least one of five sources (same as PVS1 and PP3_Strong gates).
- LP4 (>=3 PM), LP5 (2 PM + >=2 PP), LP6 (1 PM + >=4 PP): these rules can be satisfied by PM1+PM2+PM4 or similar combinations without any gene-disease evidence. Gate prevents LP classification in genes without known Mendelian disease mechanism.
- LP1, LP2, LP3 not gated: they require PVS or PS criteria, which already have their own disease association gates.
- Subtractive change: can only downgrade LP to VUS, cannot upgrade.
v3.16.6 March 2026
- PVS1: ClinGen AD Definitive/Strong genes bypass pLI/LOEUF constraint gate. Genes with ClinGen Gene-Disease Validity rating of Definitive or Strong for autosomal dominant inheritance now qualify for PVS1 regardless of population constraint metrics.
- PVS1: This addresses genes like TUBB1 (pLI near 0, LOEUF=1.586) with ClinGen AD Definitive that were excluded from PVS1 by the population constraint gate.
v3.16.0 March 2026
- ClinGen Gene-Disease Validity integration: ar_lof_genes_clingen and disease_associated_genes_clingen reference tables added to reference_db. Replaces in-memory Python constants for AR LoF gene list and disease association checks.
- PVS1 AR gene bypass now queries refdb.ar_lof_genes_clingen at classification time instead of Python AR_LOF_GENES constant.
- Disease association check now queries refdb.disease_associated_genes_clingen for PVS1, PP3_Strong, LP disease gates, and ClinVar override gate.
- Reference database count: 14 (previously 13).
v3.15.0 March 2026
- PM1: Critical functional domains migrated from Python constant (CRITICAL_PFAM_DOMAINS) to reference database table (refdb.interpro_pfam_domains). HELIX-CR-2026-012.
- PM1: Full InterPro Pfam catalog (27,481 entries) loaded into reference_db. 49 domains marked as critical (is_critical=true) across 14 categories.
- PM1: Pre-build approach - critical accessions loaded from refdb at classification time, OR LIKE clauses generated in Python. Zero performance regression vs v3.14.0.
- PM1: Adding a new critical domain is now a database UPDATE, not a code change + rebuild + deploy.
- InterPro / Pfam added as 13th reference database.
- Loader script: scripts/load_interpro_pfam.py with InterPro API fetch, JSON cache for offline fallback, and hard-fail validation.
- Reference: HELIX-CR-2026-012.
v3.14.0 March 2026
- PM1: CRITICAL_PFAM_DOMAINS converted from Pfam short names to InterPro-verified Pfam accession numbers (HELIX-CR-2026-011). VEP annotation records domain data as Pfam:PFxxxxx accession numbers, but PM1 SQL searched for short names (RyR, Pkinase, Ion_trans). Result: 0 legitimate PM1 applications since curated list introduction. Fix: all 49 domains now use verified accession numbers.
- PM1: Removed Rho (covered by PF00071 Ras superfamily, VEP confirms RAC1 annotated with PF00071) and Peptidase_S1 (covered by PF00089 Trypsin, VEP confirms F10 annotated with PF00089).
- PM1: Corrected accessions - Dynamin_N: PF01031 (Dynamin_M) corrected to PF00350 (Dynamin_N). PH: PF00168 (C2 domain) corrected to PF00169 (PH domain). FBN_EGF: mapped to PF00008 (EGF-like domain).
- PM1: All 49 accessions verified against InterPro REST API (ebi.ac.uk/interpro) on 2026-03-16. Domain coverage validated against patient VCF data (RYR1 PF01365, KRAS PF00071, KCNH2 PF00520, CDH23 PF00028, MYO7A PF00063, COL1A1 PF01391, CFTR PF00005, EGFR PF07714).
- PM1: Post-fix validation on HG2023_206: PM1 count increased from 2 (false positive) to 1,139 (legitimate). RYR1 c.6635T>C correctly receives PM1. CRYGD false positives eliminated. Zero B/LB classification changes.
- Clinical trigger: HG2023_206 RYR1 c.6635T>C (p.Val2212Ala) in RyR domain (Pfam:PF01365) classified VUS [PM2,PP3_Supporting] without PM1. CG classified LP. PM1 now applied: VUS [PM1,PM2,PP3_Supporting].
- Reference: HELIX-CR-2026-011, Richards et al. 2015 PMID:25741868, ClinGen SVI PM1 guidance.
v3.13.0 March 2026
- PVS1: ClinVar LOF evidence added as fifth constraint gate path (HELIX-CR-2026-010). Small autosomal dominant genes with ClinVar P/LP LOF variants (2+ review stars) but uninformative gnomAD constraint (pLI < 0.9, LOEUF >= 0.35, HI != 3) now qualify for PVS1 via CLINVAR_LOF_AD_GENES curated list.
- gene_has_disease_association: CLINVAR_LOF_AD_GENES added as sixth source in the disease association check.
- Clinical trigger: GCM2 p.Arg131Ter (stop_gained) in PD2025_098. GCM2 is a 5-exon AD gene for Familial hypoparathyroidism with ClinVar Pathogenic p.Tyr136Ter (2 stars) but pLI=0.0, LOEUF=1.205, HI=NULL. Now LP [PVS1,PM2].
- Reference: Abou Tayoun 2018 PMID:30192042, HELIX-CR-2026-010.
v3.12.1 March 2026
- PP3/BP4: Missense consequence guard (HELIX-CR-2026-009). PP3 BayesDel path (Strong/Moderate/Supporting) and BP4 BayesDel path (Moderate/Supporting) now require consequence = missense_variant. BayesDel_noAF is calibrated exclusively for missense variants (Pejaver et al. 2022).
- Clinical trigger: GCM2 p.Arg131Ter (stop_gained) in PD2025_098 received PP3_Strong (BayesDel=0.66) + PM2 = LP. Correct: PM2 only = VUS. 2 non-missense variants with false PP3_Strong eliminated.
- Subtractive change: can only downgrade (LP to VUS), cannot upgrade. PP3_splice (SpliceAI) unaffected.
- Reference: Pejaver 2022 PMID:36413997, Walker 2023 PMID:37352859, HELIX-CR-2026-009.
v3.12.0 March 2026
- HELIX-REF-001: Multi-Source HPO Enrichment Pipeline. HPO annotation expanded from single-source (HPO Consortium, 5,173 genes, 320K records) to multi-source enriched view (6 sources, 5,688 genes, 927K records, +10% gene coverage).
- HPO sources: HPO Consortium (primary), Orphanet disease-to-HPO via Orphadata XML (3,176 genes with frequency data), DECIPHER G2P from 7 clinical panels (2,125 genes), Monarch Initiative (4,791 genes), ClinVar-MedGen chain mapping from P/LP variants (5,258 genes), and manual clinical curation for common-disease genes.
- Enrichment depth: 3,292 existing genes received additional HPO terms. 515 net new genes added. 1,505 genes covered by all 5 automated sources.
- Source priority ordering: Orphanet > G2P > HPO Consortium > Monarch > ClinVar-MedGen. Low-confidence entries filtered from clinical view.
- PP4 criterion evaluates against enriched HPO set. No ACMG criteria logic changes.
- Data integrity validated: 0 NULLs, 0 invalid HPO IDs, 0 duplicate PKs across 926,930 records.
v3.11.4 March 2026
- PVS1: Non-canonical splice site variants excluded from PVS1 application. VEP consequence terms splice_donor_5th_base_variant (position +5), splice_donor_region_variant (positions +3 to +8), and splice_acceptor_5th_base_variant (acceptor position -5) now block PVS1. ClinGen PVS1 Decision Tree (Abou Tayoun et al. 2018) specifies canonical +/-1,2 splice sites only. Extended splice region variants retain access to PP3_splice (SpliceAI >= 0.2) at Supporting evidence level.
- Clinical trigger: PDE6B c.1401+4_1401+48del in PD2025_061 - VUS vs ClinVar Benign (2 stars). VEP annotates as splice_donor_5th_base_variant. PVS1 created false conflicting evidence (PVS1 vs BS1+BS2+BP6 = VUS).
- Impact: 2 of 18 splice_donor PVS1 variants affected in PD2025_061. 0 P/LP variants affected. Subtractive change (cannot create false P/LP).
- Log strings now use CLASSIFIER_VERSION constant instead of hardcoded version.
- Reference: HELIX-CR-2026-008, Abou Tayoun 2018 PMID:30192042.
v3.11.3 March 2026
- PVS1: NMD_transcript_variant no longer blocks PVS1. VEP NMD_transcript_variant annotation means the transcript UNDERGOES nonsense-mediated decay, which is additional evidence FOR loss-of-function, not against it. The exclusion now targets NMD_escaping_variant (transcripts that escape NMD - truncated protein expressed, may retain partial function).
- Previous behavior: frameshift_variant,NMD_transcript_variant -> PVS1 blocked. New behavior: frameshift_variant,NMD_transcript_variant -> PVS1 applies. frameshift_variant,NMD_escaping_variant -> PVS1 blocked (correct).
- Clinical trigger: HG2022_057 review (KCNN2 frameshift).
- Reference: VEP documentation (Ensembl), Abou Tayoun et al. 2018 PMID:30192042.
v3.11.0 March 2026
- De novo projection (PS2): Prospective computation of whether adding PS2 (confirmed de novo, Strong) would upgrade a VUS to LP or P. Identifies variants where trio confirmation has the highest clinical impact.
- Excludes: non-VUS, non-heterozygous, BA1 active, compound het candidates (biologically exclusive with de novo), and conflicting evidence variants.
- Minimum evidence guard (v3.11.2): PM2 alone is insufficient for de novo candidacy - requires at least one additional pathogenic criterion beyond PM2.
- De novo projection is informational only - it does not change the current ACMG classification. It flags variants where parental testing would resolve VUS status.
v3.10.1 March 2026
- disease_genes_clinvar CTE: Added review_stars >= 2 filter. The shared CTE that provides gene-level ClinVar disease association evidence (Source 1 of 5 in the disease association check) previously included all ClinVar P/LP genes without a star minimum. This created a circular reference: a 1-star ClinVar P/LP submission simultaneously triggered the ClinVar override AND provided the disease association evidence that allowed the v3.10.0 gate to pass. SHROOM2 (1-star Likely Pathogenic, zero disease association from Orphanet/ClinGen/AR_LOF_GENES/VCEP) self-validated its own gate.
- The review_stars >= 2 threshold is consistent with PS1 (ps1_min_stars = 2): a single unreviewed ClinVar submission does not constitute gene-level disease association evidence. Two or more independent submitters with concordant interpretation (2+ stars) provide meaningful gene-disease validation.
- Impact on PVS1 and PP3_Strong disease gates: minimal. Both gates reference the same CTE, but genes where PVS1 or PP3_Strong activate have redundant disease evidence from Orphanet, ClinGen haploinsufficiency, curated AR LoF gene list (~150 genes), or VCEP coverage. A gene relying solely on 1-star ClinVar for disease evidence is exactly the type of gene that should not receive PVS1 (Very Strong) or PP3_Strong.
- Clinical trigger: SHROOM2 p.Gly211Ser remained LP after v3.10.0 deployment. Now correctly classified as VUS (ClinVar_gated).
v3.10.0 March 2026
- ClinVar override: P/LP override with review_stars = 1 (single submitter, no peer review) now requires the gene to have an established disease association via at least one of five sources: ClinVar P/LP (gene-level), Orphanet entry, ClinGen haploinsufficiency score, curated AR LoF gene list (~150 genes), or VCEP coverage. Without this gate, a single ClinVar submission in a gene without any known Mendelian disease mechanism could promote a variant to Likely Pathogenic.
- ClinVar override: P/LP assertions with 2+ review stars (multiple concordant submitters) bypass the disease association gate. Benign and Likely Benign ClinVar overrides are not affected by this gate.
- ClinVar override: Gated variants (1-star P/LP in genes without disease association) fall through to standard ACMG scoring, which typically produces VUS when no ACMG criteria are activated.
- gene_has_disease_association: New boolean column in the criteria evaluation stage. Reuses the same 5-source disease check as the PVS1 gate (v3.8.0) and PP3_Strong gate (v3.8.1). Zero new data sources, zero performance overhead.
- clinvar_disease_gated: New audit trail boolean. The criteria string shows the "ClinVar_gated" prefix when the disease-association gate blocks a ClinVar P/LP override, identifying the gated variant for review.
- Clinical trigger: SHROOM2 p.Gly211Ser in PD2025_060 - classified as LP from a single ClinVar submission (1 star) in a gene with zero disease association across all five reference sources (Orphanet, ClinGen, HPO, VCEP, AR_LOF_GENES), zero constraint data (pLI/LOEUF/mis_z all NULL), and population frequency of 0.23%.
- Reference: HELIX-CR-2026-007, Ghosh et al. 2018 PMID:30311383 (limitations of single-submitter ClinVar entries).
v3.9.0 March 2026
- PVS1: Last-exon truncating variants downgraded from Very Strong (+8 points) to Strong (+4 points). ClinGen PVS1 Decision Tree (Abou Tayoun et al. 2018) Figure 1, Node 3: truncating variants in the last exon escape nonsense-mediated mRNA decay (NMD) because there is no downstream exon-exon junction complex.
- PVS1: is_last_exon detection uses VEP exon_number column (format X/Y, last exon when X == Y). NULL or missing format defaults to Very Strong (conservative fallback). Single-exon genes (1/1) correctly receive PVS1_Strong.
- PVS1: Criteria string shows PVS1 for Very Strong (non-last-exon) and PVS1_Strong for last-exon downgraded variants. Consistent with PP3_Strong, BP4_Moderate naming convention.
- PVS1: PP3_splice double-counting guard unchanged - uses original PVS1 condition covering both Very Strong and Strong paths.
- PVS1: AR_LOF_GENES last-exon downgrade applies regardless of inheritance mode. NMD escape is a biophysical mechanism independent of AD/AR inheritance.
- PVS1: Classification impact - PVS1_Strong + PM2 = 6 points = Likely Pathogenic (most common minimal combination unaffected). PVS1_Strong + PS1 drops from Pathogenic to Likely Pathogenic.
- Clinical trigger: PD2025_058 review (MYO6 p.Lys972GlufsTer5, exon 27/32 - NOT last exon, unaffected by change).
- Reference: Abou Tayoun et al. 2018 PMID:30192042.
v3.8.2 March 2026
- BP1: ClinVar pathogenic missense guard - BP1 no longer applies when ClinVar has a Pathogenic missense variant in the same gene. If ClinVar confirms pathogenic missense, the BP1 premise is falsified. Uses session-local ClinVar data.
- BP1: Clinical trigger - MCCC2 p.Val339Met in PD2025_099 received BP1 despite being ClinVar Pathogenic missense. AR enzyme genes have low mis_z but pathogenic homozygous missense.
- PP3_Supporting: Criteria string label changed from PP3 to PP3_Supporting for consistency with PP3_Strong and PP3_Moderate. ClinGen SVI (Walker et al. 2023) recommends explicit evidence strength labeling.
v3.8.1 March 2026
- PP3_Strong: Disease association gate - PP3_Strong (BayesDel >= 0.518, Strong evidence level) now requires the gene to have an established disease association via at least one of five sources: ClinVar P/LP (gene-level), Orphanet gene-disease entry, ClinGen haploinsufficiency score, curated AR LoF gene list, or VCEP coverage.
- PP3_Strong: Without this gate, PP3_Strong + PM2 could reach Likely Pathogenic in genes without any known Mendelian disease mechanism. Clinical trigger: PD2025_056 identified POU2F2 p.Val324Met and RYR3 p.Arg285Gln as false LP classifications.
- PP3_Strong: Scientific basis - Pejaver et al. 2022 calibrated BayesDel on ClinVar variants in disease genes; Walker et al. 2023 (ClinGen SVI) Section 3.2 restricts PP3/BP4 calibrated thresholds to genes with established disease mechanism.
- PP3_Moderate and PP3_Supporting are NOT gated by disease association (cannot reach LP alone).
- PP3_splice path unchanged (splice prediction independent of missense disease mechanism).
v3.8.0 March 2026
- PVS1: Disease association gate - PVS1 now requires the gene to have an established disease association via at least one of five sources: ClinVar P/LP (gene-level), Orphanet gene-disease entry, ClinGen haploinsufficiency score, curated AR LoF gene list (~150 genes), or VCEP coverage. Population constraint (pLI/LOEUF) alone is no longer sufficient.
- PVS1: Genes with high pLI but no known Mendelian disease mechanism (e.g., TTC28, ABL1, HTT) are correctly excluded from PVS1. Gene audit across 5 validation cases: 250 genes blocked, 92 genes pass, ACMG SF v3.2 safety check passed (0 of 81 genes affected).
- PVS1: Reference: Abou Tayoun et al. 2018 (ClinGen PVS1 Decision Tree) - "PVS1 is applicable when loss of function is a known mechanism of disease." HELIX-CR-2026-003.
- BS1: Constraint-implied AD fallback - when no inheritance mode data exists from any source (Orphanet, ClinGen, HPO) but the gene is LoF-constrained (pLI > 0.9 or LOEUF < 0.35), BS1 uses the AD threshold (0.1%) instead of the AR default (5%). Resolves the architectural inconsistency where PVS1 treated a gene as AD-haploinsufficient but BS1 treated it as AR.
- BS1: Explicit Orphanet AR and curated AR_LOF_GENES added as cascade priorities 5-6, so explicit AR evidence overrides constraint-implied AD inference.
- BS1: Cascade expanded from 5 to 8 levels: VCEP > ClinGen HI > Orphanet AD > Orphanet XLD > HPO AD > Orphanet AR > AR_LOF_GENES > Constraint-implied AD > Default AR.
- BS1: Criteria string now includes BS1[Orphanet-AR] and BS1[constraint-AD] source annotations.
- Reference databases: Orphanet/Orphadata and VCEP gene-specific specifications now documented as separate reference databases (9 total, previously 7).
- Clinical validation: TTC28 p.Arg21Ter (PD2025_100) changes from LP to VUS. Discordance analysis: frequency_conflict CRITICAL findings reduced from 1 to 0.
v3.7.0 March 2026
- BS1: Orphanet inheritance cascade for inheritance-aware frequency thresholds. 5-level priority: VCEP gene-specific -> ClinGen HI score=3 -> Orphanet AD -> Orphanet XLD -> HPO AD fallback -> AR default.
- BS1: Orphanet AD coverage: 1,694 autosomal dominant genes (4x improvement over ~400 from ClinGen HI alone).
- BS1: XLD/XLR distinction: X-linked dominant genes (90) receive AD threshold (0.1%); X-linked recessive genes (663) receive AR default (0.05%). XL-unspecified mapped to XLR as conservative default.
- BS1: HPO AD fallback (HP:0000006) as Priority 4 catches AD genes not in Orphanet (e.g. KCNQ1, KCNH2 - also covered by ClinGen HI).
- BS1: Criteria string now includes inheritance source annotation: BS1[ClinGen-HI], BS1[Orphanet-AD], BS1[Orphanet-XLD], BS1[HPO-AD].
- Reference DB: orphanet_gene_inheritance table added (3,197 genes, LEFT JOIN at classification time, <2ms overhead).
- Data source: Orphadata CC-BY-4.0 (INSERM, France). Covers ~6,100 rare diseases with gene-disease-inheritance annotations.
- Validation: 9/10 known AD genes confirmed (KCNJ11, GCK, HNF1A, FGFR3, SCN1A, etc.), 4/4 AR-only genes unchanged, 5/5 dual-inheritance genes correct.
v3.6.8 March 2026
- BP1: Added mis_z < 2.0 guard to prevent BP1 in missense-constrained genes.
- BP1: Genes with pLI < 0.1 AND mis_z >= 2.0 (e.g. KCNJ11, ABCC8, KCNQ1, SCN1A) have missense as primary disease mechanism - BP1 ("primarily truncating variants cause disease") is factually incorrect for these genes.
- Clinical validation: KCNJ11 p.Leu270Val correctly loses BP1 after this change.
v3.6.7 March 2026
- PP3: Missense relevance guard added. PP3 (BayesDel) now requires evidence that missense variants cause disease in the gene before applying computational prediction.
- PP3 guard has four gates: (1) mis_z > 2.0 for direct missense constraint, (2) mis_z > 1.5 AND pLI < 0.5 for AR genes like BRCA2/MLH1/ATM, (3) mis_z > 0.5 AND pLI > 0.9 for strong AD LoF-intolerant genes, (4) NULL fallback for genes without mis_z but with pLI > 0.5.
- PP3_splice path unchanged (splice prediction independent of missense mechanism).
- BP4 unchanged (benign prediction consistent with low missense constraint).
- Scientific basis: Pejaver et al. 2022 calibrated BayesDel on missense-mechanism genes; applying to LoF-mechanism genes is methodologically incorrect.
v3.6.6 March 2026
- PM1: Curated CRITICAL_PFAM_DOMAINS list (~60 domains) replaces generic Pfam presence check.
- PM1: Only well-characterized functional domains with documented pathogenic variant clustering trigger PM1.
- PM1: Excluded domains: generic structural (Caveolin, coiled-coil), DUF, transmembrane helices, repetitive regions.
- PM1: Included domains: kinase catalytic (PF00069, PF07714), DNA-binding (PF00870, PF00505, PF00096, PF00533), ion channel pores (PF00520, PF00029, PF01365), GTPase (PF00071), enzyme active sites (PF00089), and ~40 additional VCEP-documented domains. Domain matching uses Pfam accession numbers since v3.14.0.
- Scientific basis: ACMG 2015 requires PM1 for "critical and well-established functional domain without benign variation" - not any Pfam annotation.
v3.6.5 March 2026
- PP2: Now requires mis_z > 2.0 in addition to pLI > 0.5.
- PP2: ACMG 2015 requires "missense variants are a common mechanism of disease" - pLI alone measures LoF constraint, not missense constraint.
- PP2: mis_z > 2.0 indicates that the gene is missense-constrained (Samocha et al. 2014, Karczewski et al. 2020).
- Genes like CAV1 (pLI=0.846, mis_z=0.86) with LoF disease mechanism no longer incorrectly trigger PP2.
v3.6.4 March 2026
- PVS1: AR LoF gene list expanded from approximately 35 to approximately 150 genes
- AR LoF gene list covers all major AR disease categories: neurodegeneration, metabolic (amino acid, fatty acid oxidation, glycogen storage, lysosomal, peroxisomal), ciliopathies, sensory (hearing, vision), immune/hematologic, cardiac, neuromuscular, connective tissue, kidney, endocrine, liver, and respiratory
- Every gene in the curated list has ClinGen Definitive/Strong validity or equivalent published evidence for biallelic LoF disease mechanism
- VCEP AR genes (CDH23, GJB2, MYO7A, PAH, SLC26A4, USH2A) handled via VCEP pvs1_applicable gate, not duplicated in curated list
v3.6.2 March 2026
- Homozygous reference genotype guard: variants with genotype hom_ref (0/0) excluded from ACMG classification
- hom_ref variants arise from multi-allelic VCF sites where the sample has 0 ALT reads for a specific allele but was included due to another ALT allele at the same position
- Without this guard, hom_ref frameshift/stop variants in LoF-intolerant genes received PVS1+PM2 = Likely Pathogenic despite not being present in the patient
- 20 false Pathogenic/Likely Pathogenic classifications eliminated in validation (Case 19)
v3.6.1 March 2026
- PM2: NULL gnomAD allele frequency now correctly satisfies PM2 (absent from controls = never observed in ~800K individuals)
- PM2: Previous behavior required frequency data to be present (non-NULL), incorrectly excluding truly absent variants
- PVS1: Curated AR LoF gene list (~35 genes) replaces broad HPO-based bypass (v3.6 used HP:0000007 with 3,127 AR genes - too broad, caused 43 false Likely Pathogenic)
- PVS1: Curated list includes only genes with Definitive/Strong evidence for biallelic LoF disease mechanism (ClinGen Gene-Disease Validity, literature)
v3.6 March 2026
- PVS1: Autosomal recessive genes bypass pLI/LOEUF constraint gate using HPO inheritance annotation (HP:0000007)
- PVS1: Aligned with ClinGen PVS1 Decision Tree (Abou Tayoun et al. 2018) which does not require population constraint for genes with established LoF disease mechanism in recessive inheritance
- HPO database now used for PVS1 AR gene identification in addition to PP4 phenotype matching
- Note: HPO-based bypass replaced by curated gene list in v3.6.1 due to excessive breadth
v3.5 February 2026
- ClinGen VCEP gene-specific specification overlay (optional, enabled by default)
- BA1, BS1, PM2: gene-specific frequency thresholds from published VCEP specifications (~50-60 genes)
- PVS1: gene-specific applicability gate (disabled for gain-of-function genes)
- VCEP audit trail: criteria string includes [VCEP:Panel vX.Y] marker when gene-specific thresholds applied
- VCEP toggle: can be enabled/disabled per case in case settings
- Source: ClinGen Criteria Specification Registry (CSpec)
v3.4 February 2026
- Bayesian point-based classification framework (Tavtigian et al. 2018, 2020) replaces sequential 18-rule evaluation
- All 18 original ACMG combining rules produce identical results under the point system (backward compatible)
- Point system fills gaps: evidence combinations not covered by original 18 rules now classified properly
- PP3/BP4: BayesDel_noAF single-tool replaces weighted 6-predictor consensus (ClinGen SVI calibration, Pejaver et al. 2022)
- PP3 evidence strength modulation: Strong (>= 0.518), Moderate (0.290-0.517), Supporting (0.130-0.289)
- BP4 evidence strength modulation: Moderate (<= -0.361), Supporting (-0.360 to -0.181)
- PM1 + PP3 double-counting guard: combined evidence capped at Strong equivalent (4 points)
- Continuous confidence scores derived from point distance to classification threshold boundary
- Conflicting evidence handled by point summation instead of a binary VUS default
- High-confidence conflict safety check: Strong/Very Strong pathogenic vs. Strong benign flagged for manual review
- Legacy predictors (SIFT, AlphaMissense, MetaSVM, DANN, PhyloP, GERP) retained as display data only
v3.3 February 2026
- SpliceAI PP3 threshold aligned to ClinGen SVI 2023 recommendation (lowered from 0.5 to 0.2)
- PP3_splice excluded when PVS1 applies (ClinGen SVI double-counting guard)
- BP7 upgraded: synonymous + not splice_region + SpliceAI <= 0.1 (Walker et al. Figure 4)
- BP7 conservation filter intentionally omitted per Walker et al. Table S13
- Evidence strength modulation (PP3_Moderate for high SpliceAI) deferred pending VCEP specifications
v3.2 February 2026
- SpliceAI integration: PP3_splice path for splice evidence
- BP4 SpliceAI guard: requires max_score < 0.1 for benign consensus
- Criteria string distinguishes PP3 (missense) from PP3_splice (splice)
v3.1 January 2026
- Maximum sensitivity approach: removed all frequency and impact pre-filtering
- All quality-passing variants proceed through classification
- Clinicians decide clinical relevance using classification + annotations
v3.0 January 2026
- Conflicting evidence priority: pathogenic + benign evidence produces VUS
- BA1 stand-alone override: allele frequency > 5% always classified Benign
- ClinVar override restricted to non-conflicting evidence only
- SQL-based classification engine (100x performance improvement)
- PM2: required non-NULL frequency data (corrected in v3.6.1)
### References
- Tavtigian SV, Greenblatt MS, Harrison SM, et al. Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. Human Mutation. 2018;39(11):1485-1492. PMID: 30311386
- Tavtigian SV, Harrison SM, Boucher KM, Biesecker LG. Fitting a naturally scaled point system to the ACMG/AMP variant classification guidelines. Human Genetics. 2020;139(8):1057-1067. PMID: 32666219
- Pejaver V, Byrne AB, Feng BJ, et al. Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3/BP4 criteria. American Journal of Human Genetics. 2022;109(12):2163-2177. PMID: 36413997
- Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. Genetics in Medicine. 2015;17(5):405-424. PMID: 25741868
- Walker LC, Hoya M, Wiggins GAR, et al. Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: Recommendations from the ClinGen SVI Splicing Subgroup. American Journal of Human Genetics. 2023;110(7):1046-1067. PMID: 37352859
- Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, et al. Predicting Splicing from Primary Sequence with Deep Learning. Cell. 2019;176(3):535-548.e24. PMID: 30661751
- Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature. 2020;581(7809):434-443. PMID: 32461654
- Cheng J, Novati G, Pan J, et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science. 2023;381(6664):eadg7492. PMID: 37733863
- McLaren W, Gil L, Hunt SE, et al. The Ensembl Variant Effect Predictor. Genome Biology. 2016;17(1):122. PMID: 27268795
- Landrum MJ, Lee JM, Benson M, et al. ClinVar: improving access to variant interpretations and supporting evidence. Nucleic Acids Research. 2018;46(D1):D1062-D1067. PMID: 29165669
- Abou Tayoun AN, Pesaran T, DiStefano MT, et al. Recommendations for interpreting the loss of function PVS1 ACMG/AMP variant criterion. Human Mutation. 2018;39(11):1517-1524. PMID: 30192042
- Josephs KS, Roberts AM, Grozeva D, et al. Beyond gene-disease validity: capturing structured data on inheritance, allelic requirement, disease-relevant variant classes, and mechanism for inherited cardiac conditions. Genome Medicine. 2023;15:86. PMID: 37884986
- Cummings BB, Karczewski KJ, Kosmicki JA, et al. Transcript expression-aware annotation improves rare variant interpretation. Nature. 2020;581(7809):452-458. PMID: 32461655
### Questions About Our Methodology?
Clinical geneticists and laboratory directors can contact us with technical questions about the classification method.
Contact Us
For Geneticists How It Works Rare Disease Newborn Screening

### Links and cited sources

- [gnomad.broadinstitute.org](https://gnomad.broadinstitute.org)
- [ncbi.nlm.nih.gov/clinvar](https://www.ncbi.nlm.nih.gov/clinvar/)
- [sites.google.com/site/jpopgen/dbNSFP](https://sites.google.com/site/jpopgen/dbNSFP)
- [Illumina / Ensembl](https://github.com/Illumina/SpliceAI)
- [NCBI / EMBL-EBI / MANE Consortium](https://www.ncbi.nlm.nih.gov/refseq/MANE/)
- [HPO Consortium + Orphanet + G2P + Monarch + ClinVar-MedGen](https://hpo.jax.org)
- [clinicalgenome.org](https://clinicalgenome.org)
- [orphadata.com (INSERM, France)](https://www.orphadata.com)
- [cspec.genome.network](https://cspec.genome.network)
- [DECIPHER / Wellcome Sanger Institute](https://www.deciphergenomics.org/gene2phenotype)
- [Monarch Initiative (monarchinitiative.org)](https://monarchinitiative.org)
- [NCBI ClinVar + MedGen](https://www.ncbi.nlm.nih.gov/medgen/)
- [UniProt Consortium (EMBL-EBI / SIB / PIR)](https://www.uniprot.org)
- [EMBL-EBI InterPro](https://www.ebi.ac.uk/interpro/)
- [clinicalgenome.org](https://clinicalgenome.org/affiliation/gene-disease-validity/)
- [National Clinical Research Center for Geriatric Disorders, Xiangya Hospital, Central South University](http://www.genemed.tech/gofcards)
- [clinicalgenome.org](https://clinicalgenome.org/working-groups/dosage-sensitivity-curation/)
- [ensembl.org](https://www.ensembl.org/info/docs/tools/vep/index.html)
- [Contact Us](https://folklore.helena.bio/contact)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)
- [How It Works](https://folklore.helena.bio/how-it-works)
- [Rare Disease](https://folklore.helena.bio/use-cases/rare-disease)
- [Newborn Screening](https://folklore.helena.bio/use-cases/newborn-screening)

---

## Page: Trio Variant Analysis Methodology | Folklore

Source: https://folklore.helena.bio/methodology/family-analysis
Canonical: https://folklore.helena.bio/methodology/family-analysis
Description: Production trio methodology covering GLnexus joint genotyping, proband ACMG classification, PLINK relationship QC, de novo candidates, compound heterozygous phasing, and segregation limits.

## Trio Variant Analysis Methodology
Family Analysis | Complete trio workflow | ClinGen SVI-informed evidence
Folklore analyses a complete proband-mother-father trio. The member gVCFs are joint-genotyped into one normalized call set, the proband is classified through the standard ACMG workflow, and family-aware evidence is computed from the same joint genotypes.
The family module records inheritance evidence alongside the proband classification. It does not replace clinical review, and the clinical-grade reporting overlay does not independently change PS2 criterion semantics.
Contents
01 Pipeline overview 02 Family compositions 03 Sample quality control (PLINK IBD) 04 De novo variant detection 05 Compound heterozygosity 06 Cosegregation scoring (LOD) 07 Evidence summary for clinical review 08 Methodology architecture 09 Distinctive Folklore features 10 Limitations 11 Version history 12 References
### Pipeline overview
Seven-stage production flow from complete-trio validation and joint genotyping to proband classification, relationship QC, inheritance analysis, and reviewable evidence delivery.
1
Validate the complete trio
The workflow requires exactly one proband, one mother, and one father with compatible gVCF inputs and recorded pedigree metadata.
2
Joint-genotype the three gVCFs
GLnexus creates one normalized multi-sample VCF. Caller-specific preset selection and contig normalization are resolved before downstream processing.
3
Extract and classify the proband
A proband-only variant VCF is extracted from the joint call set and sent through Variant Analysis. ACMG classification and variant annotations are produced for the proband only.
4
Build the trio evidence store
Every joint-VCF site is written to the trio DuckDB with proband, father, and mother genotypes. Proband annotation and ACMG fields are joined from the classified proband database.
5
Run relationship quality control
PLINK computes pairwise Identity-by-Descent estimates to evaluate declared parent-child relationships, parental relatedness, and possible duplicate samples.
6
Compute inheritance evidence
De novo confidence, compound heterozygous phase, segregation scores, and categorical inheritance patterns are derived from the trio genotypes.
7
Apply the reporting overlay
Quality, population-frequency, consequence, and splice-signal filters identify the clinically focused reporting subset. Detector outputs and ACMG criterion semantics remain separately auditable.
### Family compositions
The current production workflow accepts one composition: a complete proband-mother-father trio.
Composition Requirement Analysis scope
Complete trio Proband + mother + father Joint genotyping, proband classification, PLINK relationship QC, de novo assessment, compound heterozygous phasing, and trio segregation context
Duo, sibling-only, extended-pedigree, and multiple-proband workflows are not part of the current production trio pipeline.
### Sample quality control (PLINK IBD)
Before inheritance analysis begins, Folklore verifies that declared family relationships are consistent with the genomic data. Sample swap detection is a critical clinical safety check: a mislabelled sample produces invalid inheritance evidence regardless of algorithmic correctness downstream.
Method
PLINK 1.9 computes pairwise Identity-by-Descent (IBD) proportions across all declared family members using biallelic single-nucleotide variants with minor allele frequency at or above 0.05 and missingness at or below 5 percent. Expected IBD for parent-child pairs is approximately 0.50; expected IBD for unrelated parents is below 0.10. Deviations beyond defined thresholds raise alerts that are surfaced to the clinical reviewer.
Alert types
critical
Sample swap
Declared parent-child pair with PI_HAT below 0.40. Pipeline fails; clinical reviewer must validate sample labelling before re-running analysis.
warning
Duplicate sample
Declared parent-child pair with PI_HAT above 0.70, indicating possible sample duplication or extreme consanguinity. Pipeline proceeds; warning surfaced.
warning
Consanguinity
Declared unrelated parents with PI_HAT above 0.125 (third-degree relatives or closer). Pipeline proceeds; warning surfaced for clinical interpretation.
### De novo variant detection
De novo variants arise in the germline of a parent or in the post-zygotic period of the proband and are absent from both parental genomes. Their identification relies on confident parental hom_ref calls at the variant position. Folklore applies a tiered confidence assessment that surfaces the strength of de novo inference rather than producing a binary call.
ACMG criteria triggered
ACMG/AMP 2015 (Richards et al., PMID: 25741868) defines PS2 (Strong) for confirmed de novo variants with verified maternity and paternity, and PM6 (Moderate) for assumed de novo variants without confirmation. The ClinGen Sequence Variant Interpretation Working Group specification (Biesecker & Harrison 2018, PMID: 29543229) refined these criteria, distinguishing confidence levels and providing application guidance.
Applicability
De novo detection requires a complete proband-mother-father trio. Both parental genotypes must be available from the same joint-genotyped call set before de novo confidence can be assessed.
Tiered confidence assessment
The detector emits high- and low-confidence candidates, plus excluded and not-applicable control states. Confidence reflects technical support for de novo origin, not pathogenicity. The clinical-grade layer is a reporting refinement and does not independently alter PS2 semantics.
high
High confidence
Both parents meet the configured genotype and coverage requirements with reference genotypes at the candidate position. The result is a technically supported de novo candidate for clinical and ACMG review.
low
Low confidence
At least one parent has sub-threshold coverage, genotype quality below recommended levels, or chromosome representation that cannot be confirmed. Variant is surfaced for manual reviewer interpretation; PS2 is not auto-applied.
excluded
Excluded
Variant present in at least one parent (parent genotype not hom_ref). De novo origin is explicitly disconfirmed.
Chromosome-presence gate
The joint multi-sample VCF provides an explicit genotype state for each trio member at every retained site. Folklore distinguishes confirmed reference genotypes, low-quality calls, no-calls, and biologically expected chromosome exceptions before assigning de novo confidence.
chrY for female parents
Female parents biologically lack chrY. Absence of chrY variants in maternal data is expected, not a coverage gap. The chromosome-presence gate accommodates this exception so chrY proband variants are not unnecessarily downgraded.
chrM for paternal contribution
Mitochondrial DNA is maternally inherited; paternal mitochondria are degraded during fertilisation. Absence of chrM variants in paternal data is expected. The gate excludes paternal chrM coverage from confidence calculations to reflect this biology.
Exclusions
The following situations are excluded from de novo interpretation or trigger downgrades to lower confidence tiers:
Variants present in at least one parent are explicitly excluded; de novo origin is disconfirmed.
Mosaic de novo variants below detection threshold in parental samples may be incorrectly classified as de novo. Clinical interpretation should consider mosaicism when phenotype suggests it.
Regions with known low-quality parental genotyping (low complexity, segmental duplications) may produce sub-threshold coverage and downgrade confidence.
Variants on parent-only positions (present in a parent but not in the proband) are not de novo candidates by definition and are excluded from de novo classification.
### Compound heterozygosity
Compound heterozygosity is a recessive mechanism in which the proband carries two distinct heterozygous pathogenic variants in the same gene, each inherited from a different parent. Trio inheritance can establish whether the variants are in trans.
Biological rationale
In autosomal recessive disorders, pathogenicity requires biallelic disruption. Two heterozygous variants in trans configuration (one from each parent) result in functional knock-out of the gene, just as a homozygous loss-of-function variant would. In solo proband whole-exome or whole-genome analysis, phase cannot be determined without parental information or long-read sequencing. Trio data resolves this gap.
ACMG criterion triggered
ACMG/AMP 2015 (Richards et al., PMID: 25741868) defines PM3 (Moderate): "Detected in trans with a pathogenic variant for recessive disorders." Compound heterozygous candidates in trans configuration with a confidently classified partner contribute to PM3 application.
Phasing methodology
If the proband is heterozygous for variant A and variant B in the same gene, and variant A is inherited from the mother (mother heterozygous, father reference) while variant B is inherited from the father (mirror configuration), then A and B are in trans configuration. This is a definitive compound heterozygous setup. When parental genotype data is incomplete (sub-threshold coverage, missing chromosome representation), Folklore surfaces the candidate at lower confidence with explicit reason rather than making a binary determination.
Variant consequence filtering
Candidate variants are restricted to clinically relevant consequence types. Intronic variants distant from splice sites, synonymous variants without splice impact, and UTR-adjacent variants are excluded from compound heterozygosity assessment.
Included consequences
missense_variant
stop_gained
stop_lost
frameshift_variant
inframe_insertion
inframe_deletion
splice_donor_variant
splice_acceptor_variant
splice_region_variant
start_lost
initiator_codon_variant
Excluded consequences
intron_variant
downstream_gene_variant
upstream_gene_variant
synonymous_variant
A pragmatic cap of 50 candidate variants per gene protects against combinatorial explosion in large genes such as TTN, MUC16, and SYNE1. Variants are ranked by ACMG severity (P > LP > VUS > LB > B) and confidence score before the cap is applied, so the highest-priority candidates are retained.
Multi-partner candidates
When a variant has multiple potential compound heterozygous partners in the same gene, Folklore selects a single primary partner deterministically and flags the multi-partner state for clinical reviewer awareness. The reviewer sees that additional partner variants exist and can examine each pairing individually. This avoids hidden ambiguity and preserves clinical decision authority.
### Cosegregation scoring (LOD)
Cosegregation of a candidate variant with phenotype across family members provides quantitative evidence for pathogenicity. Folklore implements the classical Jarvik & Browning 2016 LOD framework (PMID: 27236918) with phenotype specificity modulation per ClinGen Sequence Variant Interpretation Working Group 2021 recommendations.
LOD formula
Per-variant LOD scores are computed using the logarithm-of-the-odds framework as described in Jarvik & Browning 2016. Each informative meiotic event consistent with the inheritance hypothesis contributes positively; each violation contributes negatively; uninformative meioses contribute zero. The mathematical formulation is described in detail in the cited reference and is not reproduced here.
Supported inheritance hypotheses
Folklore computes LOD scores for the following modes of inheritance. When the inheritance hypothesis is unspecified, segregation scoring is not performed and per-variant strength is reported as not-applicable; this is documented per Option A of the implementation specification.
AD_de_novo Autosomal dominant de novo AR_compound_het Autosomal recessive (compound heterozygous) AR_homozygous Autosomal recessive (homozygous) XLR X-linked recessive XLD X-linked dominant mitochondrial Mitochondrial (maternal transmission)
Five-band evidence strength scale
The continuous LOD score is mapped to one of five evidence strength bands per ClinGen SVI 2021 recommendations. Each band corresponds to an ACMG modifier strength applied to the PP1 (cosegregation) criterion.
Strength band LOD range ACMG points ACMG strength
indeterminate [0.0, 0.5) 0 --
supporting [0.5, 2.0) 1 PP1_Supporting
moderate [2.0, 3.0) 2 PP1_Moderate
strong [3.0, 5.0) 4 PP1_Strong
very_strong [5.0, inf) 8 PP1_VeryStrong
Phenotype specificity multiplier
Per ClinGen SVI 2021, the LOD score is dampened by a multiplier reflecting phenotype specificity. Highly specific phenotypes (syndromic, characteristic of a single gene) receive full weight; broad phenotypes (common, multiple genes possible) receive half weight; unspecified phenotypes receive quarter weight as a conservative default.
specific
x 1.0
broad
x 0.5
unspecified
x 0.25
Trio-only mathematical ceiling
A complete trio contributes at most two informative meioses and has a theoretical LOD ceiling near 0.3. The configured Supporting band begins at 0.5, so trio-only results normally remain indeterminate. Moderate, strong, and very strong PP1 evidence require an extended pedigree, outside the current production scope.
### Evidence summary for clinical review
After all phases complete, Folklore delivers a structured evidence summary alongside the per-variant inheritance annotations. The clinical reviewer receives the following information for each completed analysis:
Number of de novo candidate variants per confidence tier (high, low, excluded counts).
Number of compound heterozygous candidate pairs evaluated, plus per-gene breakdown of variants with assigned partners and multi-partner cases.
LOD distribution across strength bands (indeterminate, supporting, moderate, strong, very strong counts), maximum LOD observed, and mean LOD across scored variants.
PLINK IBD quality control results when joint VCF was provided: per-pair PI_HAT values, raised alerts (sample swap, duplicate sample, consanguinity), and PLINK version for audit.
Algorithm versions used for this specific analysis run, recorded immutably so the clinical reviewer always knows which exact version produced each result.
### Methodology architecture
Two principles shape the production workflow: one normalized joint call set for classification and inheritance, and immutable algorithm-version recording for traceability.
Joint genotyping before inheritance analysis
Folklore joint-genotypes the three member gVCFs with GLnexus. The proband-only classification input and the trio inheritance table derive from the same normalized multi-sample call set. This prevents independent call-set differences from being interpreted as inheritance.
Write-once algorithm versioning
Algorithm version labels for the joint caller, de novo algorithm, segregation algorithm, and sample QC algorithm are recorded when an analysis is created and cannot be modified afterward. A completed analysis therefore identifies the algorithm version used for each result, supporting reconstruction and audit after platform upgrades.
Reference tools and databases
GLnexus
Recorded per analysis
Joint genotyping of the complete trio gVCF set
Source: DNAnexus
PLINK
Recorded per analysis
Identity-by-Descent relationship quality control
Source: Purcell Lab, Harvard
### Distinctive Folklore features
What Folklore adds beyond the public standards. Each feature is publicly documented; the value is in disciplined integration, not in algorithmic novelty hidden from clinical scrutiny.
Shared call set across classification and inheritance
The proband ACMG classification and parental genotype evidence derive from one normalized joint VCF, preserving positional and allelic consistency across the complete trio.
Tiered de novo confidence assessment
De novo confidence records the technical strength of parental genotype evidence. The clinical-grade subset is a reporting overlay; final ACMG criterion handling remains explicit and reviewable.
Chromosome-aware safeguards
The chromosome-presence gate distinguishes confident hom_ref inference from suspicious absence by verifying parental data coverage on the candidate chromosome. Sex-aware exceptions (chrY for female parents, chrM for paternal contribution) prevent false confidence downgrades arising from biologically expected absence rather than coverage gaps.
LOD ceiling disclosure
The complete trio has a theoretical LOD ceiling near 0.3, below the configured Supporting threshold of 0.5. Extended pedigrees are required before PP1 evidence strength can be reached.
Write-once algorithm versioning for regulatory reproducibility
Algorithm version labels are recorded immutably at analysis creation. Subsequent platform upgrades do not retroactively modify completed analyses. A clinical reviewer always knows which exact version produced a given result, supporting both scientific reproducibility and accreditation audit requirements.
### Limitations
The following limitations bound the interpretation of Family Analysis results.
The production workflow requires a complete proband-mother-father trio. Duo, sibling-only, and extended-pedigree workflows are not currently supported.
Trio-only segregation has a theoretical LOD ceiling near 0.3, below the configured Supporting threshold of 0.5. PP1 evidence requires an extended pedigree.
Mosaic de novo variants below the detection threshold in a parental sample may appear absent and require clinical review.
Compound heterozygous findings depend on reliable parental genotypes and establish candidate phase, not the clinical significance of both alleles.
The clinical-grade reporting layer narrows technically supported candidates for review but does not independently change PS2 trigger semantics.
Structural variants and copy-number variants are outside the current small-variant trio workflow.
Results must be interpreted with phenotype, family history, and other clinical evidence by a qualified genetics professional.
### Methodology scope
The public methodology describes the current production workflow without binding the page to a transient software release number.
Production trio workflow Current scope Continuously maintained
Complete proband-mother-father trio input.
GLnexus joint genotyping and normalized joint VCF.
Proband ACMG classification from the joint call set.
PLINK Identity-by-Descent relationship quality control.
De novo and compound heterozygous candidate detection.
Explicit trio-only segregation ceiling.
Separate clinical-grade reporting overlay.
### References
Richards S, Aziz N, Bale S, Bick D, Das S, Gastier-Foster J, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology.
Genetics in Medicine. 2015;17(5):405-424.
PMID: 25741868
Biesecker LG, Harrison SM; ClinGen Sequence Variant Interpretation Working Group. The ACMG/AMP reputable source criteria for the interpretation of sequence variants.
Genetics in Medicine. 2018;20(12):1687-1688.
PMID: 29543229
Abou Tayoun AN, Pesaran T, DiStefano MT, Oza A, Rehm HL, Biesecker LG, Harrison SM. Recommendations for interpreting the loss of function PVS1 ACMG/AMP variant criterion.
Human Mutation. 2018;39(11):1517-1524.
PMID: 30192042
Jarvik GP, Browning BL. Consideration of cosegregation in the pathogenicity classification of genomic variants.
American Journal of Human Genetics. 2016;98(6):1077-1081.
PMID: 27236918
Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, et al. Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: Recommendations from the ClinGen SVI Splicing Subgroup.
American Journal of Human Genetics. 2023;110(7):1046-1067.
PMID: 37352859
Veltman JA, Brunner HG. De novo mutations in human genetic disease.
Nature Reviews Genetics. 2012;13(8):565-575.
PMID: 22781750
Kolesnikov A, Goel S, Nattestad M, Yun T, Baid G, Yang H, et al. DeepTrio: variant calling in families using deep learning.
bioRxiv. 2021.
DOI: 10.1101/2021.04.05.438434
Purcell S, Neale B, Todd-Brown K, Thomas L, Ferreira MA, Bender D, et al. PLINK: a tool set for whole-genome association and population-based linkage analyses.
American Journal of Human Genetics. 2007;81(3):559-575.
PMID: 17701901
Manichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM. Robust relationship inference in genome-wide association studies.
Bioinformatics. 2010;26(22):2867-2873.
PMID: 20926424
McCormick EM, Lott MT, Dulik MC, Shen L, Attimonelli M, Vitale O, et al. Specifications of the ACMG/AMP standards and guidelines for mitochondrial DNA variant interpretation.
Human Mutation. 2020;41(12):2028-2057.
PMID: 33058415
### Bring family-aware variant interpretation into your clinical workflow
Folklore's Family Analysis module integrates with your existing Variant Analysis pipeline. Schedule a discussion with our scientific team to assess the fit for your laboratory.
Contact us Nuclear ACMG methodology mtDNA methodology For geneticists How it works

### Links and cited sources

- [DNAnexus](https://github.com/dnanexus-rnd/GLnexus)
- [Purcell Lab, Harvard](https://www.cog-genomics.org/plink/)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)
- [PMID: 29543229](https://pubmed.ncbi.nlm.nih.gov/29543229/)
- [PMID: 30192042](https://pubmed.ncbi.nlm.nih.gov/30192042/)
- [PMID: 27236918](https://pubmed.ncbi.nlm.nih.gov/27236918/)
- [PMID: 37352859](https://pubmed.ncbi.nlm.nih.gov/37352859/)
- [PMID: 22781750](https://pubmed.ncbi.nlm.nih.gov/22781750/)
- [DOI: 10.1101/2021.04.05.438434](https://doi.org/10.1101/2021.04.05.438434)
- [PMID: 17701901](https://pubmed.ncbi.nlm.nih.gov/17701901/)
- [PMID: 20926424](https://pubmed.ncbi.nlm.nih.gov/20926424/)
- [PMID: 33058415](https://pubmed.ncbi.nlm.nih.gov/33058415/)
- [Contact us](https://folklore.helena.bio/contact)
- [Nuclear ACMG methodology](https://folklore.helena.bio/methodology)
- [mtDNA methodology](https://folklore.helena.bio/methodology/mtdna)
- [For geneticists](https://folklore.helena.bio/for-geneticists)
- [How it works](https://folklore.helena.bio/how-it-works)

---

## Page: mtDNA Variant Interpretation under MMDWG 2020 | Folklore

Source: https://folklore.helena.bio/methodology/mtdna
Canonical: https://folklore.helena.bio/methodology/mtdna
Description: Folklore's mtDNA method: MMDWG 2020 criteria, heteroplasmy and haplogroup-aware frequency rules, APOGEE 2, MitoTIP, PON-mt-tRNA, NUMT flags, and curated evidence.

## Mitochondrial DNA Variant Interpretation
MMDWG Module v 1.0.0 | Classifier v 3.30.0 | Updated May 2026
Folklore classifies mitochondrial variants under the ClinGen Mitochondrial Disease Working Group 2020 specification ( McCormick EM et al., Hum Mutat. 2020;41(12):2028-2057. PMID: 33058415. DOI: 10.1002/humu.24107 ). The method records the thresholds, reference datasets, automated criteria, and reviewer-curated evidence used for each result.
Nuclear and mitochondrial variants follow separate classifiers. mtDNA uses MMDWG 2020; nuclear variants use ACMG/AMP 2015. Each result records its framework, module version, criteria, and evidence strengths.
Contents
01 Pipeline Overview 02 Folklore Differentiators 03 Reference Databases 04 MMDWG Classification 05 Automated Criteria (11) 06 Curated Criteria (10) 07 Excluded Criteria (7) 08 Computational Predictors 09 Combining Rules 10 Mitochondrial Genome 11 Limitations 12 Version History 13 References
### Pipeline Overview
Seven-stage mtDNA classification pipeline. mtDNA variants are routed to the MMDWG framework, annotated against the curated mitochondrial reference set, evaluated against the eleven automated criteria, integrated with reviewer-submitted curation evidence, and resolved through Richards 2015 combining rules with mtDNA-specific overrides.
1
mtDNA Variant Identification
Variants residing on the mitochondrial chromosome are identified during pipeline ingestion. Six standard mtDNA chromosome aliases are accepted (chrM, M, MT, chrMT, NC_012920.1, J01415.2) and harmonized to a single canonical reference. Nuclear pseudogene candidates (NUMT regions known from published catalogs) are flagged for visibility without blocking classification.
2
Two-Framework Architecture
Folklore applies two independent classification frameworks in the same analysis: ACMG/AMP 2015 (Richards) for nuclear variants, and ClinGen MMDWG 2020 (McCormick) for mtDNA variants. Each variant carries explicit framework provenance in its output, so reviewing geneticists know which interpretive framework produced the classification.
3
mtDNA Reference Annotation
mtDNA variants are annotated against the curated mitochondrial reference set: ClinVar mtDNA assertions, gnomAD v3.1 mtDNA population frequencies with heteroplasmy distribution, APOGEE 2 missense scores, MitoTIP and PON-mt-tRNA tRNA scores, mt-Phylotree haplogroup-defining annotations, and Ensembl mtDNA gene coordinates with biotype.
4
MMDWG Criteria Evaluation
Eleven automated criteria are evaluated per variant per McCormick 2020: PVS1, PS1, PM2_Supporting, PM4, PM5_Moderate, PM5_Supporting, PP3, BA1, BS1, BP2_Supporting, BP4, BP7. Strength tiers follow ClinGen mtDNA VCEP v1.0.0 conventions. Tier suffixes appear on criteria emitted at non-default ACMG strength.
5
Curated Evidence Integration
Ten manual-curation criteria (PS2, PS3, PS4, PM6, PP1, PP4, BS2, BS3, BS4, BP5) accept geneticist input through an audit-trailed curation interface with explicit strength tiers (Very Strong, Strong, Moderate, Supporting). Submitted curation contributes to the same combining-rule evaluation as the automated criteria.
6
Combining Rules and Override
Richards 2015 combining rules (eighteen rules, retained unchanged per McCormick Section 5.2) determine the final five-tier classification (P, LP, VUS, LB, B). BA1 acts as stand-alone benign override. ClinVar P/LP/B/LB assertions at two or more review stars provide an override path (mtDNA-specific stricter threshold per McCormick Section 5.1). Conflicting pathogenic and benign evidence at moderate or above produces VUS.
7
Persistence and Audit Trail
The variant record stores the classification, criteria string with strength tiers, framework provenance (mmdwg_2020), MMDWG module version, and classification confidence. Categorized notes record non-default decisions such as PP1 disqualification, BA1 conservative fallback, and manual curation contribution.
### Folklore Differentiators
The ClinGen MMDWG 2020 specification is the standard. The choices below describe how Folklore addresses interpretive challenges that the specification raises but does not prescribe a single implementation for. Each differentiator is described at the conceptual level; the specific implementation is part of the Folklore platform.
Personalized BA1 via Patient Haplogroup Ancestry
McCormick 2020 Section 5.2.10 requires BA1 to fire either when the variant exceeds 1% allele frequency or when the variant is haplogroup-defining for the patient haplogroup. A simple "haplogroup-defining anywhere" flag is insufficient because the clinical interpretation depends on whether the patient actually carries that haplogroup. Folklore infers the patient haplogroup from the VCF and evaluates each candidate variant against the patient ancestral haplogroup branch (not against the global population). When haplogroup inference is uncertain, Folklore falls back to a conservative evaluation that applies BA1 only to the high-frequency path, preserving sensitivity for rare variants. Illustrative example: m.3394T>C is haplogroup-defining for branch M9a but pathogenic in branch B4c (Ji 2012 PMID:22577229; Kang 2016 PMID:27220472). The same variant correctly receives BA1 in an M9a patient and remains pathogenic-eligible in a B4c patient.
NUMT-Aware Classification with Visibility Flags
Nuclear mitochondrial DNA segments (NUMTs) are stretches of mtDNA-like sequence integrated into the nuclear genome (Calabrese 2017 PMID:24723423). Variants overlapping known NUMT regions can produce false-positive mtDNA calls when sequencing depth or coverage characteristics suggest nuclear origin. Folklore maintains a curated catalog of published NUMT regions and surfaces NUMT-overlap and depth-anomaly flags for the reviewing geneticist. The flags are advisory rather than blocking: classification still proceeds under MMDWG, and the geneticist makes the final pseudogene determination informed by the visibility signal.
Two-Framework Architecture with Explicit Provenance
Mitochondrial and nuclear genetics differ in inheritance, mutation rate, heteroplasmy, and allele frequency interpretation. McCormick 2020 modifies, downgrades, or excludes ACMG criteria for mtDNA. Nuclear variants use ACMG/AMP 2015, mtDNA variants use ClinGen MMDWG 2020, and the classification_framework field identifies the module used for each result. The modules do not share thresholds or criteria sets.
Manual Curation Workflow with Strength Tiers
Ten of the twenty-five ACMG criteria require evidence not derivable from the VCF and reference data alone: parental segregation (PS2, PM6, PP1, BS4), functional studies (PS3, BS3), case-control statistics (PS4), phenotype specificity (PP4), independent healthy-individual observation (BS2), and alternative-diagnosis context (BP5). Folklore exposes these criteria through an audit-trailed curation interface with explicit ClinGen mtDNA VCEP v1.0.0 strength tiers: four-tier (Very Strong / Strong / Moderate / Supporting) for PS2 and PM6 per McCormick Section 5.2.3 and 5.2.6; two-tier (Strong / Supporting) for BS2 per McCormick Section 5.2.10; specification-defined tiers for PS3, PS4, PM6, PP1, PP4, BS3, BS4, and BP5. PS4 and PP1 strength can also be derived inline from proband counts and haplogroup-diversity statistics where these are available.
Categorized Transparency Audit Trail
A categorized audit note records each non-default classification decision, including PVS1 NMD branch removal per Abou Tayoun (mRNA path only), BA1 conservative fallback for uncertain haplogroup inference, PM5_Supporting tRNA or rRNA application per McCormick Section 5.2.2, named-reviewer curation, and ClinVar override gating. These notes support review and reconstruction of the decision path.
### Reference Databases
All mtDNA reference data is stored locally on EU-based infrastructure. No variant data is sent to external services during processing. Database versions are fixed per deployment and documented here.
mt_clinvar_variants
NCBI ClinVar (mtDNA subset, GRCh38) 3,122 mtDNA records
ClinVar clinical significance assertions filtered to chrM. Distribution: 78 Pathogenic, 109 Likely Pathogenic, 1,269 VUS, 612 Likely Benign, 1,017 Benign, 34 Not_provided, 3 Other. Review stars: 119 zero-star, 2,452 one-star, 238 two-star, 313 three-star. Folklore retains both the canonical clinical significance code (P, LP, VUS, LB, B) and the original ClinVar string for audit transparency.
Used by: PS1 (same amino acid change as known pathogenic, mRNA path only), PM5_Moderate (different missense at same residue, mRNA), PM5_Supporting (same nucleotide position, tRNA / rRNA, McCormick Section 5.2.2), BP2_Supporting (another mtDNA variant in the same session is ClinVar P/LP), and the ClinVar override gate at two or more review stars (stricter than the nuclear one-star override per McCormick Section 5.1).
Source: NCBI ClinVar
mt_population_frequencies
gnomAD v3.1 mtDNA call set 18,164 records (56,434-allele baseline)
Per-variant homoplasmic and heteroplasmic allele frequencies, per-individual maximum heteroplasmy, haplogroup-defining flag, common-low-heteroplasmy flag, MitoTIP raw score, MitoTIP tRNA prediction class, PON-mt-tRNA prediction class, and PON-mt-tRNA pathogenicity probability. Distribution: 124 variants > 5% allele frequency, 429 variants > 1%, 4,299 haplogroup-defining variants, 374 common-low-heteroplasmy variants.
Used by: BA1 (allele frequency > 1% stand-alone benign, or haplogroup-defining for the patient ancestral branch, McCormick Section 5.2.10), BS1 (allele frequency 0.5-0.99% Strong benign, McCormick Section 5.2.5), PM2_Supporting (allele frequency < 0.00002 absent-from-controls, downgraded from Moderate per McCormick Section 5.2.5), PP3 tRNA path (MitoTIP > 50th percentile combined with PON-mt-tRNA > 0.5), BP4 tRNA path (mirror inverse).
Source: gnomAD
mt_apogee2_scores
MitImpact 3.1.3 24,190 missense scores across 13 protein-coding genes
APOGEE 2 numeric pathogenicity scores for every possible mtDNA missense substitution across the 13 protein-coding genes. Score range 0.0 to 0.981 with associated probability and class label. Auxiliary predictor scores (AlphaMissense, MitoClass1, PolyPhen2, SIFT, FATHMM, CADD) and MitoMap disease status are retained alongside for clinical reference. APOGEE 1 score is retained as a legacy column for backward compatibility but is not used by the production PP3 / BP4 mRNA path.
Used by: PP3 mRNA path (APOGEE 2 score > 0.5 triggers Supporting pathogenic per McCormick Figure 3), BP4 mRNA path (APOGEE 2 score <= 0.5 triggers Supporting benign, mirror inverse).
Source: MitImpact
mt_gene_coordinates
Ensembl mtDNA annotation 37 genes (13 protein-coding + 22 Mt_tRNA + 2 Mt_rRNA)
Coordinate boundaries and biotype classification for every mtDNA gene. Biotype distribution: 22 Mt_tRNA, 13 protein_coding, 2 Mt_rRNA (MT-RNR1, MT-RNR2). Strand information is retained to handle the reverse-strand light-chain genes (MT-TQ, MT-TA, MT-TN, MT-TC, MT-TY, MT-TS1, MT-ND6, MT-TE, MT-TP).
Used by: Gene-class dispatch for criteria with biotype-specific behavior: PVS1, PS1, PM4, PM5_Moderate, PM5_Supporting, PP3, BP4, and BP7 each have biotype-aware logic per McCormick 2020.
Source: Ensembl
mt_haplogroups
PhyloTree build 17 (FU1a release) 17,439 records, 6,355 distinct haplogroups, 2,844 parents, 621 back-mutations
Maternal-haplogroup phylogenetic tree with per-edge variant assignments, parent-haplogroup links, back-mutation flags, insertion flags, and variant modifier annotations. Provides the data structure required to evaluate whether a candidate variant is haplogroup-defining for any ancestor of the patient haplogroup.
Used by: BA1 personalized path (variant is haplogroup-defining for the patient maternal lineage). When haplogroup inference is unavailable for a sample, Folklore falls back to a conservative BA1 evaluation that does not over-apply the stand-alone benign override.
Source: PhyloTree
mt_helix_curated
Bootstrap (manual curation table) Reserved for clinical curation overrides
Reserved table for Folklore clinical-curation overrides of ClinVar assertions. Each entry records the canonical Folklore classification, the rationale, and a flag indicating whether the entry supersedes a conflicting ClinVar record. The table is empty in the public reference set and is populated only by named reviewers through the audit-trailed curation interface.
Used by: Curation override layer applied before the ClinVar override gate.
Source: Helena Bioinformatics
nuclear_numt_known_regions
Lareau master catalog 784 records across 23 chromosomes
Catalog of nuclear mitochondrial DNA (NUMT) regions with chromosome, start, end, length, and source attribution. Used to flag candidate pseudogene calls during chromosome routing.
Used by: NUMT visibility flags surfaced for reviewer audit. NUMT-overlap and depth-anomaly flags are advisory; classification proceeds under MMDWG and the reviewing geneticist makes the final pseudogene determination.
Source: Published NUMT catalogs
### MMDWG Classification
Variant classification follows the ClinGen Mitochondrial Disease Working Group 2020 specification with twenty-five evidence criteria evaluated systematically. Eleven criteria are fully automated; ten require manual curation by the reviewing geneticist; seven are excluded per McCormick Section 5.3.
Classification Priority Order
Classification logic is applied in strict priority order. Higher-priority rules are evaluated first, and the first matching rule determines the final classification:
1
BA1 Stand-alone
Allele frequency above 1% homoplasmic, or variant is haplogroup-defining for the patient ancestral branch (Folklore personalized BA1 path). BA1 is the only stand-alone criterion in the MMDWG framework and cannot be overridden by any other evidence, including ClinVar assertions.
2
Conflicting Evidence
If a variant has pathogenic evidence at moderate strength or above (PVS, PS, or PM criteria triggered) AND strong benign evidence (BS criteria triggered), the variant is classified as VUS regardless of the individual evidence strength. This is a conservative approach that prioritizes clinical safety.
3
ClinVar Override
ClinVar P/LP/B/LB classification is applied only when no conflicting computational evidence exists. Requires two or more review stars (stricter than the nuclear one-star default per McCormick Section 5.1). ClinVar VUS does not override computational classification.
4
Combining Rules
The Richards 2015 eighteen combining rules (Table 5) are evaluated against the criteria triggered for the variant. Rules are retained unchanged for mtDNA per McCormick Section 5.2.
5
Default
Variants that do not meet any of the above criteria are classified as Uncertain Significance (VUS).
Classification Output
Each variant receives one of five standard classifications, a list of all criteria triggered with explicit strength tiers (e.g. "PVS1, PM2_Supporting, PP3"), an explicit framework provenance label (mmdwg_2020), the MMDWG module version, and a continuous classification confidence score:
Pathogenic
Likely Pathogenic
VUS
Likely Benign
Benign
### Automated Criteria (11 of 25)
These criteria are evaluated automatically for every mtDNA variant. Conditions and thresholds follow the McCormick 2020 specification. Each criterion lists the reference data sources it depends on and any biotype-specific behavior.
PVS1 Very Strong / Strong
Null variant in a gene where loss-of-function is a known disease mechanism
Conditions
Applies to mtDNA protein-coding (mRNA) genes only.
Truncating consequence: stop_gained or frameshift.
Strength modulation per Abou Tayoun 2018 (PMID:30192042) decision tree adapted for mtDNA: McCormick 2020 retains the Abou Tayoun framework with the explicit removal of the NMD branch (mtDNA mRNAs are not subject to canonical nuclear NMD).
Exclusions
tRNA and rRNA genes (PVS1 not defined under McCormick for these biotypes).
Stop-retained and stop-lost variants.
Non-canonical splice variants (mtDNA does not undergo nuclear-style splicing).
Reference data: Ensembl mtDNA gene coordinates (biotype dispatch), VEP (consequence)
Per McCormick Section 5.2.1, the nuclear NMD branch of the Abou Tayoun decision tree is removed for mtDNA. mtDNA truncating variants are evaluated for their effect on the mature protein only.
PS1 Strong
Same amino acid change as an established pathogenic variant
Conditions
Applies to mtDNA protein-coding (mRNA) genes only.
ClinVar Pathogenic or Likely Pathogenic at two or more review stars at the same protein residue with the same amino acid substitution.
Exclusions
tRNA and rRNA genes (PS1 not applicable for non-coding biotypes).
Reference data: mt_clinvar_variants
McCormick Section 5.2.1 retains PS1 with the standard amino-acid match definition for mRNA genes. The two-star ClinVar threshold is stricter than the nuclear default and aligns with the elevated-evidence requirement specified in McCormick Section 5.1.
PM2_Supporting Supporting
Absent from population databases
Conditions
gnomAD v3.1 mtDNA homoplasmic allele frequency below the absent-from-controls threshold (default 0.00002), or variant absent from gnomAD entirely.
Reference data: mt_population_frequencies
McCormick Section 5.2.5 explicitly downgrades PM2 from Moderate (the nuclear default) to Supporting for mtDNA, reflecting the reduced statistical power of population-frequency arguments at extremely low frequencies in mtDNA.
PM4 Moderate
Protein length change in a non-repetitive region
Conditions
Applies to mtDNA protein-coding (mRNA) genes only.
In-frame insertion or in-frame deletion that alters protein length in a region not annotated as repetitive or low-complexity.
Exclusions
tRNA and rRNA genes (PM4 not applicable, length changes are evaluated through PM5_Supporting and PP3 tRNA path).
Reference data: Ensembl mtDNA gene coordinates (biotype dispatch), VEP (consequence)
PM5_Moderate Moderate
Different missense change at a residue with established pathogenic missense
Conditions
Applies to mtDNA protein-coding (mRNA) genes only.
ClinVar Pathogenic or Likely Pathogenic at two or more review stars at the same protein residue but with a different amino acid substitution.
Exclusions
tRNA and rRNA genes (handled by PM5_Supporting under McCormick Section 5.2.2).
Reference data: mt_clinvar_variants
PM5_Supporting Supporting
Different nucleotide change at a position with established pathogenic variant in tRNA / rRNA
Conditions
Applies to mtDNA tRNA and rRNA genes only.
ClinVar Pathogenic or Likely Pathogenic at two or more review stars at the same nucleotide position but with a different alternate allele.
Exclusions
Protein-coding (mRNA) genes (covered by PM5_Moderate at residue level).
Reference data: mt_clinvar_variants, Ensembl mtDNA gene coordinates (biotype dispatch)
McCormick Section 5.2.2 introduces PM5_Supporting specifically for tRNA and rRNA biotypes where amino-acid-level matching is not applicable. Same-nucleotide-different-allele matching at Supporting strength.
PP3 Supporting
Computational evidence supports a deleterious effect
Conditions
mRNA path: APOGEE 2 score above the pathogenic threshold (default 0.5). Per McCormick Figure 3, APOGEE is the recommended in silico tool for mtDNA missense variants.
tRNA path: MitoTIP score above the 50th percentile combined with PON-mt-tRNA pathogenicity probability above 0.5.
rRNA path: not evaluated. McCormick 2020 does not endorse a calibrated in silico predictor for rRNA, so PP3 is not applied for the two rRNA genes.
Exclusions
rRNA biotype (PP3 not applied per McCormick).
PVS1 active for the same variant (no double-counting per ClinGen SVI 2023).
Reference data: mt_apogee2_scores, mt_population_frequencies (MitoTIP, PON-mt-tRNA), Ensembl mtDNA gene coordinates (biotype dispatch)
BA1 Stand-alone
Allele frequency consistent with stand-alone benign
Conditions
gnomAD v3.1 mtDNA homoplasmic allele frequency above 1% (McCormick Section 5.2.10), OR the variant is haplogroup-defining for the patient ancestral haplogroup branch (Folklore personalized BA1 path).
Reference data: mt_population_frequencies, mt_haplogroups
McCormick Section 5.2.10 lowers the BA1 threshold from the nuclear 5% to 1% for mtDNA. Folklore evaluates the haplogroup-defining clause against the patient ancestral branch (not against any branch globally) so that variants haplogroup-defining for an unrelated branch do not silently receive BA1 in the patient. When haplogroup inference is uncertain, Folklore falls back to a conservative path that applies BA1 only to the > 1% high-frequency arm.
BS1 Strong
Allele frequency above expected for the disorder
Conditions
gnomAD v3.1 mtDNA homoplasmic allele frequency between 0.5% and 0.99%.
Exclusions
Variant satisfies BA1 (BA1 takes precedence as stand-alone benign).
Reference data: mt_population_frequencies
McCormick Section 5.2.5 specifies the 0.5% Strong-benign threshold for mtDNA, separate from the 1% BA1 threshold. The two together replace the nuclear 5% / 1% pair.
BP2_Supporting Supporting
Observed alongside an established pathogenic mtDNA variant in the same individual
Conditions
Another variant in the same patient session is ClinVar Pathogenic or Likely Pathogenic in mtDNA at two or more review stars.
Reference data: mt_clinvar_variants
mtDNA-specific reframing of BP2: in the absence of nuclear-style trans / cis configuration, the presence of an established pathogenic mtDNA variant in the same individual provides Supporting benign evidence for an additional candidate variant in the same mitochondrial genome.
BP4 Supporting
Computational evidence suggests no impact
Conditions
mRNA path: APOGEE 2 score below the benign threshold (default 0.5). Mirror inverse of PP3 mRNA path.
tRNA path: MitoTIP score at or below the 50th percentile, combined with PON-mt-tRNA pathogenicity probability at or below 0.5.
Exclusions
rRNA biotype (BP4 not applied per McCormick, no calibrated rRNA predictor).
PVS1 active for the same variant.
Reference data: mt_apogee2_scores, mt_population_frequencies, Ensembl mtDNA gene coordinates (biotype dispatch)
BP7 Supporting
Synonymous variant without predicted impact
Conditions
Applies to mtDNA protein-coding (mRNA) genes only.
Synonymous consequence (no amino acid change).
No predicted splice impact (mtDNA does not undergo nuclear-style splicing; the splice clause from the nuclear ACMG version is not evaluated).
Exclusions
tRNA and rRNA genes (synonymous concept not applicable).
Reference data: Ensembl mtDNA gene coordinates (biotype dispatch), VEP (consequence)
### Curated Criteria (10 of 25)
These criteria require evidence not derivable from VCF and reference data alone: maternal segregation, functional assays, case-control statistics, phenotype specificity, healthy-individual observation, and alternative-diagnosis context. Reviewing geneticists submit curation evidence through the Folklore audit-trailed interface with explicit ClinGen mtDNA VCEP v1.0.0 strength tiers.
PS2 Very Strong / Strong / Moderate / Supporting (4-tier per McCormick Section 5.2.3)
De novo variant with confirmed maternity
mtDNA is maternally inherited. PS2 strength is determined by the number of independent maternally-confirmed de novo observations and the segregation evidence supporting de novo origin. The four-tier strength scale is selected by the reviewing geneticist through the curation interface.
PS3 Supporting only (per McCormick Section 5.2.4)
Functional studies show a deleterious effect
McCormick caps PS3 at Supporting strength for mtDNA pending future calibration of mitochondrial functional assays. Reviewing geneticists submit published functional evidence through the curation interface and cite the supporting PMID.
PS4 Strong / Moderate / Supporting (3-tier, can be derived inline from proband and haplogroup-diversity counts)
Increased prevalence in affected versus controls
PS4 strength is selected by the reviewing geneticist or, where proband-count and haplogroup-diversity statistics are available in the session metadata, derived inline by the curation interface. Curated values always supersede inline derivations and are recorded with the derivation source for audit transparency.
PM6 Very Strong / Strong / Moderate / Supporting (4-tier per McCormick Section 5.2.6)
Assumed de novo without maternity confirmation
Counterpart to PS2 when maternal confirmation is not available. The four-tier scale matches PS2.
PP1 Strong / Moderate / Supporting (3-tier, can be derived inline from maternal-member count)
Cosegregation with disease in maternal relatives
McCormick Section 5.2.7 specifies a homoplasmic-in-all-maternal-members disqualifier: if the variant is homoplasmic in every maternal relative regardless of phenotype, PP1 cannot fire because cosegregation is not informative. Folklore enforces this disqualifier automatically when the curated maternal-member counts indicate uniform homoplasmy. The audit trail records the disqualification when it applies.
PP4 Supporting (specification default)
Patient phenotype highly specific for a mitochondrial disorder
Reviewing geneticist confirms phenotype specificity through the curation interface. mtDNA-specific phenotype patterns (Leber hereditary optic neuropathy, MELAS, MERRF, NARP, Leigh syndrome) provide the typical evidence base.
BS2 Strong / Supporting (2-tier per McCormick Section 5.2.10)
Observed in a healthy maternal relative
McCormick omits the Very Strong and Moderate tiers from BS2 for mtDNA. The two-tier scale is selected by the reviewing geneticist.
BS3 Supporting only (mirror of PS3 cap)
Functional studies show no deleterious effect
Strength capped at Supporting per McCormick. Reviewing geneticists submit published benign functional evidence with the supporting PMID.
BS4 Strong (specification default)
Lack of segregation in affected maternal relatives
Strong benign evidence when affected maternal relatives do not carry the variant. Selected by the reviewing geneticist through the curation interface.
BP5 Supporting (specification default)
Variant found in a case with an alternate molecular basis
Reviewing geneticist confirms the alternate diagnosis through the curation interface, with the alternate diagnosis recorded in the audit trail.
### Excluded Criteria (7 of 25)
McCormick 2020 Section 5.3 excludes seven nuclear ACMG criteria from mtDNA interpretation because their underlying assumptions do not hold for the mitochondrial genome. Folklore does not apply any of these criteria to mtDNA variants.
PM1
Functional-domain hot-spot evidence for mtDNA tRNA variants is captured through MitoTIP (PP3 tRNA path), which is itself a domain-aware computational tool. Adding PM1 on top would double-count the same biological argument. McCormick Section 5.3 excludes PM1 to prevent this double-counting.
PM3
PM3 (in trans with a pathogenic variant for recessive disorders) presupposes biparental inheritance and a recessive autosomal-recessive mechanism. mtDNA is maternally inherited and does not have an autosomal-recessive analog, so PM3 is biologically inapplicable.
PP2
PP2 requires that missense variants be a common mechanism of disease in a constraint-quantified gene. mtDNA exhibits high sequence variability across haplotypes, making gene-level missense constraint statistics uninformative. McCormick Section 5.3 excludes PP2.
PP5
ClinGen Sequence Variant Interpretation Working Group (Biesecker 2018, PMID:29543229) recommended retiring PP5 as a redundant criterion subsumed by PS1, PM5, and the ClinVar override gate. McCormick adopts this recommendation for mtDNA.
BP1
BP1 applies to genes for which primarily truncating variants cause disease, used to argue against missense pathogenicity. The mtDNA disease catalog includes substantial pathogenic missense burden across the protein-coding genes, so the BP1 premise does not hold.
BP3
BP3 covers in-frame indels in repetitive or low-complexity nuclear regions (HVR-style). mtDNA HVR (D-loop) regions are clinically excluded from interpretation entirely, so BP3 is not needed as a separate Supporting benign criterion.
BP6
Same ClinGen SVI removal rationale as PP5: subsumed by the ClinVar override gate at two or more review stars and not needed as a separate Supporting benign criterion.
### Computational Predictors (PP3 / BP4)
PP3 and BP4 dispatch by gene biotype. mtDNA protein-coding genes use APOGEE 2 (Bianco 2023). mtDNA tRNA genes use a combined MitoTIP plus PON-mt-tRNA evaluation per McCormick Figure 3. mtDNA rRNA genes have no calibrated predictor and PP3 / BP4 are not applied for the two rRNA genes.
APOGEE 2 (mRNA Path)
APOGEE 2 (Bianco 2023, PMID:PMC10439926) is a multi-layer machine-learning model calibrated for the prediction of mtDNA missense variant pathogenicity. Folklore uses precomputed APOGEE 2 scores from the MitImpact 3.1.3 distribution covering all 24,190 possible missense substitutions across the 13 protein-coding genes. Score above the pathogenic threshold (default 0.5) triggers PP3 Supporting; score at or below the threshold triggers BP4 Supporting. APOGEE 1 is retained as a legacy column for backward compatibility but is not used by the production PP3 / BP4 logic.
MitoTIP + PON-mt-tRNA (tRNA Path)
McCormick Figure 3 specifies a combined MitoTIP (Sonney 2017, PMID:28809476) plus PON-mt-tRNA (Niroula and Vihinen 2019, PMID:31504522) evaluation for tRNA variants. PP3 fires when MitoTIP exceeds the 50th percentile and PON-mt-tRNA pathogenicity probability exceeds 0.5. BP4 is the mirror inverse. Both tools are calibrated specifically for mtDNA tRNA variants and provide complementary domain-aware and probability-based evaluations.
rRNA Path: Not Evaluated
McCormick 2020 does not endorse a calibrated in silico predictor for the two mtDNA rRNA genes (MT-RNR1, MT-RNR2). PP3 and BP4 are not applied for rRNA biotype variants. Pathogenicity assessment for rRNA variants relies on PVS1 (where applicable), PS1, PM2_Supporting, BA1, BS1, the audit-trailed curation criteria, and the ClinVar override gate.
### Combining Rules
The Richards 2015 eighteen combining rules (Table 5) are retained unchanged for mtDNA per McCormick Section 5.2. Strength tiers from McCormick (PM2_Supporting, PM5_Supporting, BS2 two-tier, PS3 / BS3 capped at Supporting) feed directly into the same combining-rule evaluation.
Pathogenic (8 rules)
P1 1 Very Strong (PVS) + >= 1 Strong (PS) P2 1 Very Strong (PVS) + >= 2 Moderate (PM) P3 1 Very Strong (PVS) + 1 Moderate (PM) + 1 Supporting (PP) P4 1 Very Strong (PVS) + >= 2 Supporting (PP) P5 >= 2 Strong (PS) P6 1 Strong (PS) + >= 3 Moderate (PM) P7 1 Strong (PS) + 2 Moderate (PM) + >= 2 Supporting (PP) P8 1 Strong (PS) + 1 Moderate (PM) + >= 4 Supporting (PP)
Likely Pathogenic (6 rules)
LP1 1 Very Strong (PVS) + 1 Moderate (PM) LP2 1 Strong (PS) + 1-2 Moderate (PM) LP3 1 Strong (PS) + >= 2 Supporting (PP) LP4 >= 3 Moderate (PM) LP5 2 Moderate (PM) + >= 2 Supporting (PP) LP6 1 Moderate (PM) + >= 4 Supporting (PP)
Benign (2 rules)
B1 1 Stand-alone (BA1) -- mtDNA threshold > 1% homoplasmic, or haplogroup-defining for patient branch B2 >= 2 Strong benign (BS1, BS2, BS3, BS4)
Likely Benign (2 rules)
LB1 1 Strong benign + 1 Supporting benign LB2 >= 2 Supporting benign
### The Mitochondrial Genome
The 16,569-base human mitochondrial genome encodes 37 genes: 13 protein-coding genes (subunits of OXPHOS complexes I, III, IV, V), 22 transfer-RNA genes (one per amino acid plus two for leucine and serine), and 2 ribosomal-RNA genes (12S and 16S). Gene biotype determines the criterion path under MMDWG.
Gene Biotype Start End Strand
MT-TF Mt_tRNA 577 647 +
MT-RNR1 Mt_rRNA 648 1,601 +
MT-TV Mt_tRNA 1,602 1,670 +
MT-RNR2 Mt_rRNA 1,671 3,229 +
MT-TL1 Mt_tRNA 3,230 3,304 +
MT-ND1 protein_coding 3,307 4,262 +
MT-TI Mt_tRNA 4,263 4,331 +
MT-TQ Mt_tRNA 4,329 4,400 -
MT-TM Mt_tRNA 4,402 4,469 +
MT-ND2 protein_coding 4,470 5,511 +
MT-TW Mt_tRNA 5,512 5,579 +
MT-TA Mt_tRNA 5,587 5,655 -
MT-TN Mt_tRNA 5,657 5,729 -
MT-TC Mt_tRNA 5,761 5,826 -
MT-TY Mt_tRNA 5,826 5,891 -
MT-CO1 protein_coding 5,904 7,445 +
MT-TS1 Mt_tRNA 7,446 7,514 -
MT-TD Mt_tRNA 7,518 7,585 +
MT-CO2 protein_coding 7,586 8,269 +
MT-TK Mt_tRNA 8,295 8,364 +
MT-ATP8 protein_coding 8,366 8,572 +
MT-ATP6 protein_coding 8,527 9,207 +
MT-CO3 protein_coding 9,207 9,990 +
MT-TG Mt_tRNA 9,991 10,058 +
MT-ND3 protein_coding 10,059 10,404 +
MT-TR Mt_tRNA 10,405 10,469 +
MT-ND4L protein_coding 10,470 10,766 +
MT-ND4 protein_coding 10,760 12,137 +
MT-TH Mt_tRNA 12,138 12,206 +
MT-TS2 Mt_tRNA 12,207 12,265 +
MT-TL2 Mt_tRNA 12,266 12,336 +
MT-ND5 protein_coding 12,337 14,148 +
MT-ND6 protein_coding 14,149 14,673 -
MT-TE Mt_tRNA 14,674 14,742 -
MT-CYB protein_coding 14,747 15,887 +
MT-TT Mt_tRNA 15,888 15,953 +
MT-TP Mt_tRNA 15,956 16,023 -
Coordinates are GRCh38 chrM (NC_012920.1, the revised Cambridge Reference Sequence). Source: Ensembl mtDNA annotation.
### Limitations and Disclaimers
Folklore is a clinical decision support tool, not a diagnostic device. Every classification requires review and confirmation by a qualified clinical geneticist.
Ten of the twenty-five MMDWG criteria require evidence not derivable from VCF and reference data alone (PS2, PS3, PS4, PM6, PP1, PP4, BS2, BS3, BS4, BP5). These criteria are entered through the audit-trailed curation interface by the reviewing geneticist with explicit ClinGen mtDNA VCEP v1.0.0 strength tiers.
Heteroplasmy interpretation is patient-specific. Homoplasmic, heteroplasmic, and tissue-restricted heteroplasmy patterns each carry different clinical implications. Folklore exposes max-heteroplasmy and per-individual heteroplasmy distribution for every variant, but the clinical interpretation of heteroplasmy levels remains a judgment call by the reviewing geneticist.
Patient haplogroup inference depends on the input VCF. When haplogroup-informative variants are absent or insufficient, Folklore falls back to a conservative BA1 evaluation rather than guessing the haplogroup. The audit trail records the fallback when it applies.
Nuclear mitochondrial DNA segments (NUMTs) can produce false-positive mtDNA calls. Folklore flags NUMT-overlap and depth-anomaly signals for reviewer audit but does not block classification. The reviewing geneticist makes the final pseudogene determination.
Population frequency data from gnomAD v3.1 mtDNA may underrepresent certain ancestral populations. Allele frequency thresholds should be interpreted in the context of the patient ancestry.
ClinVar mtDNA assertions vary in quality and currency. The two-or-more review-star threshold for the override gate (stricter than the nuclear default per McCormick Section 5.1) reduces the impact of unreviewed single-submitter assertions but cannot eliminate it entirely.
PS3 and BS3 are capped at Supporting strength per McCormick 2020 pending future calibration of mitochondrial functional assays. Functional evidence that would qualify for Strong or Moderate strength under the nuclear ACMG framework is currently down-tiered.
Haplogrep3-based haplogroup inference and HmtVAR external annotation are reserved for a future deployment phase. Production currently relies on the maternal-haplogroup phylogenetic tree (PhyloTree FU1a) and the curated reference set described above.
Results should be interpreted in the context of the patient's clinical presentation, family history, available mitochondrial functional studies, and other clinical information.
### Version History
The version history records changes to the methodology. The MMDWG module version is independent of the main classifier engine version.
MMDWG v1.0.0 Current May 2026
Production deployment of the ClinGen MMDWG 2020 mtDNA classifier (McCormick 2020 PMID:33058415). Eleven automated criteria, ten audit-trailed curation criteria with explicit ClinGen mtDNA VCEP v1.0.0 strength tiers, seven excluded per McCormick Section 5.3.
Two-framework architecture: nuclear ACMG/AMP 2015 (Richards) and mtDNA MMDWG 2020 (McCormick) operate as independent classification modules with explicit framework provenance on every variant in the output.
Personalized BA1 path: the haplogroup-defining clause from McCormick Section 5.2.10 is evaluated against the patient ancestral haplogroup branch using PhyloTree build 17 (FU1a release). Conservative fallback when haplogroup inference is uncertain.
NUMT-aware classification: a curated catalog of 784 NUMT regions across 23 chromosomes feeds visibility flags surfaced for reviewer audit. Classification proceeds under MMDWG; pseudogene determination remains a geneticist judgment call.
Audit-trailed manual curation interface for the ten curated criteria (PS2, PS3, PS4, PM6, PP1, PP4, BS2, BS3, BS4, BP5) with strength tier selection and PMID citation capture.
Categorized transparency notes: every non-default decision (PVS1 NMD branch removal, BA1 conservative fallback, PP1 disqualification, PM5_Supporting tRNA / rRNA, manual curation contribution) is recorded for reviewer audit.
### References
McCormick EM, Lott MT, Dulik MC, Shen L, Attimonelli M, Vitale O, et al. Specifications of the ACMG/AMP standards and guidelines for mitochondrial DNA variant interpretation.
Human Mutation. 2020;41(12):2028-2057.
PMID: 33058415
Richards S, Aziz N, Bale S, Bick D, Das S, Gastier-Foster J, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology.
Genetics in Medicine. 2015;17(5):405-424.
PMID: 25741868
Abou Tayoun AN, Pesaran T, DiStefano MT, Oza A, Rehm HL, Biesecker LG, Harrison SM. Recommendations for interpreting the loss of function PVS1 ACMG/AMP variant criterion.
Human Mutation. 2018;39(11):1517-1524.
PMID: 30192042
Bianco SD, Parca L, Petrizzelli F, Biagini T, Giovannetti A, Liorni N, et al. APOGEE 2: multi-layer machine-learning model for the interpretable prediction of mitochondrial missense variants.
Nature Communications. 2023;14:5562.
PMID: PMC10439926
Sonney S, Leipzig J, Lott MT, Zhang S, Procaccio V, Wallace DC, Sondheimer N. Predicting the pathogenicity of novel variants in mitochondrial tRNA with MitoTIP.
PLoS Computational Biology. 2017;13(12):e1005867.
PMID: 28809476
Niroula A, Vihinen M. PON-mt-tRNA: a multifactorial probability-based method for classification of mitochondrial tRNA variations.
Nucleic Acids Research. 2016;44(5):2020-2027.
PMID: 31504522
Schoenherr S, Weissensteiner H, Kronenberg F, Forer L. Haplogrep 3 - phylogenetic analysis and quality control for mitochondrial DNA variants.
Nucleic Acids Research. 2023;51(W1):W263-W268.
PMID: PMC10320189
Calabrese FM, Simone D, Attimonelli M. Primates and mouse NumtS in the UCSC Genome Browser.
BMC Bioinformatics. 2012;13(Suppl 4):S15.
PMID: 24723423
Ji Y, Liang M, Zhang J, Zhang M, Zhu J, Meng X, et al. Mitochondrial haplotypes may modulate the phenotypic manifestation of the LHON-associated m.11778G>A mutation.
Mitochondrion. 2012;12(5):597-602.
PMID: 22577229
Kang X, Wei X, Jiang L, Niu C, Zhang J, Chen S, Meng D. Composition and variation analysis of the 16S rDNA in mitochondrial haplogroups.
Mitochondrion. 2016;30:60-65.
PMID: 27220472
Walker LC, Hoya M, Wiggins GAR, Lindy A, Vincent LM, Parsons MT, et al. Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: Recommendations from the ClinGen SVI Splicing Subgroup.
American Journal of Human Genetics. 2023;110(7):1046-1067.
PMID: 37352859
Biesecker LG, Harrison SM; ClinGen Sequence Variant Interpretation Working Group. The ACMG/AMP reputable source criteria for the interpretation of sequence variants.
Genetics in Medicine. 2018;20(12):1687-1688.
PMID: 29543229
### Questions About Our mtDNA Methodology?
Clinical geneticists, mitochondrial-disease specialists, and laboratory directors can contact us with technical questions about the mtDNA method.
Contact Us Nuclear ACMG Methodology For Geneticists How It Works

### Links and cited sources

- [NCBI ClinVar](https://www.ncbi.nlm.nih.gov/clinvar/)
- [gnomAD](https://gnomad.broadinstitute.org)
- [MitImpact](https://mitimpact.css-mendel.it)
- [Ensembl](https://www.ensembl.org)
- [PhyloTree](https://www.phylotree.org)
- [Helena Bioinformatics](https://www.helena.bio)
- [Published NUMT catalogs](https://www.ncbi.nlm.nih.gov/pubmed/24723423)
- [PMID: 33058415](https://pubmed.ncbi.nlm.nih.gov/33058415/)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)
- [PMID: 30192042](https://pubmed.ncbi.nlm.nih.gov/30192042/)
- [PMID: PMC10439926](https://pubmed.ncbi.nlm.nih.gov/10439926/)
- [PMID: 28809476](https://pubmed.ncbi.nlm.nih.gov/28809476/)
- [PMID: 31504522](https://pubmed.ncbi.nlm.nih.gov/31504522/)
- [PMID: PMC10320189](https://pubmed.ncbi.nlm.nih.gov/10320189/)
- [PMID: 24723423](https://pubmed.ncbi.nlm.nih.gov/24723423/)
- [PMID: 22577229](https://pubmed.ncbi.nlm.nih.gov/22577229/)
- [PMID: 27220472](https://pubmed.ncbi.nlm.nih.gov/27220472/)
- [PMID: 37352859](https://pubmed.ncbi.nlm.nih.gov/37352859/)
- [PMID: 29543229](https://pubmed.ncbi.nlm.nih.gov/29543229/)
- [Contact Us](https://folklore.helena.bio/contact)
- [Nuclear ACMG Methodology](https://folklore.helena.bio/methodology)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)
- [How It Works](https://folklore.helena.bio/how-it-works)

---

## Page: SV/CNV Methodology - Riggs 2020 ClinGen/ACMG Standard | Folklore

Source: https://folklore.helena.bio/methodology/sv
Canonical: https://folklore.helena.bio/methodology/sv
Description: Documentation of Folklore's mandatory structural-variant and copy-number-variant (SV/CNV) processing under the semiquantitative Riggs 2020 ClinGen/ACMG standard: the constitutional loss metric (Table 1), gain metric (Table 2), dosage-sensitivity map (Collins 2022, ClinGen Dosage), five-tier classification, and explicit non-applicability boundaries.

## SV/CNV Methodology
Framework Riggs 2020 | Updated June 2026
Documentation of how Folklore evaluates structural variants and copy-number variants (SV/CNV) under the Riggs 2020 joint ACMG and ClinGen technical standard. It describes the point-based loss metric (Table 1) and gain metric (Table 2), the dosage-sensitivity evidence, and the five-tier classification. This documentation is intended for clinical geneticists, laboratory directors, and accreditation auditors.
Contents
01 Where the SV pass sits 02 Loss metric (Table 1) 03 Gain metric (Table 2) 04 Dosage-sensitivity evidence 05 Classification thresholds 06 Reference data 07 Limitations and documented hooks 08 Version History 09 References
### Where the SV pass sits
The SV/CNV pass is the third classification framework over the variants table, alongside nuclear ACMG/AMP 2015 and mitochondrial MMDWG 2020. It scores only SV-typed rows: losses (DEL, or copy number below 2) are scored under Table 1, and gains (DUP, or copy number above 2) under Table 2; the two row sets are mutually exclusive, so no row is scored twice.
1
SV row routing
SV-typed rows are separated from nuclear and mitochondrial rows. Losses and gains are partitioned by mutually exclusive conditions (loss takes precedence on a contradictory row), so each SV row passes through exactly one of the two point metrics.
2
Per-section Riggs scoring
For each row the Riggs per-section points (genomic content, dosage sensitivity, gene number, population frequency, inheritance) are summed into a total score. The point values and thresholds are read from a single source of configuration constants; there are no inline clinical numbers in the code.
3
Five-tier verdict assembly
For applicable constitutional loss and gain rows, the total score is mapped through the single five-tier ladder into a label (P/LP/VUS/LB/B) and a per-section audit breakdown. The framework label is cnv_riggs_2020. Non-applicable typed SV rows retain NULL verdict fields.
### Loss metric (Riggs 2020, Table 1)
A semiquantitative point metric for copy-number losses. The point values below are the Riggs "suggested points per case" column values, not the maximum-score cap. The sections active at launch are scored from the available evidence columns; hooks that the standard defines but the current schema cannot establish are scored 0 and explicitly flagged.
Section 1 (1B)
Genomic content
-0.60 Active
1B (-0.60): the loss overlaps neither a protein-coding gene nor an exon (zero coding content). Any coding overlap is 1A (continue, 0.00).
Section 2 - 2A
Established haploinsufficiency (HI)
+1.00 Active
2A (+1.00): an established-HI region token (ClinGen score 3) with complete overlap - the loss completely overlaps the established-HI region.
Section 2 - 2H
Predicted haploinsufficiency
+0.15 Active
2H (+0.15): a predicted-HI signal - Collins pHaplo >= 0.86. A missing (unannotated) pHaplo value does not fire 2H.
Section 2 - 2E
Intragenic PVS1 decomposition
+0.30 / +0.15 Partially active
A single-gene intragenic loss (one gene with exon overlap) is scored by the ClinGen SVI PVS1 ladder. Active from the available geometry: Moderate +0.30 (>= 2 exons) and Supporting +0.15 (1 exon). The VeryStrong +0.90 and Strong +0.45 tiers are documented hooks - unreachable from the current SV schema (no NMD-competence / exon-resolution column), so they are never selected at launch (conservative under-call).
Section 2 - 2F
Established-benign containment
0.00 Documented hook (not active)
2F (-1.00 by standard): a documented hook scored 0 at launch. The shipped complete-overlap token encodes the opposite geometry (the SV contains the region, not a variant within a region), so awarding -1.00 would fabricate a strong benign downgrade; scoring 0 is the clinically safe under-call. Must be described as not yet active, not as active scoring.
Section 3 (3A/3B/3C)
Gene number
0.00 / +0.45 / +0.90 Active
Banded on gene number: 0-24 -> 0.00 (3A); 25-34 -> +0.45 (3B); 35+ -> +0.90 (3C).
Section 4 (4O)
Common population variation
-1.00 Active
4O (-1.00): population frequency above 1%. The predicate structurally excludes the absent sentinel (-1.0) and the missing value, so neither reads as "common". 4A-4N have no shipped per-case column and are scored 0.
Section 5 (5F)
Inheritance
0.00 Active (5F); 5G/5H unreachable
At launch the contribution is 5F = 0.00 - no shipped SV column carries a de novo / segregation signal. 5G (+0.10) and 5H (+0.15, the erratum value) are carried in config for a future family-side signal but are unreachable at launch (documented, not active).
### Gain metric (Riggs 2020, Table 2)
A clean mirror of the loss metric over the gain rows and gain constants, with the differences specific to Table 2. Where Table 2 awards 0 points by definition (unlike Table 1), this is flagged as "0 by table" so it is visible that it is not an omission.
Section 1 (1B)
Genomic content
-0.60 Active
1B (-0.60): the gain overlaps neither a gene nor an exon - the same zero-coding-content predicate as the loss.
Section 2 - 2A
Established triplosensitivity (TS)
+1.00 Active
2A (+1.00): an established-TS region token (ClinGen score 3 on the TS axis) with complete overlap.
Section 2 - predicted TS
Predicted triplosensitivity
0.00 0 by table
Collins pTriplo >= 0.94 carries a predicted-TS signal, but Table 2 has no +point row for predicted TS (unlike loss 2H). The launch contribution is therefore 0.00; the threshold is read and the zero surfaces in the breakdown so it is visible. Must be described as 0 by table, not as +points.
Section 2 - 2H
HI gene contained within the gain
0.00 0 by table
2H under Table 2 is 0.00 (continue) - explicitly NOT the loss 2H +0.15. The +0.15 value from the loss must never be copied onto this row.
Section 2 - 2D
Smaller than an established-benign gain
0.00 Documented hook (not active)
2D (-1.00 by standard): a documented hook scored 0. The shipped token carries no benign-region-containment signal, so the launch contribution is 0. Must be described as not yet active.
Section 2 - 2K
Breakpoint in an HI gene + specific phenotype
0.00 Documented hook (not active)
2K (+0.45 by standard): a documented hook scored 0. The shipped token carries no breakpoint-in-gene + phenotype-specificity signal, so the launch contribution is 0. Must be described as not yet active.
Section 2 - 2I
Intragenic PVS1 decomposition
+0.30 / +0.15 Partially active
The same ClinGen SVI PVS1 ladder as loss 2E - one SVI specification applied to the intragenic category of each table. Active: Moderate +0.30 / Supporting +0.15; VeryStrong +0.90 / Strong +0.45 are documented hooks, unreachable from the current geometry.
Section 3 (3A/3B/3C)
Gene number (gain bands)
0.00 / +0.45 / +0.90 Active
Gain bands, different from the loss: 0-34 -> 0.00 (3A); 35-49 -> +0.45 (3B); 50+ -> +0.90 (3C).
Section 4 (4O)
Common population variation
-1.00 Active
4O (-1.00): population frequency above 1%; the same predicate as the loss, excluding the absent sentinel and the missing value.
Section 5 (5F)
Inheritance
0.00 Active (5F); 5G/5H unreachable
The launch contribution is 5F = 0.00. 5G (+0.10) and 5H (+0.15, as published in Table 2 - not the loss erratum) are carried in config but are unreachable at launch.
### Dosage-sensitivity evidence
The point metrics read predicted and established dosage-sensitivity signals from two named standards.
Collins 2022 (pHaplo / pTriplo)
pHaplo >= 0.86 is the threshold for predicted haploinsufficiency (fires loss 2H, +0.15). pTriplo >= 0.94 is the threshold for predicted triplosensitivity (0 by table for the gain).
ClinGen Dosage Sensitivity Map
Score 3 = established HI/TS (fires 2A +1.00 on the matching axis). Score 40 = dosage-sensitivity-unlikely / established benign (named in config for the future 2F wiring; not active scoring at launch). Score 30 = autosomal recessive. Score -1 = not evaluated.
### Classification thresholds
The total score is mapped to a five-tier label by the exact cutoffs of the single ladder (shared between loss and gain). These are the SV metric cutoffs - NOT the point bands of the nuclear ACMG page, which are a different metric.
Tier Score range
Pathogenic (P) total >= 0.99
Likely Pathogenic (LP) 0.90 <= total <= 0.98
Uncertain Significance (VUS) -0.89 <= total <= 0.89
Likely Benign (LB) -0.98 <= total <= -0.90
Benign (B) total <= -0.99
### Reference data
The SV pass joins no separate reference tables during scoring; it reads the SV evidence columns already materialised on the row by the SV annotation stage. The assets listed are the sources those evidence columns depend on.
ClinGen Dosage Sensitivity Map
Scores for established HI/TS regions plus the dosage-unlikely / autosomal-recessive / not-evaluated sentinels.
Collins 2022 pHaplo / pTriplo
Predicted dosage-sensitivity scores (predicted HI and TS).
gnomAD-SV v4 / gnomAD CNV v4 population frequency
The basis for the structural-variant population frequency used by Section 4O.
Materialised SV evidence columns
Gene/exon overlap counts, HI/TS region tokens, pHaplo/pTriplo, and population frequency, produced by the SV annotation stage and read directly from the row.
### Limitations and documented hooks
SV processing participates on every case. Riggs verdict assembly applies only to the shipped constitutional copy-number loss and gain predicates; unsupported or ambiguous SV types remain visibly unclassified rather than being routed into another framework. If the shipped schema cannot establish a contribution defined by the standard, that section scores 0 and is listed as inactive.
Mandatory processing does not imply universal Riggs applicability: typed INS, INV, BND, ambiguous CNV, and other non-applicable SV rows retain NULL acmg_class, acmg_criteria, confidence_score, and classification_framework fields.
Loss 2F (established-benign containment, -1.00 by standard) is a documented hook scored 0: the shipped token encodes the opposite geometry, so awarding -1.00 would be a fabricated benign downgrade.
Gain 2D (-1.00 by standard) and 2K (+0.45 by standard) are documented hooks scored 0: the shipped token carries no benign-region-containment signal, nor a breakpoint-in-gene with phenotype specificity.
Predicted triplosensitivity (pTriplo >= 0.94) and gain 2H are 0 by table: Table 2 awards them no +points. The loss 2H +0.15 value is not copied onto the gain.
The VeryStrong (+0.90) and Strong (+0.45) tiers of the intragenic PVS1 decomposition (loss 2E / gain 2I) are documented hooks - unreachable from the current SV schema, which carries no NMD-competence or exon-resolution column; only Moderate and Supporting are active.
The 5G and 5H inheritance tiers are carried in config but are unreachable at launch - no shipped SV column carries a de novo / segregation signal or phenotype specificity.
Repeat expansions are not classified by this framework; they remain outside the scope of the SV/CNV pass.
### Version History
Build and documentation milestones of the SV/CNV framework, including the mandatory deterministic processing cutover and its explicit applicability boundaries.
Riggs 2020 Current June 2026
Implemented and documented the Table 1 loss point metric (HEL-CR-2026-SV-LOSS-SCORING-RIGGS-TABLE1-001).
Implemented and documented the Table 2 gain point metric (HEL-CR-2026-SV-GAIN-SCORING-RIGGS-TABLE2-001).
Implemented the intragenic PVS1 decomposition for loss 2E and gain 2I (HEL-CR-2026-SV-INTRAGENIC-PVS1-DECOMPOSITION-001).
Made SV role routing, sequence-resolved DEL/INS normalization, evidence annotation, and Riggs dispatch mandatory for every case; only applicable constitutional loss/gain rows receive cnv_riggs_2020 verdicts, while INS/INV/BND/ambiguous CNV rows remain explicitly unclassified (CR-2026-VA-MANDATORY-DETERMINISTIC-SV-PROCESSING-001).
### References
Riggs ER, Andersen EF, Cherry AM, Kantarci S, Kearney H, Patel A, et al. Technical standards for the interpretation and reporting of constitutional copy-number variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics (ACMG) and the Clinical Genome Resource (ClinGen).
Genetics in Medicine. 2020;22(2):245-257.
PMID: 31690835
Riggs ER, Andersen EF, Cherry AM, Kantarci S, Kearney H, Patel A, et al. Erratum: Technical standards for the interpretation and reporting of constitutional copy-number variants (item 5H default 0.15).
Genetics in Medicine. 2021;23(11):2230.
PMID: 33731880
Collins RL, Glessner JT, Porcu E, Lepamets M, Brandon R, Lauricella C, et al. A cross-disorder dosage sensitivity map of the human genome.
Cell. 2022;185(16):3041-3055.
PMID: 35917817
Riggs 2020 ( cnv_riggs_2020 ) -- Riggs ER et al., Genet Med. 2020;22(2):245-257. PMID: 31690835 ; Collins RL et al., Cell. 2022;185(16):3041-3055. PMID: 35917817 ; ClinGen Dosage Sensitivity Map (clinicalgenome.org)
### Bring structural-variant classification into your clinical workflow
Folklore's SV/CNV framework extends the Variant Analysis Service. Schedule a discussion with our scientific team to assess the fit for your laboratory.
Contact us Nuclear ACMG methodology mtDNA methodology For geneticists

### Links and cited sources

- [PMID: 31690835](https://pubmed.ncbi.nlm.nih.gov/31690835/)
- [PMID: 33731880](https://pubmed.ncbi.nlm.nih.gov/33731880/)
- [PMID: 35917817](https://pubmed.ncbi.nlm.nih.gov/35917817/)
- [Contact us](https://folklore.helena.bio/contact)
- [Nuclear ACMG methodology](https://folklore.helena.bio/methodology)
- [mtDNA methodology](https://folklore.helena.bio/methodology/mtdna)
- [For geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Clinical Screening, Context-Aware Variant Prioritization | Folklore

Source: https://folklore.helena.bio/platform/clinical-screening
Canonical: https://folklore.helena.bio/platform/clinical-screening
Description: Multi-component, age and context-aware variant scoring with diagnostic and screening modes. Tier system, actionability levels, ACMG SF v3.2, gene panels, inheritance-aware demotion.

Analysis Workflow
## Clinical Screening
Engine v2.1.0 | 7-component scoring | 9 clinical boosts
ACMG classification assigns the variant class. Phenotype Matching scores relevance to the patient's presentation. Clinical Screening uses age, sex, ethnicity, family history, sample structure, and selected gene panels to prioritize review and actionability.
The Screening Service applies different weights to neonatal, adult, and pre-conception carrier contexts. The output records each component and caps contextual boosts so they cannot override the underlying biological score.
Contents
01 Clinical Positioning 02 Two Operational Modes 03 The Seven Scoring Components 04 Context-Aware Weighting 05 Clinical Context Boosts 06 Inheritance-Aware Demotion 07 Tiers and Actionability 08 Gene Panels 09 Inputs and Outputs 10 Standards and Boundaries
### Clinical Positioning
Variant analysis produces classifications. Phenotype matching identifies which variants align with a patient symptoms. But several real-world clinical questions remain unanswered.
Four Clinical Questions
A neonate is in ICU with no usable phenotype yet. Which variants on the genome should the team look at first ?
A 45-year-old presents for proactive health screening with family history of breast cancer. Which variants matter for her , and which actionability level applies to each?
Two parents are planning pregnancy. Which carrier findings need a reproductive counseling conversation?
A trio sample with both parents sequenced. Which de novo variants deserve elevated priority?
Each scenario uses different prioritisation logic. The service changes weights for age of onset, ancestry-specific founder variants, and inheritance context, then records the score components and resulting tier.
### Two Operational Modes
One engine supports a phenotype-driven diagnostic mode and five proactive screening modes with different weight distributions.
Diagnostic Mode
When: Patient has structured phenotype data (HPO terms)
Phenotype match dominates the score. Other components support or qualify it. Age relevance contributes nothing because phenotype provides direct evidence. Used for routine diagnostic cases referred for a specific clinical indication.
Neonatal Screening
When: Newborn or infant, no usable phenotype yet
Age relevance, gene constraint, and dosage sensitivity dominate. Optimised for early-onset, treatable, time-critical conditions. Phenotype contributes only as gene-disease burden.
Pediatric Screening
When: Child, proactive or non-specific indication
Similar weight profile to neonatal but tuned to childhood-onset disease genes.
Proactive Adult Screening
When: Adult, ACMG SF v3.2 or hereditary cancer / cardiac questions
Deleteriousness and age relevance dominate. ACMG SF gene set, hereditary cancer panels, and hereditary cardiac conditions are the primary signal sources.
Carrier Screening
When: Pre-conceptional, reproductive counseling
Recessive carrier findings prioritised for reproductive counselling. Compound heterozygote detection elevated. Inheritance-aware demotion still applies for non-reproductive contexts within the same case.
Pharmacogenomics
When: Drug response and metabolism (when enabled)
Pharmacogene variants prioritised. Run alongside other modes when relevant.
### The Seven Scoring Components
Seven component scores are combined with context-specific weights. The output records each component and the total for review.
1
Gene Constraint
How intolerant the gene is to loss-of-function and missense variation. Genes that have been depleted of damaging variants by selection are more likely to cause disease when damaged. Coding and splicing variants exploit this signal; non-coding variants are heavily discounted regardless of gene constraint.
2
Deleteriousness
Several in silico predictors cover protein structure, evolutionary conservation, splice effects, and ClinGen-calibrated meta-prediction. If one value is unavailable, its weight is redistributed across the remaining predictors.
3
Phenotype
In diagnostic mode, the overlap between patient HPO terms and the gene phenotypic spectrum. In screening mode, a discounted gene-disease burden signal capped well below the diagnostic ceiling, full phenotype credit is reserved for cases where actual HPO matching has been done.
4
Dosage Sensitivity
Whether the gene is haploinsufficient (one functional copy is not enough) or triplosensitive (an extra copy causes disease). Combined with consequence type, a loss-of-function variant in a haploinsufficient gene is far more concerning than the same variant in a dosage-tolerant gene.
5
Consequence Severity
A graded score reflecting the predicted impact on protein function, from stop-gained and frameshift down through missense and synonymous. Non-protein-coding consequences receive minimal credit.
6
Compound Heterozygote Detection
Identifies pairs of heterozygous coding/splicing variants in the same gene that may form a biallelic loss of function, the diagnostic answer for many recessive conditions. Intronic, synonymous, and regulatory variants are correctly excluded from compound het candidacy.
7
Age-Appropriate Disease Context
Whether the gene is known to cause disease at the patient life stage. A pediatric-onset metabolic disease gene scores high in a neonate, low in an adult; an adult-onset cancer predisposition gene scores high in a 50-year-old and low in a newborn. When a curated gene panel is in use, panel metadata (age relevance, ClinGen evidence level) takes precedence over generic gene lists.
### Context-Aware Weighting
The engine selects component weights from the clinical question. The same variant can therefore receive different priority scores for the same patient under different screening modes.
Mode Weight Behaviour
Diagnostic case (phenotype present) Phenotype dominates. Age relevance contributes nothing.
Neonatal screening Age relevance, gene constraint, and dosage sensitivity dominate. Phenotype contributes only as gene-disease burden.
Pediatric screening Similar to neonatal but tuned to childhood-onset disease genes.
Proactive adult screening Deleteriousness and age relevance (ACMG SF cancer/cardiac genes) dominate.
Elderly screening Age relevance dominates more aggressively, since most clinically actionable variants at this age are a small set of high-evidence cardiac and cancer genes.
### Clinical Context Boosts
After the seven-component score is computed, bounded boosts apply submitted clinical context. No single contextual boost can override the underlying biological score.
ACMG class
Pathogenic and Likely Pathogenic variants receive the largest boost. VUS with PVS1 (strong null variant) evidence receive a smaller boost reflecting strong supporting evidence even at uncertain classification.
Phenotype tier
When the Phenotype Matching Service has been run, its Tier 1 and Tier 2 assignments translate into boosts that elevate phenotype-relevant variants in the screening output.
Ethnicity-aware
Targeted increases for variants in genes with established founder mutations in the patient ancestry: Ashkenazi Jewish, African, East Asian, South Asian, European-specific. Used carefully and modestly, never as a substitute for the underlying evidence.
Family history
When a family history is reported, variants in genes consistent with that history receive a measured boost. Bounded so it cannot, on its own, elevate a weak signal to a top tier.
Sex-linked
X-linked variants receive different boosts in male and female patients, reflecting the difference between hemizygous (affected) and heterozygous (often carrier) status.
Consanguinity
When consanguinity is reported, homozygous variants receive elevated weighting consistent with autosomal recessive disease in consanguineous families.
De novo
For trio or duo samples with parental sequencing, variants in constrained genes receive a boost reflecting the disproportionate clinical impact of de novo mutations in such genes.
Pregnancy and family planning
When the patient is pregnant or in family planning, variants in recessive genes commonly screened pre-conceptionally and prenatally receive a boost focused on reproductive actionability.
Gene panel
Membership in the chosen gene panel, modulated by ClinGen evidence level (Definitive, Strong, Moderate, Limited, Disputed, Refuted). Variants in strongly evidenced panel genes are favoured over those in genes with weaker evidence within the same panel.
### Inheritance-Aware Demotion
Pathogenic-or-likely-pathogenic alone does not earn Tier 1 in a recessive context.
A heterozygous P/LP variant in an autosomal recessive gene, without a compound heterozygous partner, is a carrier finding, not the diagnostic answer, and is demoted out of Tier 1 to reflect that.
The exception is a confirmed PVS1 (strong null variant) call, which is treated as a safety net for cases where upstream classification has independently confirmed a dominant loss-of-function mechanism.
Why this matters: This rule keeps heterozygous carrier findings in recessive genes out of the diagnostic Tier 1 group.
### Tiers and Actionability
Screening tier answers how strongly does the evidence point to this variant? Clinical actionability answers what should be done about it, on what timeline? A Tier 1 VUS in an actionable gene can still carry Monitoring rather than Immediate actionability because the axes are scored independently.
Tier Label Clinical Meaning
Tier 1 High Priority Immediate review. Strong combined evidence across multiple components. The variants the geneticist should see first.
Tier 2 Moderate Priority Monitor and follow up. Substantive evidence but not conclusive on its own. Worth active surveillance and additional confirmation.
Tier 3 Low Priority Future consideration. Weak evidence in this clinical context. Worth revisiting if the clinical picture changes or new evidence emerges.
Tier 4 Very Low Priority Likely benign or clinically irrelevant in this context.
Clinical Actionability
Immediate
Known pathogenic in an ACMG SF or otherwise actionable gene. Should trigger an urgent clinical pathway.
Monitoring
Significant evidence. Appropriate for active surveillance and confirmatory testing.
Future
Moderate evidence. Revisit as the clinical picture evolves or as new evidence emerges.
Research
Exploratory. Not currently actionable but may become so.
### Gene Panels
Gene panels are first-class objects in the Screening Service. They scope the analysis to a clinically meaningful set of genes, reduce noise, and let panel metadata (disease association, age relevance, ClinGen evidence level, display labels) drive scoring decisions.
Built-in Panels
Built-in panels ship with the platform and are maintained by Folklore. The neurology panels are derived from a doctoral validation study of WGS/WES in 76 patients with rare neurological diseases (51.3% diagnostic yield, 18 novel pathogenic variants reported to ClinVar).
Diabetes
MODY 1-14 (monogenic diabetes subtypes)
KATP channel (neonatal diabetes and congenital hyperinsulinism)
Neonatal diabetes (non-MODY, non-KATP)
Wolfram syndrome
Insulin resistance and lipodystrophy
Syndromic diabetes
T2D common variants
Comprehensive diabetes (all subtypes)
Neonatal Screening
Early-onset, treatable, time-critical conditions
Neurology (validation study, 76 patients, 51.3% diagnostic yield)
Intellectual disability
Epilepsy
Neuromuscular disease
Neurometabolic disease
Mitochondrial disease
Movement disorders
Custom Panels
Organisation-specific. Laboratory administrators create, edit, and assign panels for their own organisation without affecting other organisations. Gene symbols are validated against HGNC at entry time, aliases normalised automatically.
Panel Suggestions
Computed from the patient age group. The system pre-selects high-relevance built-in panels and explains why each suggestion was made, with an explicit auto-select recommendation that the geneticist can override.
Custom Genes
Can be added on top of any panel for ad-hoc, case-specific screening, with user-defined priority and age relevance.
### Inputs and Outputs
Classified variants and submitted clinical context produce tiered scores, actionability labels, and per-component rationale.
Inputs from the Pipeline
Completed variant analysis session
ACMG/AMP classification and applied criteria
Gene constraint, dosage sensitivity, inheritance mode
In silico predictors, conservation scores, splice predictions
Population frequencies (gnomAD), ClinVar context
Optional phenotype matching results
Inputs from the Geneticist
Patient demographics: age (in days or years), sex
Recommended context: ethnicity, indication for testing, family history, consanguinity
Optional context: HPO terms, free-text notes, sample type (singleton/duo/trio/quad), parental samples, reproductive context, secondary-findings preferences
Gene panel selection: built-in panel IDs, custom panels, ad-hoc custom genes
Filtering preferences: maximum Tier 1 results, minimum total score, inclusion of VUS / Pathogenic / Tier 4
Outputs for the Geneticist
Four-tier ranking with summary counts (Pathogenic, Likely Pathogenic, VUS, per-tier)
Per-variant: total score, per-component breakdown, all applied boosts, tier, clinical actionability, justification, panel context
Gene-level summaries with best tier and best actionability per gene
Panel header metadata for reporting (panel names, ClinGen status, full gene list)
Outputs for Downstream Services
Persistent screening results consumed by the AI Service for the screening report
Per-case data available to cohort-level analytics for population work
### Standards and Boundaries
The sources below define the data inputs, interpretation limits, and review boundary.
ACMG/AMP
Variant classification follows ACMG/AMP 2015 with subsequent ClinGen specifications. Performed upstream by the Variant Analysis Service. The Screening Service consumes that classification, it does not reclassify.
Reference: Richards et al., Genetics in Medicine, 2015, PMID: 25741868
ACMG SF v3.2
Secondary findings handling aligns with the current ACMG recommendations on reportable secondary findings. The proactive adult mode uses the ACMG SF gene set as a primary signal source.
Reference: Miller et al., ACMG SF v3.2, Genet Med. 2023, PMID: 37347242
ClinGen Gene-Disease Validity
Gene-disease validity is encoded as panel metadata and modulates the panel boost. Genes with Definitive or Strong evidence are favoured over Limited or Disputed within the same panel.
HGNC
Every gene symbol added to a panel is validated against the HGNC approved-symbols set, with automatic normalisation of aliases.
HPO
Phenotype input uses the Human Phenotype Ontology. Gene-disease burden in screening mode is also derived from HPO-mapped associations.
Reference: Kohler et al., Nucleic Acids Research, 2021, PMID: 33264411
Reporting Boundary
Clinical Screening ranks variants and exposes the scores behind each tier. It prioritizes review; it does not make a diagnostic decision.
Data Residency
The service runs within the Folklore platform on EU-based infrastructure compliant with GDPR Article 9 and 1+MG technical requirements.
### What Sets It Apart
Context-specific weights, bounded boosts, carrier safeguards, and recorded score components define the prioritization workflow.
Context-aware scoring, not flat scoring
The same variant in different patients, or different clinical questions, receives different priorities. Clinical relevance is a function of the question being asked, not a property of the variant alone.
Two clinically distinct modes in one engine
Diagnostic (phenotype-driven) and screening (proactive, with five sub-modes covering neonatal, pediatric, adult, carrier, and pharmacogenomic use cases).
Seven independent scoring components
The output records each component, boost, and total score for reviewer inspection.
Nine clinical-context boosts
Layered on top of the biological evidence, each bounded so no single boost can override the underlying signal.
Inheritance-aware tier assignment
Heterozygous carriers in recessive genes are not surfaced as diagnoses. PVS1 safety net preserved for confirmed dominant loss-of-function mechanisms.
First-class gene panel support
Built-in clinical panels (diabetes, neonatal, neurology), organisation-specific custom panels, ClinGen-modulated scoring, age-aware panel suggestions, HGNC validation.
Two independent output axes
Tier (how strong is the evidence) and actionability (what should be done, on what timeline). The combination drives the clinical conversation, not either alone.
Designed for diagnostic and proactive workflows
Pre-conceptional carrier screening, neonatal genomic screening, adult ACMG SF screening, and trio/family analysis all supported by the same engine with different mode selections.
### See Clinical Screening in Practice
Compare diagnostic and proactive rankings with per-variant component scores and clinical-context boosts.
Contact Us Technical Documentation Screening Methodology ACMG Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [Technical Documentation](https://folklore.helena.bio/docs/screening)
- [Screening Methodology](https://folklore.helena.bio/screening-methodology)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Cohort Analysis (Prism), Population-Level Genomic Analytics | Folklore

Source: https://folklore.helena.bio/platform/cohort-analysis
Canonical: https://folklore.helena.bio/platform/cohort-analysis
Description: Cohort analysis on classified WGS/WES samples, including Fisher, CMC, and SKAT-O burden tests, pathway enrichment, pLoF analysis, GWAS replication, polygenic risk scores, and weighted candidate-gene nomination.

Analysis Workflow
## Cohort Analysis
Engine v1.0 | Six statistical analyses | Internal codename Prism
Cohort analysis aggregates classified samples into a unified matrix for rare-variant burden tests, pathway analysis, GWAS replication, compound heterozygous review, and candidate-gene nomination.
The service consumes pre-classified samples, preserves their ACMG context, and adds population-level statistics without variant re-calling or re-annotation.
Contents
01 Clinical and Research Positioning 02 Pipeline Architecture 03 Cohort Variant Matrix 04 Cohort Quality Control 05 Gene-Level Burden Testing 06 Pathway Enrichment 07 pLoF and Frequency Analysis 08 GWAS Replication and Polygenic Scores 09 Candidate Gene Nomination 10 Inputs and Outputs 11 Standards and Boundaries
### Clinical and Research Positioning
Cohort analysis bridges per-patient ACMG classification and population-level disease genetics. Concrete examples illustrate where it changes the answer.
Three Question Types
Gene burden in a disease cohort. A monogenic diabetes cohort of 176 samples is analysed for rare variant enrichment in established susceptibility genes. Burden testing against a gnomAD reference identifies which genes carry significantly more carriers than expected, with FDR correction across all tested genes and per-gene power analysis on the side.
Pathway-level enrichment. When several burden-significant genes share a common biological pathway, that pathway is more likely to be causally implicated than any single gene in isolation. Pathway enrichment surfaces these convergent signals as additional evidence beyond gene-level p-values.
Cross-sample compound heterozygotes. A pair of heterozygous variants in a recessive gene is invisible at the per-patient level if neither variant is independently flagged. Cohort analysis identifies compound het candidates per sample with phasing context derived from the pre-classified data.
### Pipeline Architecture
Six sequential phases convert cohort metadata and per-sample VCFs into the analytical outputs. Each phase persists state so the pipeline can resume or run again with different parameters without restarting from scratch.
1
Bulk Classification
For samples not yet classified, the service dispatches the standard variant analysis pipeline per sample. A distributed concurrency limit ensures the cluster is not overwhelmed when ingesting hundreds of samples at once.
2
Quality Control
Per-sample metrics are computed from each classified DuckDB: variant count, transitions over transversions, heterozygous over homozygous ratio, and mean read depth. Cohort-level mean and standard deviation are calculated, and samples deviating beyond a configurable threshold on any metric are flagged as outliers with the specific deviating metric recorded.
3
Matrix Construction
A unified cohort variant matrix is built using deduplicated variant catalog and sparse genotype representation. Each sample is attached read-only and ingested into the cohort store. Cohort-wide allele frequencies, carrier counts, and ACMG consensus across samples are computed in a final pass.
4
Compound Heterozygous Detection
When a gene panel is provided, the service identifies pairs of heterozygous variants in the same gene per sample, restricted to coding and splicing consequences. A noise filter excludes genes with an excessive number of heterozygous variants per sample, which typically indicate common low-penetrance variation rather than disease-causing biallelic loss of function.
5
Statistical Analysis
Burden testing, pathway enrichment, pLoF, frequency analysis, GWAS replication, and polygenic risk scoring run as separate API-driven analyses on the completed matrix. Each analysis writes its own results table keyed by analysis run identifier.
6
Candidate Nomination
A weighted scoring engine integrates evidence across the prior phases and produces a ranked list of candidate genes with per-component score breakdowns and human-readable evidence summaries.
### Cohort Variant Matrix
A unified sparse matrix supports the downstream analyses while retaining per-variant and per-sample values across hundreds of WGS samples.
Deduplicated Variant Catalog
Each unique chromosome, position, reference, alternate combination is stored once across the entire cohort. This is the foundation for cohort-wide allele frequency calculation and for cross-sample variant comparison.
Sparse Genotype Matrix
Only non-reference genotypes are stored, keyed by variant identifier and sample identifier. For typical disease cohorts where each sample carries variants at a tiny fraction of the genome, this is far more efficient than dense representation.
ACMG Consensus Resolution
When the same variant is classified differently across samples, the cohort matrix records the most severe classification along with a discordance flag. A variant called Pathogenic in one sample and Likely Benign in another is treated as Pathogenic at the cohort level for burden testing, with the discordance flag preserving the disagreement for review.
Cohort-Wide Statistics
Cohort allele frequency, carrier count, and homozygote count are computed once and stored on the variant catalog. Statistical analyses query these pre-computed values rather than re-aggregating across the full genotype matrix.
### Cohort Quality Control
Outlier samples corrupt cohort-level statistics. Quality control runs before matrix construction with multiple metrics evaluated independently. A sample flagged on any single metric is sufficient for outlier review.
Variant Count
Total variant calls per sample. Outliers may indicate library preparation issues, alignment problems, or sample contamination.
Transitions over Transversions
Ti/Tv ratio is a classical sequencing quality metric. Substantial deviation from expected values may reflect base-call quality issues.
Heterozygous over Homozygous Ratio
Het/hom ratio reflects ancestry and consanguinity but extreme deviations may indicate sample-swap or contamination.
Mean Read Depth
Average depth across all variants. Low depth correlates with reduced calling sensitivity and lower confidence on variant interpretation.
### Gene-Level Burden Testing
The core question: which genes carry excess rare variant burden in the cohort relative to a reference population? Multiple complementary methods are run on every gene with FDR correction across the full set of tested genes.
Fisher Exact Test
Two-by-two contingency table comparing cohort carriers against gnomAD-expected carriers per gene. Two-sided test capturing both enrichment and depletion. The default method for gene-level burden testing.
CMC (Combined Multivariate and Collapsing)
Binary carrier collapsing per gene. At the gene level, this is equivalent to Fisher exact and provides a methodological cross-check.
SKAT-O (Optional)
Sequence Kernel Association Test, optimised. Optional R integration for variance-component testing. Falls back to CMC when not available.
Multiple Testing Correction
Benjamini-Hochberg false discovery rate (FDR) is applied across all tested genes. Bonferroni-corrected p-values are also reported. Genes are flagged significant when FDR is below the configured threshold.
Statistical Power Analysis
Minimum detectable odds ratio at standard 80% power is computed per gene. This is essential context for null results: a non-significant gene with low power is not the same as a non-significant gene with high power.
Method Discordance Flag
When two methods disagree on significance, the gene is flagged for manual review. Concordant signals across methods are stronger than single-method calls.
### Pathway Enrichment
When burden-significant genes share a common biological pathway, that pathway is more likely to be causally implicated. Pathway enrichment is a complementary signal, not a substitute for gene-level evidence.
Method. Fisher exact test, one-sided alternative, per pathway. The contingency table compares burden-significant genes that are members of the pathway against burden-significant genes that are not.
Background. All genes tested in the burden phase, not all genes in the genome. This is critical for correctness: testing a pathway against the full genome inflates significance because cohort capture varies by gene.
Pathway sources. KEGG, Reactome, and Gene Ontology biological process. The researcher provides pathway definitions; the engine runs the test and applies multiple testing correction across all evaluated pathways.
Output. Each pathway result includes its p-value and the significant genes contributing to the enrichment.
### pLoF and Frequency Analysis
Two complementary single-variant analyses targeting different signal classes.
Predicted Loss-of-Function
Variants with frameshift, stop-gained, or canonical splice site consequences are the strongest single-variant signals for haploinsufficiency mechanisms. The pLoF analysis aggregates these per gene, surfacing pLoF carrier counts alongside gene constraint metrics and ClinVar context.
Cohort vs gnomAD Frequency
Per-variant binomial test against the configured gnomAD reference, with FDR correction across all tested variants. Allele count is computed correctly using heterozygous and homozygous genotypes, and direction (enriched or depleted) is reported alongside the p-value.
### GWAS Replication and Polygenic Risk Scores
Common-variant complement to the rare variant analyses. Tests how the cohort behaves at known signals and at population-derived risk score scales.
GWAS Signal Replication
Per known GWAS signal (typically by rs-identifier), the cohort allele frequency is tested against the gnomAD reference. Output includes the cohort frequency, p-value, and a comparison to the published odds ratio for the trait.
Polygenic Risk Scoring
Polygenic risk scores are computed per sample using PGS Catalog weight files. Coverage of expected variants is reported alongside the score. Below a defined threshold, the score is marked as directional only and whole-genome data is recommended before clinical use. A within-cohort percentile supports relative comparison.
### Candidate Gene Nomination
A single weighted scoring engine integrates evidence across all prior analyses and produces a ranked list of candidate genes with per-component breakdowns. Researchers see both the top candidate and the specific reasons it ranks where it does.
1
Burden
Gene reaches significance in the cohort burden test against the chosen control population.
2
Pathway
Gene contributes to a pathway that is significantly enriched among burden-significant genes.
3
pLoF Carriers
Gene has predicted loss-of-function carriers in the cohort. pLoF variants are the strongest single-variant evidence for haploinsufficiency mechanisms.
4
GWAS Overlap
Gene region overlaps a known GWAS signal for the disease or related phenotype.
5
Disease Association
Gene has prior disease association in ClinVar or OMIM. Established disease genes are weighted differently from novel candidates.
6
Constraint
Gene is constrained against loss-of-function or missense variation in the general population (high pLI or low LOEUF). Constraint is a strong prior on gene-level disease relevance.
7
Compound Heterozygous Carriers
Gene has cohort samples with paired heterozygous variants suggesting biallelic loss of function. Normalised by the number of carriers detected.
Output. A ranked list of candidate genes with combined score, per-component score breakdown, and a human-readable evidence summary. Rankings are reproducible: the same cohort with the same parameters produces the same ranked output.
### Inputs and Outputs
Inputs come from upstream analysis and the researcher; outputs support review and downstream services.
Inputs from the Pipeline
N classified per-sample DuckDB files from variant analysis (one per cohort sample)
ACMG/AMP classification, criteria, and supporting evidence already applied per variant
Per-sample annotations: gene symbol, consequence, in silico predictors, gnomAD frequency, ClinVar context, HPO
Optional gene panel for compound heterozygous detection scoping
Inputs from the Researcher
Cohort name, sequencing type (WGS, WES, or CES), control population for burden comparison
Per-sample metadata: sample identifier, sex, age, optional HPO terms, optional clinical subgroup
Qualifying variant criteria: maximum allele frequency, minimum impact, consequence filter, classification inclusion (VUS, P/LP), collapsing strategy
Optional gene panel and pathway definitions (KEGG, Reactome, GO biological process)
Optional GWAS signals for replication, optional PGS Catalog weight files for polygenic risk scoring
Outputs for the Researcher
Per-gene burden results: Fisher, CMC, SKAT-O p-values, FDR q-value, Bonferroni p-value, minimum detectable odds ratio
Per-pathway enrichment results: Fisher p-value, FDR q-value, contributing significant genes
Per-gene pLoF summary: variant counts, carrier counts, constraint metrics, ClinVar context, HPO
Per-variant frequency analysis: cohort allele count vs gnomAD with binomial p-value and direction (enriched or depleted)
Per-signal GWAS replication: cohort allele frequency, p-value, comparison to published odds ratio
Per-sample polygenic scores with coverage warnings and within-cohort percentile
Per-gene candidate ranking: combined score, per-component breakdown, evidence summary
Compound heterozygous candidate pairs per sample per gene
### Standards and Boundaries
The references below govern burden testing, enrichment, scoring, and the research-to-clinical boundary.
ACMG/AMP
Variant classification follows ACMG/AMP 2015 with subsequent ClinGen specifications. Performed upstream by the Variant Analysis Service per cohort sample. The Cohort Service consumes that classification as input and does not reclassify.
Reference: Richards et al., Genetics in Medicine, 2015, PMID: 25741868
CMC Method
Combined Multivariate and Collapsing method for gene-level rare-variant burden testing.
Reference: Li and Leal, AJHG, 2008, PMID: 18691683 (CMC method)
SKAT-O
Sequence Kernel Association Test, optimised. Variance-component test for rare variant association. Optional R integration with documented fallback to CMC.
Reference: Lee et al., AJHG, 2012, PMID: 22863193 (SKAT-O methodology)
Benjamini-Hochberg FDR
False discovery rate correction applied across all tested genes for burden, pathways for enrichment, and variants for frequency analysis. The standard multiple testing approach for high-dimensional genomic studies.
Reference: Benjamini and Hochberg, JRSSB, 1995 (FDR control)
PGS Catalog
Polygenic risk score computation uses weight files from the PGS Catalog. Coverage of expected variants is reported alongside the score, and below a defined threshold the score is flagged as directional only.
Reference: PGS Catalog, Lambert et al., Nature Genetics, 2021, PMID: 33692572
gnomAD
Default control population for burden testing and frequency analysis. Population stratification is configurable per analysis run (NFE is the default).
Reporting Boundary
Cohort Analysis produces statistical results, candidate-gene rankings, and per-sample evidence for research review. Individual diagnostic conclusions require separate clinical assessment.
Data Residency
The service runs within the Folklore platform on EU-based infrastructure compliant with GDPR Article 9 and 1+MG technical requirements. Cohort genomic data does not leave the platform during analysis.
### What Sets It Apart
A shared cohort matrix feeds the statistical analyses and candidate-gene nomination.
Maximum-fidelity input
Operates on pre-classified samples from variant analysis. ACMG context is preserved and consumed as evidence rather than recomputed. No information loss between per-sample classification and cohort analytics.
Six analyses on one matrix
Burden testing, pathway enrichment, pLoF analysis, frequency analysis, GWAS replication, and polygenic risk scoring all share the same cohort matrix. Re-running an analysis with different parameters does not require re-ingesting samples.
Power-aware results
Every burden result includes the minimum detectable odds ratio at standard power. Null results are interpretable: not detecting a signal in an underpowered gene is not the same as not detecting a signal in a well-powered gene.
Method cross-check by design
Fisher and CMC run on every burden test. SKAT-O runs when available. Method discordance is surfaced explicitly so manually reviewed signals are concordant signals.
ACMG consensus across cohort
When the same variant is classified differently across samples, the most severe classification is recorded with an explicit discordance flag. Conservative for burden testing, transparent for review.
Weighted candidate ranking
A single ranked list integrates evidence from all six analyses with per-component breakdowns and human-readable evidence summaries. The researcher sees both the top candidate and why it is the top candidate.
Sparse matrix architecture
Deduplicated variant catalog plus sparse genotype storage scales to hundreds of WGS samples without dense N times M memory cost.
Reproducible runs
Each analysis run records the classifier version, qualifying variant criteria, control population, and gene panel. Re-running the same cohort with a new classifier version produces a new run rather than overwriting prior results.
### See Cohort Analysis in Practice
Review a cohort from per-sample classification through statistical analysis and candidate-gene ranking.
Contact Us ACMG Methodology Screening Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Screening Methodology](https://folklore.helena.bio/screening-methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Family and Trio Analysis, Inheritance-Aware Variant Interpretation | Folklore

Source: https://folklore.helena.bio/platform/family-trio
Canonical: https://folklore.helena.bio/platform/family-trio
Description: Trio variant analysis using GLnexus joint genotyping, proband ACMG classification, PLINK relationship QC, de novo detection, compound heterozygous phasing, and segregation assessment.

Analysis Workflow
## Family and Trio Analysis
Workflow Joint-genotyped trio | 3 inheritance analyses | ClinGen SVI guidance
Family context can change the evidence available in a rare Mendelian case. Folklore analyses a complete proband-mother-father trio to identify de novo candidates, phase compound heterozygous pairs, and assess segregation within the limits of a three-member pedigree.
The three member gVCFs are joint-genotyped into one normalized call set. The proband is classified through the standard ACMG workflow, and the family module records inheritance evidence without replacing the upstream classification.
Contents
01 Clinical Positioning 02 Supported Family Compositions 03 Pipeline Architecture 04 Sample Quality Control 05 De Novo Detection 06 Compound Heterozygous Phasing 07 Segregation Scoring 08 Feasibility Planning 09 Inputs and Outputs 10 Standards and Boundaries
### Clinical Positioning
Family analysis adds evidence in three common clinical scenarios.
Three Scenarios
VUS in proband, de novo confirmed by parents. A variant may remain a VUS after proband classification. Trio-supported de novo evidence identifies it for criterion review; final ACMG handling remains explicit and reviewable.
Two heterozygous variants in a recessive gene, phased in trans. A pair of variants in a recessive disease gene, one inherited from each parent, can form a biallelic loss-of-function finding. Parental phasing distinguishes it from an in-cis pair.
Variant segregates with disease. Segregation strength depends on informative meioses. A complete trio has a theoretical LOD ceiling near 0.3, below the configured Supporting threshold of 0.5. Extended pedigrees are required for PP1 evidence.
Each of these is a different inheritance algorithm with different data requirements. The Family Service runs all three on the same trio data and presents the combined evidence to the geneticist alongside the upstream ACMG classification.
### Supported Composition
The current production workflow requires a complete trio: proband, mother, and father.
#### Complete Trio
Proband, mother, and father represented in the same joint-genotyped call set. This supports parental-origin assessment, relationship QC, de novo detection, and in-trans compound heterozygous phasing.
Capabilities: De novo detection, compound heterozygous phasing, segregation context, and PLINK relationship QC.
### Pipeline Architecture
The workflow runs sequentially from trio joint genotyping to proband classification and inheritance analysis. Runtime depends on input size, sequencing mode, and available compute resources.
1
Stage 1
Joint Genotyping
GLnexus joint-genotypes the proband, mother, and father gVCFs into one normalized multi-sample VCF. The same call set becomes the source for both proband classification and trio inheritance analysis.
2
Stage 2
Proband Classification
A proband-only VCF is extracted from the normalized joint VCF and processed through the standard Variant Analysis workflow. ACMG classification belongs to the proband.
3
Stage 3
Trio Evidence Construction and QC
Joint-VCF genotypes are combined with the proband classified DuckDB to build the trio evidence store. PLINK evaluates declared parent-child relationships, possible sample swaps, duplicate samples, and parental relatedness.
4
Stage 4
Inheritance Analysis and Reporting
De novo detection, compound heterozygous phasing, segregation scoring, inheritance-pattern derivation, and the clinical-grade reporting overlay run on the trio evidence store. Results remain separate from the proband ACMG classification.
### Sample Quality Control
Before any inheritance analysis, the service verifies that the family relationships in the metadata match the genetic relationships in the data. Sample-swap, accidental duplication, and unreported consanguinity all corrupt downstream inheritance calls if undetected.
PLINK Identity-by-Descent Analysis
PLINK analyses the joint multi-sample VCF produced by the trio workflow. Biallelic SNPs provide pairwise Identity-by-Descent estimates used to check the declared parent-child relationships and identify unexpected relatedness or duplicate samples.
Sample Swap
When the IBD between proband and a declared parent does not match the expected parent-offspring relationship.
Consanguinity
When the IBD between the two declared parents is elevated above the expected unrelated-individuals baseline.
Duplicate Sample
When two declared family members are genetically identical, typically an upload error rather than a real biological scenario.
Alerts surface in the case summary alongside the inheritance evidence. The geneticist sees QC results before reading variant calls.
### De Novo Detection
Identifies variants present in the proband but absent in both parents, with explicit confidence tiers reflecting the strength of supporting evidence.
Explicit Confidence States
Technically supported candidates are reported as high or low confidence. Excluded and not-applicable are control states, while an intermediate tier is reserved for future use. Confidence describes the de novo inference, not pathogenicity.
Chromosome-Presence Gate
When a parent genotype is NULL, the system distinguishes confident inferred homozygous-reference (the chromosome was sequenced and called as reference) from suspicious absence (the chromosome may not have been adequately covered). This is a key clinical safety mechanism preventing false de novo calls from coverage gaps.
Gender-Aware Chromosome Exceptions
Biologically expected NULLs are recognised as such: chromosome Y in a female parent, mitochondrial DNA in a father. These do not trigger confidence downgrades. Gender of each family member is part of the input.
Clean ACMG Classification Preserved
The de novo annotation is added on top of the upstream ACMG classification. The original Pathogenic / Likely Pathogenic / VUS / Likely Benign / Benign assignment from variant analysis is preserved verbatim. Geneticists see both: the classification, and the new family-aware evidence.
### Compound Heterozygous Phasing
Identifies pairs of heterozygous coding or splicing variants in the same gene that may form biallelic loss of function. Phasing requires parental origin information and distinguishes in-trans (biallelic) from in-cis (single allele) configurations.
Phasing from Parental Origin
When parental genotypes are available, the service determines whether a pair of heterozygous variants in the same gene came from different parents (in trans, biallelic) or from the same parent (in cis, single-allele). Only in-trans pairs constitute genuine biallelic loss-of-function.
Multi-Partner Detection
A variant may participate in multiple candidate compound heterozygous pairs. The service records all candidate pairs in the gene for geneticist review.
Coding and Splicing Restriction
Compound heterozygous candidacy is restricted to coding and splicing consequences. Synonymous, intronic, and regulatory variants are correctly excluded, they do not constitute biallelic loss of function regardless of zygosity.
Symmetric Annotation
When variants A and B form a compound heterozygous pair, both rows in the trio table are annotated. The geneticist can land on either variant and immediately see the partner.
### Segregation Scoring
Per-variant likelihood-ratio LOD scores using the Jarvik framework as specified by ClinGen SVI 2021. Mapped to ClinGen-aligned evidence bands with explicit acknowledgement of trio data limits.
ClinGen SVI 2021 Framework
Per-variant LOD scores are computed using the Jarvik likelihood-ratio framework as specified by the ClinGen Sequence Variant Interpretation Working Group. The methodology is published, peer-reviewed, and explicitly designed for clinical use.
LOD Evidence Bands
Computed LOD values map to indeterminate, supporting, moderate, strong, or very strong bands. Not-applicable is a separate state used when no segregation score is computed.
Trio LOD Ceiling
A complete trio contributes at most two informative meioses and has a theoretical LOD ceiling near 0.3. Because the configured Supporting band begins at 0.5, trio-only segregation normally remains indeterminate. Extended pedigrees are required for PP1 evidence.
Hypothesis-Aware
When the inheritance hypothesis is unknown, the segregation phase explicitly writes not_applicable rather than emitting a spurious LOD = 0 result. This prevents misinterpretation of "no evidence calculated" as "evidence against".
### Feasibility Planning
Before the pipeline runs, the service evaluates which inheritance phases are feasible given the family composition and inheritance hypothesis. Feasibility flags and a free-form rationale are persisted alongside the trio analysis record, providing clinical audit trail.
Flag Meaning
de_novo_feasible True when both parents are sequenced and an inheritance hypothesis consistent with de novo (autosomal dominant or sporadic) is plausible. False otherwise, for example, a duo without the affected parent.
compound_het_feasible True when at least one parent is sequenced and the inheritance hypothesis is autosomal recessive. Phasing requires parental origin information.
segregation_feasible True when at least two family members with affected status are present. Trio with affected proband and one affected parent qualifies; singleton does not.
plan_rationale Free-form structured explanation of why each phase was planned as feasible or not. Provides clinical audit trail, the geneticist can verify which inheritance evidence was looked for and which was not, and why.
### Inputs and Outputs
Three member gVCFs and pedigree metadata produce a normalized joint VCF, a proband classification, and a separate trio inheritance evidence store.
Inputs from the Pipeline
Proband, mother, and father gVCFs from a compatible calling workflow
One normalized joint multi-sample VCF generated by GLnexus
Proband classified_variants.duckdb generated from the extracted proband VCF
Reference genome and contig mapping required for normalization
Inputs from the Geneticist
Complete trio composition: proband, mother, and father
Per-member sample identifier, sex, family role, and affected status
Inheritance hypothesis and phenotype specificity when available
Outputs for the Geneticist
De novo annotation per variant: confidence tier and supporting evidence
Compound heterozygous pairs: phase determination, partner variants, multi-partner flags
Segregation evidence: per-variant LOD score and ClinGen-aligned band
Sample QC summary: IBD analysis results, alerts for sample-swap or consanguinity
Feasibility and rationale: which phases were planned as feasible and why
Evidence summary JSON denormalised on the trio analysis record for fast retrieval
Outputs for Downstream Services
Persistent trio_variants DuckDB consumed by the AI Service for the family analysis report
Per-case data available to cohort-level analytics for population work
### Standards and Boundaries
The methods below govern de novo analysis, phasing, segregation, sample quality control, and clinical review.
ACMG/AMP
Variant classification follows ACMG/AMP 2015 with subsequent ClinGen specifications. Performed upstream by the Variant Analysis Service. The Family Service consumes that classification and adds family-aware evidence on top, it does not reclassify.
Reference: Richards et al., Genetics in Medicine, 2015, PMID: 25741868
ClinGen SVI 2021 Segregation
Segregation LOD scoring follows the ClinGen Sequence Variant Interpretation Working Group framework for clinical use. The Jarvik likelihood-ratio methodology is published and validated.
Reference: ClinGen Sequence Variant Interpretation Working Group, 2021
Jarvik Likelihood Framework
Per-variant likelihood-based segregation LOD computation.
Reference: Jarvik and Browning, AJHG, 2016, PMID: 27374771
PLINK 1.9
Sample identity-by-descent analysis for sample-swap, consanguinity, and duplicate-sample detection.
Reference: Chang et al., GigaScience, 2015, PMID: 25722852
Joint-Genotyped Trio Workflow
The member gVCFs are joint-genotyped inside Folklore. The proband classification and trio inheritance table derive from the same normalized call set, preventing cross-sample call-set drift.
Reporting Boundary
Family and Trio Analysis adds inheritance annotations, confidence tiers, evidence bands, and feasibility rationale to the upstream classification. A qualified clinical geneticist decides how those results affect the case.
Data Residency
The service runs within the Folklore platform on EU-based infrastructure compliant with GDPR Article 9 and 1+MG technical requirements. No family genomic data leaves the platform during analysis.
### What Sets It Apart
Inheritance logic, feasibility checks, algorithm versioning, and sample quality control define the family workflow.
One call set for classification and inheritance
The proband classification and parental genotype evidence derive from the same normalized joint VCF. Independently called missing rows are not treated as confirmed reference genotypes.
Three inheritance algorithms in one pipeline
De novo detection, compound heterozygous phasing, and segregation scoring run sequentially on the same trio data. The geneticist sees the full inheritance picture in one place.
Trio segregation limit stated explicitly
The theoretical trio-only LOD ceiling is near 0.3, below the configured Supporting threshold of 0.5. Folklore presents this as limited segregation context rather than PP1-strength evidence.
Chromosome-presence gate
Distinguishes confident inferred homozygous-reference from suspicious absence when parent genotype is NULL. A clinical safety mechanism that prevents false de novo calls from coverage gaps.
Gender-aware exceptions
chrY NULL for a female parent and chrM NULL for a father are biologically expected, not coverage gaps. The system recognises this and does not penalise confidence.
Feasibility flags with rationale
Each analysis records which phases were feasible, which ran, and why, allowing the geneticist to review the planned evidence search.
Atomic phase semantics
If a phase fails, later phases do not run. The service records the failure stage and preserves partial state for operator inspection.
Sample QC built in
PLINK IBD analysis runs as part of the standard pipeline when joint VCF is available. Sample-swap, consanguinity, and duplicate-sample alerts surface before clinical interpretation.
### See Family Analysis in Practice
Follow de novo detection, compound heterozygous phasing, and segregation scoring for a trio.
Contact Us ACMG Methodology Screening Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Screening Methodology](https://folklore.helena.bio/screening-methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Literature Evidence, Local PubMed Mining for Variant Interpretation | Folklore

Source: https://folklore.helena.bio/platform/literature-evidence
Canonical: https://folklore.helena.bio/platform/literature-evidence
Description: Local PubMed database with genetics-filtered publications and per-publication relevance scoring. Sub-second clinical search aligned to ACMG evidence categories.

Analysis Workflow
## Literature Evidence
Service Local PubMed Mirror | Six-component scoring | ACMG-aligned
The Literature Evidence service maintains a local, genetics-filtered PubMed mirror with pre-extracted variants, genes, and phenotypes. A six-component model ranks publications for the case and maps results to ACMG evidence categories.
Typical searches return in under a second with ranked publications, per-component score breakdowns, and ACMG-aligned strength labels inside the variant interpretation workflow.
Contents
01 Why Local PubMed Mining 02 Ingestion Pipeline 03 Genetics Relevance Filtering 04 Variant, Gene, and Phenotype Extraction 05 Clinical Search Workflow 06 Relevance Scoring 07 ACMG-Aligned Evidence Strength 08 Inputs and Outputs 09 Standards and Boundaries
### Why Local PubMed Mining
Using public PubMed directly introduces three constraints for case-level variant review.
Three Problems
Latency. Structured queries across several genes, variants, and phenotype terms require repeated public PubMed requests. A local index supports those queries within the case-review interface.
Inconsistency. Search results include component scores and source links for later review.
Data residency. Sending case context, including HPO terms and gene panels, to public PubMed transmits potentially sensitive context outside the platform. A local mirror keeps clinical search inside the EU-resident Folklore infrastructure.
### Ingestion Pipeline
Six pipeline stages convert raw PubMed XML into a clinically searchable local database. Each file completes all stages before the next begins, so processing can restart after failure without repeating completed files.
1
Download
PubMed baseline and daily update files are pulled from the NCBI FTP source with parallel transfers and integrity verification. The full PubMed baseline is approximately 1,300 compressed XML files representing tens of millions of articles.
2
Parsing
Compressed PubMed XML is parsed with a streaming XML reader, producing structured publication records. Memory-efficient parsing handles the largest files without loading them fully into memory.
3
Filtering
Publications are evaluated against curated MeSH descriptors that signal genetics relevance, against accepted publication types, and against a publication date floor. The filtering ratio reduces the input to a manageable genetics-focused subset of approximately seven to eight percent of the input volume.
4
Extraction
For each retained publication, the service extracts variant notations, gene symbols, and phenotype mentions. Extracted entities are validated against authoritative reference data before persistence.
5
Loading
Extracted records are inserted into the local literature database in batches with conflict-aware upsert semantics. Re-processing a file is idempotent: existing records are updated rather than duplicated.
6
Cleanup
Downloaded compressed XML files are removed after successful processing to reclaim disk space. Cleanup policy is configurable and can preserve files for audit or troubleshooting.
### Genetics Relevance Filtering
The filtering layer selects a genetics-focused subset using three independent dimensions.
Genetics MeSH Descriptors
A curated set of Medical Subject Headings descriptors signals that a publication concerns genetics or genomics. MeSH classifications are the most reliable filtering signal because they are professionally indexed at the source.
Publication Type
Case reports, clinical trials, original research, and review articles are accepted. Editorial pieces, news items, and similar non-research formats are excluded from the indexed dataset.
Publication Date Floor
Publications older than the configured date floor are excluded. The default floor reflects the period over which contemporary genetic nomenclature and reporting standards stabilised.
### Variant, Gene, and Phenotype Extraction
During ingestion, the extraction layer records variants, genes, and phenotypes from publication text and metadata. Search then queries those precomputed entities.
Variant Mentions
HGVS cDNA and protein notations alongside legacy notations are extracted from publication text. Notation is normalised so a single variant referenced in different formats across different publications is recognised consistently in search.
Gene Mentions
Candidate gene symbols are validated against the human protein-coding gene reference. This eliminates false positives from common abbreviations that overlap with non-gene acronyms. Mention counts per publication are preserved as evidence of gene centrality to the paper.
Phenotype Mentions
Phenotype names are mapped to HPO, OMIM, and MeSH identifiers when available. Stemmed morphological matching ensures that variations such as plural and adjectival forms link to the same underlying phenotype.
### Clinical Search Workflow
Six steps convert a case clinical context into a ranked, persisted, stream-ready set of literature evidence. The full workflow targets sub-second response for typical queries.
Step 1
Two-Source Publication Discovery
For each query gene, the service finds candidate publications through two complementary sources. The first is a direct lookup against extracted gene mentions, providing publications where the gene is identified as a meaningful subject. The second is a text search across titles and abstracts, catching mentions that may not have been formally extracted. The two sources are merged into a unique candidate set.
Step 2
Publication Enrichment
Each candidate publication is enriched with full metadata: title, abstract, journal, publication date, authors, DOI and PMC identifiers, MeSH descriptors, and publication types. Gene mention counts, variant mentions for query genes, and phenotype mentions are attached to support downstream relevance scoring.
Step 3
Parallel Relevance Scoring
Enriched publications are scored across multiple components. Scoring runs in parallel across worker processes, bypassing Python concurrency limits to deliver sub-second results even when hundreds of candidate publications are evaluated.
Step 4
Filter and Rank
Publications below a minimum relevance threshold are discarded. Remaining publications are sorted by total score in descending order and capped at the requested result limit.
Step 5
Persist to Session Store
Top results are written into the session output store alongside the variant classification data. The geneticist consumes both data types from a single per-session file.
Step 6
Stream-Ready Export
A compressed newline-delimited JSON file is produced for the frontend. Metadata is emitted on the first line, results follow one per line, and a completion marker closes the stream. This format allows progressive display in the frontend as results are received rather than waiting for the full payload.
### Relevance Scoring
Six weighted components produce a relevance score per publication. The result includes each component alongside the total.
1
Phenotype Match
How well the publication phenotype mentions overlap the patient HPO terms. Stemmed morphological matching ensures that small lexical variations do not break the match. The dominant signal in the score because phenotype alignment is the strongest single indicator that a publication is relevant to the case at hand.
2
Publication Type
Case reports and clinical trial reports are weighted highest. Original research follows. General journal articles and review articles are weighted lower. The hierarchy reflects the relative value of each publication type as evidence in ACMG variant classification.
3
Gene Centrality
How frequently the query gene is mentioned in the publication. A paper that mentions the gene of interest dozens of times in the abstract and main text is more centrally about that gene than one that mentions it once in passing. Centrality is bounded so a small number of publications mentioning a gene many times does not dominate the ranking.
4
Functional Data
Whether the publication describes functional studies relevant to ACMG functional evidence criteria. Indicators include MeSH terms for animal models, knockout studies, cell line experiments, and molecular biology techniques. Functional data is a key prerequisite for ACMG functional evidence.
5
Variant Match
Exact variant notation match between the query and the extracted variant mentions in the publication scores highest. Same-gene different-variant scores lower. This component captures the difference between a paper about the patient exact variant and a paper about a different variant in the same gene.
6
Recency
Publications decay linearly over a recent window. A current-year publication scores at the top of this component; older publications score progressively lower. Recency is a relatively small contribution because older landmark papers can remain authoritative.
### ACMG-Aligned Evidence Strength
Each ranked publication receives an evidence-strength label mapped to ACMG/AMP categories, producing a structured pool for criteria review.
Strength Description and ACMG Alignment
Strong Publication describes the exact variant and includes functional studies. Candidate evidence for ACMG PS3 (well-established functional studies showing damaging effect) or PP3 functional component context.
Moderate Publication describes the exact variant OR functional data, but not both. Candidate context for moderate-weight ACMG criteria.
Supporting Publication describes the gene with phenotype overlap to the case. Candidate context for ACMG PP4 (phenotype highly specific for a single genetic etiology) or related supporting evidence.
Weak Gene is mentioned but no exact variant or phenotype-specific context is present. Background reference material for completeness rather than direct ACMG evidence.
### Inputs and Outputs
Query genes, variants, and HPO terms produce a ranked publication set with component scores and source links.
Inputs from the Pipeline
Patient query genes from the upstream variant analysis or phenotype matching results
Patient HPO terms from the case clinical context
Optional exact variant notations from the case variant analysis output
Inputs from the Geneticist
Search initiation as part of the standard variant interpretation workflow
Optional gene panel scoping when only specific genes are of interest
Optional result limit configuration
Outputs for the Geneticist
Ranked publication list with overall relevance score for the case
Per-publication score breakdown across all six relevance components
Per-publication evidence strength label aligned to ACMG categories
Direct links to PubMed identifier, DOI, and PMC where available
Highlighted variant matches, gene mention counts, and phenotype matches per publication
Persisted session results consumable from the same data store as variant classifications
Stream-ready compressed JSON for instant frontend rendering
### Standards and Boundaries
The references below define the evidence categories, vocabularies, and review boundary.
ACMG/AMP
Relevance scoring components and evidence-strength labels map to ACMG/AMP categories. The four levels are Strong, Moderate, Supporting, and Weak; a geneticist reviews the source before assigning a criterion.
Reference: Richards et al., Genetics in Medicine, 2015, PMID: 25741868
PubMed
The PubMed baseline and daily update streams are the source of the literature data. PubMed is maintained by the U.S. National Library of Medicine at NCBI. The local mirror applies those updates on its configured schedule.
Reference: PubMed, U.S. National Library of Medicine, NCBI
MeSH
Medical Subject Headings are the controlled vocabulary maintained by NLM for indexing biomedical literature. MeSH descriptors support genetics filtering and the functional-data score, and NLM curators assign the source indexing.
Reference: Medical Subject Headings (MeSH), U.S. National Library of Medicine
HPO
The Human Phenotype Ontology provides the structured phenotype vocabulary used for matching publication phenotype mentions against patient HPO terms. Stemmed morphological matching extends HPO matching to lexical variants without requiring exact term lookup.
Reference: Kohler et al., Nucleic Acids Research, 2021, PMID: 33264411
HGNC
Gene symbols extracted from publications are validated against the HGNC approved-symbols set. This eliminates false positives from non-gene acronyms and ensures that gene mentions normalise consistently across publications using older gene symbol versions.
Reference: HUGO Gene Nomenclature Committee, hgnc.symbolreport
Reporting Boundary
Literature Evidence ranks publications and exposes component scores and evidence-strength labels. It does not assign pathogenicity; the geneticist reads the sources and decides whether they support a case criterion.
Data Residency
The service runs within the Folklore platform on EU-based infrastructure compliant with GDPR Article 9 and 1+MG technical requirements. The local PubMed mirror is hosted within the same EU infrastructure, so clinical search does not transmit case data outside the platform.
### What Sets It Apart
Local indexing, entity extraction, relevance scoring, evidence labels, and session integration define the service.
Local mirror, sub-second search
The PubMed baseline is mirrored locally and indexed for retrieval. Typical clinical queries return ranked results in well under a second inside the variant interpretation workflow.
Genetics-focused dataset
A curated MeSH-based filter reduces millions of articles to a focused genetics-relevant subset, removing noise that would otherwise dilute search results. The filtering ratio is conservative and biased toward inclusion when relevance is plausible.
Pre-extracted variants, genes, and phenotypes
Variant notations, validated gene symbols, and HPO-mapped phenotype mentions are extracted at ingestion time, not at search time. The geneticist sees publications already enriched with the entities that drive ACMG decisions.
Validated gene symbols
Every candidate gene symbol is validated against the HGNC approved-symbols set before storage. Common abbreviation collisions, the dominant source of false positives in gene mention extraction, are eliminated upstream.
Six-component relevance scoring
A weighted model combines phenotype match, publication type, gene centrality, functional-data signal, variant match, and recency. The result includes each component alongside the total score.
ACMG-aligned strength labels
Each ranked publication carries a strength label mapped to ACMG evidence categories, creating a categorized pool for geneticist review.
Session-integrated storage
Search results persist in the same per-session data store as variant classifications. The frontend loads both data types from a single source, simplifying architecture and ensuring consistency across the case lifetime.
EU data residency
The local PubMed mirror runs within Folklore's EU infrastructure, so a clinical search does not transmit case data outside the platform.
### See Literature Evidence in Practice
Enter query genes and HPO terms, then inspect the ranked publications, component scores, and ACMG-aligned strength labels.
Contact Us ACMG Methodology Screening Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Screening Methodology](https://folklore.helena.bio/screening-methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Mitochondrial DNA Analysis, MMDWG 2020 Variant Interpretation | Folklore

Source: https://folklore.helena.bio/platform/mitochondrial-dna
Canonical: https://folklore.helena.bio/platform/mitochondrial-dna
Description: Dedicated mtDNA classification engine following the McCormick 2020 ACMG/AMP specifications from the ClinGen Mitochondrial Disease Variant Curation Expert Panel. Heteroplasmy-aware, haplogroup-aware, conservation-aware.

Analysis Workflow
## Mitochondrial DNA Analysis
Framework MMDWG 2020 | ClinGen Expert Panel | Heteroplasmy aware
Maternal inheritance, heteroplasmy, lack of recombination, and haplogroup structure require different evidence rules for mitochondrial variants. Folklore routes them through a dedicated classifier that implements the MMDWG 2020 specifications.
MMDWG 2020 evaluates all 28 ACMG/AMP criteria for mtDNA, excludes seven with a stated biological rationale, and defines fourteen with mtDNA-specific thresholds. Folklore preserves the Richards 2015 combining rules and adds no platform-specific criteria.
Contents
01 Why mtDNA Needs Its Own Framework 02 Unique Mitochondrial Biology 03 Dual Classifier Routing 04 Criteria Applied to mtDNA 05 Criteria Explicitly Excluded 06 PVS1 for mtDNA Variants 07 In Silico Predictors by Gene Type 08 Haplogroup-Aware Frequency 09 Manual Evidence Placeholders 10 Inputs and Outputs 11 Standards and Boundaries
### Why mtDNA Needs Its Own Framework
Treating a mitochondrial variant through a nuclear classifier produces clinically incorrect results. The MMDWG 2020 specifications exist precisely because mtDNA biology is sufficiently different that several ACMG criteria are inapplicable and others require mtDNA-specific thresholds.
A Worked Example
The classical MELAS variant m.3243A>G has over a thousand published cases, a 4-star ClinVar pathogenic assertion, and is the textbook mitochondrial encephalomyopathy variant. Run through a nuclear ACMG classifier without mtDNA awareness, the criteria string is empty, the confidence is zero, and the assignment is VUS. Run through the MMDWG 2020 classifier, the same variant correctly receives the criteria that apply to it and is classified as Pathogenic with full evidence trace.
Nuclear and mitochondrial variants use different criteria and frequency assumptions. Folklore routes mtDNA variants to the MMDWG classifier and nuclear variants to the ACMG/AMP classifier, recording the framework used for each result.
### Unique Mitochondrial Biology
Five biological features distinguish mtDNA from the nuclear genome and shape the MMDWG 2020 specifications.
Maternal Inheritance
Mitochondrial DNA is exclusively maternally inherited through the female germline. There is no paternal contribution and no autosomal recessive pattern. Several ACMG criteria designed for autosomal genetics are inapplicable as a result.
Heteroplasmy
A cell contains many mitochondria, and each mitochondrion contains many copies of mtDNA. A pathogenic variant may be present at any heteroplasmy level from a few percent to homoplasmy. Phenotypic expression often follows a tissue-specific threshold effect: a variant may be silent at low heteroplasmy and severe at high heteroplasmy in the same individual.
No Splicing, No Recombination
mtDNA genes do not undergo splicing, and the mtDNA genome does not undergo homologous recombination. Splice prediction is irrelevant. Variants accumulate without recombination smoothing, producing high natural variability across haplogroups.
Haplogroup Structure
Because mtDNA does not recombine, variants cluster into fixed lineages called haplogroups. Some variants are top-level haplogroup defining and reach high allele frequency in their lineage while remaining absent elsewhere. A pathogenic variant in one haplogroup may be a benign defining variant in another, requiring haplogroup-aware interpretation.
Threshold Effect
A phenotype may only manifest when a variant reaches a particular heteroplasmy threshold in a given tissue. Below the threshold the carrier is asymptomatic; above it, the disease manifests. This shapes the clinical evidence: variant heteroplasmy correlates with disease severity in segregation studies.
### Dual Classifier Routing
Every variant in a session is routed to either the nuclear or the mitochondrial classifier based on chromosome and gene symbol context. Edge cases including D-loop annotation gaps and NUMT pseudogene candidates are handled explicitly with informational flags rather than fail-hard errors.
1
Mitochondrial Path
When: Mitochondrial chromosome with a mitochondrial gene symbol
Variant is routed to the MMDWG 2020 classifier. Criteria specifications are mtDNA-aware. Combining rules from Richards 2015 are preserved unchanged. Output is labelled with the mtDNA framework so downstream services and reports treat it as such.
2
Nuclear Path
When: Nuclear chromosome with a non-mitochondrial gene symbol
Variant is routed to the standard ACMG 2015 classifier. Unchanged. The vast majority of variants in a typical WGS or WES session follow this path.
3
D-Loop and Annotation Gaps
When: Mitochondrial chromosome but no mitochondrial gene symbol assigned by upstream annotation
Routed to the mtDNA classifier with an explicit annotation flag. The mtDNA D-loop is a regulatory region without protein-coding genes; some annotation pipelines do not assign gene symbols to D-loop variants. The classifier handles this case explicitly.
4
NUMT Pseudogene Candidates
When: Nuclear chromosome with a mitochondrial-prefixed gene symbol
NUMTs are nuclear copies of mitochondrial DNA segments. They are a known biological reality and a recognised cause of false-positive mtDNA calls. The variant is routed to the nuclear classifier with an explicit NUMT pseudogene candidate flag, plus additional informational markers when low heteroplasmy, low depth, or known NUMT-prone region indicators are present. Decision remains with the geneticist.
### Criteria Applied to mtDNA
Twenty ACMG/AMP criteria are applicable to mtDNA variants per MMDWG 2020. Several have mtDNA-specific weight or threshold differences from the nuclear classifier. Several require manual evidence input from the geneticist.
PVS1
Very Strong evidence for pathogenicity. For mtDNA, a large heteroplasmic deletion encompassing at least one full gene is treated as Very Strong. Smaller protein-coding truncations follow the Abou Tayoun decision tree, which produces Strong, Moderate, or Supporting weight depending on transcript impact and truncated fraction. Not applicable to single nucleotide changes in tRNA or rRNA.
PS1
Strong. Same nucleotide change as a previously established pathogenic variant in a protein-coding mitochondrial gene. Restricted to the protein-coding mtDNA gene set.
PS3
Functional studies. MMDWG 2020 caps PS3 at Supporting evidence weight for mtDNA, citing the absence of standard or universal parameters for objectively analysing cybrid studies. Available as a manual evidence placeholder for geneticist annotation.
PS4
Case prevalence in unrelated probands across diverse top-level haplogroups. Manual evidence placeholder. Affected proband definition follows MMDWG 2020 exactly: classic mitochondrial disease syndromes or red flag features per Haas 2007 and 2008.
PM2_Supporting
Population frequency below 1 in 50,000 from controls in reliable mitochondrial databases. MMDWG 2020 specifies Supporting weight for this criterion in mtDNA, a single-tier downgrade from the nuclear PM2 Moderate. gnomAD does not contain mtDNA frequencies; MITOMAP and HmtDB are the primary frequency sources.
PM4
In-frame insertion or deletion in a non-repeat region or stop-loss. Restricted to protein-coding mtDNA genes.
PM5
Different change at the same position as a known pathogenic variant. Moderate weight in protein-coding genes; Supporting weight in tRNA and rRNA.
PM6
Assumed de novo without confirmation of maternity by full mtDNA sequencing. Manual evidence placeholder.
PP1
Segregation in maternal family members with heteroplasmy correlation. Cannot be applied when the variant is homoplasmic in all family members. Manual evidence placeholder. Heteroplasmy must segregate with disease severity.
PP3
In silico predictors. Tool selection depends on gene type: APOGEE2 for protein-coding mRNA genes, MitoTIP and HmtVAR for tRNA genes. Discordant predictions yield no contribution. rRNA in silico prediction is not currently supported.
PP4
Phenotype-specific evidence. For mtDNA, requires decreased electron transport chain enzyme activity below 20 percent of control mean from a CLIA-approved laboratory in muscle, liver, or fibroblasts. Manual evidence placeholder.
BA1
Stand-alone benign evidence. Allele frequency above 1 percent in mitochondrial population databases, OR a top-level haplogroup defining variant in an individual whose predicted haplogroup matches. The haplogroup-defining condition is dual: a variant is a defining marker for haplogroup X but pathogenic in haplogroup Y for an unrelated condition. Conflict detection prevents incorrect application.
BS1
Strong benign. Allele frequency between 0.5 and 0.99 percent in mitochondrial population databases. Strict range, not an open-ended threshold.
BS2
Heteroplasmy comparison. Variant observed at higher heteroplasmy in a healthy adult than in the same tissue of an affected individual. Manual evidence placeholder.
BP2_Supporting
Another mtDNA variant in the same individual is independently confirmed as pathogenic. Symmetrical: both variants receive this annotation against each other.
BP4
In silico predictors point to benign. Mirror of PP3 with the same gene-type-specific tool selection.
BP5
Alternate molecular cause. Variant found in an individual with a confirmed nuclear-DNA-related mitochondrial disease. Manual evidence placeholder.
BP7
Synonymous variant in a protein-coding mtDNA gene. MMDWG 2020 specifies that mtDNA synonymous variants are supporting benign evidence because cryptic splice site disruption, the rationale for the nuclear BP7 conservation check, does not apply to mtDNA.
BS3
Functional studies showing no effect. Manual evidence placeholder.
BS4
Lack of segregation in affected family members, or paternal segregation. Paternal segregation is itself disqualifying for an mtDNA pathogenic variant. Manual evidence placeholder.
### Criteria Explicitly Excluded
Seven ACMG criteria are excluded for mtDNA per MMDWG 2020. Each exclusion has biological rationale published verbatim in the framework. Folklore preserves these exclusions strictly.
PM1
Conservation, domain, and structural information is incorporated into the in silico predictors used in PP3 (APOGEE2, MitoTIP, HmtVAR). Including PM1 separately would double-count the same evidence.
PM3
Compound heterozygous and recessive inheritance does not apply to mtDNA. The genome is maternally inherited as a single haploid unit.
PP2
Low rate of benign missense in the gene of interest does not apply to mtDNA. High variability across haplogroups is well documented; mtDNA tolerates high missense rates because it lacks recombination and has a relatively high mutation rate without histone protection.
PP5
ClinGen Sequence Variant Interpretation guidance: any variant observed in a database should be assessed using its underlying evidence, not relying on the assertion alone. This is general SVI policy, applied to both nuclear and mitochondrial paths.
BP1
Most variants in protein-coding mtDNA genes are missense, not truncating. The premise that missense variants in a gene where loss of function is the main mechanism are likely benign does not apply.
BP3
In-frame indels in repetitive regions of mtDNA are concentrated in two well-known hypervariable regions and are routinely excluded from clinical interpretation pre-classification.
BP6
Same rationale as PP5: assertions in databases should not be relied on without underlying evidence per ClinGen SVI policy.
### PVS1 for mtDNA Variants
PVS1 is the strongest single criterion in the ACMG framework. For mtDNA, MMDWG 2020 specifies three implementation paths reflecting the biological differences between large deletion syndromes, protein-coding truncations, and tRNA or rRNA variants.
Path 1, Large Deletion (Very Strong)
A heteroplasmic mtDNA deletion that encompasses at least one full mitochondrial gene receives PVS1 at Very Strong weight. This corresponds clinically to single large mtDNA deletion syndromes including KSS, Pearson, and CPEO. The detection criterion is functional, encompassing a full gene rather than a base-pair threshold.
Path 2, Protein-Coding Truncation (Abou Tayoun decision tree)
Frameshift, nonsense, and small deletion variants in protein-coding mtDNA genes follow the Abou Tayoun 2018 PVS1 decision tree. Nonsense-mediated decay does not occur for mtDNA, so the truncated fraction of the protein governs the assigned weight. Variants leaving more than ninety percent of the coding region intact typically receive PVS1 at Moderate weight.
Path 3, tRNA and rRNA
Single nucleotide variants in tRNA or rRNA genes are not eligible for PVS1 unless they are part of a deletion encompassing a full gene as in Path 1. MMDWG 2020 is explicit that, with the exception of large deletions, this criterion cannot be applied to tRNA or rRNA variants.
### In Silico Predictors by Gene Type
MMDWG 2020 specifies different predictor sets for different mtDNA gene types. Tools optimised for nuclear protein-coding variants are not transferable to mitochondrial tRNA structures.
Protein-Coding (mRNA)
APOGEE2 is the recommended predictor. A pathogenic prediction supports PP3; a neutral prediction supports BP4. Discordant ensemble predictions yield no contribution.
Transfer RNA (tRNA)
MitoTIP and HmtVAR are evaluated jointly. Concordant pathogenic prediction across both tools supports PP3. Concordant benign prediction supports BP4. Discordant predictions yield no contribution.
Ribosomal RNA (rRNA)
No standardised in silico predictors for rRNA are currently in use. PP3 and BP4 are not evaluated for rRNA variants in this framework. Future predictors based on rRNA secondary structure may be incorporated as the field matures.
### Haplogroup-Aware Frequency
Because mtDNA does not recombine, allele frequency aggregated globally can mask haplogroup-specific patterns. A variant may reach high frequency in one haplogroup while remaining absent in others. The classifier handles this explicitly.
BA1 Stand-Alone Benign, Two Conditions
BA1 fires when global allele frequency exceeds one percent in mitochondrial population databases. It also fires when the variant is a top-level haplogroup-defining marker and the patient predicted haplogroup matches. Both pathways are tracked in the evidence trace.
Conflict Detection
A variant defining haplogroup X may be pathogenic when it appears outside that lineage. MMDWG 2020 explicitly notes that BA1 must not fire when there is conflicting evidence of pathogenicity in a different haplogroup context. The classifier checks this and surfaces a haplogroup conflict warning to the geneticist.
Frequency Sources
MITOMAP and HmtDB are the primary mtDNA population frequency databases per MMDWG 2020. gnomAD does not provide mtDNA frequencies. PM2 supporting evidence requires frequency below one in fifty thousand, BS1 strong benign requires the strict range from 0.5 to 0.99 percent, and BA1 stand-alone benign requires above one percent.
### Manual Evidence Placeholders
Several MMDWG 2020 criteria require evidence not directly extractable from variant annotation. Examples include functional studies, segregation in maternal family members, ETC enzyme activity, and confirmed maternity by full mtDNA sequencing. The classifier exposes these as manual evidence placeholders for geneticist annotation.
The placeholders include PS3 (functional studies, capped at Supporting weight), PS4 (case prevalence with haplogroup diversity), PM6 (assumed de novo without confirmation), PP1 (segregation with heteroplasmy correlation), PP4 (phenotype-specific ETC deficiency in CLIA-approved laboratory testing), BS2 (heteroplasmy comparison between healthy and affected family members), BS3 (functional studies showing no effect), BS4 (lack of segregation or paternal segregation), and BP5 (alternate molecular cause from nuclear DNA).
Geneticist annotations are recorded alongside the automated criteria with timestamp and user identity. Re-running the classifier preserves manual annotations and re-evaluates only the automated criteria.
### Inputs and Outputs
Inputs come from upstream analysis and the geneticist; outputs support clinical review.
Inputs from the Pipeline
Pre-annotated variants from the upstream variant analysis pipeline
Per-variant gene symbol, consequence, transcript context, and population frequency
In silico predictor scores when applicable to the gene type
ClinVar context and review status when available
Heteroplasmy and read depth when present in the source VCF
Inputs from the Geneticist
Manual evidence annotations for criteria requiring external data: PS3, PS4, PM6, PP1, PP4, BS2, BS3, BS4, BP5
Optional inheritance hypothesis context (sporadic, maternally inherited, unknown)
Optional haplogroup designation when established by external testing
Outputs for the Geneticist
ACMG class assignment: Pathogenic, Likely Pathogenic, VUS, Likely Benign, or Benign
Applied criteria string with explicit framework label distinguishing mtDNA from nuclear classifications
Per-criterion evidence trace including triggering data points
Confidence score reflecting evidence strength and consistency
NUMT pseudogene candidate flags when applicable for variants on nuclear chromosomes with mitochondrial-prefixed gene symbols
Haplogroup conflict warnings when a variant is haplogroup-defining for one lineage but appears in a different lineage
### Standards and Boundaries
The classifier operates against published standards and within explicit clinical boundaries.
MMDWG 2020 (McCormick et al.)
The mitochondrial DNA classification framework produced by the ClinGen Mitochondrial Disease Variant Curation Expert Panel and the Mitochondrial Disease Sequence Data Resource Consortium. ClinGen SVI and the Clinical Domain Working Group Oversight Committee approved it.
Reference: McCormick et al., Hum Mutat. 2020;41(12):2028-2057, PMID: 33058415
ACMG/AMP 2015
Original ACMG/AMP classification framework. MMDWG 2020 preserves Richards 2015 combining rules unchanged and specifies which of the original 28 criteria apply to mtDNA, which are excluded with verbatim rationale, and which require mtDNA-specific thresholds.
Reference: Richards et al., Genetics in Medicine, 2015, PMID: 25741868
Abou Tayoun 2018 (PVS1)
Decision tree for PVS1 application to truncating variants. MMDWG 2020 references this for protein-coding mtDNA frameshift, nonsense, and small deletion variants. Used unchanged for the protein-coding path.
Reference: Abou Tayoun et al., Hum Mutat. 2018, PMID: 30192042 (PVS1 decision tree)
APOGEE2
In silico pathogenicity predictor for mitochondrial protein-coding variants, used per the MMDWG 2020 PP3 specification.
Reference: Castellana et al., MitImpact / APOGEE2 (in silico mtDNA pathogenicity predictor)
MitoTIP and HmtVAR
In silico pathogenicity predictors for mitochondrial tRNA variants, used per the MMDWG 2020 PP3 specification with concordance requirements.
Reference: Sonney et al., PLoS Comput Biol. 2017, PMID: 28732077 (MitoTIP for tRNA variants)
MITOMAP and HmtDB
Primary mitochondrial population frequency databases. The MMDWG 2020 specification explicitly identifies these as the reliable sources for mtDNA allele frequency. gnomAD is not used for mtDNA frequency.
Scope Boundary
Folklore mtDNA analysis is for primary mitochondrial disease, matching the MMDWG 2020 scope statement verbatim. The framework is not designed for and is not used for cancer somatic mtDNA variants, longevity studies, or complex disease predisposition contexts.
Reporting Boundary
The mtDNA module records the classification, applied criteria, and supporting evidence. A qualified clinical geneticist reviews the result before clinical use.
Data Residency
The service runs within the Folklore platform on EU-based infrastructure compliant with GDPR Article 9 and 1+MG technical requirements.
### What Sets It Apart
Framework selection, evidence rules, curation, and audit records remain separate from the nuclear classifier.
Dedicated mtDNA classifier
Mitochondrial variants are routed to a separate classification engine that follows the MMDWG 2020 specifications. They are not processed through nuclear ACMG with its inapplicable criteria.
Haplogroup-aware BA1
A variant defining one haplogroup may cause disease in another lineage. The classifier checks both global frequency and haplogroup-defining status against the patient predicted haplogroup, and detects conflicts where the same variant has dual evidence.
Heteroplasmy in the workflow
Heteroplasmy is captured per variant where present in the source data and is presented to the geneticist alongside the classification. Segregation evidence (PP1) requires heteroplasmy correlation with disease severity.
NUMT awareness
Nuclear copies of mitochondrial DNA segments are a known cause of false-positive mtDNA calls. The classifier detects nuclear-chromosome variants with mitochondrial-prefixed gene symbols and surfaces NUMT candidate flags with low-heteroplasmy, low-depth, and known-region indicators.
Gene-type-specific predictors
APOGEE2 for protein-coding mRNA genes. MitoTIP and HmtVAR for tRNA genes. rRNA in silico prediction is explicitly not supported, matching the MMDWG 2020 scope.
Verbatim exclusion rationale
Seven ACMG criteria are excluded for mtDNA. Each exclusion carries the MMDWG 2020 rationale verbatim, preserved in the system documentation and exposed in the Standards reference. The geneticist sees why a criterion did not fire.
No Folklore-specific extensions
The classifier ships strict MMDWG 2020 compliance. Combining rules from Richards 2015 are preserved unchanged. No platform-specific thresholds, no novel criteria combinations, no informal rules layered on top of the published specification.
Manual evidence with audit trail
Criteria requiring external data (functional studies, segregation, ETC enzyme activity, alternate molecular cause) are exposed as manual evidence placeholders. Geneticist annotations are recorded with timestamp and user identity for audit.
### See mtDNA Analysis in Practice
Trace mitochondrial variants through MMDWG 2020 criteria, haplogroup checks, and NUMT detection.
Contact Us mtDNA Methodology Detail ACMG Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [mtDNA Methodology Detail](https://folklore.helena.bio/methodology/mtdna)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Phenotype Matching, Phenotype-First Variant Prioritization | Folklore

Source: https://folklore.helena.bio/platform/phenotype-matching
Canonical: https://folklore.helena.bio/platform/phenotype-matching
Description: Whole-genome phenotype-to-gene matching with HPO semantic similarity, a five-tier priority hierarchy, and decomposable scoring.

Analysis Workflow
## Phenotype Matching
Standard HPO | Whole-genome scale | Second-level latency
ACMG classification tells you how pathogenic a variant is. It does not tell you which variant explains your patient . Phenotype Matching answers that second question, turning clinical relevance into a reproducible score and a five-level priority hierarchy aligned with how a geneticist actually reads a case.
Phenotype-first prioritization can rank a VUS with strong phenotypic relevance above a Pathogenic variant for an unrelated condition. The tier rules apply that ordering before manual review.
Contents
01 Clinical Positioning 02 Workflow for the Geneticist 03 The Five-Tier System 04 Why Phenotype-First Matters 05 Scale and Performance 06 Inputs and Outputs 07 Standards and Boundaries 08 What Sets It Apart
### Clinical Positioning
After ACMG classification of a whole-genome or exome case, the geneticist is left with hundreds to thousands of variants flagged as Pathogenic, Likely Pathogenic, or VUS. Most are irrelevant to the referral indication.
The Real Diagnostic Question
A Pathogenic BRCA1 variant has no diagnostic value for a child referred for refractory epilepsy. Manually correlating each candidate against the patient's clinical presentation traditionally consumes the better part of a working week per case, and that effort scales poorly across cohorts and trio analyses.
Phenotype Matching reads the patient's clinical presentation as HPO terms, compares them with the phenotypic spectrum of each candidate gene, and assigns each variant a match score, priority score, and clinical tier. The results are stored with the variant data for screening, reporting, and AI-assisted summaries.
Built-in clinical principle: a P/LP variant without phenotype match is reported as an Incidental Finding, not as Tier 1. A VUS with strong phenotype relevance can outrank a P/LP variant for an unrelated disease.
### Workflow for the Geneticist
The same four-step workflow applies across Folklore access paths.
1
Phenotype Capture
The geneticist captures the patient's clinical findings as HPO terms. Three entry modes are supported: direct HPO ontology search with name and synonym autocomplete; AI-assisted extraction from free-text referral letters or discharge summaries with automatic exclusion of negated phrases; and direct lookup by HP:ID. Free-text clinical notes can be attached to the case.
2
Run Matching
The service reads the patient's HPO terms together with the variant data already produced by upstream classification. Phenotype relevance is computed for every variant in scope, tiering rules are applied, and results are written back to the case alongside the existing variant annotations.
3
Review
Results are presented gene-first. Each gene card shows the best tier achieved by any variant in that gene, the strongest phenotype match score, the count of variants per tier, and which patient HPO terms the gene matched on. Expanding a gene reveals every variant with full annotation context and a per-term breakdown of how each patient HPO term aligned with the gene's phenotypic spectrum.
4
Report
A branded PDF report lists Tier 1, Tier 2, and Incidental Findings in detail and gives summary counts for Tier 3 and Tier 4. It contains phenotype-matching data without AI-generated interpretation and requires clinical geneticist review before clinical use.
### The Five-Tier System
Each variant is placed into exactly one tier. Tier ranges do not overlap, so a variant's score alone determines its tier.
Tier Label Score Range Clinical Meaning
Tier 1 Actionable 80.00 - 99.99 Pathogenic or Likely Pathogenic with strong phenotype match. The variant most likely explaining the patient's presentation.
Tier 2 Potentially Actionable 60.00 - 79.99 VUS with strong supporting evidence (high impact, strong phenotype match, rare). Or P/LP variants where inheritance context warrants further investigation before calling them diagnostic.
IF Incidental Finding 40.00 - 59.99 Pathogenic or Likely Pathogenic, but not relevant to the referral phenotype. Clinically important as a secondary finding, but separated from the primary diagnostic search.
Tier 3 Uncertain 20.00 - 39.99 VUS with moderate evidence. Worth tracking; insufficient for clinical action.
Tier 4 Unlikely 0.00 - 19.99 Benign, Likely Benign, common, or with poor phenotype relevance.
Sorting Within a Tier
Within each tier, finer-grained sorting is driven primarily by phenotype match strength, with ACMG classification and population frequency contributing secondary signal. The geneticist always sees the strongest candidates first within each clinical category.
Conservative Tier 1 and Tier 2
The service is deliberately conservative about the top tiers. A variant cannot reach Tier 1 or Tier 2 on impact alone. Phenotype relevance is required, and benign evidence, ClinVar Benign with curator review, or strong benign ACMG criteria, blocks elevation regardless of phenotype score.
### Why Phenotype-First Matters Clinically
Most variant prioritization tools start from the variant and ask: is this pathogenic? They sort the patient's variants by ACMG class and leave clinical relevance to manual filtering. Tier 1 can therefore include pathogenic variants in genes unrelated to the referral.
The Reversed Question
Phenotype Matching reverses the question. It starts from the patient and asks: given this clinical picture, which variants deserve attention first? The service records that prioritization as scores, tiers, and per-term contributions for reviewer inspection.
#### A relevant VUS is not buried
A VUS in a gene whose phenotypic spectrum closely matches the patient's presentation appears in Tier 2 instead of being buried in a flat pathogenicity-ordered list.
#### An unrelated pathogenic variant is not noise
A Pathogenic variant for a condition unrelated to the referral indication is preserved as an Incidental Finding, visible, scored, reportable as a secondary finding when clinically appropriate, but not competing for attention with the variant that actually explains the case.
#### Inheritance is respected
A heterozygous pathogenic variant in a gene that causes disease only in the homozygous or compound heterozygous state is treated differently from one in a dominant gene. Carrier status is not a diagnosis. Dual-inheritance genes are flagged so the geneticist can assess context.
### Scale and Performance
A typical WGS case can carry hundreds of thousands of HPO-annotated variants. The matching pipeline returns per-variant phenotype scores, tiers, and match details.
Pre-warmed Parallelism
Worker processes are initialized at service startup with the HPO ontology already loaded and caches primed. Incoming requests do not pay the cold-start cost.
Scale-aware Computation
Variants in the same gene share phenotypic context, so the pipeline computes that context once and reuses it across per-variant results.
Optimized Frontend Rendering
Gene-level summaries are exported as a small compressed payload that streams in near-instantly when the case is opened. The heavier per-gene variant detail is fetched on demand only when the geneticist expands a gene. The interface stays responsive even on cases with hundreds of thousands of variants.
### Inputs and Outputs
The service combines classified variants with patient HPO terms and returns phenotype scores, tiers, and match details.
Inputs from the Pipeline
Completed variant analysis session
ACMG/AMP classification per variant
Gene and HPO annotation
Population frequencies (gnomAD)
ClinVar context and review status
Inheritance context (Orphanet AD/AR/XLD/XLR, ClinGen dosage)
Inputs from the Geneticist
List of HPO terms representing the patient's clinical presentation
Optional free-text clinical notes that travel with the case
Outputs for the Geneticist
Phenotype match score (0-100) and clinical tier for every variant with HPO data
Gene-level summary view ranked by clinical priority
Per-variant breakdown of which patient HPO terms matched, against what, and how strongly
Inheritance context notes (e.g., dual-inheritance genes, recessive carrier flags)
Downloadable branded PDF report with Tier 1, Tier 2, and Incidental Findings detail
Outputs for Downstream Services
Phenotype tier and score data consumed by the Screening Service for tier-aware variant scoring
Phenotype context consumed by the AI Service for clinical interpretation reports
### Standards and Boundaries
HPO semantics, the scoring model, and the clinical-review boundary are documented below.
HPO
The Human Phenotype Ontology is the international standard for structured clinical phenotype representation in genetics. The service operates entirely in HPO terms with full ontology hierarchy support.
Reference: Kohler et al., Nucleic Acids Research, 2021, PMID: 33264411
ACMG/AMP
Variant classification follows the ACMG/AMP 2015 guidelines with subsequent ClinGen specifications. Classification itself is performed upstream by the Variant Analysis Service. Phenotype Matching consumes that classification, it does not reclassify.
Reference: Richards et al., Genetics in Medicine, 2015, PMID: 25741868
Reporting Boundary
Phenotype Matching supplies scores and tiers for clinical review. It does not make a diagnosis or authorize clinical use; sign-off remains with a qualified clinical geneticist.
Data Residency
The service runs within the Folklore platform on EU-based infrastructure compliant with GDPR Article 9 (special category data) and 1+MG technical requirements. Patient phenotype data and variant data remain inside the platform.
### What Sets It Apart
Explicit controls translate phenotype relevance into scores, tiers, and review order.
Phenotype-first by design
The tier system gives referral-phenotype relevance more weight than pathogenicity in an unrelated gene. A relevant VUS can therefore outrank an unrelated Pathogenic variant.
Whole-genome scale, second-level latency
For a typical WGS case, the output includes phenotype scores, tiers, and per-term breakdowns across the annotated variants.
Transparent, decomposable scoring
Scores decompose into phenotype, ACMG, and frequency components, and matches decompose into per-HPO-term contributions for reviewer inspection.
Inheritance-aware throughout
Heterozygous findings in recessive genes, dual-inheritance genes, and compound heterozygous candidates are handled distinctly rather than collapsed into a single pathogenic-or-not flag.
Conservative on Tier 1 and Tier 2
Strong benign evidence, a curator-reviewed ClinVar Benign assertion, BA1, or strong benign ACMG criteria blocks tier elevation regardless of phenotype score.
Clinically reviewable output
Reports separate Tier 1, Tier 2, and Incidental Findings so secondary findings remain distinct from the primary diagnostic search.
### See Phenotype Matching in Practice
Start with HPO terms and review the resulting shortlist ranked by relevance to the patient's presentation.
Contact Us Technical Documentation ACMG Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)
- [Technical Documentation](https://folklore.helena.bio/docs/phenotype-matching)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Variant Analysis & Classification | Folklore Platform

Source: https://folklore.helena.bio/platform/variant-classification
Canonical: https://folklore.helena.bio/platform/variant-classification
Description: VCF processing with Ensembl VEP annotation and rule-based ACMG/AMP classification across all 28 evidence criteria, with recorded data sources and thresholds.

Analysis Workflow
## Variant Analysis & Classification
Folklore annotates VCF variants through Ensembl VEP and applies the 2015 ACMG/AMP guidelines through deterministic rules. AI output does not enter the class calculation.
The pipeline evaluates all 28 evidence criteria against population data, clinical archives, computational predictions, and gene-level constraint metrics and records the evidence behind each applied criterion.
### Core Capabilities
The workflow has four parts: VCF processing, annotation, ACMG/AMP classification, and evidence recording.
#### VCF Processing
Standard VCF 4.1 and 4.2 parsing from any sequencing platform. Whole genome, whole exome, and targeted panels supported. Quality filtering preserves clinically significant ClinVar variants regardless of quality scores.
#### Multi-Source Annotation
Ensembl VEP for consequence prediction, protein impact, and functional domain mapping. Multi-source enrichment from ClinVar, gnomAD, dbNSFP, ClinGen, and 12+ in silico predictors. Over 60 annotations per variant.
#### Rule-Based ACMG/AMP Classification
All 28 evidence criteria from the 2015 ACMG/AMP guidelines (Richards et al.) systematically evaluated. No AI determines pathogenicity. Five-tier output: Pathogenic, Likely Pathogenic, VUS, Likely Benign, Benign.
#### Audit Trail
Each applied criterion is recorded per variant with its evidence source, threshold, and classifier version.
### Technical Foundation
The classification pipeline queries the genomic and clinical resources listed below for each variant.
Ensembl VEP
Variant effect prediction, consequence, protein impact
ClinVar
Clinical variant interpretations and review status
gnomAD
Population allele frequencies (global and population-specific)
ClinGen
Gene-disease validity and dosage sensitivity
dbNSFP
Aggregated functional predictions across 12+ tools
AlphaMissense
Missense pathogenicity prediction (DeepMind)
SpliceAI
Splice site impact prediction
BayesDel
Meta-predictor with ClinGen SVI-calibrated thresholds
GERP++ / PhyloP / PhastCons
Evolutionary conservation
pLI / LOEUF / o/e
Gene-level constraint metrics
### ACMG/AMP Criteria Coverage
All 28 evidence criteria from the 2015 ACMG/AMP guidelines are systematically evaluated. Pathogenic and benign evidence is weighted according to standard combining rules.
PVS1
Very Strong Pathogenic
- PVS1, Null variant in gene with established LOF mechanism
PS1–PS4
Strong Pathogenic
- PS1, Same amino acid change as established pathogenic
- PS2, De novo (paternity and maternity confirmed)
- PS3, Functional studies showing damaging effect
- PS4, Significantly increased prevalence in affected individuals
PM1–PM6
Moderate Pathogenic
- PM1, Mutational hot spot or critical functional domain
- PM2, Absent or extremely low frequency in population databases
- PM3, Detected in trans with a pathogenic variant (recessive)
- PM4, Protein length changes (in-frame indels, stop-loss)
- PM5, Novel missense at residue with different pathogenic change
- PM6, Assumed de novo without confirmation of paternity/maternity
PP1–PP5
Supporting Pathogenic
- PP1, Cosegregation with disease in multiple affected family members
- PP2, Missense in gene with low rate of benign missense variants
- PP3, Multiple computational lines of evidence support deleterious effect
- PP4, Patient phenotype highly specific for the gene
- PP5, Reputable source reports as pathogenic
BA1
Stand-alone Benign
- BA1, Allele frequency above 5% in population databases
BS1–BS4
Strong Benign
- BS1, Allele frequency greater than expected for disease
- BS2, Observed in healthy adults (recessive, dominant, X-linked)
- BS3, Functional studies showing no damaging effect
- BS4, Lack of segregation in affected family members
BP1–BP7
Supporting Benign
- BP1, Missense in gene where only truncating cause disease
- BP2, Observed in trans with pathogenic (dominant) or in cis
- BP3, In-frame indel in repetitive region without known function
- BP4, Multiple computational lines of evidence suggest no impact
- BP5, Variant found in case with alternate molecular basis
- BP6, Reputable source reports as benign
- BP7, Synonymous with no predicted splice impact and conserved
Detailed criteria reference
### What You Receive
The output is a classified variant set with evidence records for clinical review and downstream phenotype matching.
#### Classified Variant Table
Every variant in the input VCF receives a five-tier ACMG classification with explicit listing of all applied criteria.
#### Evidence Detail Per Variant
Population frequencies, ClinVar status, in silico predictions, conservation scores, and gene constraint metrics in one unified view.
#### Criteria Audit Trail
For each criterion applied (e.g., PVS1, PM2_supporting), the underlying evidence and threshold used is documented and traceable.
#### Confidence Scoring
A numerical confidence score reflects the strength and number of supporting criteria and can be used to order review.
### See Variant Classification in Practice
Inspect the path from VCF upload to a classified variant table with recorded evidence.
Contact Us Technical Documentation ACMG Methodology Full Pipeline For Geneticists

### Links and cited sources

- [Detailed criteria reference](https://folklore.helena.bio/docs/classification/criteria-reference)
- [Contact Us](https://folklore.helena.bio/contact)
- [Technical Documentation](https://folklore.helena.bio/docs/classification)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Full Pipeline](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Screening Methodology - Multi-Dimensional Variant Prioritization | Folklore

Source: https://folklore.helena.bio/screening-methodology
Canonical: https://folklore.helena.bio/screening-methodology
Description: How Folklore prioritizes classified variants for clinical review using seven scoring dimensions, age-aware weights, clinical profile boosts, and four-tier ranking.

## Screening Methodology
Screening Engine v 2.1.0 | Updated March 2026
This page documents how Folklore prioritizes classified variants for clinical review. ACMG classification determines the variant class; screening determines review order from patient-specific clinical relevance.
The screening algorithm evaluates each variant across seven dimensions, applies patient-specific clinical profile boosts, and produces a four-tier priority ranking. Results show the component scores used for each variant. Screening completes in under one second for typical cases.
Contents
01 Overview 02 Screening Pipeline 03 Scoring Components 04 Deleteriousness Ensemble 05 Weight System 06 Clinical Profile Boosts 07 Tier System 08 Gene Lists 09 Screening Modes 10 Limitations 11 Version History 12 References
### Overview
Classification and screening solve different problems. Both are necessary for efficient clinical review.
Classification
Determines what a variant IS. Assigns one of five ACMG categories based on evidence criteria. Documented on the Classification Methodology page.
Screening
Determines which variants to REVIEW FIRST. Ranks all classified variants by clinical relevance to this specific patient using a multi-dimensional scoring algorithm adapted to patient demographics and clinical context.
Why Screening Is Necessary
After classification, a clinician typically faces 50-200+ variants requiring review. Manual review of every variant is impractical. Screening reduces this to 3-20 high-priority candidates (Tier 1) by incorporating clinical context that classification alone does not consider: patient age, sex, ethnicity, family history, clinical phenotype, and the specific screening strategy.
A VUS in a highly constrained gene with strong phenotype match and family history of genetic disease may be more clinically relevant than a Pathogenic variant in an unrelated gene. Screening captures these contextual relationships.
### Screening Pipeline
Single-pass pipeline over classified variants. Total processing time under one second for typical cases (100-500 variants).
1
Load Classified Variants
Pre-classified variants (Pathogenic, Likely Pathogenic, VUS) are loaded with full annotation data. Phenotype tier data is joined when available. Gene panel filters are applied if selected.
2
Calculate Base Component Scores
Each variant is scored across seven independent dimensions (constraint, deleteriousness, phenotype, dosage, consequence, compound heterozygote, age relevance). Each score is normalized to [0.0, 1.0].
3
Apply Scoring Weights
Component scores are combined using weights appropriate to the patient's age group and screening mode. All weight sets sum to exactly 1.0 and are validated at runtime.
4
Calculate Clinical Boosts
Patient-specific context (ACMG class, phenotype match, ethnicity, family history, sex, consanguinity, pregnancy, gene panel) adds additional priority. Total score capped at 1.0.
5
Assign Priority Tiers
Each variant is assigned to one of four tiers based on boosted score and clinical context. All Pathogenic/Likely Pathogenic variants are guaranteed Tier 1 regardless of base scores.
6
Export Results
Tiered results are persisted for downstream consumption. Gene-level summaries are exported for summary-first clinical review.
### Scoring Components
Each variant is evaluated across seven independent dimensions. All scores are normalized to [0.0, 1.0] for direct comparability.
Gene Constraint
Measures the gene's intolerance to variation using gnomAD constraint metrics. Strategy varies by consequence type: loss-of-function variants are scored using pLI and LOEUF jointly, missense variants use mis_z combined with pLI, and non-coding variants receive a heavily discounted score regardless of gene constraint.
Loss-of-function variants in genes with high pLI and low LOEUF receive the maximum score
Missense variants in genes with high missense constraint (mis_z) and high pLI receive elevated scores
Non-coding consequences (intron, upstream, downstream, synonymous) receive minimal scores - gene constraint is clinically irrelevant for variants that do not affect protein function
Data sources: gnomAD v4.1.0 (pLI, LOEUF, mis_z)
Deleteriousness
Weighted aggregate of eight computational predictors for variant deleteriousness. BayesDel_noAF is the primary ClinGen SVI-calibrated predictor; SpliceAI covers splice impact, and six additional predictors cover missing values.
BayesDel_noAF is the primary signal, normalized from its native range to [0, 1]
SpliceAI provides orthogonal splice impact prediction independent of missense predictors
AlphaMissense contributes independent protein structure signal derived from AlphaFold
When BayesDel is unavailable, its weight is redistributed proportionally among remaining predictors
Conservation scores (PhyloP, GERP) are normalized from their native ranges to [0, 1]
Data sources: dbNSFP 4.9c (BayesDel_noAF, DANN, SIFT, AlphaMissense, MetaSVM, PhyloP, GERP), Ensembl (SpliceAI)
Phenotype Relevance
Evaluates how well the variant's gene matches the patient's clinical presentation. Operates in two modes depending on whether patient HPO terms are available.
Diagnostic mode (HPO terms provided): computes overlap between patient HPO terms and gene HPO associations
Screening mode (no phenotype): uses gene-disease burden as proxy, capped at 0.5 (full 1.0 is reserved for actual HPO phenotype match). Non-coding variants receive additionally discounted scores.
Data sources: HPO gene-phenotype associations
Dosage Sensitivity
Based on ClinGen haploinsufficiency evidence, applied only to loss-of-function variants.
Only evaluated for loss-of-function consequences (frameshift, stop_gained, splice_donor, splice_acceptor)
Three-tier scoring based on ClinGen evidence level: sufficient, emerging, or limited evidence
Non-LoF variants and genes without dosage data receive zero
Data sources: ClinGen dosage sensitivity (haploinsufficiency_score)
Consequence Severity
Hierarchical severity ranking based on VEP consequence type and transcript biotype.
Non-protein-coding biotypes receive minimal scores regardless of consequence
Hierarchy: frameshift/stop_gained/splice_donor/splice_acceptor > start_lost/stop_lost > inframe > missense > splice_region > UTR > synonymous > intron
Data sources: Ensembl VEP (consequence, biotype)
Compound Heterozygote
Detects potential compound heterozygotes for autosomal recessive conditions by identifying multiple heterozygous coding variants in the same gene.
Only true coding/splicing consequences qualify - intronic, synonymous, and splice_region+intron_variant combinations are excluded (splice_region alone without coding impact is not a true coding consequence)
Variants pre-flagged by the upstream classification pipeline receive the maximum score
Same-gene heterozygous coding variant pairs are detected via efficient pre-grouped gene lookup
Data sources: Pipeline-internal genotype analysis
Age Relevance
Prioritizes age-appropriate disease genes using curated gene lists. Different gene categories are emphasized depending on the patient's age group.
Neonatal: early-onset disease genes and treatable metabolic conditions receive highest priority
Pediatric: childhood-onset genes receive highest priority, early-onset genes slightly lower
Adult: cancer predisposition and cardiac genes receive highest priority
Elderly: cardiac genes receive highest priority; cancer gene priority is reduced
ACMG Secondary Findings genes receive elevated priority across all age groups
Panel genes with age_group_relevance="all" receive minimum 0.6 score regardless of patient age
ClinGen gene-disease validity modulates panel gene scores: Definitive/Strong=1.0, Moderate=0.85, Limited=0.7
Data sources: ACMG Secondary Findings v3.2, curated pediatric and adult gene lists, gene panel metadata (when panel selected)
### Deleteriousness Ensemble
Eight-predictor weighted ensemble with BayesDel_noAF as the primary signal. Approved via clinical review HELIX-CR-2026-002.
Predictor Signal Type Contribution Source
BayesDel_noAF ClinGen SVI calibrated missense Primary dbNSFP 4.9c
SpliceAI Splice impact (4 delta scores) Major (orthogonal) Ensembl MANE
AlphaMissense Protein structure (AlphaFold) Significant dbNSFP 4.9c
DANN Deep learning pathogenicity Supporting dbNSFP 4.9c
SIFT Sequence homology Supporting dbNSFP 4.9c
MetaSVM Ensemble meta-predictor Supporting dbNSFP 4.9c
PhyloP (100-way) Per-site conservation Minor dbNSFP 4.9c
GERP++ Per-element conservation Minor dbNSFP 4.9c
BayesDel_noAF Normalization
BayesDel_noAF scores are linearly normalized from their observed range to [0, 1] before weighting. ClinGen SVI evidence tiers are used for justification text but do not directly determine the screening score - the normalized continuous value is used instead for maximum discrimination.
NULL Handling
When BayesDel_noAF is unavailable for a variant (NULL or NaN), its weight is redistributed proportionally among the remaining seven predictors. Each predictor uses type conversion with NaN detection, and the score records the available predictor set.
### Weight System
Different clinical contexts require different scoring emphasis. The weight system selects appropriate component weights based on patient age group, screening mode, and phenotype availability. All weight sets sum to exactly 1.0.
Diagnostic
Patient has HPO phenotype terms (any age)
Constraint
Standard
Deleteriousness
Standard
Phenotype
Dominant
Dosage
Low
Consequence
Minimal
Compound Het
Minimal
Age Relevance
None
Rationale: Phenotype matching drives prioritization when clinical presentation is available. Age relevance is unnecessary because the phenotype itself guides variant selection.
Neonatal / Pediatric
0-18 years, no phenotype
Constraint
High
Deleteriousness
Standard
Phenotype
Low
Dosage
Elevated
Consequence
Low
Compound Het
Minimal
Age Relevance
Elevated
Rationale: Gene constraint is critical in neonates and children. Dosage sensitivity is elevated because LoF in haploinsufficient genes is urgent. Age relevance captures actionable early-onset conditions.
Adult Proactive
18-65 years, no phenotype
Constraint
Standard
Deleteriousness
Elevated
Phenotype
Low
Dosage
Low
Consequence
Low
Compound Het
Minimal
Age Relevance
Elevated
Rationale: Deleteriousness predictions are elevated because missense predictions are critical for cancer and cardiac risk. Age relevance captures adult-actionable conditions.
Elderly
65+ years, no phenotype
Constraint
Reduced
Deleteriousness
Standard
Phenotype
Low
Dosage
Low
Consequence
Low
Compound Het
Minimal
Age Relevance
Highest
Rationale: Age relevance receives the highest weight because older patients benefit most from narrowly actionable findings. Constraint is reduced because many genes have already been expressed without clinical phenotype.
Design Principle
Age relevance weight increases with patient age: Elevated for neonatal/pediatric, Elevated for adult, Highest for elderly. This reflects the clinical reality that older patients benefit most from narrowly actionable findings, while younger patients warrant broader screening.
### Clinical Profile Boosts
Nine boost categories use the clinical profile submitted with each screening request. Boosts are added to the base weighted score and can promote variants to higher tiers. The total score is capped at 1.0.
ACMG Classification
Pathogenic and Likely Pathogenic variants receive priority boosts. VUS variants with strong null variant evidence (PVS1 criterion) also receive an elevated boost.
Activated when: ACMG class is P, LP, or VUS with PVS1
Phenotype Match Tier
Variants in genes with strong phenotype correlation receive priority boosts proportional to the phenotype matching tier.
Activated when: Phenotype Matching Service has been executed
Ethnicity
Population-specific founder mutations receive elevated priority. Supports Ashkenazi Jewish, African, East/South Asian, and European founder variant lists.
Activated when: Patient ethnicity is provided
Family History
Cancer and cardiac predisposition genes receive priority boosts when family history is reported. Additional boost when indication specifically references family history.
Activated when: Family history flag is set
Sex-Linked Inheritance
X-linked disease genes receive priority boosts based on patient sex. Males receive a larger boost due to hemizygous expression.
Activated when: Variant on chromosome X in a recognized X-linked gene
Consanguinity
Homozygous variants receive elevated priority in consanguineous families, reflecting increased likelihood of identical-by-descent inheritance.
Activated when: Consanguinity flag is set
De Novo Proxy
Variants in highly constrained genes receive a proxy de novo boost when trio/duo data is available. This is a prioritization heuristic, not confirmed de novo status.
Activated when: Sample type is trio or duo
Pregnancy / Family Planning
Prenatal actionable genes receive elevated priority for pregnant patients. Carrier screening genes receive priority for family planning.
Activated when: Pregnancy or family planning flag is set
Gene Panel
Panel genes receive priority boosts modulated by ClinGen gene-disease validity level and age-group relevance matching.
Activated when: A gene panel is selected
Boost Interaction
Multiple boosts can apply simultaneously. A variant in BRCA1 for an Ashkenazi Jewish patient with family history of cancer will receive ACMG classification boost, ethnicity boost, and family history boost concurrently. The total boosted score is capped at 1.0 to prevent score inflation.
### Tier System
Four-tier priority ranking with two-stage assignment: base tier from component scores, then final tier after applying all clinical boosts.
Tier 1 : High Priority - Immediate Review
Variants requiring immediate clinical attention. Includes all Pathogenic and Likely Pathogenic variants regardless of base score, strong phenotype matches, and variants exceeding the high-priority score threshold. Capped at a configurable maximum (default: 20 variants).
Tier 2 : Moderate Priority - Monitor
Variants with moderate clinical relevance. Includes Tier 2 phenotype matches and variants with intermediate boosted scores. Warrant review but are less urgent than Tier 1.
Tier 3 : Low Priority - Future Consideration
Variants with low but non-trivial clinical relevance. May become significant with additional clinical information or future variant reclassification.
Tier 4 : Very Low Priority - Likely Benign
Variants with minimal clinical relevance under current evidence. Excluded from results by default; included only when explicitly requested.
Base Tier Assignment
The base tier uses both the total weighted score and individual component peaks. A variant with an exceptional signal in a single component (e.g., a highly constrained gene or very high deleteriousness) can be promoted to Tier 1 even if other components are moderate. This prevents clinically significant variants from being buried by low scores in irrelevant dimensions.
Final Tier Determination (Post-Boost)
After clinical boosts are applied, the final tier may differ from the base tier. Pathogenic and Likely Pathogenic variants, VUS with strong null-variant evidence (PVS1), and strong phenotype matches appear in Tier 1 regardless of their base component scores. ACMG classification takes priority over component-level scoring.
### Gene Lists
Curated gene lists drive age relevance scoring, clinical actionability assessment, and ethnicity-specific prioritization.
ACMG Secondary Findings v3.2 (81 genes)
Three categories: Cancer Predisposition (25 genes including APC, BRCA1, BRCA2, MLH1, MSH2, TP53, VHL), Cardiac (34 genes including KCNH2, KCNQ1, MYBPC3, MYH7, SCN5A, LMNA), and Metabolic (8 genes including LDLR, BTD, RPE65). These genes receive elevated age relevance scores across all age groups and are used for clinical actionability assessment.
Source: Miller DT et al. Genet Med. 2023;25(1):100726
Pediatric Gene Lists
Three curated categories: Early-Onset Disease Genes (CFTR, SMN1, GAA, and other genes causing conditions diagnosable at birth), Treatable Metabolic Conditions (PAH, GALT, and other genes where early intervention changes outcomes), and Childhood-Onset Genes (NF1, PKD1, and other genes causing conditions typically presenting in childhood).
Adult-Onset Gene Lists
Two curated categories: Cancer High-Risk (BRCA1, BRCA2, MLH1, MSH2, TP53, PTEN, and other high-penetrance cancer predisposition genes) and Cardiac (KCNH2, MYBPC3, SCN5A, LMNA, and other genes associated with sudden cardiac death or cardiomyopathy).
Population-Specific Founder Variants
Curated founder mutation gene lists for ethnicity-aware prioritization: Ashkenazi Jewish (BRCA1/2, GBA, HEXA, FANCC, BLM, MSH2, MSH6), African ancestry (HBB, G6PD), East/South Asian (ALDH2, CYP2C19, HBA1, HBA2, HBB), and European (BRCA1/2, CFTR).
### Screening Modes
Six screening modes determine the weight profile and prioritization strategy.
Diagnostic
Patient has HPO phenotype terms. Phenotype matching dominates scoring. Used when clinical presentation guides interpretation.
Neonatal Screening
Newborn screening (0-28 days). Prioritizes early-onset disease genes, treatable metabolic conditions, and haploinsufficient genes.
Pediatric Screening
Child and adolescent screening (1-18 years). Prioritizes childhood-onset conditions with emphasis on gene constraint and age-relevant genes.
Proactive Adult
Adult health screening (18-65 years). Emphasizes cancer predisposition, cardiac risk genes, and computational deleteriousness predictions.
Carrier Screening
Recessive carrier identification for reproductive risk assessment.
Pharmacogenomics
Drug response screening for medication safety and efficacy.
Age Group Determination
Patient age is converted to one of six age groups: Neonatal (0-28 days), Infant (29 days - 1 year), Child (1-12 years), Adolescent (12-18 years), Adult (18-65 years), Elderly (65+ years). Day-level precision is used for neonatal/infant boundary. Age can be provided in days, years, or both.
### Limitations
Screening is a prioritization tool, not a diagnostic tool. It determines review order, not variant pathogenicity. ACMG classification determines pathogenicity.
Compound heterozygote detection is inferred from genotype data without formal phasing. Trio data or long-read sequencing provides definitive confirmation.
De novo boost is a proxy based on gene constraint, not confirmed de novo status. Parental genotype comparison is required for confirmation.
Ethnicity-based boosts use curated founder mutation lists. Population-specific variants outside these lists do not receive ethnicity boosts.
Phenotype scoring without HPO terms uses gene-disease burden as a proxy, which favors well-characterized genes over recently described disease associations.
Age relevance gene lists are curated and may not include all relevant disease genes for each age group. Lists are updated periodically.
Carrier screening and pharmacogenomics modes are not yet weight-differentiated from adult screening. Dedicated weight profiles are planned.
Gene panel boosts are modulated by ClinGen gene-disease validity, which may not be available for all panel genes.
All scores are normalized to [0.0, 1.0]. Raw metric magnitudes are compressed into a relative scale.
Screening results should always be interpreted by a qualified clinical geneticist in the context of the patient's clinical presentation and family history.
### Version History
The version history records changes to the methodology.
v2.1.0 Current March 2026
Panel-aware age relevance scoring: panel genes with age_group_relevance="all" receive minimum 0.6 score regardless of patient age
ClinGen gene-disease validity modulates panel gene scores: Definitive/Strong=1.0, Moderate=0.85, Limited=0.7
Compound het detection: splice_region + intron_variant combinations excluded from coding consequences (splice_region alone is not true coding impact)
Phenotype scoring: gene-disease burden capped at 0.5 in screening mode (full 1.0 reserved for actual HPO phenotype match)
Non-coding variants receive discounted phenotype scores proportional to their inability to exploit gene-disease burden
Panel metadata corrections: INS/NEUROD1/LMNA age_group_relevance set to "all", GATA4 ClinGen status corrected to Strong
New sub-panels added: 5 diabetes sub-panels + Comprehensive Diabetes Panel (47 genes)
v2.0.0 February 2026
Initial production release with 7-component scoring system
BayesDel_noAF as primary deleteriousness predictor with ClinGen SVI calibration
8-predictor weighted ensemble for deleteriousness scoring
Age-aware weight profiles: Diagnostic, Neonatal, Pediatric, Adult, Elderly
Clinical profile boosts: ethnicity, family history, sex-linked, consanguinity, pregnancy
Four-tier priority ranking with ACMG classification guarantees
Gene panel boost with ClinGen modulation
### References
Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants.
Genetics in Medicine. 2015;17(5):405-424.
PMID: 25741868
Pejaver V, Byrne AB, Feng BJ, et al. Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3/BP4 criteria.
Am J Hum Genet. 2022;109(12):2163-2177.
PMID: 36413997
Karczewski KJ, Francioli LC, Tiao G, et al. The mutational constraint spectrum quantified from variation in 141,456 humans.
Nature. 2020;581(7809):434-443.
PMID: 32461654
Cheng J, Novati G, Pan J, et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense.
Science. 2023;381(6664):eadg7492.
PMID: 37733863
Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, et al. Predicting Splicing from Primary Sequence with Deep Learning.
Cell. 2019;176(3):535-548.e24.
PMID: 30661751
Miller DT, Lee K, Abul-Husn NS, et al. ACMG SF v3.2: Reducing noise and improving clinical actionability.
Genet Med. 2023;25(1):100726.
PMID: 36344267
Kohler S, Gargano M, Matentzoglu N, et al. The Human Phenotype Ontology in 2021.
Nucleic Acids Research. 2021;49(D1):D1207-D1217.
PMID: 33264411
### Questions About Our Screening Methodology?
Clinical geneticists and laboratory directors can contact us with technical questions about the screening method.
Contact Us Classification Methodology Screening Documentation Phenotype Matching Newborn Screening

### Links and cited sources

- [Classification Methodology](https://folklore.helena.bio/methodology)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)
- [PMID: 36413997](https://pubmed.ncbi.nlm.nih.gov/36413997/)
- [PMID: 32461654](https://pubmed.ncbi.nlm.nih.gov/32461654/)
- [PMID: 37733863](https://pubmed.ncbi.nlm.nih.gov/37733863/)
- [PMID: 30661751](https://pubmed.ncbi.nlm.nih.gov/30661751/)
- [PMID: 36344267](https://pubmed.ncbi.nlm.nih.gov/36344267/)
- [PMID: 33264411](https://pubmed.ncbi.nlm.nih.gov/33264411/)
- [Contact Us](https://folklore.helena.bio/contact)
- [Screening Documentation](https://folklore.helena.bio/docs/screening)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)
- [Newborn Screening](https://folklore.helena.bio/use-cases/newborn-screening)

---

## Page: Folklore Clinical Evidence and Performance Studies

Source: https://folklore.helena.bio/studies
Canonical: https://folklore.helena.bio/studies
Description: Study design, case selection, cohort reports, methodological divergence, limitations, and real-world performance evidence for the Folklore classifier.

Clinical Evidence
## Validation and Real-World Performance Studies
This page reports the study designs, cohort composition, comparator methods, observed divergences, and stated limitations used to evaluate Folklore. Detailed case-level reports are available under appropriate confidentiality terms.
Preprint under peer review
The methodology described on this page has been formalized as a scientific manuscript. The paper is openly available as a preprint and has been submitted for peer review at Frontiers in Genetics (Human and Medical Genomics). Read the preprint: "A Methodological Framework for Real-World Performance Studies of Clinical Variant Classification Platforms at Early Organizational Stages" (Mitev, 2026). DOI: 10.5281/zenodo.21105027.
DOI: 10.5281/zenodo.21105027
197
Cases evaluated
across the validation and real-world cohorts
24
Clinical domains
in the comparative domain table
0
Opposite-direction errors
no P/LP-vs-B/LB discordance in any cohort
100%
Cohort 1 acceptance
pre-registered threshold met
Contents
01 Overview 02 Case Selection Methodology 03 Methodology Framework 04 Methodological Divergence 05 Conflict of Interest Disclosure 06 Limitations 07 Phased Pathway 08 References
Master Methodology
Folklore Real-World Performance Framework v1.0
The methodology document for Folklore performance studies defines the three-layer performance model, multi-source ground truth construction, divergence characterization, conflict-of-interest controls, reporting standards, and the phased path toward regulatory submission.
Read the framework
### Overview
Folklore conducts two complementary types of clinical evidence work: pre-registered validation studies that evaluate platform performance against defined acceptance criteria, and real-world performance studies that characterize platform behavior on real clinical cases under a multi-source ground truth methodology.
Qualified laboratories report 25 to 35 percent disagreement under the ACMG/AMP framework, so a single laboratory cannot serve as definitive ground truth. The method constructs multi-source ground truth, assigns platform-versus-comparator disagreements to six disposition categories, and distinguishes classifier defects from design behavior, feature gaps, and upstream pipeline failures. The preprint is available at DOI: 10.5281/zenodo.21105027.
Inter-laboratory variant classification studies under the ACMG/AMP 2015 framework consistently document 25 to 35 percent legitimate methodological disagreement between qualified laboratories (Amendola et al. 2016; Harrison et al. 2017). Real-world platform performance characterization for clinical bioinformatics software therefore requires multi-source ground truth construction rather than single-source comparison.
The Real-World Performance Framework (HEL-RWP-FRAMEWORK-v1.0, April 2026) defines the methodology for ongoing Folklore studies. It specifies a three-layer performance model, multi-source ground truth construction, divergence characterization, conflict-of-interest disclosure, and a phased path toward future regulatory submission. It replaces prior internal validation-framework documents and limits current claims to real-world performance characterization.
Validation Study
Cohort 1: Pathogenic Variant Concordance Study
Pre-registered validation study with binary acceptance criteria: pathogenic variant concordance against an established peer reference comparator.
Read the full study
Real-World Performance Study
Cohort 2: Multi-Source Performance Characterization
Real-world performance characterization against multi-source ground truth, with a six-category methodological-disposition taxonomy.
Read the full study
Real-World Performance Study
Cohort 3: Real-World Performance Study
The study processed 57 real-world cases under classifier v3.39.1 and re-adjudicated platform-versus-comparator differences per variant. It reports zero confirmed classifier-logic defects.
Read the full study
Deterministic Classification Performance Study
Cohort 4: Deterministic Classification Performance Study
All 100 real-world cases completed under classifier v3.39.1. The study reports 67.0% FULL plus CLINICAL agreement and a 2.2% discordant set; forensic review converted bounded findings into a defined correction programme.
Read the full study
### Case Selection Methodology
Both cohorts used performance-blind selection with pre-registered diversity criteria. No case was added or excluded after Folklore output was observed. The dated selection record identifies the criteria and rationale used.
Cohort 1 Selection
Pre-Registered Validation Study Selection
Source registry Established clinical reference laboratory patient registry Selection date February 2026 (prior to any Folklore processing) Selection bias control Performance-blind: cases chosen before any Folklore platform output was observed Pre-registration Diversity criteria documented in writing with date stamp before processing Inclusion criteria VCF availability in reference laboratory storage; clinically reported pathogenic or likely pathogenic outcome; complete patient phenotype documentation Diversity coverage 11 clinical domains, 6 variant types, 3 zygosity categories Domain composition Cardiology, oncology, ophthalmology, nephrology, neurology, hearing, syndromic, metabolic, rare disease, neuromuscular, reproductive Operator independence External operator (no role in classifier development), independent of Folklore R&D
Cohort 2 Selection
Real-World Performance Study Selection
Source registry Same established clinical reference laboratory patient registry as Cohort 1 Selection date March 2026 (prior to any Folklore processing) Selection bias control Performance-blind: cases chosen before any Folklore platform output was observed Funnel Approximately 171 available VCF files in registry; approximately 32 eligible candidates after applying inclusion criteria; 20 selected per pre-specified diversity axes Pre-registration Five diversity axes documented in writing before selection: clinical domain, mutation type, zygosity, complexity, pathogenic mechanism Leading principle Complementarity, not repetition: cases selected to characterize platform behavior in clinical domains and mutation types not covered by Cohort 1, and to increase case complexity Inclusion criteria VCF availability in reference laboratory storage; clinically reported pathogenic or likely pathogenic outcome; complete patient phenotype documentation Diversity coverage 13 clinical domains, 7 variant types, 8 inheritance patterns Complexity expansion Compound heterozygous representation quadrupled relative to Cohort 1 (4 cases versus 1); 2 multi-gene cases (up to 3 variants in 3 distinct genes); novel chromosomal contexts including first Y-chromosome variant evaluation
Five Diversity Axes (Cohort 2)
Selection structure for Cohort 2 was deliberately complementary to Cohort 1, not repetitive. Five orthogonal diversity axes were pre-specified to maximize coverage of clinical scenarios and platform challenge cases.
1
Clinical domain
Seven entirely new clinical domains relative to Cohort 1 (neuromuscular, autoinflammatory, connective tissue, craniofacial, ciliopathy, Y-linked reproductive, pulmonology), plus expansion of existing domains (nephrology, neurology). The two cohorts together cover 24 distinct clinical domains.
2
Mutation type
Balanced mix across SNVs (8 cases), deletions (7), duplications (1), complex rearrangements (1), and mixed compound heterozygous configurations (4 cases with two variants of different types in one gene). Tests platform behavior across the full spectrum of variant calling outputs.
3
Zygosity
14 heterozygous (autosomal dominant or autosomal recessive carrier), 1 homozygous (autosomal recessive biallelic), 2 hemizygous (X-linked and Y-linked). The Y-chromosome case constitutes the first hemizygous Y-linked evaluation for the platform.
4
Case complexity
14 simple cases (one variant in one gene), 4 compound heterozygous (two variants in one gene), 2 multi-gene cases (two to three variants in distinct genes). This complexity distribution quadruples compound heterozygous representation relative to Cohort 1, reflecting the real clinical scenario in autosomal recessive diseases.
5
Pathogenic mechanism
Coverage spans loss-of-function (frameshift, splice), gain-of-function, enzymatic deficiency, channelopathy, structural defect, and pharmacogenomic significance. Distinct pathogenic mechanisms test classifier behavior across the diverse molecular pathologies encountered in clinical practice.
Clinical Domain Coverage Across Cohorts
The two cohorts together cover 17 distinct clinical domains. Cohort 2 adds seven entirely new domains while extending coverage in nine domains explored in Cohort 1.
Clinical domain Cohort 1 Cohort 2 Status
Cardiology 3 0 Covered in Cohort 1
Oncology 3 0 Covered in Cohort 1
Neurology and movement disorders 2 2 Expanded
Nephrology 2 2 Expanded
Ophthalmology 1 0 Covered in Cohort 1
Hearing disorders 1 0 Covered in Cohort 1
Metabolic 1 0 Covered in Cohort 1
Syndromic / RASopathy 1 1 Expanded
Reproductive / Carrier screening 1 1 Expanded
Neuromuscular 0 4 NEW in Cohort 2
Autoinflammatory 0 1 NEW in Cohort 2
Connective tissue / Ehlers-Danlos 0 1 NEW in Cohort 2
Craniofacial 0 1 NEW in Cohort 2
Ciliopathy 0 2 NEW in Cohort 2
Endocrinology / Diabetes 0 1 NEW in Cohort 2
Y-linked / Reproductive 0 1 NEW in Cohort 2
Pulmonology 0 1 NEW in Cohort 2
Selection Limitations
The following limitations apply to current Phase 1 selection procedures and inform future cohort design.
Single-source patient data: all 40 cases across both cohorts originate from one peer reference laboratory. Multi-laboratory selection is not available within the current organizational stage. Phase 2 Pivotal Cohorts will introduce multi-laboratory case sourcing as standard.
Positive-case enrichment: 100% of selected cases have a clinically reported pathogenic or likely pathogenic outcome. No negative controls (population samples, healthy individuals, or established benign references) are included. This limits clinical specificity claims and reflects the founder-stage focus on detection and classification correctness for known pathogenic variants.
No pre-specified VCEP-curated subset: at selection time, no explicit attempt was made to enrich either cohort with Variant Curation Expert Panel curated variants for higher-tier ground truth comparison. Multi-source ground truth is constructed retrospectively for each case where evidence is available; this is appropriate to current organizational stage but not optimal for inferential claims.
Sample size: n=20 per cohort yields a Wilson 95% confidence interval of approximately 20 percentage points at 70% concordance. This supports descriptive characterization but limits inferential claims. Future Phase 2 Pivotal Cohorts with at least n=30-50 would narrow the confidence interval.
Convenience sampling within constraints: selection from the eligible candidate pool was guided by pre-specified diversity criteria but is not strict random sampling. Bias minimization is achieved through explicit performance-blind selection and pre-registered diversity axes, not through randomization.
### Methodology Framework
Folklore conducts performance studies under HEL-RWP-FRAMEWORK-v1.0 (April 2026), a master methodology document establishing study design, ground truth construction, performance metrics, quality controls, conflict of interest management, and reporting standards.
Multi-source ground truth
Inter-laboratory concordance studies document 25-35% methodological disagreement between qualified laboratories, which limits a single laboratory as definitive ground truth. The Folklore Real-World Performance Framework triangulates ClinGen Variant Curation Expert Panels (highest weight where available), ClinVar 2-star+ aggregated submissions, functional studies with clinical follow-up, population evidence, calibrated predictors per ClinGen SVI 2022, published case reports and segregation, and disease-specific consortium databases.
Three-layer performance model
Performance characterization is structured as three independent and complementary layers. Layer 1 (analytical performance) addresses variant detection, annotation correctness, repeatability, and reproducibility. Layer 2 (classification performance) addresses ACMG class concordance with multi-source ground truth and methodological characterization of all divergences. Layer 3 (clinical performance) addresses tier placement, phenotype match score, screening rank, AI report quality, and AI faithfulness.
Methodological characterization, not adjudication
When Folklore and a reference source diverge, the framework records which ACMG criteria each platform applied or withheld, whether Folklore reflects ClinGen SVI 2020+ specifications, whether independent sources support either classification, and how the difference compares with published inter-laboratory disagreement. Future Pivotal Cohorts reserve formal adjudication for external panels.
Pre-registration and transparency
Cohort composition and selection criteria are dated before processing begins, and selection is performance-blind. Pre-registered Validation Studies define acceptance criteria before processing. Real-World Performance Studies report values across the metric families, compare them with published benchmarks, characterize divergences, record platform defects and remediation status, and state limitations without a binary pass/fail disposition.
### Understanding Methodological Divergence
Inter-laboratory disagreement of 25 to 35 percent is a documented baseline across the field (Amendola 2016, Harrison 2017). Where Folklore and the peer comparator differ, every divergence is methodologically characterized. No opposite-direction (P/LP-vs-B/LB) discordance has been observed in any cohort.
Folklore applies stricter ClinGen SVI 2020+ specifications
Where Folklore and the peer comparator differ, the difference frequently reflects methodological evolution, not platform defect. Folklore implements Pejaver 2022 calibrated PP3/BP4 thresholds, Walker 2023 SpliceAI thresholds aligned to ClinGen SVI 2023, ClinGen SVI PP5 deprecation, BS2 inheritance-aware thresholds, and PVS1 expression-aware guards (gnomAD pext v4.1, GTEx v10). Many peer laboratories continue to apply historical ACMG/AMP 2015 criteria without subsequent ClinGen SVI updates.
Conservative guards operating per design
Classification guards limit overclassification. A computational-only LP guard requires at least one observed criterion (PS1, PM3, PM4, PP4, or PP5) before a variant reaches Likely Pathogenic when its only Strong evidence is computational. An autosomal recessive carrier guard blocks Likely Pathogenic for heterozygous loss-of-function variants in AR-only LoF genes without a compound heterozygous partner. Dual-mechanism bypasses retain Likely Pathogenic eligibility for genes with both monoallelic and biallelic pathways. Outcomes from these guards are categorized as design behavior.
Documented feature-coverage gaps with active roadmap
Three feature gaps were identified as recurring sources of LP/VUS boundary disagreement: PP3 tool divergence, PM1 critical-domain coverage scope, and PP2 not yet implemented. Each is roadmap-tracked. The PM1 coverage gap was directly addressed by the recent UniProt residue-level evidence integration (HELIX-CR-2026-082, classifier v3.30.0, April 2026), which expanded PM1 evidence from approximately 50 hardcoded Pfam domains to 113,126 expert-curated UniProt residues - a 2,262-fold increase in evidence density.
Inherent VCF-only automation limitations
Classification criteria that require manual-curation evidence (PS4 case-control literature, PP1 family cosegregation, PP5 single-submitter evaluations under deprecation review) cannot be fully automated from VCF input alone. Where the peer comparator applies these criteria via manual curation, Folklore conservatively declines to fire them - this is an inherent scope limitation of automated VCF-only classification, transparently disclosed.
No opposite-direction discordance
No opposite-direction (P/LP-vs-B/LB) discordance has been observed in any cohort: Cohorts 1 and 2 recorded no DISCORDANT outcomes, and every DISCORDANT outcome in Cohort 3 is a both-defensible adjacent-category call with no opposite-direction error. The CSER Consortium (Amendola 2020) reports 11% of variants exhibit discordance affecting clinical recommendations across nine laboratories; in Folklore evaluation, opposite-direction discordance is 0%. All differences fall within the P/LP boundary or the LP/VUS boundary, neither of which alters clinical actionability under ClinGen SVI guidance.
### Conflict of Interest Disclosure
Helena Bioinformatics maintains a transparent conflict of interest disclosure framework aligned with academic publication standards (CIOMS, ICH GCP, Helsinki Declaration) and consistent with downstream regulatory requirements. Disclosure is mandatory; mitigations are tailored to study type and organizational stage.
Founder and Chief Executive Officer
Financial (sole owner), Employment
Disclosure: Sole owner of Helena Bioinformatics. Performs all internal Folklore roles in the current single-founder organizational stage, including technical development, scientific direction, business operations, and clinical validation oversight.
Mitigation: Disclosed in all Real-World Performance Study Reports. The multi-source ground truth methodology mitigates single-author classification framing. Future organizational growth will introduce role separation.
Scientific Advisor
Family relationship
Disclosure: Family relationship to the Folklore founder.
Mitigation: Excluded from formal adjudication. Role limited to Scientific Advisor (methodology consultation, framework review, scientific guidance). Family relationship disclosed in any report where input is referenced.
Cohort 2 case operator
Employment
Disclosure: Employed by the peer reference laboratory used in Cohort 2.
Mitigation: Disclosed in the Cohort 2 Real-World Performance Study Report. Operator role limited to technical processing and comparison documentation; not classification authority. The multi-source ground truth methodology mitigates single-source dependency. Future cohorts will introduce independent operator structure with no peer-laboratory affiliations.
External clinical signatory
Professional affiliation
Disclosure: Has signed peer reference laboratory clinical reports as external clinical signatory, including reports referenced in Cohort 2. Not an employee of the peer reference laboratory.
Mitigation: Cannot serve as adjudicator on cases previously signed. May serve as Folklore Scientific Advisor or methodology consultant for non-overlapping cohorts. Disclosed in any report where input is referenced.
### Limitations and Caveats
The following limitations apply to current Phase 1 evidence and inform future cohort design.
The 40 cases across two cohorts limit inferential statistical power. At 70% concordance with n=20, the Wilson 95% confidence interval is approximately 20 percentage points. The results support descriptive trend characterization, while future Phase 2 Pivotal Cohorts with at least n=30-50 would narrow the interval.
Single peer reference comparator. All peer comparisons in both cohorts are with one established clinical reference laboratory. Inter-laboratory variation inherent in ACMG classification means that disagreement with any single laboratory does not necessarily indicate platform error. The multi-source ground truth methodology mitigates single-source dependency but does not replace multi-laboratory comparison. Phase 2 Pivotal Cohorts will introduce multi-laboratory comparison as standard.
Cohort selection enriched for known pathogenic and likely pathogenic variants. Performance on Variant of Uncertain Significance classification, benign variant filtering, and novel variants without prior literature is not directly assessed in these cohorts. Future Methodology Validation Studies (database-derived against ClinVar 3-star+ benchmarks and GIAB truth sets) will address these performance dimensions at scale.
Genome build and transcript namespace differences. The peer reference comparator uses GRCh37 with RefSeq NM transcripts; Folklore uses GRCh38 with Ensembl ENST canonical transcripts. While liftover and cross-namespace mapping are generally reliable for SNVs and small indels, documented edge cases include transcript-isoform divergence with N-terminal exon usage differences and occasional position-offset inconsistencies. These are documented methodological observations, not platform defects.
Methodological characterization without external adjudication. Per the Real-World Performance Framework, methodological characterization of PARTIAL outcomes is performed by Folklore R&D rather than by external multi-adjudicator panels. This is appropriate to the current Phase 1 organizational stage but does not provide the same level of independence as multi-adjudicator panel review. Phase 2 Pivotal Cohorts will introduce external adjudication panels with explicit conflict-of-interest controls.
Operator structure for Cohort 2. The Cohort 2 case operator is employed by the peer reference comparator entity. The conflict of interest is disclosed with mitigations applied (operator role limited to technical processing; multi-source ground truth methodology mitigates single-source dependency). Phase 2 Pivotal Cohorts will introduce independent operator structure.
VCEP coverage. Only one of twenty Cohort 2 cases had a Variant Curation Expert Panel curated reference available, reflecting the state-of-the-field limitation that VCEP coverage across rare-disease genes remains sparse (approximately 250 genes with active VCEP curation as of April 2026 per ClinGen).
Results should be interpreted in the context of the patient's clinical presentation, family history, and other available clinical information. Folklore is a clinical decision support tool, not a diagnostic device. A qualified clinical geneticist must review and confirm each classification.
### Phased Pathway to Regulatory Submission
Folklore operates within Phase 1 of an explicitly phased pathway toward eventual regulatory submission. Phase progression is conditional on organizational growth and is not pursued at the expense of current scientific work.
Phase 1 Foundation (Current) Active
Solo founder organizational stage. Real-World Performance Studies and Methodology Validation Studies. Foundational scientific evidence accumulation.
- Real-World Performance Studies (Cohorts 1 and 2 completed)
- Methodology Validation Studies (database-derived against ClinVar 3-star+ benchmarks and GIAB truth sets, in planning)
- Scientific publications and conference abstracts
- Foundational evidence carryforward to Phase 2 and Phase 3
Phase 2 Growth Future
Small team. Series A or strategic partnership. Dedicated Quality and Regulatory functions. First Pivotal Cohort possible. Multi-cohort program execution.
- Phase 1 activities continue
- First Pivotal Cohort (n=30-50, formal acceptance criteria)
- Multi-adjudicator panel introduction
- Multi-laboratory comparison as standard
- Independent operator structure
Phase 3 Pre-submission Future
Established team. Mature Quality Management System. Notified Body engagement. Pivotal Cohorts at full composition. Multi-site validation.
- Pivotal Cohorts at full composition (60/30/10 positive/negative/edge)
- Multi-site validation across independent laboratories
- Pre-submission Notified Body consultations
- IVDR (EU 2017/746) conformity assessment submission
- Post-market performance follow-up planning
### References
Peer-reviewed literature underlying the methodology framework, ACMG/AMP standards, ClinGen SVI specifications, and inter-laboratory benchmark studies.
Amendola LM, Jarvik GP, Leo MC, et al. Performance of ACMG-AMP variant-interpretation guidelines among nine laboratories in the Clinical Sequencing Exploratory Research Consortium.
American Journal of Human Genetics. 2016;98(6):1067-1076.
PMID: 27181684
Harrison SM, Dolinsky JS, Knight Johnson AE, et al. Clinical laboratories collaborate to resolve differences in variant interpretations submitted to ClinVar.
Genetics in Medicine. 2017;19(10):1096-1104.
PMID: 28301460
Richards S, Aziz N, Bale S, et al. Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology.
Genetics in Medicine. 2015;17(5):405-424.
PMID: 25741868
Tavtigian SV, Greenblatt MS, Harrison SM, et al. Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework.
Human Mutation. 2018;39(11):1485-1492.
PMID: 30311386
Pejaver V, Byrne AB, Feng BJ, et al. Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3/BP4 criteria.
American Journal of Human Genetics. 2022;109(12):2163-2177.
PMID: 36413997
Walker LC, Hoya M, Wiggins GAR, et al. Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing: Recommendations from the ClinGen SVI Splicing Subgroup.
American Journal of Human Genetics. 2023;110(7):1046-1067.
PMID: 37352859
Abou Tayoun AN, Pesaran T, DiStefano MT, et al. Recommendations for interpreting the loss of function PVS1 ACMG/AMP variant criterion.
Human Mutation. 2018;39(11):1517-1524.
PMID: 30192042
Roy S, Coldren C, Karunamurthy A, et al. Standards and guidelines for validating next-generation sequencing bioinformatics pipelines: a joint recommendation of the Association for Molecular Pathology and the College of American Pathologists.
Journal of Molecular Diagnostics. 2018;20(1):4-27.
PMID: 29154853
Amendola LM, Muenzen K, Biesecker LG, et al. Variant classification concordance using the ACMG-AMP variant interpretation guidelines across nine genomic implementation research studies.
American Journal of Human Genetics. 2020;107(5):932-941.
PMID: 33108757
Cummings BB, Karczewski KJ, Kosmicki JA, et al. Transcript expression-aware annotation improves rare variant interpretation.
Nature. 2020;581(7809):452-458.
PMID: 32461655
Mitev V. A Methodological Framework for Real-World Performance Studies of Clinical Variant Classification Platforms at Early Organizational Stages.
Zenodo [preprint]. 2026. doi: 10.5281/zenodo.21105027.
10.5281/zenodo.21105027
### Request Detailed Cohort Reports
Detailed cohort reports, per-case performance documentation, and the full Real-World Performance Framework are available to clinical partners, peer reviewers, investors, and regulatory consultants under appropriate confidentiality terms.
Contact Us ACMG Methodology Founding Clinical Partner How It Works For Geneticists

### Links and cited sources

- [DOI: 10.5281/zenodo.21105027](https://doi.org/10.5281/zenodo.21105027)
- [Master Methodology Folklore Real-World Performance Framework v1.0 The methodology document for Folklore performance studies defines the three-layer performance model, multi-source ground truth construction, divergence characterization, conflict-of-interest controls, reporting standards, and the phased path toward regulatory submission. Read the framework](https://folklore.helena.bio/studies/framework)
- [Validation Study Cohort 1: Pathogenic Variant Concordance Study Pre-registered validation study with binary acceptance criteria: pathogenic variant concordance against an established peer reference comparator. Read the full study](https://folklore.helena.bio/studies/cohort-1)
- [Real-World Performance Study Cohort 2: Multi-Source Performance Characterization Real-world performance characterization against multi-source ground truth, with a six-category methodological-disposition taxonomy. Read the full study](https://folklore.helena.bio/studies/cohort-2)
- [Real-World Performance Study Cohort 3: Real-World Performance Study The study processed 57 real-world cases under classifier v3.39.1 and re-adjudicated platform-versus-comparator differences per variant. It reports zero confirmed classifier-logic defects. Read the full study](https://folklore.helena.bio/studies/cohort-3)
- [Deterministic Classification Performance Study Cohort 4: Deterministic Classification Performance Study All 100 real-world cases completed under classifier v3.39.1. The study reports 67.0% FULL plus CLINICAL agreement and a 2.2% discordant set; forensic review converted bounded findings into a defined correction programme. Read the full study](https://folklore.helena.bio/studies/cohort-4)
- [PMID: 27181684](https://pubmed.ncbi.nlm.nih.gov/27181684/)
- [PMID: 28301460](https://pubmed.ncbi.nlm.nih.gov/28301460/)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)
- [PMID: 30311386](https://pubmed.ncbi.nlm.nih.gov/30311386/)
- [PMID: 36413997](https://pubmed.ncbi.nlm.nih.gov/36413997/)
- [PMID: 37352859](https://pubmed.ncbi.nlm.nih.gov/37352859/)
- [PMID: 30192042](https://pubmed.ncbi.nlm.nih.gov/30192042/)
- [PMID: 29154853](https://pubmed.ncbi.nlm.nih.gov/29154853/)
- [PMID: 33108757](https://pubmed.ncbi.nlm.nih.gov/33108757/)
- [PMID: 32461655](https://pubmed.ncbi.nlm.nih.gov/32461655/)
- [10.5281/zenodo.21105027](https://zenodo.org/records/21105027)
- [Contact Us](https://folklore.helena.bio/contact)
- [ACMG Methodology](https://folklore.helena.bio/methodology)
- [Founding Clinical Partner](https://folklore.helena.bio/partners/cellgenetics)
- [How It Works](https://folklore.helena.bio/how-it-works)
- [For Geneticists](https://folklore.helena.bio/for-geneticists)

---

## Page: Cohort 1: Pathogenic Variant Concordance Study | Folklore

Source: https://folklore.helena.bio/studies/cohort-1
Canonical: https://folklore.helena.bio/studies/cohort-1
Description: Pre-registered validation study of the Folklore variant classification platform: pathogenic variant concordance against an established peer reference comparator, with binary acceptance criteria.

Back to Studies
Validation Study
## Cohort 1: Pathogenic Variant Concordance Study
Pre-registered validation study with binary acceptance criteria. Twenty clinical cases with known pathogenic or likely pathogenic variants evaluated against an established clinical reference laboratory.
Public scientific report
### Cite this study
The archived report, persistent DOI and full metadata record are available through OSF.
Suggested citation
Mitev V. Cohort 1: Pathogenic Variant Concordance Study. Helena Bioinformatics public scientific report, 1.0, August 2026. doi: 10.17605/OSF.IO/KVJ8E . DOI 10.17605/OSF.IO/KVJ8E
Full record on OSF
Study specification
Document HEL-VAL-C1-2026-001 v1.0 Study type Pre-registered validation study Period February - March 2026 Classifier version v3.6.4 (final) Reference comparator Established clinical reference laboratory Sample size 20 cases Clinical domains 11 (rare disease, oncology, cardiology, neurology, metabolic, ophthalmology, nephrology, neuromuscular, hearing, syndromic, reproductive) Variant types 6 (missense, nonsense, frameshift, splice, compound heterozygous, multi-gene carrier) Zygosity coverage 3 (heterozygous, homozygous, compound heterozygous) Acceptance threshold at least 18/20 (90%) for all primary criteria Standards applied ISO 15189:2022, EU IVDR (2017/746), AMP/CAP 2018, ACMG/AMP 2015
Results against pre-registered acceptance criteria
Metric Result
P1 - Variant detection 20/20 (100%)
P2 - Gene assignment 20/20 (100%)
P3 - ACMG concordance (FULL) 14/20 (70%)
P3 - ACMG concordance (CLINICAL, one-tier P/LP) 6/20 (30%)
P3 - PARTIAL or DISCORDANT 0/20 (0%)
P3 - Overall (FULL + CLINICAL) 20/20 (100%)
S1 - HGVS notation match 20/20 (100%)
S2 - Consequence match 20/20 (100%)
S3 - Zygosity match 20/20 (100%)
All primary criteria met 20/20 (100%)
Validation decision: GO
All 20 cases (100%) met all three primary concordance criteria (variant detection, gene assignment, ACMG concordance), exceeding the pre-defined GO threshold of 90% (>= 18/20). No PARTIAL or DISCORDANT results were observed. The six CLINICAL (one-tier P/LP) differences are within the expected range of inter-laboratory ACMG classification variability as documented in the peer-reviewed literature (Amendola 2016, Harrison 2017). The six CLINICAL differences fall in two patterns: cases where Folklore applied more conservative ClinGen SVI 2023 thresholds (SpliceAI < 0.2 for PP3_splice), and cases where Folklore applied ClinVar 2-star+ override or VCEP gene-specific rules to elevate Likely Pathogenic to Pathogenic.
The validation process additionally identified and resolved four classifier defects through iterative improvement (v3.6.0 to v3.6.4). The most significant was the systematic autosomal recessive loss-of-function gene curation expansion (approximately 150 genes with Definitive or Strong ClinGen Gene-Disease Validity for biallelic LoF disease mechanism), which prevents PVS1 misclassification of loss-of-function variants in autosomal recessive disease genes whose population constraint metrics are statistically underpowered.
Back to Studies

### Links and cited sources

- [Back to Studies](https://folklore.helena.bio/studies)
- [doi: 10.17605/OSF.IO/KVJ8E](https://doi.org/10.17605/OSF.IO/KVJ8E)
- [Full record on OSF](https://osf.io/kvj8e/)

---

## Page: Cohort 2: Multi-Source Performance Characterization | Folklore

Source: https://folklore.helena.bio/studies/cohort-2
Canonical: https://folklore.helena.bio/studies/cohort-2
Description: Real-world performance characterization of the Folklore variant classification platform against multi-source ground truth, with a six-category methodological-disposition taxonomy.

Back to Studies
Real-World Performance Study
## Cohort 2: Multi-Source Performance Characterization
Real-World Performance Study under HEL-RWP-FRAMEWORK-v1.0. Twenty cases evaluated under a three-layer performance model with multi-source ground truth construction. Methodological characterization of all divergences. Not a binary pass/fail study; foundational scientific evidence at the current organizational stage.
Public scientific report
### Cite this study
The archived report, persistent DOI and full metadata record are available through OSF.
Suggested citation
Mitev V. Cohort 2: Multi-Source Performance Characterization. Helena Bioinformatics public scientific report, 1.0, August 2026. doi: 10.17605/OSF.IO/RTBK8 . DOI 10.17605/OSF.IO/RTBK8
Full record on OSF
Study specification
Document HEL-RWP-C2-2026-001 v1.0 Study type Real-World Performance Study under HEL-RWP-FRAMEWORK-v1.0 Period March - April 2026 Classifier version v3.28.0 (final at cohort closure) Reference sources Multi-source ground truth: VCEP curations (where available), ClinVar 2-star+ aggregated submissions, peer reference comparator Sample size 20 cases Clinical domains 13 (neurology, autoinflammatory, nephrology X-linked, RASopathy, neuro-metabolic, endocrine, neuromuscular, lysosomal storage, skeletal dysplasia, connective tissue, respiratory, craniofacial, ciliopathy, Y-linked reproductive, mitochondrial) Variant types 7 (missense, nonsense, frameshift, splice region, polypyrimidine tract, complex indel, large recurrent indel) Inheritance patterns 8 (AD reduced penetrance, AD haploinsufficient, AR, AR carrier, AR compound heterozygous, X-linked recessive, Y-linked LoF, dual AD/AR mechanism) Performance model Three-layer (analytical, classification, clinical workflow) Disposition methodology Methodological characterization (not binary pass/fail), six-category disposition taxonomy L1
Layer 1 - Analytical performance
Variant detection (P1) 19/20 (95%)
Single non-detection attributed to upstream sequencing pipeline failure via 7-step audit trail; Folklore platform integrity verified via 21:21 input-to-database record correspondence; Folklore attribution CLEARED
Gene assignment (P2) 19/19 (100% where evaluable) HGVS / consequence / zygosity (S1-S3) All PASS where evaluable L2
Layer 2 - Classification performance
FULL concordance 13 components
Identical ACMG class with the peer reference comparator. Anchored variously in ClinVar 2-3 star, ClinGen Hearing Loss VCEP v1.0, KDIGO 2025 ADPKD guideline, or independent multi-criterion convergence.
PARTIAL components 6 components (across 4 cases)
Methodologically characterized: 2 validated design behavior (computational-only LP guard and AR carrier guard operating per design), 3 feature-gap accumulation at LP/VUS boundary (PP3 tool divergence, PM1 coverage, PP2 implementation), 1 manual-criteria-dependent inherent automation limitation
DISCORDANT outcomes 0
No opposite-direction classifications observed
NOT EVALUABLE 1
Upstream variant calling pipeline failure (verified non-Folklore attribution)
Inter-laboratory benchmark Within published range
Amendola 2016 (66%), Harrison 2017 (72-76%), Bergquist 2025 (70-75%) inter-laboratory FULL concordance benchmarks
L3
Layer 3 - Clinical workflow performance
Q1 - Tier placement PASS in diagnostic indication cases
Target P/LP variants placed in Tier 1 of phenotype matching output
Q2 - Phenotype match score HIGH in cases with appropriate HPO input
at least 50% match score
Q4 - AI report quality GOOD in majority
Variant mentioned, interpretation correct, overall quality good
AI Faithfulness 14/20 (70%) PASS
Seven consecutive PASS at cohort closure (Cases 14-20). Improvement trajectory documented across the cohort.
Change Request validations 10 deployed CRs validated
Trigger-configuration validation of conservative-guard architecture across diverse real-world clinical scenarios
Six-category methodological-disposition taxonomy
Cohort 2 assigns each observed concordance pattern to a methodological-disposition category. The taxonomy records why results differ and informs future cohort design without reducing the study to a binary pass/fail result.
Tier 1 ground-truth concordance
13 components
Folklore and the peer comparator both reach the same classification, anchored in VCEP curation, ClinVar 2-star or higher, or established clinical practice guidelines. Multiple ACMG criteria converge on the same class.
Disposition: No action. System operates correctly. Multiple validation milestones documented.
Validated design behavior
2 components
Folklore classification reflects a documented classifier guard operating per ClinGen-aligned conservative design principles. Disagreement with the peer comparator does not reflect platform error; the guard is the explicit intended output for genotype-context-aware classification.
Disposition: No remediation. Update Change Request registry status to validated with cohort cases as evidence anchor.
Feature-gap accumulation at LP/VUS boundary
3 components
Multiple documented feature gaps in Folklore automated classification scope collectively prevent reaching the LP combining-rule threshold. Three identified gaps: PP3 computational predictor tool divergence, PM1 critical-domain coverage scope, and PP2 not yet implemented in the automated classifier. Each gap is independently characterized and roadmap-tracked. Distinct from defect: the gaps are documented limitations, not errors.
Disposition: Roadmap-tracked. PM1 domain coverage expansion in flight (HELIX-CR-2026-082, deployed April 2026 with UniProt residue-level evidence integration). PP2 implementation roadmap formalization. PP3 tool-choice divergence: no classifier change recommended (defensible methodological choice per ClinGen SVI).
Manual-criteria-dependent inherent automation limitation
1 component
Classification depends primarily on manual-curation evidence (case-control / case-series literature, family cosegregation, individual ClinVar submitter evaluation) that is not available from VCF input. Distinct from feature gap: not a classifier roadmap item, but an inherent scope limitation of automated VCF-only classification.
Disposition: Long-term automation roadmap items only: literature mining for PS4 case-control automation, external trio data ingestion for PP1 segregation, gene-aware PP2 thresholds for small genes. Not classifier defect remediation.
Input data layer upstream pipeline failure
1 case (NOT EVALUABLE)
Target variants absent from input VCF due to upstream sequencing or variant calling pipeline. Folklore platform integrity verified via 21:21 input-to-database record correspondence at correct genomic target coordinates. Failure is upstream, not Folklore.
Disposition: No Folklore remediation required (attribution cleared). Upstream pipeline investigation pending sequencing facility action.
Back to Studies

### Links and cited sources

- [Back to Studies](https://folklore.helena.bio/studies)
- [doi: 10.17605/OSF.IO/RTBK8](https://doi.org/10.17605/OSF.IO/RTBK8)
- [Full record on OSF](https://osf.io/rtbk8/)

---

## Page: Cohort 3: Real-World Performance Study | Folklore

Source: https://folklore.helena.bio/studies/cohort-3
Canonical: https://folklore.helena.bio/studies/cohort-3
Description: Real-world performance characterization of the Folklore variant classification platform (classifier v3.39.1) on 57 real-world clinical cases against a multi-source ground truth and the CellGenetics peer comparator, with per-variant forensic re-adjudication and a six-category methodological-disposition taxonomy.

Back to Studies
Real-World Performance Study
## Cohort 3: Real-World Performance Study
Cohort 3 includes 57 scored real-world clinical cases under HEL-RWP-FRAMEWORK-v1.0: 56 are scoreable per variant, 3 awaiting-comparator cases are excluded, and one pharmacogenetic re-analysis is reported separately. Cases were processed with locked classifier version v3.39.1 against multi-source ground truth and the CellGenetics peer comparator. Platform-versus-comparator differences received independent per-variant forensic re-adjudication. This Phase 1 study reports foundational evidence, not a binary pass/fail result.
Public scientific report
### Cite this study
The archived report, persistent DOI and full metadata record are available through OSF.
Suggested citation
Mitev V. Cohort 3: Real-World Performance Study. Helena Bioinformatics public scientific report, 1.0, August 2026. doi: 10.17605/OSF.IO/K8UWF . DOI 10.17605/OSF.IO/K8UWF
Full record on OSF
Study specification
Document HEL-RWP-C3-2026-001 v5_0 (DRAFT, forensic-adjudicated) Study type Real-World Performance Study under HEL-RWP-FRAMEWORK-v1.0 Period Concordance validation under v3.39.1; cohort closed June 2026 Classifier version v3.39.1 (final at cohort closure) Reference sources Multi-source ground truth: VCEP curations, ClinVar 2-star+ aggregated submissions, calibrated computational predictors, and the CellGenetics peer comparator (comparator only, not ground truth) Cohort size 60 manifest cases; 57 scored (3 awaiting-comparator excluded); 56 per-variant scoreable (one TPMT-only pharmacogenetic re-analysis segregated); 184 shared-variant denominator Clinical domains 15+ (cardiology, nephrology, neurology, neuromuscular, skeletal dysplasia, ophthalmology, oncology, metabolic, hematology, endocrinology, RASopathy, deafness, reproductive, thrombophilia, carrier screening) Variant types Full coding range: missense, nonsense, frameshift/indel, canonical and non-canonical splice, synonymous with RNA effect, in-frame deletion, start-loss, 5-prime UTR, complex delins (out-of-scope SV/CNV surfaced by the comparator are reported separately) Inheritance patterns 8+ (autosomal dominant, autosomal recessive incl. compound-het and carrier states, X-linked recessive and dominant, Y-linked, dual AD/AR mechanism, low-penetrance risk alleles) Performance model Three-layer (analytical, classification, clinical) with an added per-variant forensic re-adjudication of every LP/VUS discordance Disposition methodology Methodological characterization (not binary pass/fail), six-category disposition taxonomy L1
Layer 1 - Analytical performance
Variant detection (P1) 184/186 in-scope shared variants (98.9%, Wilson 95% CI 96.2-99.7%)
Two in-scope absences attributed under Category 6 (Case 38 P4HA1; Case 44 SERAC1); 11 further absences are out-of-scope SV/CNV events excluded from the detection denominator per framework Section 2.1. Benchmark: CLSI MM09-A2 at least 99%.
Gene / HGVS / consequence / zygosity (P2, S1-S3) Concordant on all detected shared variants
Namespace-aware (Ensembl ENST vs RefSeq NM reconciled per variant)
Repeatability / reproducibility Deterministic where examined
Identical class on re-resolution within session
L2
Layer 2 - Classification performance
FULL concordance vs peer 102/184 (55.4%, Wilson 95% CI 48.2-62.4%)
Within the published 60-75% inter-laboratory benchmark range (Amendola 2016; Harrison 2017; Bergquist 2025).
FULL + CLINICAL (clinical-equivalent) 121/184 (65.8%, Wilson 95% CI 58.6-72.2%)
Clinically equivalent agreement per ClinGen SVI; within the published inter-laboratory range.
PARTIAL components 56 across 34 cases
Every PARTIAL is an LP-vs-VUS or LB-vs-VUS adjacent-category divergence -- the most frequent inter-laboratory locus (Bergquist 2025). No opposite-direction PARTIAL.
DISCORDANT outcomes 7 (all both-defensible)
Low-penetrance / risk-allele boundary calls (e.g. HFE, CHEK2, SERPINA1 Z, WARS2, AMPD1). None is an opposite-direction (P/LP-vs-B/LB) error; on two of them the evidence favours Folklore.
NOT EVALUABLE 13 (reported separately)
11 out-of-scope SV/CNV scope-boundary absences; 2 in-scope variant-calling gaps (Cases 38, 44). Excluded from the concordance fractions per framework Section 2.1.
Forensic re-adjudication 0 confirmed classifier-logic defects
The LP/VUS discordance set was independently re-grounded from the session databases and traced through the full classifier path; each resolves to a validated guard, a correct eligibility or combining outcome, or a comparator over-call on a single carrier allele. No Category 1 defect confirmed.
L3
Layer 3 - Clinical workflow performance
Q1 - Tier placement PASS on detected, P/LP, phenotype-matched diagnostic anchors
N/A on genotype-only carrier screens with no phenotype input.
Q2 - Phenotype match HIGH on matched diagnostic cases
For example NPR2 85.4%, USP9Y 94.3%, PKD1 100%.
Q4 - AI report quality AI interpretation in 49/57 cases
Three material items escalated to the AI service; the remaining are polish items.
AI faithfulness 49/49 AI-bearing cases free of gene-disease name hallucination
Clean cohort-wide; the earlier-cohort hallucination pattern did not recur.
Six-category methodological-disposition taxonomy
Cohort 3 assigns non-FULL outcomes to six methodological-disposition categories. LP/VUS differences also received independent per-variant forensic re-adjudication.
Classifier defect
None confirmed
Output inconsistent with ACMG/AMP 2015 or ClinGen SVI. This category is reserved for the human clinical-scientific gate and is never auto-assigned. The re-adjudicated LP/VUS discordance set was traced end to end and none was a classifier-logic defect.
Disposition: No Category 1 remediation owed by this cohort. Two engineering findings the study surfaced (PM5 re-enablement; comp-het splice-partner detection) were nonetheless fixed and are live.
Validated design behavior
Dominant for LoF-null discordances
A documented conservative guard operating per design -- most often the autosomal-recessive carrier guard correctly holding a single heterozygous null allele at carrier VUS, and the PM3 comp-het partner-validation guard withholding auto-fire where no ClinVar-anchored partner exists in trans.
Disposition: No remediation. Change Request registry updated to validated with the cohort cases as evidence anchor.
Feature-gap accumulation at LP/VUS boundary
Roadmap-tracked
Documented, roadmap-tracked feature gaps or both-defensible low-penetrance boundaries prevent reaching LP. Distinct from a defect: these are documented limitations, including the low-penetrance risk-allele frequency-convention boundary.
Disposition: Roadmap registration; no specification breach.
Manual-criteria-dependent automation limitation
Inherent scope limit
The comparator classification rests primarily on evidence unavailable from VCF input alone -- de-novo trio, functional, segregation, or literature evidence. An inherent limit of automated VCF-only classification, not a defect.
Disposition: Long-term automation roadmap only.
Tier 1 ground-truth concordance
Anchors the 102 FULL outcomes
Folklore and the comparator reach the same class, anchored in Tier 1 ground truth (VCEP curations such as ENIGMA BRCA1/2, ClinGen Hearing Loss GJB2, RASopathy, LGMD, Platelet, Monogenic Diabetes, Hemoglobinopathy, Lysosomal), often via different criteria.
Disposition: No action. System operates correctly.
Input data layer upstream pipeline failure
13 NOT EVALUABLE
The target was absent from the input VCF in 11 out-of-scope SV/CNV events and 2 in-scope variant-calling gaps. Separate checks found no generated in-scope variant used to fill a gap.
Disposition: Out-of-scope absences require no platform remediation; the two in-scope gaps trigger the upstream-attribution audit.
Back to Studies

### Links and cited sources

- [Back to Studies](https://folklore.helena.bio/studies)
- [doi: 10.17605/OSF.IO/K8UWF](https://doi.org/10.17605/OSF.IO/K8UWF)
- [Full record on OSF](https://osf.io/k8uwf/)

---

## Page: Cohort 4: Deterministic Classification Performance Study | Folklore

Source: https://folklore.helena.bio/studies/cohort-4
Canonical: https://folklore.helena.bio/studies/cohort-4
Description: Large-scale fixed-version evaluation of Folklore deterministic ACMG classification across 100 completed real-world cases, with session-level forensic review and a defined correction programme.

Back to Studies
Deterministic Classification Performance Study
## Cohort 4: Deterministic Classification Performance Study
Cohort 4 is the largest Folklore study to date: 100 of 100 real-world cases completed under locked classifier version v3.39.1. The primary analysis covers 367 evaluable outcomes from 395 reported outcomes, followed by session-level forensic review of 19 priority variants across 16 cases. The study separates deterministic classifier defects, reference-data gaps, manual evidence limitations and legitimate inter-laboratory differences, then converts confirmed findings into a defined correction and full-cohort regression programme.
Public scientific report
### Cite this study
The archived report, persistent DOI and full metadata record are available through OSF.
Suggested citation
Mitev V. Cohort 4: Deterministic Classification Performance Study. Helena Bioinformatics public scientific report, 1.0, August 2026. doi: 10.17605/OSF.IO/NPHW8 . DOI 10.17605/OSF.IO/NPHW8
Full record on OSF
Study specification
Document HEL-RWP-C4-2026-001-PUBLIC v1.0 Study type Large-scale fixed-version real-world performance study under HEL-RWP-FRAMEWORK-v1.0 Period Cohort closed July 2026; forensic review completed August 2026 Classifier version Folklore v3.39.1 at study closure Execution 100/100 cases completed; no failed session Primary denominator 367 evaluable outcomes from 395 reported outcomes (92.9%) Forensic review set 19 priority variants across 16 cases Comparator role CellGenetics peer comparison, evaluated against multi-source evidence rather than treated as single-source ground truth Privacy Aggregate results only; no patient, case, session, batch or variant identifiers are published L1
Layer 1 - Cohort execution and evaluability
Completed cases 100/100 (100%)
All planned real-world cases completed under the locked study version; no session failed.
Evaluable outcomes 367/395 (92.9%)
The denominator includes outcomes for which a platform-versus-comparator classification assessment was methodologically supportable.
Not evaluable 28 outcomes
Reported separately and excluded from concordance fractions rather than silently treated as disagreement.
L2
Layer 2 - Deterministic classification performance
FULL concordance 223/367 (60.8%; Wilson 95% CI 55.6-65.6%)
Exact five-tier ACMG/AMP agreement.
FULL + CLINICAL 246/367 (67.0%; Wilson 95% CI 62.1-71.6%)
Exact or clinically equivalent adjacent-category agreement.
PARTIAL outcomes 113/367 (30.8%)
Adjacent-boundary methodological differences, concentrated at Pathogenic/Likely Pathogenic/VUS and VUS/Likely Benign boundaries.
DISCORDANT outcomes 8/367 (2.2%)
A small, reviewable set concentrated in risk-allele and mechanism-specific boundaries.
Forensic deep dive 19 variants across 16 cases
Each priority difference was traced from session DuckDB evidence through deterministic ACMG criteria and available reference data.
L3
Layer 3 - Findings and correction programme
Confirmed deterministic defects 4 mechanisms plus 1 family-architecture finding
Disease-aware BS1 selection, dual-mechanism/PVS1 handling, BP4_Moderate strength handling, carrier/class separation, and a separate family-level architecture correction.
Reference-data gaps Confirmed as a distinct source category
Some differences depend on evidence absent from the current automated reference layer; these are assigned to source-ingestion work, not mislabelled as combining-logic defects.
Manual evidence limitations Expected scope boundary
Functional, segregation, de-novo and case-specific literature evidence cannot always be derived from variant input alone and is not counted as a classifier defect.
Correction and regression Defined and cohort-wide
Confirmed defects will be corrected, relevant reference sources will be added, and all 100 cases will be reprocessed in shadow mode against the locked baseline.
Study conclusion Successful
The cohort met its scientific purpose: it measured performance at scale, identified bounded defects and converted them into a verifiable improvement programme.
Principal findings and action programme
The forensic review assigns material differences to the source that can actually change them. This preserves a clean distinction between software defects, missing reference evidence, non-automatable clinical evidence and legitimate peer-method differences.
Deterministic classifier defects
Confirmed and bounded
Four deterministic mechanisms require correction: disease-aware BS1 source selection, dual-mechanism PVS1 eligibility, BP4_Moderate strength handling, and separation of carrier context from variant class. A separate family-architecture finding is tracked with the same regression programme.
Action: Implement the corrections and prove each one with targeted regression fixtures plus a 100-case shadow re-run.
Reference-data coverage gaps
Source-ingestion programme
Some comparator evidence was not available in the automated reference layer at study time. The gap is attributable to missing source coverage rather than faulty ACMG combining logic.
Action: Add the authoritative sources to the reference-data lifecycle, preserve provenance and versioning, then re-evaluate affected outcomes.
Manual clinical evidence
Recognised automation boundary
Segregation, de-novo status, functional assays and case-specific literature interpretation may require human evidence not present in a VCF or reference snapshot.
Action: Keep these outcomes in specialist review and do not count the absence of non-automatable evidence as a classifier defect.
Inter-laboratory classification differences
Expected and reviewable
Qualified laboratories can reach adjacent ACMG/AMP categories through different admissible evidence sets and strength assignments. CellGenetics manual calls are therefore evaluated, not accepted as ground truth by default.
Action: Retain both-defensible differences as methodological variation; correct Folklore only where the evidence trace demonstrates a platform defect or data gap.
Back to Studies

### Links and cited sources

- [Back to Studies](https://folklore.helena.bio/studies)
- [doi: 10.17605/OSF.IO/NPHW8](https://doi.org/10.17605/OSF.IO/NPHW8)
- [Full record on OSF](https://osf.io/nphw8/)

---

## Page: Real-World Performance Framework v1.0 | Folklore

Source: https://folklore.helena.bio/studies/framework
Canonical: https://folklore.helena.bio/studies/framework
Description: Folklore Real-World Performance Framework v1.0 - master methodology document for Real-World Performance Studies of the Folklore variant classification platform. Multi-source ground truth, three-layer performance model, methodological characterization, transparent COI disclosure, phased pathway to regulatory submission.

Back to Studies
## Folklore Real-World Performance Framework
Methodology for Real-World Performance Studies of Folklore
Version 1.0 - April 2026
This page presents Folklore Real-World Performance Framework v1.0 for clinical partners, peer reviewers, investors, strategic partners, and regulatory consultants. It describes Helena Bioinformatics's current Phase 1 work. The resulting evidence may support a future regulatory submission but is not itself a regulatory submission.
Preprint under peer review
The methodology described on this page has been formalized as a scientific manuscript. The paper is openly available as a preprint and has been submitted for peer review at Frontiers in Genetics (Human and Medical Genomics). Read the preprint: "A Methodological Framework for Real-World Performance Studies of Clinical Variant Classification Platforms at Early Organizational Stages" (Mitev, 2026). DOI: 10.5281/zenodo.21105027.
DOI: 10.5281/zenodo.21105027
Document ID HEL-RWP-FRAMEWORK-v1.0 Version 1.0 (initial release) Status Active Document type Real-World Performance Study Master Methodology (Living Document) Authoring entity Helena Bioinformatics Effective date April 2026 Applicable to All real-world performance characterization activities for Folklore, including ongoing Cohort 2 study and prospective cohorts. Primary audience Scientific community, peer reviewers, investors, strategic partners, and as foundational evidence for future regulatory submission. Eventual regulatory path IVDR (EU 2017/746), UKCA, IVDR-CH - anticipated submission Q3 2027 or later, subject to organizational maturity and partnership development. Current Framework supports preparatory evidence accumulation, not immediate submission. Methodology standards ACMG/AMP 2015 (Richards et al.), ClinGen SVI specifications (Pejaver 2022, Walker 2023, Abou Tayoun 2018, Biesecker & Harrison 2020). CIOMS conflict of interest disclosure. CONSORT-style reporting where applicable.
Contents
1. Framework Purpose and Positioning 2. Study Architecture 3. Performance Metrics 4. Cohort Design 5. Methodology 6. Reference Source Standards 7. Reporting Standards 8. Conflict of Interest Management 9. Path to Future Regulatory-Grade Validation 10. Document Control
### 1. Framework Purpose and Positioning
#### 1.1. Purpose
This document defines how Helena Bioinformatics conducts Real-World Performance Studies of the Folklore variant classification platform. It covers study design, ground truth construction, performance metrics, quality controls, conflict-of-interest management, and reporting standards for Folklore's current organizational stage.
This Framework replaces the archived Folklore Clinical Validation Framework v1.0 and v1.1. It limits current work to real-world performance characterization at Folklore's present organizational stage; regulatory-validation studies require mature quality infrastructure and are reserved for later phases.
#### 1.2. Why Reframing
The prior framework treated a single peer reference laboratory as ground truth and structured study activities around regulatory validation conventions. Three observations during early use motivated reframing:
Inter-laboratory concordance studies (Amendola 2016, Harrison 2017, Bergquist 2025) document 25 to 35 percent methodological disagreement between qualified laboratories. A single peer laboratory therefore cannot serve as definitive ground truth.
The Folklore classifier applies ClinGen SVI 2020+ specifications that may differ from peer laboratories using the historical 2015 ACMG framework. Comparator differences must be characterized against the criteria and evidence used by each system before they are attributed to a classifier defect.
Helena Bioinformatics operates in Phase 1 without immediate Notified Body engagement. Current studies support scientific performance claims; regulatory dossiers are reserved for a later organizational phase.
Near-term work is classified as Real-World Performance Studies and may be submitted for scientific peer review. Formal regulatory submission remains a later-phase activity.
#### 1.3. Helena Bioinformatics Current Position
Helena Bioinformatics currently operates within Phase 1 of an explicitly phased pathway (Section 9.1). Phase 1 activities reflect the organizational scope appropriate to this stage. The platform processes clinical genomic data through partnership arrangements with established peer reference laboratory partners.
The current stage admits the following considerations, which this Framework explicitly accommodates:
Phase 1 organizational scope - activities reliant on external collaboration for clinical and reference functions where appropriate.
Partnership-based patient data access - VCF files originating from collaborating peer reference laboratories under appropriate agreements.
Phased pathway to Notified Body engagement - IVDR submission anticipated Q3 2027 or later, conditional on funding and organizational growth.
Within Phase 1, this Framework defines the evidence that current studies can produce and the additional controls required before later regulatory work.
#### 1.4. Intended Audience for Performance Study Outputs
Outputs of studies conducted under this Framework are intended for the following audiences:
Scientific community: Peer-reviewed publications. Conference abstracts. Methodology demonstration.
Investors and strategic partners: Clinical credibility evidence. Capability demonstration. Due diligence material.
Potential collaborators: Platform capability evidence supporting partnership discussions with laboratories and clinical centers.
Future Notified Body (downstream): Foundational evidence retained for citation in eventual IVDR submission. Not pivotal evidence under current Framework - designated as foundational methodology and capability demonstration.
Internal Helena Bioinformatics R&D: Defect identification, classifier improvement targets, methodology refinement.
### 2. Study Architecture
#### 2.1. Three-Layer Performance Model
Performance characterization is structured as three independent and complementary layers. Each layer addresses a distinct aspect of platform behavior and produces evidence directly informative to clinical credibility.
Layer 1: Analytical Performance
Does the platform detect, annotate, and process variants with required technical performance? Evidence of correct read-in, correct annotation, repeatability, reproducibility. Output: technical performance metrics, with comparison to standard NGS pipeline benchmarks (CLSI MM09-A2).
Layer 2: Classification Performance
Does the platform classify variants in concordance with established ACMG/AMP standards, ClinGen SVI specifications, and multi-source reference evidence? Evidence of agreement with multi-source ground truth across the available evidence hierarchy. Output: classification concordance metrics with explicit ground truth source labeling. Methodological characterization of disagreements.
Layer 3: Clinical Performance
Does the platform produce clinically meaningful, prioritized, and interpretable results that support clinical decision-making? Evidence of diagnostic prioritization, phenotype matching, and interpretation quality. Output: clinical workflow metrics: tier placement, screening rank, AI report quality, AI faithfulness.
#### 2.2. Multi-Source Ground Truth Construction
For each variant in a Real-World Performance Study, the Framework constructs ground truth from the available evidence sources and weights them by evidence strength.
VCEP-curated classifications
Multi-expert consensus through ClinGen Variant Curation Expert Panels. ClinVar 3-star or 4-star entries.
Weighting: Highest weight. Where available, treated as definitive ground truth.
Functional studies with clinical follow-up
Published experimental evidence demonstrating variant effect, supplemented by reported clinical phenotype.
Weighting: High weight. Particularly for missense variants in genes with established functional assays.
ClinVar aggregated submissions (2-star)
Multiple submitters, criteria provided, no conflicts. Aggregated independent classifications.
Weighting: Moderate weight. Useful for triangulation.
Population genetics evidence
gnomAD allele frequencies, constraint metrics (pLI, oe LoF, LOEUF, missense Z), regional constraint, pext expression.
Weighting: Moderate weight. Strong for ruling out pathogenicity (high frequency or LoF tolerance) and supporting (LoF intolerance for null variants).
Computational predictors (calibrated)
BayesDel noAF, AlphaMissense, REVEL with ClinGen SVI 2022 calibrated thresholds for evidence strength.
Weighting: Supporting weight. Used per ClinGen Pejaver 2022 calibration thresholds.
Published case reports and segregation
Peer-reviewed literature reporting variant in affected individuals, family segregation evidence.
Weighting: Variable weight depending on case quality, segregation strength, and publication rigor.
Disease-specific consortium databases
LOVD, gene-specific locus databases, registry data.
Weighting: Moderate weight. Domain-specific value where available.
Single peer reference laboratory
Single-laboratory clinical classification.
Weighting: Comparator only, not ground truth. Inter-laboratory disagreement of 25 to 35 percent is expected per published benchmarks; disagreement does not establish error.
#### 2.3. Concordance Levels
Comparisons between Folklore classification and any reference source are recorded at the following concordance levels:
FULL
Identical ACMG classification (P=P, LP=LP, VUS=VUS, LB=LB, B=B).
Direct agreement on classification.
CLINICAL
P/LP boundary, or LB/B boundary. Same clinical management implication.
Clinically equivalent. Per ClinGen SVI guidance, P and LP do not differ in clinical management.
PARTIAL
Adjacent categories that differ in clinical management. P/LP versus VUS, LB versus VUS.
Methodological divergence requiring characterization. Not automatic platform error.
DISCORDANT
Opposite-direction classifications. P/LP versus B/LB.
Significant divergence requiring detailed investigation, regardless of source.
#### 2.4. Characterization, Not Adjudication
For a PARTIAL or DISCORDANT result, the Framework records the following methodological characteristics:
Which ACMG criteria each platform applied and which were withheld.
Whether Folklore choice reflects ClinGen SVI 2020+ specifications (deprecation of PP5, PS4 single-case caution, calibrated computational thresholds, etc.).
Whether independent multi-source evidence supports one classification over another.
Whether the divergence falls within published inter-laboratory disagreement benchmarks (60 to 75 percent FULL concordance per Amendola 2016, Harrison 2017, Bergquist 2025).
Formal adjudication by external expert panel is reserved for Pivotal Cohorts in future organizational stages (see Section 9.2). Real-World Performance Studies in Phase 1 (current Folklore stage) employ methodological characterization, supplemented by optional academic peer review of the final report.
### 3. Performance Metrics
Performance is measured across three metric families corresponding to the three architectural layers. Metrics in this Framework are descriptive and benchmarked against published literature, not gated against pre-set thresholds. Pass/fail thresholds are reserved for future Pivotal Cohorts under expanded Framework versions.
#### 3.1. Family A - Analytical Performance
Detection Sensitivity
Fraction of variants present in input VCF correctly detected and presented in classified output.
Benchmark: Compared to CLSI MM09-A2 benchmarks for diagnostic NGS pipelines (typical at least 99 percent).
Annotation Concordance
Fraction of variants where Folklore annotation (gene, consequence, transcript) matches authoritative reference (Ensembl, MANE Select).
Benchmark: Compared to Ensembl VEP standard (typical at least 99 percent).
Repeatability
Identical classification when same VCF processed twice on same classifier version.
Benchmark: Deterministic classifier expectation: 100 percent.
Reproducibility
Identical classification on different infrastructure with same classifier version.
Benchmark: Deterministic classifier expectation: 100 percent.
Upstream Pipeline Failure Rate
Fraction of cases where target variant is absent from input VCF, attributable to upstream sequencing or variant calling pipeline.
Benchmark: Reported separately. Excluded from Folklore Detection Sensitivity calculation.
#### 3.2. Family B - Classification Performance
FULL Concordance with VCEP Reference
Identical ACMG class with VCEP-curated classification (where available).
Benchmark: Higher expectation given expert consensus reference. Reported per available cases.
FULL Concordance with Peer Reference Laboratory
Identical ACMG class with peer laboratory comparator.
Benchmark: Compared to published inter-laboratory benchmarks: Amendola 2016 (66 percent), Harrison 2017 (72 to 76 percent), Bergquist 2025 (70 to 75 percent). Folklore results reported in this published context.
FULL+CLINICAL Combined
Identical or P/LP boundary classification.
Benchmark: Reported as clinical-equivalent agreement.
PARTIAL Rate with Methodological Characterization
Fraction of cases at PARTIAL with documented root cause analysis.
Benchmark: Reported with full root cause table. Target: 100 percent characterization coverage.
DISCORDANT Rate
Fraction of cases at DISCORDANT.
Benchmark: Reported with detailed investigation. Target: low, but reported transparently regardless of value.
ClinGen SVI 2020+ Stricter Application Rate
Fraction of PARTIAL cases attributable to Folklore application of ClinGen SVI specifications more recent than peer laboratory reference framework.
Benchmark: Documented as methodological divergence reflecting field evolution, not platform defect.
Multi-Source Agreement Rate
For PARTIAL/DISCORDANT cases: fraction where multi-source evidence (VCEP, ClinVar consensus, functional, population) supports Folklore classification.
Benchmark: Reported as evidence triangulation outcome.
#### 3.3. Family C - Clinical Performance
Q1 - Tier Placement
Fraction of target P/LP variants placed in Tier 1 of Phenotype Matching output.
Benchmark: Reported as PASS (Tier 1), ACCEPTABLE (Tier 2), FAIL (Tier 3 or 4).
Q2 - Phenotype Match Score
Fraction of cases where target gene phenotype match score is at least 50 percent.
Benchmark: Reported as HIGH (at least 50 percent), MODERATE (20 to 49 percent), LOW (less than 20 percent).
Q3 - Screening Rank
Ordinal rank of target variant in Screening service output.
Benchmark: Reported as TOP-5, TOP-10, TOP-20, OUTSIDE-20.
Q4 - AI Report Quality
Quality rating of AI-generated clinical report.
Benchmark: Reported as GOOD, ACCEPTABLE, POOR per Q4a (variant mentioned), Q4b (interpretation correct), Q4c (overall quality).
AI Faithfulness Rate
Fraction of AI reports free of factual hallucinations.
Benchmark: Target: 100 percent. Any hallucination triggers separate AI Service investigation Change Request.
Diagnostic Yield (cohort-level)
Fraction of cases with at least one P/LP variant addressing primary clinical indication.
Benchmark: Compared to published WGS yield (Lionel 2018: 41 percent; Marshall 2020: 43 percent).
#### 3.4. Reporting Disposition Logic
Real-World Performance Studies do not produce binary pass/fail dispositions. Instead, study outputs report:
Performance values across all metric families with reference to published benchmarks.
Methodological characterization of all PARTIAL and DISCORDANT cases.
Identified platform defects with linked Change Requests and remediation status.
Limitations and caveats explicitly acknowledged.
Conclusions appropriate to study scope: descriptive performance characterization, not regulatory clearance.
Future Pivotal Cohorts under expanded Framework versions will introduce binary pass/fail thresholds. Such thresholds are not applicable at the current organizational stage.
### 4. Cohort Design
#### 4.1. Study Type Designations
Each cohort conducted under this Framework is designated, prior to initiation, as one of three study types. Designation determines applicable study design, sample size, and intended use of evidence.
Real-World Performance Study
Purpose: Characterize Folklore behavior on real clinical cases. Document inter-laboratory concordance and methodology divergence. Identify defects.
Typical size: 10 to 30 cases
Intended output: Performance characterization report. Peer-reviewable scientific document. Foundational evidence for downstream regulatory work.
Methodology Validation Study
Purpose: Statistical validation against established benchmark datasets (ClinVar 3-star+, GIAB, VCEP curated sets). Demonstrate classifier accuracy at scale.
Typical size: Hundreds to thousands of cases (database-derived)
Intended output: Methodology validation paper. Statistical performance characterization. Suitable for journal submission (JMG, Genet Med, AJHG, etc.).
Pivotal Cohort Study (future)
Purpose: Formal validation evidence for IVDR/UKCA conformity assessment. Multi-site, multi-adjudicator, full bias controls.
Typical size: 30 to 50+ cases minimum
Intended output: Notified Body submission evidence. Reserved for Phase 2/3 organizational stage.
#### 4.2. Cohort Selection Criteria for Real-World Performance Studies
Cohort selection must be documented and reproducible. Selection criteria are pre-specified in the Cohort Design Document before any case is processed. Selection cannot be performance-aware.
For Real-World Performance Studies (current Folklore stage), the following selection guidance applies:
Pre-registration: Cohort composition and selection criteria documented in writing with date stamp before processing begins.
Diversity coverage: Each cohort spans multiple clinical domains (minimum 5 distinct domains), inheritance patterns (AD, AR, X-linked, Y-linked, mitochondrial where applicable), and variant types (SNV, indel, splice, structural where in scope).
Source diversity (where feasible): Where multiple source laboratories or sources can contribute, cases are selected to represent multiple sources rather than concentrating on a single source.
Multi-source ground truth feasibility: Selection should preferentially include variants where multi-source evidence is achievable (VCEP-curated cases, ClinVar 3-star+ cases, well-characterized variants in published literature).
Demographic balance: Sex ratio between 30:70 and 70:30 where data permits. Age range spanning pediatric and adult cases where data permits.
Composition flexibility: Real-World Performance Studies may operate at flexible composition (positive cases, negative controls, edge cases in proportions tailored to study question). Strict 60/30/10 composition is reserved for future Pivotal Cohorts.
#### 4.3. Sample Size Guidance
Sample size for Real-World Performance Studies is selected to support the study descriptive purpose and the feasibility of execution under current organizational stage:
Real-World Performance Study (initial)
Minimum n: 10
Recommended n: 15-20
Rationale: Sufficient to identify methodology gaps, document inter-laboratory concordance trend, and identify platform defects. Statistical power is descriptive, not inferential.
Real-World Performance Study (expanded)
Minimum n: 20
Recommended n: 25-30
Rationale: Wilson 95% CI at 70% concordance: ±18 pp at n=25. Sufficient for trend characterization and pre-publication preparation.
Methodology Validation Study
Minimum n: Hundreds-thousands
Recommended n: Thousands+
Rationale: Database-derived. Statistical power for inferential claims.
Pivotal Cohort (future)
Minimum n: 30
Recommended n: 50+
Rationale: Wilson 95% CI narrows to ±13 pp at n=50. Required for IVDR Class C pivotal evidence.
### 5. Methodology
#### 5.1. Detection Assessment (Layer 1)
Detection assessment evaluates whether Folklore correctly identifies variants present in input VCF.
Input VCF processed through Folklore pipeline using production classifier version under study.
For each target variant, the classified_variants DuckDB queried by genomic position (GRCh38), HGVS coding change, protein change, and ClinVar identifier. Detection confirmed if any query returns a match.
For variants not detected, root cause determined by examining: input VCF presence, liftover output, quality filter status, annotation pipeline logs.
Detection failures categorized: (a) Folklore pipeline failure - counts against Detection Sensitivity; (b) upstream pipeline failure - reported separately as upstream failure rate; (c) quality filter failure - recorded as quality-flagged detection.
#### 5.2. Classification Assessment (Layer 2)
Classification assessment evaluates Folklore ACMG class output against multi-source ground truth.
For each target variant, all available reference sources are documented (peer laboratory class, ClinVar entry, VCEP curation if any, functional studies in literature, population evidence).
Folklore ACMG class extracted from classified_variants database for the target variant.
Concordance computed independently against each reference source per Section 2.3 levels.
For PARTIAL and DISCORDANT cases (Folklore vs any reference), ACMG criteria differences documented in detail. Methodology root cause categorized: (a) ClinGen SVI 2020+ stricter application by Folklore, (b) evidence access difference, (c) classifier defect, (d) reference data limitation.
Cases categorized as classifier defect escalated as Change Requests with full traceability.
Multi-source agreement determined: where Folklore disagrees with peer laboratory but agrees with multi-source evidence (VCEP, ClinVar consensus, functional), this is documented as Folklore-aligned-with-broader-evidence.
#### 5.3. Clinical Performance Assessment (Layer 3)
Clinical Performance assessment evaluates platform behavior on clinical workflow dimensions. Conducted by clinical reviewer (with COI disclosure per Section 8 if applicable).
Q1 (Tier Placement): Reviewer records Phenotype Matching tier assignment for target variant. Rated PASS (Tier 1), ACCEPTABLE (Tier 2), FAIL (Tier 3 or 4).
Q2 (Phenotype Match Score): Reviewer records percentage match score. Rated HIGH (at least 50 percent), MODERATE (20 to 49 percent), LOW (less than 20 percent).
Q3 (Screening Rank): Reviewer records ordinal position of target variant in screening service output. Rated TOP-5, TOP-10, TOP-20, OUTSIDE-20.
Q4 (AI Report Quality): Reviewer reads AI-generated clinical report. Rated Q4a (variant mentioned: YES/NO), Q4b (interpretation correct: YES/PARTIAL/NO), Q4c (overall quality: GOOD/ACCEPTABLE/POOR). Any factual hallucination triggers AI Faithfulness flag.
#### 5.4. Methodological Characterization of Divergences
For all PARTIAL and DISCORDANT cases, the following characterization is mandatory and constitutes the substantive scientific output of the study:
Side-by-side ACMG criteria table (Folklore vs each reference source).
Identification of which Folklore-applied or Folklore-withheld criteria reflect ClinGen SVI 2020+ specifications (PP5 deprecation, PS4 single-case caution, calibrated computational thresholds, BS2 inheritance-aware thresholds, PVS1 pext/expression evidence requirement, etc.).
Independent multi-source evidence summary (VCEP, ClinVar consensus, functional, population, segregation).
Statement on whether disagreement falls within published inter-laboratory disagreement benchmarks.
Disposition: methodological divergence (Folklore and reference apply different but defensible methodologies), classifier defect (Folklore error to be remediated), reference limitation (Folklore correct, reference applies less stringent standard), or evidence-genuinely-uncertain (current evidence does not unambiguously support one classification).
### 6. Reference Source Standards
#### 6.1. VCEP-Curated Variants
Highest-strength reference source. Sourcing procedure:
ClinVar database queried for variants with review_status of reviewed by expert panel (3-star) or practice guideline (4-star).
Variants filtered by gene of interest and disease context.
Selected variants documented with sourcing date and ClinVar Variation ID.
Periodic re-verification: At each cohort initiation, all included VCEP variants re-verified for current 3-star+ status. ClinVar VCEP classifications can change over time.
#### 6.2. ClinVar Aggregated Submissions
Useful supplementary reference. Used at 2-star (multiple submitters, criteria provided, no conflicts) where 3-star+ unavailable.
Disagreement among ClinVar submitters (conflicting interpretations) is recorded and is not used as definitive ground truth.
Star rating preserved in all reports for transparency.
#### 6.3. Peer Reference Laboratory
In Real-World Performance Studies under this Framework, the peer reference laboratory is the comparator. The relationship has the following structure:
Peer laboratory classifications are designated as peer reference comparator in all study documentation. The term reference standard is reserved for VCEP and adjudicated multi-expert sources.
Inter-laboratory concordance is reported as an independent metric and benchmarked against published multi-laboratory studies (Amendola 2016, Harrison 2017, Bergquist 2025).
Disagreements between Folklore and the peer laboratory are characterized methodologically (Section 5.4). Where independent multi-source evidence supports Folklore classification, this is documented; where multi-source evidence supports the peer laboratory, this is also documented.
Where peer laboratory personnel serve in Folklore study roles (e.g., as case operator), conflicts of interest are managed under Section 8 with appropriate disclosure and methodological independence safeguards.
Communication with the peer laboratory regarding individual case disagreements remains the responsibility of Helena Bioinformatics. Methodological characterization documents may be shared with the peer laboratory for their own quality program.
#### 6.4. Functional and Population Evidence
Triangulation evidence sources used in characterization but not as standalone ground truth:
Functional studies: PMID-cited published functional assay results. Used in PVS1, PS3, BS3 evidence assessment.
Population frequency: gnomAD v4.1 allele frequencies, popmax frequencies. Used in PM2, BA1, BS1 evidence assessment.
Constraint metrics: pLI, oe LoF, LOEUF, missense Z, regional pext. Used in PVS1 evidence strength modulation.
Segregation: published or available family segregation data. Used in PP1 evidence assessment when source data is reliable.
### 7. Reporting Standards
#### 7.1. Per-Case Performance Report
Every case in every Real-World Performance Study produces a Per-Case Performance Report following the structure below. This structure is consistent with the Per-Case Concordance Reports already produced for Cohort 1 and Cohort 2; section additions reflect this Framework emphasis on multi-source characterization.
Case Header and Performance Summary: Patient ID (anonymized), gene, variant identifiers, clinical diagnosis, HPO terms, session ID, classifier version, report date. Performance summary statement and concordance outcomes against each reference source.
Multi-Source Comparison Table: Side-by-side: Folklore classification vs each available reference (peer laboratory, ClinVar with star rating, VCEP if any, published literature). ACMG criteria, HGVS, consequence, zygosity, gnomAD AF documented.
Evidence Review: ClinVar entries with review_status detail, in silico predictions with calibrated thresholds, population frequency analysis, gene constraint metrics, sequencing quality metrics, published case reports if any, functional evidence if any.
Layer 1 Assessment (Detection): P1, P2 results. Detection failure root cause if applicable.
Layer 2 Assessment (Classification): P3 result against each reference source. Methodological characterization for any PARTIAL/DISCORDANT outcome. Specific identification of ClinGen SVI 2020+ application differences.
Layer 3 Assessment (Clinical Performance): Q1, Q2, Q3, Q4 results with rationale. AI Faithfulness check explicitly noted.
Diagnostic Relevance Assessment: Comparative clinical utility analysis: Folklore versus reference. Documented per intended use.
Platform Issues Identified: Numbered list of any defects, limitations, or improvement opportunities. Each linked to existing or proposed Change Request.
COI Statement: Disclosure of any conflict of interest applicable to operator, reviewer, or other contributor to this case. Mitigation applied.
Conclusions and Methodological Disposition: Final disposition: methodological divergence / classifier defect / reference limitation / evidence-genuinely-uncertain. Action items if any. Responsible authority sign-off.
#### 7.2. Cohort-Level Real-World Performance Study Report
At cohort completion, Helena Bioinformatics produces a Real-World Performance Study Report. This is the primary scientific output of the study and is structured to be peer-publishable.
Title and Abstract: Cohort identifier, classifier version, key findings.
Introduction: Background on Folklore, classifier methodology (ACMG/AMP 2015 with ClinGen SVI specifications), study purpose.
Methods: Cohort selection criteria, processing pipeline, classifier version, reference sources used, comparison methodology, COI disclosure.
Results - Layer 1: Detection performance with attention to upstream pipeline failures.
Results - Layer 2: Classification concordance against each reference source. Inter-laboratory concordance with peer laboratory benchmarked against published literature. Multi-source agreement analysis.
Results - Layer 3: Clinical workflow metrics.
Discussion: Methodological characterization of PARTIAL/DISCORDANT cases. Identified defects and remediation status. Limitations. Comparison to published benchmarks.
Conclusions: Study scope, findings, limitations, and status as foundational evidence for later regulatory work.
Appendices: All Per-Case Performance Reports, classifier version manifest, reference database snapshots, COI Disclosure Registry.
#### 7.3. Identified Defects and Change Request Tracking
All identified defects, methodology gaps, and improvement opportunities are entered into a centralized Change Request log. Log entries contain: discovery date, source case(s), description, severity, assigned Change Request identifier, status, target resolution, post-remediation re-validation requirement.
#### 7.4. Optional Academic Peer Review
Real-World Performance Study Reports may be submitted to independent academic clinical geneticists for pre-publication scientific review. This review does not constitute formal regulatory adjudication. The final report records the reviewers, their conflicts of interest under Section 8, and the disposition of their feedback.
### 8. Conflict of Interest Management
#### 8.1. Position
Helena Bioinformatics maintains a transparent COI disclosure framework aligned with academic publication standards (CIOMS, ICH GCP, Helsinki Declaration) and consistent with downstream regulatory requirements (ISO 15189:2022 Section 7.2, IVDR conformity assessment expectations). Disclosure is mandatory; mitigations are tailored to study type and organizational stage.
#### 8.2. COI Categories
Family
Personal, marital, or family relationship between individual and Folklore leadership.
Default management: Disclosure mandatory. Excluded from formal adjudication. Permitted in advisory or scientific consultation roles.
Employment
Current or recent (within 24 months) employment with Helena Bioinformatics or with reference laboratories used in study.
Default management: Disclosure mandatory. Specific role restrictions in study type and cohort context.
Professional Affiliation
Past collaboration, consulting, signatory, or mentorship relationship with reference laboratory or Helena Bioinformatics, falling short of formal employment.
Default management: Disclosure mandatory. Case-by-case mitigation. Includes individuals serving as external clinical signatories on reference laboratory reports.
Financial
Equity ownership, options, royalties, advisory fees, or material financial interest in Helena Bioinformatics or in any organization with material business relationship to study outcome.
Default management: Disclosure mandatory. Does not categorically disqualify but is disclosed in all reports.
Business
Material business relationship between individual primary affiliation and Helena Bioinformatics.
Default management: Disclosure mandatory. Role restrictions may apply.
#### 8.3. Disclosure Procedure
All individuals serving in the following roles complete a COI Disclosure Form before commencing role activities:
Case operator (Real-World Performance Study)
Clinical reviewer (Layer 3 evaluation)
Academic peer reviewer (pre-publication)
Scientific Advisor
External methodology consultant
The COI Disclosure Form (HEL-COI-DISCLOSURE-FORM-v1.0) is maintained as a separate template document. Completed forms retained in Folklore document repository for the duration of engagement plus 15 years.
#### 8.4. Disclosure Registry - Current State
As of v1.0 effective date, the following roles have disclosed conflicts of interest. Disclosures and mitigations apply to all studies conducted under this Framework. Individual names are documented in the internal Disclosure Registry and disclosed in any per-case report or publication where the individual contribution is referenced.
Founder and Chief Executive Officer
Sole technical and scientific lead in current organizational stage.
Category: Financial (sole owner), Employment
Disclosure: Owns 100 percent of Helena Bioinformatics. Performs all internal Folklore roles in current organizational stage.
Mitigation: Disclosed in all Real-World Performance Study Reports. The multi-source ground truth methodology mitigates single-author classification framing. Future organizational growth will introduce role separation.
Scientific Advisor
Methodology consultation and scientific guidance.
Category: Family (relationship to founder)
Disclosure: Family relationship to Folklore founder.
Mitigation: Excluded from formal adjudication. Role limited to Scientific Advisor (methodology consultation, framework review, scientific guidance). Family relationship disclosed in any report where input is referenced.
Cohort 2 case operator
Real-World Performance Study case processing.
Category: Employment (peer reference laboratory)
Disclosure: Employed by the peer reference comparator laboratory used in Cohort 2.
Mitigation: Disclosed in Cohort 2 Real-World Performance Study Report. Operator role limited to technical processing and comparison documentation; not classification authority. Multi-source ground truth methodology mitigates single-source dependency. Future cohorts will introduce independent operator structure.
External clinical signatory
Historical signatory on peer reference laboratory clinical reports referenced in Cohort 2. Potential future Folklore Scientific Advisor or methodology consultant.
Category: Professional Affiliation
Disclosure: Has signed peer reference laboratory clinical reports (including those referenced in Cohort 2) as external clinical signatory. Not a peer reference laboratory employee.
Mitigation: Cannot serve as adjudicator on cases previously signed. May serve as Folklore Scientific Advisor or methodology consultant for non-overlapping cohorts. Disclosed in any report where input is referenced.
#### 8.5. Mitigation Strategies
For each disclosed COI, one or more mitigations apply:
Role exclusion: Individual excluded from specific roles where COI creates unmitigatable bias.
Multi-source triangulation: Single-source dependency mitigated by multi-source ground truth methodology (Section 2.2).
Transparent disclosure: COI disclosed in all relevant documentation.
Procedural separation: Individual operates under documented procedural controls limiting influence.
Role limitation: Individual restricted to roles where COI does not create methodological dependency on outcome.
Periodic review: COI status reviewed annually.
#### 8.6. Annual Review
Disclosure Registry reviewed annually. Off-cycle review triggered by:
New individual engagement requiring COI disclosure.
Material change in individual affiliations, financial interests, or family circumstances.
Material change in Helena Bioinformatics organizational structure.
External inquiry regarding COI.
### 9. Path to Future Regulatory-Grade Validation
#### 9.1. Phased Pathway
This Framework operates within Phase 1 of an explicitly phased pathway toward eventual regulatory submission. Phase progression is conditional on organizational growth and is not pursued at the expense of current activities.
Phase 1 Foundation (Current)
Organizational stage: Current Phase 1 organizational stage. No internal Quality, Regulatory, or Clinical functions; activities reliant on external collaboration where required.
Permitted activities:
- Real-World Performance Studies
- Methodology Validation Studies (database-derived)
- Scientific publications
- Foundational evidence accumulation
Phase 2 Growth
Organizational stage: Small team (3 to 10). Series A or strategic partnership. Dedicated Quality and Regulatory functions.
Permitted activities:
- Phase 1 activities continue
- First Pivotal Cohort possible
- Multi-cohort program execution
- Multi-adjudicator panel introduced
Phase 3 Pre-submission
Organizational stage: Established team (10+). Mature QMS. Notified Body engagement.
Permitted activities:
- Pivotal Cohorts at full composition
- Multi-site validation
- Notified Body conformity assessment submission
- Post-market performance follow-up planning
#### 9.2. Differences in Future Phases
Activities reserved for Phase 2 and Phase 3 (and not pursued in current Phase 1) include:
Formal external adjudication panels (multi-adjudicator, paid engagement)
Independent operator structure (no peer laboratory affiliations)
Pivotal Cohort composition (60/30/10 positive/negative/edge)
Multi-site validation across independent laboratories
Pre-submission Notified Body consultations
Risk Management File at full ISO 14971 detail
Quality Management System at full ISO 13485 detail
#### 9.3. Foundational Evidence Carryforward
The following evidence from Real-World Performance Studies can support future Pivotal submissions:
Performance characterization data is cited as preliminary evidence supporting platform readiness.
Defect and remediation records provide a history of quality-management actions.
Published methodology sources document the rationale for classifier choices.
Multi-source agreement analysis measures whether Folklore classifications agree with broader expert consensus when the peer reference differs.
These records can be cited as foundational evidence, subject to the scope and limitations of each study.
### 10. Document Control
#### 10.1. Version History
v1.0
Initial release of reframed framework. Replaces archived Folklore Clinical Validation Framework v1.0 and v1.1. Reframes scope from regulatory validation to Real-World Performance Studies appropriate to current organizational stage. Establishes multi-source ground truth methodology, methodological characterization (rather than formal adjudication), academic peer review, and explicit phased pathway to future regulatory submission.
Author: Helena Bioinformatics
#### 10.2. Change Control
Living document. Revision triggered by:
Methodology refinement during cohort execution.
Updates to applicable scientific standards (ACMG, ClinGen SVI, ClinVar review tiers).
External peer review feedback on Framework or Study Reports.
Material change in organizational stage (Phase 1 to 2 to 3 transition).
Material change in COI Disclosure Registry.
Internal periodic review (minimum annual).
Minor revisions issued as v1.x. Major revisions (e.g., transition to Pivotal Cohort framework) issued as v2.0+.
#### 10.3. Document Retention
All versions retained indefinitely as part of Helena Bioinformatics records. Minimum 15 years retention applied for downstream regulatory traceability.
#### 10.4. Approval
This Framework v1.0 is approved by Helena Bioinformatics on the date stated in Section 10.1.
Approved by Helena Bioinformatics Date of approval April 2026 Document ID HEL-RWP-FRAMEWORK-v1.0 Distribution Internal - Helena Bioinformatics. Provided to academic peer reviewers, scientific collaborators, and prospective regulatory consultants upon engagement.
End of Framework v1.0
Back to Studies Contact Us

### Links and cited sources

- [Back to Studies](https://folklore.helena.bio/studies)
- [DOI: 10.5281/zenodo.21105027](https://doi.org/10.5281/zenodo.21105027)
- [Contact Us](https://folklore.helena.bio/contact)

---

## Page: Carrier Screening Variant Interpretation | Folklore

Source: https://folklore.helena.bio/use-cases/carrier-screening
Canonical: https://folklore.helena.bio/use-cases/carrier-screening
Description: Automated ACMG classification for carrier screening panels. Population-specific frequency analysis, structured clinical reports, and transparent evidence chains for reproductive genetics.

## Carrier Screening Variant Interpretation
Expanded carrier screening panels now cover dozens to hundreds of conditions. As panel size increases, more variants require systematic evidence review. Folklore applies ACMG classification and population-aware frequency analysis across those panel sizes.
### Carrier Screening at Scale
Expanded panels increase the number of conditions, variants, and population-frequency contexts that a reviewer must assess.
#### Variant Volume at Scale
Expanded carrier screening panels test for hundreds of conditions simultaneously, producing dozens to hundreds of candidate variants per individual. Each requires systematic evidence evaluation against population databases, clinical repositories, and functional predictions.
#### Population-Specific Frequencies
Carrier frequencies vary across populations. A variant common in Ashkenazi Jewish populations may be absent in East Asian populations. Interpretation therefore uses ancestry-specific allele frequencies alongside global averages.
#### Classification Nuance
Carrier screening operates in a different clinical context than diagnostic testing. Variants classified as VUS in a diagnostic setting may still be reportable in a carrier context depending on the condition, population, and clinical guidelines adopted by the laboratory.
#### Reporting Complexity
Carrier screening results can affect reproductive counseling. Reports therefore present the classification, population data, and supporting evidence in a structured form.
### Carrier Screening Workflow
1
Upload carrier screening VCF
Standard VCF from any sequencing platform - targeted panel, exome, or genome.
2
Automated annotation and classification
Population frequencies, clinical significance, functional predictions, and ACMG criteria applied to every variant.
3
Clinical review of classified variants
The geneticist reviews Pathogenic and Likely Pathogenic findings with the applied criteria and evidence sources.
4
Generate carrier screening report
Structured report with classification, population data, and evidence attribution.
5
Genetic counseling
Report supports informed reproductive counseling with transparent, traceable findings.
### Platform Capabilities for Carrier Screening
#### Full ACMG Classification for Every Variant
Carrier screening variants receive ACMG/AMP classification using 19 automated criteria, a Bayesian point framework, BayesDel ClinGen SVI-calibrated thresholds, and VCEP gene-specific specifications where available.
#### Population-Aware Frequency Analysis
gnomAD v4.1 provides global and population-specific allele frequencies across multiple ancestry groups. Folklore displays global frequency and population-maximum frequency (popmax) for each variant so the geneticist can assess it against the patient's reported ancestry.
#### Founder Variant Recognition
Established founder variants - including Ashkenazi Jewish founder mutations in HEXA (Tay-Sachs), BRCA1/2, CFTR, and others - are annotated through ClinVar integration with clinical significance assertions and review star quality levels. Disease associations from OMIM and ClinGen provide additional clinical context.
#### Gene-Specific VCEP Thresholds
For approximately 50-60 genes with published ClinGen VCEP specifications, gene-specific frequency thresholds and criterion modifications are applied automatically. This includes clinically important carrier screening genes such as CFTR, PAH, and BRCA1/BRCA2 (ENIGMA specifications).
#### Structured Reporting
Clinical reports present findings in structured sections with ACMG classification, evidence summary, population frequency data, and literature references. Reports are generated in PDF and DOCX formats suitable for inclusion in patient records and genetic counseling documentation.
#### Transparent Evidence Chains
Each classification records the applied ACMG criteria, database versions, and computational thresholds for review and reporting.
### See Carrier Screening Classification
Review carrier-screening classification with ACMG evidence and population-specific frequency data.
Contact Us

### Links and cited sources

- [Contact Us](https://folklore.helena.bio/contact)

---

## Page: Neonatal Genomic Screening Software | Folklore

Source: https://folklore.helena.bio/use-cases/newborn-screening
Canonical: https://folklore.helena.bio/use-cases/newborn-screening
Description: Age-aware variant prioritization for neonatal and pediatric genomic screening. Curated neonatal gene lists, time-critical interpretation, and phenotype-agnostic whole-genome analysis.

## Neonatal Genomic Screening
Genomic newborn screening can extend targeted biochemical assays with whole-genome analysis. Review then begins with millions of variants, limited phenotype information, and time-sensitive clinical questions.
Under 60 minutes
Full genome interpretation
No phenotype required
Pathogenicity-first analysis
### The Neonatal Interpretation Challenge
Neonatal intensive care combines short decision windows, incomplete phenotypes, and a broad diagnostic scope.
#### Time-Critical Decisions
In neonatal intensive care, variant review may need to inform time-sensitive decisions for presentations such as unexplained seizures or metabolic crisis.
#### Non-Specific Presentations
Newborns present with non-specific symptoms - respiratory distress, hypotonia, seizures, metabolic acidosis. These presentations overlap across hundreds of genetic conditions, making phenotype-guided gene panel selection unreliable.
#### Evolving Clinical Picture
A newborn's phenotype changes rapidly in the first days and weeks of life. Features that could narrow the differential diagnosis may not yet be apparent. Interpretation must work with incomplete and evolving clinical information.
#### Genome-Wide Scope
Targeted gene panels risk missing diagnoses outside their scope. For critically ill neonates, whole genome or exome analysis is increasingly the first-line approach - but it produces millions of variants that need systematic evaluation.
### Phenotype-Agnostic Analysis
When a neonatal phenotype is incomplete, Folklore first ranks variants by pathogenicity evidence and then adds phenotype relevance as clinical information becomes available.
#### No Phenotype Required to Start
Folklore performs genome-wide pathogenicity-first prioritization. Every variant is classified and scored based on population frequency, functional predictions, gene constraint, and clinical databases - independent of phenotypic input. When HPO terms are available, phenotype matching adds a correlation layer. When they are not yet available (common in NICU), the system still produces a clinically useful prioritized variant list.
#### Phenotype Matching as Evolving Layer
As the newborn's clinical picture develops, HPO terms can be updated and the case reprocessed. New phenotype information immediately re-ranks variants based on genotype-phenotype correlation, surfacing candidates that were classified but not initially prioritized. The underlying classification does not change - only the clinical prioritization adapts.
#### Unexpected Diagnoses Surface
Phenotype-agnostic review retains pathogenic and likely pathogenic variants in genes outside the initial differential diagnosis. This matters when the neonatal presentation is broad or involves several systems.
### Age-Aware Screening Profiles
Variant prioritization adapts to the clinical context. A newborn in the NICU and a child in a developmental genetics clinic have different urgent gene lists and different clinical priorities.
#### Neonatal
0 - 28 days
Highest priority for conditions with neonatal onset and available treatment. Gene weighting emphasizes disorders where early intervention changes outcomes - metabolic conditions (PKU, galactosemia, MCAD deficiency), cystic fibrosis, spinal muscular atrophy, congenital adrenal hyperplasia, and early-onset epilepsies.
Key genes: CFTR, SMN1, GAA, GBA, HEXA, PAH, GALT, ACADM, CYP21A2, and curated neonatal-onset gene lists
#### Pediatric
29 days - 18 years
Broader scope including childhood-onset conditions. Weight profiles shift toward developmental disorders, childhood-onset metabolic conditions, immunodeficiencies, and genetic epilepsies. Lower urgency than neonatal but still requires efficient evidence gathering for timely clinical decisions.
Key genes: Expanded panels including developmental delay, intellectual disability, epilepsy, and immunodeficiency gene lists
#### Proactive
Any age
Population-scale screening for actionable genetic conditions regardless of current clinical presentation. Focuses on conditions where early knowledge enables preventive action - hereditary cancer syndromes, cardiac conditions (LQTS, HCM), pharmacogenomic variants, and carrier status.
Key genes: ACMG SF v3.2 secondary findings list, pharmacogenomic panels, carrier screening panels
Full documentation on age-aware prioritization
### Platform Capabilities
#### Processing Speed
Gene panel VCF: 1-2 minutes. Whole exome: 2-5 minutes. Whole genome: under 15 minutes. Classification, annotation, and phenotype matching run in parallel across the full variant set.
#### Neonatal Gene Curation
Age-aware scoring profiles weight neonatal-onset conditions with available treatment higher than adult-onset conditions. Gene lists are curated from OMIM, GeneReviews, and newborn screening program publications.
#### ACMG Classification
Each variant receives ACMG/AMP classification under the Bayesian point framework. The geneticist can inspect the evidence recorded for each finding.
#### Tiered Clinical Output
Results are presented in clinical tiers: Tier 1 for findings requiring immediate review, Tier 2 for findings requiring further evaluation, and a separate incidental-findings group.
#### Reprocessing Without Re-Upload
As clinical information evolves, cases can be reprocessed with updated HPO terms or different screening profiles without re-uploading the VCF. Updated reference databases are applied automatically, and previous results are preserved for comparison.
#### GDPR-Native EU Infrastructure
Neonatal genomic data is among the most sensitive categories of personal data. All processing occurs on dedicated infrastructure in Helsinki, Finland. No variant data leaves the EU. No external API calls are made during analysis.
### See Neonatal Screening in Action
Trace a neonatal case from whole-genome VCF to age-aware clinical prioritization.
Contact Us

### Links and cited sources

- [Full documentation on age-aware prioritization](https://folklore.helena.bio/docs/screening/age-aware-prioritization)
- [Contact Us](https://folklore.helena.bio/contact)

---

## Page: Rare Disease Genome Interpretation and HPO Matching | Folklore

Source: https://folklore.helena.bio/use-cases/rare-disease
Canonical: https://folklore.helena.bio/use-cases/rare-disease
Description: Rare-disease genome review with ACMG classification, HPO semantic similarity, literature retrieval, variant prioritization, and structured evidence for geneticist review.

## Rare Disease Genome Interpretation
Folklore classifies candidate variants, compares their genes with the patient's HPO terms, retrieves linked literature, and ranks the resulting evidence for geneticist review.
5-7 years
Average time to rare disease diagnosis
3-5
Specialists consulted before diagnosis
41%
Patients receiving at least one VUS
7.3%
VUS that are ever reclassified
### The Interpretation Bottleneck
Rare-disease sequencing produces more candidate variants than a reviewer can assess through ad hoc database and literature searches.
#### Variant Volume
Whole exome sequencing produces 20,000-30,000 variants per patient. Whole genome sequencing produces 4-5 million. Each pathogenic candidate requires cross-referencing multiple databases, literature sources, and phenotype associations.
#### Manual Evidence Gathering
For each candidate variant, a geneticist must search ClinVar, gnomAD, PubMed, OMIM, and functional prediction tools individually. A single rare disease case with dozens of candidate variants can consume 5-10 days of expert time.
#### VUS Accumulation
Rare-disease cases often contain Variants of Uncertain Significance that require additional population, functional, phenotype, and literature evidence before reclassification.
#### Phenotype Complexity
Rare diseases often present with overlapping phenotypes, incomplete penetrance, and variable expressivity. Connecting a patient's specific clinical presentation to the correct gene-disease association requires structured phenotype matching, not keyword searches.
### The Rare-Disease Review Path
Six stages connect variant annotation, ACMG classification, phenotype matching, literature retrieval, evidence synthesis, and report preparation.
1
#### Variant Annotation
Variants are annotated against eight reference sources: gnomAD population frequencies, ClinVar clinical significance, dbNSFP functional predictions, SpliceAI splice impact, gnomAD gene constraint, HPO phenotype associations, ClinGen dosage sensitivity, and Ensembl VEP consequence prediction.
2
#### Phenotype-First Prioritization
Patient HPO terms are matched against gene-disease phenotype profiles using semantic similarity analysis that accounts for ontology hierarchy and information content. A variant in a gene associated with the patient's specific phenotype is prioritized over an equally classified variant in an unrelated gene. This is not keyword matching - it is structured ontological reasoning.
3
#### Bayesian ACMG Classification
The pipeline evaluates 19 automatable ACMG criteria using the Tavtigian Bayesian point framework and BayesDel ClinGen SVI-calibrated thresholds. VCEP gene-specific specifications apply for approximately 50-60 genes, and the output includes a classification confidence score.
4
#### Automated Literature Evidence
For every candidate gene and variant, Folklore searches a local database of millions of genetics-relevant PubMed publications with pre-extracted gene mentions, variant mentions, and phenotype associations. Evidence is ranked by clinical relevance and returned with full PMID attribution.
5
#### Clinical Evidence Synthesis
An on-premise AI assistant summarizes classification results, phenotype correlations, and linked literature in a structured narrative for geneticist review.
6
#### Structured Clinical Report
A tiered report presents Tier 1 and Tier 2 variants with the ACMG classification, phenotype match score, supporting literature, population frequency, and computational predictions for clinical review.
### Rare Disease Analysis Controls
Rare-disease review combines variant classification with phenotype matching, literature search, and evidence traceability.
#### Phenotype Matching That Understands Ontology
HPO semantic similarity uses information content and ontological hierarchy to identify gene-phenotype relationships that exact keyword matching would miss. "Seizures" and "Epilepsy" are recognized as related, not treated as different terms.
How semantic similarity works
#### VUS Evidence Aggregation
Each VUS is presented with population frequencies, functional predictions, literature citations, and phenotype correlations to support later review or reclassification.
Understanding confidence scores
#### Whole Genome Processing
Folklore processes whole-genome VCF files containing 4-5 million variants in under an hour. The pipeline does not pre-filter variants by frequency or predicted impact before classification.
See the full pipeline
#### Transparent Evidence Chains
Each classification records the applied ACMG criteria, database versions, computational thresholds, and evidence strength for reviewer inspection.
ACMG classification methodology
### The Geneticist Remains Central
Rare-disease diagnosis requires a geneticist to interpret incomplete phenotypes, recognize atypical presentations, integrate family history, and communicate uncertainty to patients.
Before that review, Folklore queries the configured databases, searches literature, matches the phenotype, and applies the ACMG criteria. The geneticist receives the resulting evidence and retains the clinical decision.
Folklore prepares the evidence. The geneticist makes the clinical decision.
### Review a Rare-Disease Case Workflow
See a rare-disease genome move from VCF upload through phenotype matching to a report prepared for review.
Contact Us

### Links and cited sources

- [How semantic similarity works](https://folklore.helena.bio/docs/phenotype-matching/semantic-similarity)
- [Understanding confidence scores](https://folklore.helena.bio/docs/classification/confidence-scores)
- [See the full pipeline](https://folklore.helena.bio/how-it-works)
- [ACMG classification methodology](https://folklore.helena.bio/docs/classification/acmg-framework)
- [Contact Us](https://folklore.helena.bio/contact)

---

## Page: Folklore Documentation

Source: https://folklore.helena.bio/docs
Canonical: https://folklore.helena.bio/docs
Description: Complete documentation for Folklore clinical genetics platform. Getting started, classification methodology, reference databases, and clinical workflows.

## Documentation
Documentation for the Folklore clinical genetics platform, including clinically relevant rules, thresholds, data sources, workflows, and limitations. This documentation is intended for clinical geneticists, laboratory directors, accreditation auditors, and bioinformaticians.
For the full classification methodology with all criteria thresholds and combining rules, see the dedicated Methodology page.
Getting Started
Upload your first VCF file, set patient phenotype, and understand your results.
Variant Analysis
ACMG/AMP framework, 28 evidence criteria, combining rules, and ClinVar integration.
Family Analysis
Complete-trio joint genotyping, relationship QC, de novo candidates, compound-heterozygous phasing, and segregation limits.
Cohort Analysis
Build Studies, compare quality across samples, and interpret cohort-level variant, gene, and carrier evidence.
Structural Variants and CNVs
Structural and copy-number variant input, evidence, framework routing, interpretation, and limitations.
Mitochondrial DNA
mtDNA classification under MMDWG 2020, including heteroplasmy, haplogroups, predictors, and NUMT flags.
Computational Predictors
BayesDel, SpliceAI, SIFT, AlphaMissense, and conservation scores used in PP3/BP4.
Reference Databases
gnomAD, ClinVar, dbNSFP, HPO, ClinGen, and Ensembl VEP. Versions and update policy.
Phenotype Matching
HPO-based semantic similarity, clinical tier assignment, and score interpretation.
Screening
Multi-dimensional variant prioritization, tier system, and clinical screening modes.
Literature Evidence
Local PubMed retrieval, case-aware relevance ranking, evidence labels, coverage, and clinical review boundaries.
AI Clinical Assistant
Natural language queries, clinical interpretation, and report generation.
Data and Privacy
EU infrastructure, GDPR compliance, data retention, and zero external API calls.
Limitations
What the platform cannot do. Honest documentation for clinical trust.
Glossary
Quick reference for genetics and platform terminology.
FAQ
Common questions from laboratory directors and clinical geneticists.
Changelog
Versioned history of all methodology and database changes.
Documentation Principles
This documentation explains what Folklore does, the evidence shown to users, and the boundaries that affect clinical interpretation. Clinically relevant thresholds and data provenance are documented where they support review and reproducibility; proprietary implementation details are not published. Every limitation is acknowledged. The platform is a clinical decision support tool -- the reviewing geneticist always has the final word.

### Links and cited sources

- [Methodology](https://folklore.helena.bio/methodology)
- [Getting Started Upload your first VCF file, set patient phenotype, and understand your results.](https://folklore.helena.bio/docs/getting-started)
- [Variant Analysis ACMG/AMP framework, 28 evidence criteria, combining rules, and ClinVar integration.](https://folklore.helena.bio/docs/classification)
- [Family Analysis Complete-trio joint genotyping, relationship QC, de novo candidates, compound-heterozygous phasing, and segregation limits.](https://folklore.helena.bio/docs/family-analysis)
- [Cohort Analysis Build Studies, compare quality across samples, and interpret cohort-level variant, gene, and carrier evidence.](https://folklore.helena.bio/docs/cohort-analysis)
- [Structural Variants and CNVs Structural and copy-number variant input, evidence, framework routing, interpretation, and limitations.](https://folklore.helena.bio/docs/structural-variants)
- [Mitochondrial DNA mtDNA classification under MMDWG 2020, including heteroplasmy, haplogroups, predictors, and NUMT flags.](https://folklore.helena.bio/methodology/mtdna)
- [Computational Predictors BayesDel, SpliceAI, SIFT, AlphaMissense, and conservation scores used in PP3/BP4.](https://folklore.helena.bio/docs/predictors)
- [Reference Databases gnomAD, ClinVar, dbNSFP, HPO, ClinGen, and Ensembl VEP. Versions and update policy.](https://folklore.helena.bio/docs/databases)
- [Phenotype Matching HPO-based semantic similarity, clinical tier assignment, and score interpretation.](https://folklore.helena.bio/docs/phenotype-matching)
- [Screening Multi-dimensional variant prioritization, tier system, and clinical screening modes.](https://folklore.helena.bio/docs/screening)
- [Literature Evidence Local PubMed retrieval, case-aware relevance ranking, evidence labels, coverage, and clinical review boundaries.](https://folklore.helena.bio/docs/literature)
- [AI Clinical Assistant Natural language queries, clinical interpretation, and report generation.](https://folklore.helena.bio/docs/ai-assistant)
- [Data and Privacy EU infrastructure, GDPR compliance, data retention, and zero external API calls.](https://folklore.helena.bio/docs/data-and-privacy)
- [Limitations What the platform cannot do. Honest documentation for clinical trust.](https://folklore.helena.bio/docs/limitations)
- [Glossary Quick reference for genetics and platform terminology.](https://folklore.helena.bio/docs/glossary)
- [FAQ Common questions from laboratory directors and clinical geneticists.](https://folklore.helena.bio/docs/faq)
- [Changelog Versioned history of all methodology and database changes.](https://folklore.helena.bio/docs/changelog)

---

## Page: AI Clinical Assistant | Folklore Documentation

Source: https://folklore.helena.bio/docs/ai-assistant
Canonical: https://folklore.helena.bio/docs/ai-assistant
Description: The Folklore AI Assistant queries case variants and local literature, explains recorded evidence, and drafts reviewable clinical interpretation reports.

Documentation / AI Clinical Assistant
## AI Clinical Assistant
The Folklore AI Assistant provides conversational access to the current case, recorded variant evidence, and the local literature index. It helps geneticists inspect results, compare findings with the phenotype, and prepare reviewable clinical interpretation drafts.
Model inference runs on dedicated infrastructure within the EU. Case data is not sent to external hosted language-model services during assistant use.
How It Works
1
Context Loading
When a conversation starts, the assistant loads the complete analysis context for the current case: ACMG classifications, pathogenic variants, phenotype matching results, screening tiers, clinical profile, and any previously generated interpretation.
2
Natural Language Interaction
The geneticist asks questions in natural language. The assistant can answer directly from its clinical knowledge, or invoke tools to query the patient database or search the literature database.
3
Tool Execution
When a question requires data, the assistant automatically translates it to SQL, executes it against the appropriate database, and incorporates the results into its response. Up to five sequential tool calls can be chained in a single interaction.
4
Streaming Response
Responses are streamed in real-time via Server-Sent Events. Text appears token by token, with structured events for query results, literature findings, and visualization suggestions.
Key Principles
The assistant only discusses genes and variants that are present in the patient's data. It does not invent findings or add textbook examples that are not in the actual results.
All responses are grounded in the analysis data. When the assistant references a variant, it has either seen it in the loaded context or queried it from the database.
The assistant uses a phenotype-first approach: it considers the patient's clinical presentation to identify candidate genes, then checks whether those genes appear in the variant data.
Clinical interpretation reports clearly separate AI-generated content from template-based sections (patient demographics, legal disclaimer). The AI never fabricates patient information.
Conversation history is maintained for 24 hours in an encrypted cache, enabling multi-turn discussions that build on previous questions and findings.
EU-Hosted Model Inference
The Folklore AI Assistant runs on dedicated GPU infrastructure within the EU and is accessed through an internal service path. This design keeps model inference within the controlled processing environment and supports the platform's data-protection controls.
In This Section
Capabilities
What the Folklore AI Assistant can do: variant queries, literature search, interpretation drafting, report generation, and visualization.
Asking Questions
How to interact with the AI assistant effectively -- question patterns, tool invocation, and multi-turn conversations.
Clinical Interpretation
AI-generated diagnostic reports with four interpretation levels, structured headers, and legal disclaimers.
Database Queries
Natural language to SQL translation against patient variant and biomedical literature databases.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Capabilities What the Folklore AI Assistant can do: variant queries, literature search, interpretation drafting, report generation, and visualization.](https://folklore.helena.bio/docs/ai-assistant/capabilities)
- [Asking Questions How to interact with the AI assistant effectively -- question patterns, tool invocation, and multi-turn conversations.](https://folklore.helena.bio/docs/ai-assistant/asking-questions)
- [Clinical Interpretation AI-generated diagnostic reports with four interpretation levels, structured headers, and legal disclaimers.](https://folklore.helena.bio/docs/ai-assistant/clinical-interpretation)
- [Database Queries Natural language to SQL translation against patient variant and biomedical literature databases.](https://folklore.helena.bio/docs/ai-assistant/database-queries)

---

## Page: Asking Questions | Folklore Documentation

Source: https://folklore.helena.bio/docs/ai-assistant/asking-questions
Canonical: https://folklore.helena.bio/docs/ai-assistant/asking-questions
Description: How to interact with Helix AI effectively -- question patterns, tool invocation, multi-turn conversations, and phenotype-driven analysis.

Documentation / AI Clinical Assistant / Asking Questions
## Asking Questions
Helix AI responds to natural language questions about the current case. It automatically decides whether to answer from its clinical knowledge, query the patient variant database, or search the biomedical literature. Understanding how the assistant makes these decisions helps you get better results.
When the Assistant Queries the Database
The assistant automatically queries the variant database when your question involves specific data about this patient's variants. It does not query for general genetics knowledge or conceptual explanations.
Triggers Database Query Answers Directly
"Show me pathogenic variants in BRCA1" "What does ACMG stand for?"
"How many VUS are on chromosome 7?" "Explain the PP3 criterion"
"List genes with pLI above 0.9" "What is autosomal dominant inheritance?"
"What is the gnomAD frequency of the TP53 variant?" "Why is population frequency important?"
Effective Question Patterns
Specific Filtering
"Show me missense variants in cardiac genes with gnomAD frequency below 0.01% and DANN score above 0.95"
Precise filters produce focused results. The assistant translates each condition into SQL WHERE clauses.
Aggregation
"How many variants per ACMG class are there on each chromosome?"
The assistant generates GROUP BY queries and suggests appropriate visualizations for the aggregated data.
Phenotype-Driven
"Given the patient's seizure phenotype, which genes should I focus on?"
The assistant uses its gene-phenotype knowledge to identify candidate genes (SCN1A, SCN2A, KCNQ2, etc.), then checks the variant data for those genes.
Follow-Up
"Now show me the literature evidence for that gene"
The assistant maintains conversation context. "That gene" resolves to the gene discussed in the previous response.
Clinical Correlation
"Are any of the Tier 1 phenotype matches in ACMG Secondary Findings genes?"
Cross-referencing different analysis modules helps identify clinically actionable findings.
Multi-Turn Conversations
The assistant maintains a conversation window of the last 20 messages. This enables multi-turn investigations where each question builds on previous findings. A typical diagnostic workflow might follow this pattern:
1
"What are the pathogenic and likely pathogenic variants in this case?"
The assistant queries the database for P/LP variants and presents them with key annotations.
2
"Tell me more about the MYBPC3 variant"
The assistant provides detailed information about the specific variant, including ACMG criteria, population frequency, and functional predictions.
3
"What does the literature say about this variant?"
The assistant searches the literature database for publications mentioning MYBPC3 and the specific variant notation.
4
"Does the phenotype match support this as the causative variant?"
The assistant checks the phenotype matching results for MYBPC3 and discusses the correlation with the patient's clinical presentation.
5
"Generate the clinical interpretation report"
The assistant produces a comprehensive diagnostic report synthesizing all findings from the conversation and analysis modules.
Chained Tool Execution
A single question can trigger up to five sequential tool calls. The assistant decides when to chain tools based on the complexity of the question. For example, asking "Find all pathogenic variants in constrained genes and check the literature for each" may trigger a variant database query followed by multiple literature searches -- all within one response.
Tips for Best Results
Be specific about what you want to see. "Show pathogenic missense variants" produces better results than "show me some interesting variants".
Use clinical terminology naturally. The assistant understands ACMG classes, HPO terms, gene symbols, HGVS notation, and genomic coordinates.
Ask follow-up questions to drill down. The assistant remembers context and resolves references like "that gene" or "those variants" from previous responses.
For complex analysis, break it into steps. First find the variants, then check the literature, then correlate with phenotype.
If the assistant misunderstands a query, rephrase with more specific criteria. Adding explicit column names or thresholds helps the SQL generator produce accurate queries.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [AI Clinical Assistant](https://folklore.helena.bio/docs/ai-assistant)

---

## Page: Capabilities | Folklore Documentation

Source: https://folklore.helena.bio/docs/ai-assistant/capabilities
Canonical: https://folklore.helena.bio/docs/ai-assistant/capabilities
Description: Folklore AI Assistant capabilities: case-specific variant queries, local literature search, evidence explanation, interpretation drafting, reports, and visualization.

Documentation / AI Clinical Assistant / Capabilities
## Capabilities
The Folklore AI Assistant works with the current case context and local literature index. It can query recorded variants, retrieve relevant publications, explain the available evidence, draft interpretation text, and suggest visualizations.
Conversational Variant Analysis
The assistant can translate a natural-language question into a read-only query against the current case database. Returned fields are limited to the data needed for the question and remain traceable to the stored analysis.
"Show me all pathogenic variants in cardiac genes"
"How many VUS have a gnomAD frequency below 0.01%?"
"List frameshift variants in genes with pLI above 0.9"
"What is the ACMG classification breakdown for chromosome 17?"
"Find compound heterozygote candidates in metabolic genes"
Literature Search
The assistant searches a local biomedical literature database containing over 1 million publications, 400,000 gene mentions, and 100,000 variant mentions. Literature queries run against this local mirror -- no external API calls are made during the search.
"What does the literature say about SCN1A and epilepsy?"
"Find recent publications mentioning BRCA1 pathogenic variants"
"Are there any case reports for this specific HGVS notation?"
"What is the evidence for this gene in cardiomyopathy?"
Clinical Interpretation
The assistant can draft a structured clinical interpretation for the current case using the recorded ACMG results, phenotype correlation, screening priorities, and literature evidence. The draft requires geneticist review. See Clinical Interpretation for details on interpretation levels and report structure.
Report Generation
Clinical interpretations can be exported as branded PDF or DOCX reports. PDF reports use Folklore branding with proper page headers, footers, page numbers, and a "CONFIDENTIAL" watermark. DOCX reports provide editable Word documents for further customization before distribution.
Intelligent Visualization
When the assistant executes a database query, it analyzes the results and suggests an appropriate visualization. The suggestion is genomics-aware: ACMG classification queries get pie charts with standard pathogenicity colors, chromosome distribution queries get genomically-sorted bar charts, and gene constraint queries can get scatter plots with clinical priority quadrants.
Query Type Visualization
ACMG classification breakdown Pie chart with pathogenicity colors
Variant impact distribution Severity-ordered bar chart (HIGH to MODIFIER)
Variants per chromosome Genomically-sorted bar chart (chr1-22, X, Y, M)
Top genes by variant count Bar chart with optional pLI constraint overlay
Gene constraint vs. frequency Scatter plot with clinical priority quadrants
Specific variant details Sortable, exportable data table
Clinical Knowledge
Even without querying databases, the assistant has extensive clinical genetics knowledge covering ACMG/AMP classification guidelines, Mendelian inheritance patterns, gene-disease associations, population genetics principles, HPO ontology, and functional predictor interpretation. It can explain why a specific ACMG criterion was triggered, discuss inheritance modes for a gene, or clarify the clinical significance of a computational prediction score.
What the Assistant Does Not Do
It does not make clinical diagnoses. The assistant provides analysis and interpretation support, but all findings require validation by a qualified clinical geneticist.
It does not modify variant classifications. ACMG classifications are determined by the automated pipeline and can only be overridden by the reviewing geneticist.
It does not access external databases during conversation. All queries run against local data that was loaded during the analysis pipeline.
It does not retain information between separate analysis sessions. Each case has its own isolated context.
It does not reclassify variants or substitute its own ACMG assessment for the pipeline's classification.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [AI Clinical Assistant](https://folklore.helena.bio/docs/ai-assistant)
- [Clinical Interpretation](https://folklore.helena.bio/docs/ai-assistant/clinical-interpretation)

---

## Page: Clinical Interpretation | Folklore Documentation

Source: https://folklore.helena.bio/docs/ai-assistant/clinical-interpretation
Canonical: https://folklore.helena.bio/docs/ai-assistant/clinical-interpretation
Description: AI-generated clinical genomic interpretation reports with four adaptive levels, structured templates, and data grounding safeguards.

Documentation / AI Clinical Assistant / Clinical Interpretation
## Clinical Interpretation
The Folklore AI Assistant can draft a structured clinical genomic interpretation from the analysis modules completed for the current case. The generated text remains a reviewable draft and does not change the recorded variant classifications.
Interpretation Levels
The system adapts the sections it can populate to the available case data. Missing modules are identified rather than filled through speculation.
Level 1 Variants Only
ACMG classification completed
Pathogenic and Likely Pathogenic variants, notable VUS with HIGH impact. Recommends enabling additional analysis modules.
Level 2 Screening-Focused
+ Clinical screening completed
Actionability tiers, prioritized variants with constraint scores, age-appropriate gene relevance. Adds clinical context from screening boosts.
Level 3 Phenotype-Focused
+ Phenotype matching completed
Genotype-phenotype correlation, semantic similarity scores, clinical tier assignments. Correlates findings with patient symptoms.
Level 4 Full Analysis
+ All modules completed
Comprehensive diagnostic synthesis integrating classification, screening, phenotype matching, and literature evidence into a cohesive clinical narrative.
Report Structure
Every report has three sections. The header and disclaimer are template-driven and never AI-generated, ensuring accuracy for patient information and legal text. Only the clinical interpretation body is produced by the AI.
Header (Template)
Report metadata, patient demographics, ethnicity, clinical indication, family history, consanguinity status, consent settings, HPO phenotype terms, analysis modules completed, and dataset summary (total variants, P/LP/VUS counts).
Clinical Interpretation (AI-Generated)
Reviewable narrative covering the available findings, recorded ACMG evidence, allele frequencies, inheritance context, genotype-phenotype correlation, and points requiring geneticist assessment.
Disclaimer (Template)
Legal statement noting the AI origin of the interpretation, requirement for validation by a qualified clinical geneticist, liability limitation, and co-signing requirement.
Data Grounding
The interpretation system enforces strict data grounding rules to prevent hallucination:
The AI only discusses genes and variants that appear in the provided analysis data. It does not invent or fabricate findings.
It does not add "textbook" pathogenic variants that are not present in the patient's results.
It does not reclassify variants. If the pipeline classified a variant as VUS, the interpretation discusses it as VUS.
Allele frequencies are reported in scientific notation with the exact values from the database.
When data for a specific analysis module is missing, the AI acknowledges the gap and recommends enabling that module rather than speculating.
Export Formats
Format Features
PDF Branded A4 layout with Folklore header, page numbers, "CONFIDENTIAL" watermark, print-friendly links, proper table formatting.
DOCX Editable Word document with formatted headings, tables, and inline styling. Suitable for further customization before distribution.
Important
The clinical interpretation is AI-generated and does not constitute a medical diagnosis. All findings must be independently validated by a qualified clinical geneticist before being used in patient care. The report includes a standard disclaimer section stating this requirement.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [AI Clinical Assistant](https://folklore.helena.bio/docs/ai-assistant)

---

## Page: Database Queries | Folklore Documentation

Source: https://folklore.helena.bio/docs/ai-assistant/database-queries
Canonical: https://folklore.helena.bio/docs/ai-assistant/database-queries
Description: How Helix AI translates natural language to SQL and queries patient variant and biomedical literature databases.

Documentation / AI Clinical Assistant / Database Queries
## Database Queries
Helix AI can query two databases through natural language: the patient's classified variant database and the biomedical literature database. The assistant automatically determines which database to query based on your question, generates SQL, executes it, and incorporates the results into its response.
Two Databases
Database Content Scale
Patient Variants Classified variants with ACMG annotations, population frequencies, functional predictions, phenotype matches, and screening scores. ~2.3M variants, 70 columns per variant
Biomedical Literature PubMed publications with gene mentions, variant mentions, abstracts, MeSH terms, and publication metadata. 1M+ publications, 400K gene mentions, 100K variant mentions
How Queries Work
1
Question Analysis
The assistant determines whether the question requires a database query. Questions about specific patient data trigger a query; general genetics knowledge is answered directly.
2
SQL Generation
The question is sent to a specialized SQL generation module that translates natural language into DuckDB SQL. The generator uses a low temperature (0.1) for precise, deterministic output and has access to the complete database schema.
3
Execution
The SQL query runs against a read-only DuckDB connection with a 30-second timeout. All queries are read-only -- the assistant cannot modify patient data.
4
Result Filtering
For detail queries (specific variants), results are filtered to 20 clinically essential columns out of 70, reducing token usage by approximately 70%. Aggregation queries preserve all columns.
5
Response Integration
The assistant receives the query results and incorporates them into its clinical response, adding visualization suggestions for chart-appropriate data.
Queryable Data
The patient variant database contains 70 columns per variant. The most commonly queried fields include:
Category Fields
Identity gene_symbol, chromosome, position, hgvs_protein, hgvs_cdna, rsid, transcript_id
Classification acmg_class, acmg_criteria, confidence_score
Consequence consequence, impact, biotype, exon_number, domains
Population Frequency gnomad_af, gnomad_popmax, gnomad_popmax_af, gnomad_hom
ClinVar clinvar_significance, clinvar_review_status, stars
Functional Predictions sift_prediction, alphamissense_prediction, metasvm_prediction, dann_score
Gene Constraint gene_pli, gene_oe_lof, gene_loeuf
Phenotype hpo_terms, hpo_count, hpo_phenotypes
Screening priority_score, priority_tier
Query Performance
Operation Typical Latency
SQL generation 1-3 seconds
Variant database query Under 200 milliseconds
Literature database query Under 500 milliseconds
Total (generation + execution) 2-4 seconds
Safety
All database access is strictly read-only. The assistant cannot insert, update, or delete any data. Query execution has a 30-second timeout to prevent runaway queries. Results are capped at a safe size limit to maintain responsive conversation flow.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [AI Clinical Assistant](https://folklore.helena.bio/docs/ai-assistant)

---

## Page: Changelog | Folklore Documentation

Source: https://folklore.helena.bio/docs/changelog
Canonical: https://folklore.helena.bio/docs/changelog
Description: Version history and release notes for the Folklore platform.

Documentation / Changelog
## Changelog
Version history and release notes for the Folklore platform.
v 1.6.0 April 2026 Classification Engine v3.28.0
Subtractive sophistication guards (v3.25.3 - v3.27.2)
Subtractive sophistication is the deliberate downgrade of Likely Pathogenic to VUS when the evidence profile is computational-only or biologically inconsistent with the gene mechanism. Six independent guards introduced across versions v3.25.3 through v3.27.2 reduce false positive LP classifications without ever creating new P/LP. The combined effect is more conservative classification in genes where the ACMG framework alone would over-call pathogenicity.
ar_biallelic_missense_guard (v3.25.3, HELIX-CR-2026-061): PP2 and LP4/LP5/LP6 blocked for het missense variants in AR/biallelic-only LoF genes without compound het partner and without ClinVar P/LP missense evidence. Richards 2015 Table 3: PP2 requires "missense variants are a common mechanism of disease" -- objectively false for AR LoF genes without documented missense pathogenicity.
computational_only_lp_guard (v3.25.4, HELIX-CR-2026-062): LP4/LP5/LP6 blocked when ALL pathogenic criteria are computational/annotation-based (PM1, PM2, PP2, PP3, PP3_splice). At least one observed/clinical criterion (PS1, PM3, PM4, PP4, PP5) required. De novo projection NOT gated -- PS2 is observed evidence.
ar_biallelic_het_lof_guard (v3.25.5, HELIX-CR-2026-063): all LP rules blocked for het LoF (HIGH impact) variants in AR/biallelic-only LoF genes without compound het partner. Het LoF in AR-only gene = carrier state, not pathogenic event. P rules NOT blocked: P1 with ClinVar PS1 is legitimate override.
LP2 PP2 bypass fix (v3.25.6, HELIX-CR-2026-064): PP2 (gene-level constraint) no longer bypasses the PP3_Strong+PM2 computational-only LP guard. PP4 (HPO match) and PP5 (ClinVar 1-star) still legitimately override.
LP2 PP3_splice_Strong extension (v3.25.9, HELIX-CR-2026-067): computational-only guard extended to block PP3_splice_Strong + PM2 without observed evidence, same as PP3_Strong + PM2.
LP2 splice bypass for AD/XL LoF genes (v3.27.2, HELIX-CR-2026-081): PP3_splice_Strong (SpliceAI >= 0.8) in AD/XL LoF-intolerant genes with disease association and splice proximity bypasses the computational-only guard. High-confidence splice disruption in established LoF-mechanism gene approaches functional LoF proxy. Six mandatory bypass conditions including MANE Select positional guard and GoF/DN exclusion.
PP3 / PP3_splice mutual exclusion (v3.25.7 - v3.25.8)
PP3_splice missense-tolerated guard (v3.25.7, HELIX-CR-2026-065): PP3_splice (all levels) blocked for pure missense variants without VEP splice consequence when BayesDel noAF data is available. ClinGen SVI Walker 2023 Section 3.2: PP3 is single criterion, not applied twice via different tools.
PP3 + PP3_splice mutual exclusion (v3.25.8, HELIX-CR-2026-066): PP3 BayesDel blocked for dual-consequence variants (missense + splice_region/donor/acceptor) where PP3_splice applies. SpliceAI is the more specific tool when VEP confirms splice proximity.
Combined effect: exactly one PP3 path per variant -- never both. Resolves the double-counting that previously inflated evidence count for dual-consequence variants.
PM3 ClinVar partner validation (v3.25.1)
PM3 ClinVar partner validation guard (HELIX-CR-2026-058): compound_het_candidate alone no longer satisfies PM3. Trans partner must have ClinVar P/LP (clinical_significance LIKE Pathogenic% OR Likely_pathogenic%, NOT LIKE %Conflicting%) with review_stars >= 2. Aligns with Richards 2015 Table 3: PM3 explicitly requires detection in trans with a pathogenic variant.
ClinVar-only guard (no acmg_class dependency): avoids circular reference where partner classification depends on PM3 which depends on partner classification.
Subtractive change: PM3 can only be removed, not added. 241 LP -> VUS expected. 5 LP unchanged (validated partner).
BS2 inheritance-aware AR threshold (v3.26.8)
BS2 inheritance-aware AR threshold (HELIX-CR-2026-074): AR-only genes (Orphanet has_ar=TRUE, has_ad=FALSE, or HI=30) with early-onset-only diseases use threshold of 10 homozygotes instead of generic 15.
Onset guard from refdb.orphanet_disease_onset: AR threshold activated only when gene has no Adult/Elderly/All ages onset disease. Aligns with Richards 2015 Table 3: BS2 "with full penetrance expected at an early age".
Criteria string marker: BS2[AR] when AR threshold path was used.
Dual-mechanism AD/AR LoF bypass (v3.27.0)
Dual-mechanism guard bypass (HELIX-CR-2026-079): ar_biallelic_het_lof_guard (CR-063) extended with bypass for genes with documented monoallelic LoF mechanism. Genes with both biallelic LoF AND monoallelic LoF have an AD pathway where het LoF is pathogenic, not carrier state.
Two bypass paths: gene_disease_mechanism monoallelic LoF (definitive/strong/moderate) OR CLINVAR_LOF_AD_GENES (IFT140, GCM2). ClinGen AD Definitive/Strong path removed in v2 review: ClinGen AD moi does not distinguish LoF from GoF/DN.
Additive change: VUS to LP for het LoF in dual-mechanism AD/AR LoF genes. Affected gene: IFT140 (3rd ADPKD gene per Senum 2022 PMID:34890546).
gnomAD LoF tolerance signal + BP_regional (v3.24.0 - v3.25.0)
gnomAD LoF population tolerance signal (v3.24.0, HELIX-CR-2026-057 Phase 1): boolean gnomad_lof_tolerant flags genes with HC LoF variants and homozygous/hemizygous carriers in healthy gnomAD individuals outside MANE CDS. Phase 1 informational only.
BP_regional Supporting benign (v3.25.0, HELIX-CR-2026-057 Phase 3): gnomad_lof_tolerant signal converted into Supporting benign criterion. Applied to HIGH impact variants outside MANE CDS in genes with documented LoF tolerance. AD genes excluded -- het LoF tolerance in AD is a different clinical argument. Compound het candidates excluded.
Reference data: gnomad_lof_tolerant_gene_summary table aggregates max_nhomalt, max_ac_hemi, lof_variant_count per gene with example_variant audit trail. Criteria string marker: BP_regional[gnomAD_LoF].
BP4_splice_benign for non-canonical splice (v3.25.2)
BP4_splice_benign (HELIX-CR-2026-060): Supporting benign for non-canonical splice region variants (splice_donor_region, splice_donor_5th_base, splice_acceptor_5th_base, splice_region, splice_polypyrimidine_tract) with SpliceAI <= 0.1. Excludes synonymous (BP7 handles) and missense (BP4 BayesDel handles).
PVS1 guard: not applied when PVS1 condition is satisfied. Threshold consistent with BP7. Adds benign evidence only -- cannot create new P/LP.
ClinGen negative evidence guard (v3.26.9)
ClinGen negative evidence guard (HELIX-CR-2026-078): genes with ClinGen Gene-Disease Validity classifications limited to Limited/Disputed/Refuted (no Definitive/Strong/Moderate) no longer accepted as having disease association via Orphanet entry alone.
clingen_negative_genes CTE: ~480 genes with ClinGen negative-only evaluation. Affects ClinVar 1-star override gate, LP4/LP5/LP6 disease gate, PVS1 disease gate, PP3_Strong disease gate, de novo projection.
Subtractive change: 205 affected genes with Orphanet entry + ClinGen negative-only evaluation no longer satisfy disease association gate.
PS1 minimum review stars correction (v3.26.7)
PS1 minimum review stars (HELIX-CR-2026-073): ps1_min_stars corrected from 1 to 2 in processing.yaml. Code default was already 2; YAML override allowed PS1 (Strong) at ClinVar 1-star, bypassing v3.10.0 disease association gate.
PP5 restoration: review_stars = 1 ClinVar P/LP now correctly receives PP5 (Supporting) instead of PS1 (Strong). PP5 range no longer empty set.
PM3 partner validation: partner.review_stars >= ps1_min_stars (shared config).
Combining rule corrections via Bug Bounty (v3.26.3 - v3.26.7)
P8 combining rule fix (v3.26.3, BB-001): ps_count >= 1 AND pm_count >= 4 corrected to ps_count >= 1 AND pm_count >= 1 AND pp_count >= 4 per Richards 2015 Table 5 rule (viii). Previous implementation was dead code (subsumed by P6) and encoded a non-existent rule.
De novo projection compound het + hemizygous (v3.26.4, BB-003): compound_het_candidate guard removed from de_novo_projection outer CASE. PS2 (origin) and PM3 (biallelic configuration) are orthogonal evidence types. Hemizygous genotype support added with chromosome guard (X/Y).
PM2 af_grpmax = 0.0 fix (v3.26.5, BB-006): af_grpmax = 0.0 treated as absent from controls, consistent with global_af path. Affects RASopathy genes with pm2_threshold=0.0.
B2 combining rule BP4_Moderate exclusion (v3.26.6, BB-007): B2 (>= 2 Strong benign) replaced with explicit (has_bs1 + has_bs2) >= 2. BP4_Moderate (Moderate benign per Pejaver 2022) excluded from B2. LB1b fallthrough added: BS1 + BP4_Moderate = LB.
De novo projection refinements (v3.26.0 - v3.26.4)
De novo AF ceiling + post-projection conflict guard (v3.26.0, HELIX-CR-2026-071): de_novo_projection CTE filters by AF (global_af > 0.001 OR af_grpmax > 0.001 blocks projection). Post-projection conflict guard: bs_count > 0 blocks de novo projection unconditionally.
Compound het + hemizygous compatibility (v3.26.4, BB-003): de novo and compound het are biologically compatible. Hemizygous males on X/Y eligible for de novo projection.
Subtractive change: removes false de novo candidates. Cannot create new P/LP or new de novo candidates.
Processing notes refactor (v3.27.1, v3.28.0)
Processing notes CASE priority fix (v3.27.1, HELIX-CR-2026-080): single first-match-wins CASE structure replaced with independent CASE WHEN concatenation. Variants matching multiple conditions now receive all applicable notes separated by semicolons.
Processing notes sentinel-guarded dedup (v3.28.0, HELIX-CR-HEL-VA-2026-001): all 14 ACMG note categories prefixed with [ACMG] sentinel. Symmetric refactor in artifact detection stage with [ARTIFACT] sentinel. Eliminates stale note preservation across reclassification runs and N-fold note duplication on repeated classifier runs.
Zero classification logic changes in either CR. Operational improvement only.
Reference databases
gnomAD pext (v4.1, GTEx v10): per-exon expression proportion across 49 GTEx tissues. 189,856 exon records for 18,923 MANE Select genes. Used by PVS1 expression-aware guard (v3.23.0).
MANE Select transcripts (v1.4): 19,354 transcripts (19,288 Select + 66 Plus Clinical) with CDS coordinates. Used by PVS1 positional guard (v3.22.0) and PM4.
gnomad_lof_tolerant_gene_summary: aggregated max_nhomalt, max_ac_hemi, lof_variant_count per gene. Source of gnomAD LoF tolerance signal (v3.24.0) and BP_regional criterion (v3.25.0).
orphanet_disease_onset (HELIX-REF-005): per-disease onset annotations (Adult/Elderly/All ages flags) from Orphadata. Used by BS2 inheritance-aware AR threshold (v3.26.8).
clinvar_missense_genes: 2,136 genes with ClinVar P/LP missense evidence (>= 2 at 2+ stars OR >= 3 at 1+ star). Used by PP3 missense relevance guard Gate 4a (v3.19.0) and BP1 reference-based ClinVar guard (v3.20.9).
gene_disease_mechanism (HELIX-REF-003): unified molecular mechanism table from G2P + GoFCards + manual curation. 2,617 records. Source for gof_genes_unified view (347 GoF/DN/GoE genes) and gof_genes_exclusive view (excluded 55 dual-mechanism genes).
v 1.5.0 March 2026 Classification Engine v3.17.0
Classification (HELIX-CR-2026-023)
PVS1: G2P molecular mechanism integration. DECIPHER Gene2Phenotype (G2P) database now provides primary GoF/DN guard for PVS1. gof_genes_g2p view contains 198 pure GoF/DN monoallelic genes where PVS1 is blocked.
PVS1: GOF_AD_GENES reduced from 36 to 14 fallback genes (genes not in G2P: 6 neurodegeneration/toxic aggregation, 3 somatic/neomorphic, 5 absent from G2P 2026-02-28 release). PRSS1 added to fallback.
PVS1: GNAS removed from GOF_AD_GENES (dual-mechanism: McCune-Albright = GoF, pseudohypoparathyroidism 1A = LoF). GNAS LoF variants now correctly receive PVS1.
PVS1: 49 dual-mechanism genes (SCN5A, LMNA, KCNH2, KCNQ1, FGFR1) excluded from GoF view, preserving PVS1 for legitimate LoF phenotypes.
PVS1: G2P confidence filter: only definitive, strong, moderate included. Limited confidence excluded.
Subtractive change: can only remove PVS1 from GoF/DN genes. Cannot create false P/LP.
Clinical trigger: PD2025_090 PRSS1 p.Gly177Ter (stop_gained in GoF gene).
Reference Databases
DECIPHER G2P: Molecular mechanism data (g2p_gene_mechanism table, 2,372 records) added to reference_db alongside existing HPO enrichment data.
gof_genes_g2p view: 198 pure GoF/DN monoallelic genes for PVS1 guard. Dual-mechanism genes excluded.
New loader script: scripts/load_g2p_mechanism.py with dry-run, force flags, and 9-point validation.
v 1.4.0 March 2026 Classification Engine v3.16.8
Classification
PVS1: GoF AD gene exclusion (HELIX-CR-2026-022). GOF_AD_GENES curated list (36 genes) blocks PVS1 for AD genes where disease mechanism is gain-of-function, dominant-negative, or toxic aggregation.
PVS1: ClinGen AD Definitive/Strong genes bypass pLI/LOEUF constraint gate (v3.16.6). Genes like TUBB1 with ClinGen AD Definitive now qualify for PVS1.
LP classification: Disease association gate for LP rules LP4, LP5, LP6 (v3.16.7, HELIX-CR-2026-021). Prevents LP in genes without known Mendelian disease mechanism.
ClinGen Gene-Disease Validity integration (v3.16.0): ar_lof_genes_clingen and disease_associated_genes_clingen reference tables replace Python constants.
Clinical triggers: PD2025_082 RAC1 frameshift (GoF gene), PD2025_090 discordance analysis.
Reference Databases
ClinGen Gene-Disease Validity: ar_lof_genes_clingen and disease_associated_genes_clingen tables added to reference_db.
Total reference databases: 14 (previously 13).
v 1.3.0 March 2026 Classification Engine v3.15.0
Classification
PM1: Critical functional domains migrated from Python constant to reference database table (refdb.interpro_pfam_domains). HELIX-CR-2026-012.
PM1: Full InterPro Pfam catalog (27,481 entries) loaded. 49 domains marked critical across 14 categories.
PM1: CRITICAL_PFAM_DOMAINS converted from short names to InterPro-verified accession numbers (v3.14.0, HELIX-CR-2026-011). Fixed 0 legitimate PM1 applications since curated list introduction.
PM1: Post-fix validation on HG2023_206: PM1 count from 2 (false positive) to 1,139 (legitimate).
PVS1: ClinVar LOF evidence as fifth constraint gate path (v3.13.0, HELIX-CR-2026-010). Small AD genes with ClinVar P/LP LOF but uninformative gnomAD constraint now qualify.
PP3/BP4: Missense consequence guard (v3.12.1, HELIX-CR-2026-009). BayesDel path now requires missense_variant consequence.
Clinical triggers: HG2023_206 RYR1 (PM1 fix), GCM2 p.Arg131Ter (PVS1 + PP3 guard).
Reference Databases
InterPro / Pfam added as 13th reference database (27,481 Pfam entries).
Loader script: scripts/load_interpro_pfam.py with InterPro API fetch and offline fallback.
Total reference databases: 13 (previously 12).
v 1.2.0 March 2026 Classification Engine v3.12.0
HPO Enrichment (HELIX-REF-001)
Multi-Source HPO Enrichment Pipeline: gene-phenotype annotations expanded from 1 source (5,173 genes) to 6 sources (5,688 genes, +10%)
HPO Consortium: 320K records, 5,173 genes (primary source, curated gene-phenotype associations)
Orphanet disease-to-HPO: 168K records, 3,176 genes with clinical frequency data (via Orphadata en_product4.xml + en_product6.xml)
DECIPHER Gene2Phenotype (G2P): 43K records, 2,125 genes from 7 clinical panels (DD, Eye, Cardiac, Skin, Skeletal, Cancer, Ear)
Monarch Initiative: 151K records, 4,791 genes (HPO Consortium redistribution via Monarch KG)
ClinVar-MedGen: 245K records, 5,258 genes (P/LP variant -> MedGen CUI -> HPO chain mapping, largest contributor of new genes)
Manual clinical curation: common-disease genes outside rare-disease HPO scope (TBC1D4 with PMID evidence)
3,292 genes with increased HPO term coverage from cross-source aggregation
Source priority ordering: Orphanet > G2P > HPO Consortium > Monarch > ClinVar-MedGen
Low-confidence filtering: Orphanet modifying/candidate and G2P limited entries excluded from clinical view
Classification
Classifier version bumped to v3.12.0
PP4 criterion evaluates against enriched HPO set (more genes trigger PP4)
No ACMG criteria logic changes -- only annotation data expanded
Reference Databases
DECIPHER G2P added as new reference database (7 clinical panels)
Monarch Initiative added as new reference database
MedGen HPO-OMIM mapping added as new reference database
HPO database entry updated to reflect multi-source enriched view
Total reference databases: 12 (previously 9)
v 1.1.0 March 2026 Classification Engine v3.11.4
Classification
PVS1 disease association gate: requires established disease association before applying PVS1 (ClinVar P/LP, Orphanet, ClinGen HI, AR LoF list, or VCEP coverage)
BS1 constraint-implied AD fallback: uses AD threshold (0.1%) for LoF-constrained genes without inheritance data
BS1 cascade expanded from 5 to 8 levels with explicit Orphanet AR and AR_LOF_GENES priorities
PVS1 non-canonical splice exclusion: splice_donor_5th_base_variant, splice_donor_region_variant, splice_acceptor_5th_base_variant blocked from PVS1 (v3.11.4)
PVS1 NMD_transcript_variant exclusion fix: NMD_transcript no longer blocks PVS1, NMD_escaping_variant now blocks instead (v3.11.3)
De novo projection: prospective PS2 upgrade computation for VUS variants to guide trio testing (v3.11.0)
ClinVar low-confidence disease association gate: 1-star P/LP requires gene disease evidence (v3.10.0)
disease_genes_clinvar CTE circular reference fix: review_stars >= 2 filter (v3.10.1)
PVS1 last-exon NMD downgrade: last-exon truncating variants receive PVS1_Strong instead of PVS1 Very Strong (v3.9.0)
BP1 ClinVar pathogenic missense guard and PP3_Supporting label (v3.8.2)
PP3_Strong disease association gate (v3.8.1)
Reference Databases
Orphanet/Orphadata documented as separate reference database (gene-disease-inheritance for 3,200 genes)
VCEP gene-specific specifications documented as separate reference database (~50-60 genes)
Total reference databases: 9 (previously 7)
v 1.0.0 February 2026 Initial Release
Classification
ACMG/AMP 2015 classification with Bayesian point-based framework (Tavtigian 2018/2020)
19 of 28 ACMG criteria automated
BayesDel_noAF with ClinGen SVI-calibrated thresholds for PP3/BP4 (Pejaver 2022)
SpliceAI integration for PP3_splice with PVS1 double-counting prevention
ClinVar override logic with review star quality filtering
Gene-specific VCEP threshold support
Reference Databases
gnomAD v4.1 (759M variants, 807K individuals)
gnomAD Constraint v4.1 (18.2K genes)
ClinVar 2025-01
dbNSFP 4.9c
SpliceAI precomputed (Ensembl MANE Release 113)
HPO gene-phenotype associations
ClinGen dosage sensitivity
Ensembl VEP Release 113 with local offline cache
Phenotype Matching
Lin semantic similarity with HPO ontology graph
Five-tier clinical priority system
Gene-level deduplication and aggregation
Automatic HPO term extraction from free-text clinical descriptions
Screening
Seven-component scoring algorithm (constraint, deleteriousness, phenotype, dosage, consequence, compound het, age relevance)
Six screening modes (diagnostic, neonatal, pediatric, proactive adult, carrier, pharmacogenomics)
Age-aware prioritization with curated gene lists
Clinical boosts for ethnicity, family history, sex-linked inheritance, consanguinity, pregnancy
Four-tier priority ranking with clinical actionability labels
AI Clinical Assistant
Conversational variant analysis with natural language database queries
Biomedical literature search (1M+ publications, local PubMed mirror)
Four-level adaptive clinical interpretation generation
PDF and DOCX report export with Folklore branding
Genomics-aware visualization suggestions
On-premise LLM inference within EU infrastructure
Infrastructure
EU-based processing (Helsinki, Finland)
All databases stored and queried locally
Zero external API calls during variant processing
GDPR-compliant data handling
DuckDB-based analytical pipeline

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)

---

## Page: Variant Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification
Canonical: https://folklore.helena.bio/docs/classification
Description: How Folklore annotates variants, selects the appropriate clinical framework, evaluates evidence, and presents reviewable classifications.

## Variant Analysis
Folklore turns normalized variant records into a reviewable clinical evidence summary. It annotates each supported variant, assigns the appropriate interpretation framework, evaluates available evidence, and records the framework, evidence codes, and resulting class.
The automated result is decision support, not a diagnosis. Patient phenotype, inheritance, family evidence, assay quality, and expert review remain essential to clinical interpretation.
Framework Selection
Nuclear sequence variants
Evaluated with the ACMG/AMP evidence framework, ClinGen guidance, and applicable gene-specific VCEP specifications.
Mitochondrial variants
Evaluated separately with mtDNA-specific specifications that account for heteroplasmy, haplogroups, and mitochondrial population evidence.
Structural variants
Constitutional copy-number loss and gain use the ClinGen/ACMG Riggs 2020 framework. Other eligible nuclear structural events can use the nuclear variant framework; unsupported records remain explicitly unclassified.
What the Result Contains
Five-tier class
Pathogenic, Likely Pathogenic, VUS, Likely Benign, or Benign when a supported framework reaches a result.
Evidence trace
The criteria and strength modifiers that contributed to the automated result.
Framework provenance
The nuclear, mitochondrial, or structural framework that owned the interpretation.
Review context
Warnings, source markers, and any geneticist reclassification remain visible alongside the pipeline result.
In This Section
ACMG/AMP Framework
The evidence model used for nuclear sequence variants.
Criteria Reference
A clinical-level guide to the 28 ACMG/AMP evidence codes.
Combining Rules
How evidence strengths are combined into the five-tier result.
ClinVar and ClinGen
How external assertions enter the evidence and decision paths.
Conflicting Evidence
How contradictory evidence is held for safe review.
Confidence Indicator
What the displayed class-level indicator means and does not mean.
Mitochondrial Variants
The dedicated mtDNA interpretation framework and its boundaries.
Review and Reclassification
How a geneticist can record, audit, and revert a reviewed class.

### Links and cited sources

- [ACMG/AMP Framework The evidence model used for nuclear sequence variants.](https://folklore.helena.bio/docs/classification/acmg-framework)
- [Criteria Reference A clinical-level guide to the 28 ACMG/AMP evidence codes.](https://folklore.helena.bio/docs/classification/criteria-reference)
- [Combining Rules How evidence strengths are combined into the five-tier result.](https://folklore.helena.bio/docs/classification/combining-rules)
- [ClinVar and ClinGen How external assertions enter the evidence and decision paths.](https://folklore.helena.bio/docs/classification/clinvar-integration)
- [Conflicting Evidence How contradictory evidence is held for safe review.](https://folklore.helena.bio/docs/classification/conflicting-evidence)
- [Confidence Indicator What the displayed class-level indicator means and does not mean.](https://folklore.helena.bio/docs/classification/confidence-scores)
- [Mitochondrial Variants The dedicated mtDNA interpretation framework and its boundaries.](https://folklore.helena.bio/docs/classification/mitochondrial-variants)
- [Review and Reclassification How a geneticist can record, audit, and revert a reviewed class.](https://folklore.helena.bio/docs/classification/reclassification)

---

## Page: ACMG/AMP Framework for Nuclear Variant Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/acmg-framework
Canonical: https://folklore.helena.bio/docs/classification/acmg-framework
Description: How Folklore represents ACMG/AMP evidence, strength modifications, gene-specific guidance, provenance, and expert-review boundaries.

Documentation / Variant Analysis / ACMG/AMP Framework
## ACMG/AMP Framework for Nuclear Variant Analysis
Folklore uses the ACMG/AMP five-tier framework for supported nuclear sequence variants. Evidence is recorded by direction and strength, then combined through published rules with additional safety checks for conflicting or incomplete evidence.
Evidence Model
The 28 ACMG/AMP criteria cover population, computational, functional, segregation, de novo, allelic, phenotype, and curated clinical evidence. Strength-modified criteria retain their original code with an explicit strength suffix, following ClinGen nomenclature.
Direction Strengths Examples
Pathogenic Very Strong, Strong, Moderate, Supporting PVS1, PS1, PM2, PP3
Benign Stand-alone, Strong, Supporting BA1, BS1, BP4
Automated and Review-Dependent Evidence
Nineteen criteria have active automated production paths. Eight criteria depend on case, family, functional, or segregation evidence that requires clinical review. PM5 has an implemented same-position evidence path but remains governed by deployment configuration. The interface distinguishes computed evidence from evidence that still needs expert assessment.
Gene-Specific Guidance
When an applicable ClinGen Variant Curation Expert Panel specification is available, Folklore can apply its gene- or disease-specific evidence refinements. The result records the panel and specification version so the reviewer can see when a general rule was modified.
Interpretation boundary
A reproducible rules engine can apply recorded evidence consistently, but it cannot replace assessment of phenotype fit, penetrance, inheritance, assay validity, or newly published evidence. The displayed class must be reviewed in the full clinical context.
Reference: Richards S, et al. Genetics in Medicine. 2015;17(5):405-424. PMID: 25741868

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)

---

## Page: ClinVar and ClinGen Evidence in Variant Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/clinvar-integration
Canonical: https://folklore.helena.bio/docs/classification/clinvar-integration
Description: How Folklore uses ClinVar assertions, review status, ClinGen expert-panel evidence, and explicit conflict boundaries.

Documentation / Variant Analysis / ClinVar and ClinGen
## ClinVar and ClinGen Evidence in Variant Analysis
Folklore treats public assertions as structured evidence with provenance, not as a single unquestioned answer. Assertion type, review status, source authority, disease context, and conflicts determine how an assertion can contribute.
Three Roles
Criterion evidence
Eligible exact-match assertions can support ACMG evidence codes. Source-quality requirements differ by evidence strength.
Bounded classification path
Eligible ClinVar classifications can determine the class only after higher-priority conflict and safety checks have been evaluated.
Expert-panel provenance
Applicable ClinGen VCEP assertions and specifications are identified separately so their authority and version remain visible.
Review Status
ClinVar review stars summarize the level of review, from submissions without assertion criteria through expert-panel and practice-guideline review. They are an authority signal, not a guarantee that an assertion is current or correct. Folklore records the review status and applies minimum quality requirements to the evidence path.
When an Assertion Does Not Decide the Class
The assertion is a VUS or another non-decisive category.
The review status does not meet the required evidence level.
Pathogenic and benign evidence conflict.
The assertion does not apply to the relevant disease or mechanism.
The record represents a risk allele rather than a conventional Mendelian classification.
A higher-authority, variant-specific expert-panel interpretation applies.
Versioned evidence
ClinVar and ClinGen content changes over time. Folklore processes governed local reference releases and exposes source identifiers and review context so a result can be interpreted against the evidence available to that classifier release.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)

---

## Page: ACMG/AMP Evidence Combining Rules | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/combining-rules
Canonical: https://folklore.helena.bio/docs/classification/combining-rules
Description: How Folklore combines pathogenic and benign evidence strengths, including the published ACMG/AMP rules and explicit safety guards.

Documentation / Variant Analysis / Combining Rules
## ACMG/AMP Evidence Combining Rules
Folklore groups applied criteria by pathogenic or benign direction and by evidence strength. It then evaluates the published ACMG/AMP combinations in a deterministic order, with documented ClinGen strength modifications where applicable.
Rule Families
Pathogenic
P1: 1 Very Strong + at least 1 Strong
P2: 1 Very Strong + at least 2 Moderate
P3: 1 Very Strong + 1 Moderate + 1 Supporting
P4: 1 Very Strong + at least 2 Supporting
P5: at least 2 Strong
P6: 1 Strong + at least 3 Moderate
P7: 1 Strong + 2 Moderate + at least 2 Supporting
P8: 1 Strong + 1 Moderate + at least 4 Supporting
Likely Pathogenic
LP1: 1 Very Strong + 1 Moderate
LP1b: PVS1 or PVS1_Strong + 1 Supporting when an applicable specification downgrades evidence strength
LP2: 1 Strong + 1 or 2 Moderate
LP3: 1 Strong + at least 2 Supporting
LP4: at least 3 Moderate
LP5: 2 Moderate + at least 2 Supporting
LP6: 1 Moderate + at least 4 Supporting
Benign and Likely Benign
B1: 1 Stand-alone benign criterion
B2: at least 2 Strong benign criteria
LB1: 1 Strong benign + 1 Supporting benign
LB2: at least 2 Supporting benign criteria
Safety Guards
A matching evidence count is necessary but not always sufficient. Folklore can hold a result at VUS when evidence is contradictory, when inheritance does not support the proposed mechanism, when a profile relies only on limited computational evidence, or when required clinical context is absent. These guards prevent a mathematically matching combination from being presented as stronger than the available evidence supports.
No hidden point total
The current nuclear classifier records criteria and strength levels and applies rule combinations. It does not expose a general per-variant Bayesian point total as the source of the final class. Calibrated predictor thresholds may modify evidence strength, but the displayed class is produced by the rule hierarchy and its safety checks.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)

---

## Page: Variant Classification Confidence Indicator | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/confidence-scores
Canonical: https://folklore.helena.bio/docs/classification/confidence-scores
Description: What the Folklore confidence indicator represents, its current class-level values, and why it is not a probability or boundary-distance score.

Documentation / Variant Analysis / Confidence Indicator
## Variant Classification Confidence Indicator
The current confidence field is a class-level display indicator. It is assigned from the resulting five-tier class and is not calculated from distance to an ACMG boundary.
Current Values
Classification Indicator Display level
Pathogenic 0.85 High
Likely Pathogenic 0.70 Medium
VUS 0.50 Low
Likely Benign 0.70 Medium
Benign 0.85 High
What It Does Not Mean
It is not the probability that the variant is pathogenic.
It is not the probability that the variant explains the patient phenotype.
It is not a continuous measure of distance from the next classification boundary.
Two variants in the same class can have the same indicator while relying on different evidence profiles.
Use the evidence trace
Clinical review should use the classification, criteria, source provenance, warnings, phenotype, inheritance, and case evidence. The confidence indicator is secondary metadata and must not be used alone for patient management.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)

---

## Page: Conflicting Evidence in Variant Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/conflicting-evidence
Canonical: https://folklore.helena.bio/docs/classification/conflicting-evidence
Description: How Folklore handles expert-panel assertions, risk alleles, stand-alone benign evidence, and opposing pathogenic and benign criteria.

Documentation / Variant Analysis / Conflicting Evidence
## Conflicting Evidence in Variant Analysis
Contradictory evidence is not reduced to an unqualified score. Folklore evaluates evidence authority, variant context, and the direction and strength of the competing criteria before a final class is displayed.
Decision Order
1
Expert-panel evidence
Applicable ClinGen expert-panel P/LP assertions are handled through an authority-specific path and remain visible in the evidence provenance.
2
Risk-allele safety path
An established risk allele is not treated as a conventional Mendelian Pathogenic result. It is held in a reviewable uncertain state with source context.
3
Stand-alone benign evidence
BA1 can support Benign when its population premise is met, but observed Strong or Very Strong pathogenic evidence prevents a silent stand-alone resolution.
4
Pathogenic versus benign conflict
Moderate-or-stronger pathogenic evidence opposed by Strong benign evidence is held at VUS for review.
5
Remaining evidence profiles
If no higher-priority path applies, the standard combining rules and additional safety guards determine the result.
Review the source evidence
A VUS produced by a conflict is an explicit safety outcome, not an absence of analysis. Reviewers should examine population evidence, assertion provenance, phenotype fit, inheritance, assay quality, and whether the evidence applies to the same disease mechanism.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)

---

## Page: ACMG/AMP Criteria Reference | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/criteria-reference
Canonical: https://folklore.helena.bio/docs/classification/criteria-reference
Description: A clinical-level reference to the 28 ACMG/AMP evidence codes and their current automated, review-dependent, or governed status in Folklore.

Documentation / Variant Analysis / Criteria Reference
## ACMG/AMP Criteria Reference
This page summarizes the clinical meaning and current execution status of each ACMG/AMP criterion. Exact application depends on variant type, disease mechanism, inheritance, source quality, applicable ClinGen guidance, and the evidence available for the case.
Status labels
Automated means the production pipeline can evaluate the criterion from governed data. Clinical review means the evidence depends on case-level judgment or information not inferred by the pipeline. Governed means an implementation exists but activation depends on validated reference data and deployment configuration.
Pathogenic Evidence
PVS1 Very Strong, with supported modifications Automated
Predicted loss of function where loss of function is an established disease mechanism
PS1 Strong Automated
Same amino acid change as an established pathogenic variant
PS2 Strong Clinical review
De novo occurrence with confirmed parental relationships
PS3 Strong Clinical review
Well-established functional evidence supports a damaging effect
PS4 Strong Clinical review
Increased prevalence in affected individuals
PM1 Moderate Automated
Location in a critical functional region without benign variation
PM2 Moderate or Supporting Automated
Absent or sufficiently rare in population reference data
PM3 Moderate Automated
Observed in trans with a pathogenic partner in a recessive disorder
PM4 Moderate Automated
Protein length change in an applicable non-repeat context
PM5 Moderate Governed
Different pathogenic missense change at the same residue
PM6 Moderate Clinical review
Assumed de novo occurrence without confirmed parental relationships
PP1 Supporting, with supported modifications Clinical review
Cosegregation with disease in affected family members
PP2 Supporting Automated
Missense change in a gene where missense variation is an established disease mechanism
PP3 Supporting, Moderate, or Strong Automated
Calibrated computational evidence supports a damaging effect
PP4 Supporting Automated
Patient phenotype is highly specific for the associated disorder or gene
PP5 Supporting Automated
A reputable clinical source reports pathogenicity
Benign Evidence
BA1 Stand-alone Automated
Population frequency is incompatible with the disorder
BS1 Strong, with supported modifications Automated
Population frequency is greater than expected for the disorder
BS2 Strong Automated
Observed in healthy individuals where full penetrance is expected
BS3 Strong Clinical review
Well-established functional evidence supports no damaging effect
BS4 Strong Clinical review
Lack of segregation with disease
BP1 Supporting Automated
Missense change in a gene where truncating variation is the established mechanism
BP2 Supporting Automated
Allelic observation is inconsistent with the proposed disease mechanism
BP3 Supporting Automated
In-frame length change in a repetitive region without known function
BP4 Supporting or Moderate Automated
Calibrated computational evidence supports no damaging effect
BP5 Supporting Clinical review
An alternate molecular basis explains the case
BP6 Supporting Automated
A reputable clinical source reports benignity
BP7 Supporting Automated
Synonymous change with no predicted splice impact
Framework boundary
This reference describes nuclear sequence-variant criteria. Mitochondrial and structural variants use dedicated specifications that include, exclude, or modify criteria for their biology and evidence model.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)

---

## Page: Mitochondrial Variant Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/mitochondrial-variants
Canonical: https://folklore.helena.bio/docs/classification/mitochondrial-variants
Description: How Folklore routes mtDNA variants to a dedicated framework that accounts for heteroplasmy, haplogroups, mitochondrial population data, and mtDNA-specific evidence.

Documentation / Variant Analysis / Mitochondrial Variants
## Mitochondrial Variant Analysis
Mitochondrial variants are not evaluated by the nuclear sequence-variant path. Folklore uses a dedicated mtDNA framework based on the Mitochondrial Disease Working Group specifications described by McCormick et al. (2020).
Why mtDNA Is Separate
Heteroplasmy
The proportion of mitochondrial genomes carrying a variant can differ across samples and tissues.
Maternal inheritance
Inheritance and segregation require mitochondrial rather than autosomal interpretation.
Haplogroup context
Population lineage context can change how a common or haplogroup-defining variant is interpreted.
Different evidence sources
mtDNA-specific population, clinical, functional, and computational resources are used where available.
Evidence and Output
The mtDNA path evaluates the subset of criteria that can be applied responsibly to mitochondrial variants, records mtDNA-specific provenance, and excludes nuclear criteria that do not translate to mitochondrial biology. Evidence that depends on functional studies, maternal segregation, tissue heteroplasmy, or expert curation remains visible as a review boundary rather than being inferred.
The result uses the familiar five-tier terminology, but its evidence trace identifies the mitochondrial framework. A nuclear and a mitochondrial result should not be interpreted as if they were produced by the same criteria set.
Clinical limitation
Blood heteroplasmy may not represent an affected tissue, and a computational result cannot establish tissue distribution, threshold effect, or causality. Review should include phenotype, maternal history, sample type, heteroplasmy, and relevant functional evidence.
Reference: McCormick EM, et al. Human Mutation. 2020;41(12):2028-2057. PMID: 33058415

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)
- [PMID: 33058415](https://pubmed.ncbi.nlm.nih.gov/33058415/)

---

## Page: Clinical Review and Variant Reclassification | Folklore Documentation

Source: https://folklore.helena.bio/docs/classification/reclassification
Canonical: https://folklore.helena.bio/docs/classification/reclassification
Description: How Folklore preserves the automated class while allowing an authorized geneticist to record a justified, auditable, and reversible reviewed classification.

Documentation / Variant Analysis / Review and Reclassification
## Clinical Review and Variant Reclassification
Folklore separates the pipeline classification from the geneticist's reviewed classification. An authorized reviewer can record a different five-tier class with a clinical justification while the original pipeline result remains preserved.
Review Workflow
1. Review the evidence
Assess the pipeline class, criteria, source provenance, patient phenotype, inheritance, family evidence, and relevant literature.
2. Record the reviewed class
Select a five-tier class and provide a clinical justification. The first pipeline class is preserved as the original result.
3. Use the effective class
While an override is active, lists and summaries use the reviewed class and expose the override metadata.
4. Revert when appropriate
The reviewed class can be removed, restoring the preserved pipeline class as the active result.
Auditability
Reclassification records the original class, reviewed class, justification, reviewer attribution, and timestamp. Reverting a reviewed class is also audited. The automated evidence is not rewritten, so reviewers can distinguish what the pipeline produced from the active clinical decision.
Important distinction
A reviewed class is a clinical decision, not a retraining event and not a hidden modification of the classifier. Future pipeline runs may change as reference data or classifier versions change, while the recorded review remains attributable and reversible.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Variant Analysis](https://folklore.helena.bio/docs/classification)

---

## Page: Cohort Analysis and Studies | Folklore Documentation

Source: https://folklore.helena.bio/docs/cohort-analysis
Canonical: https://folklore.helena.bio/docs/cohort-analysis
Description: How Folklore Studies organize samples, run cohort-relative quality control, build a cohort variant matrix, and present current gene- and variant-level results.

Documentation / Cohort Analysis
## Cohort Analysis (Studies)
A Study is Folklore's workspace for analyzing a group of samples together. It preserves the individual Variant Analysis result for each sample while adding cohort-level quality, carrier, frequency, and gene evidence.
Studies are living cohorts. Samples can be added, removed, or retried, and the cohort matrix and downstream results should be regenerated when membership changes.
Individual Classification Remains Authoritative
Cohort Analysis does not replace or silently rewrite a sample's Variant Analysis result. The cohort layer aggregates classifications and genotypes for comparison. Where classifications differ across samples, Folklore records the most clinically severe observed class for cohort display and marks the variant as discordant for review.
Required Study Context
A defined sequencing type: WES, WGS, or CES.
A screening panel or explicit gene list for the analysis in scope.
Unique sample identifiers and compatible VCF or completed-analysis inputs.
Comparable reference build, calling approach, and analysis provenance across samples.
Production Workflow
1
Create the Study
Name the Study, select the sequencing type, and define the clinical profile or screening panel that determines the genes in scope.
2
Add and Process Samples
Upload VCF files, import from connected storage, reuse completed analyses, or start from a batch manifest. Each sample is classified through Variant Analysis before cohort aggregation unless a compatible completed analysis is reused.
3
Run Cohort Quality Control
Folklore reports per-sample variant count, Ti/Tv ratio, heterozygous-to-homozygous ratio, and mean depth, then flags cohort-relative outliers.
4
Build the Cohort Matrix
A deduplicated variant catalog and sparse sample-genotype matrix provide carrier counts, homozygous counts, cohort allele frequency, and classification-discordance flags.
5
Generate Study Results
The current workflow produces actionable findings, burden results, pLoF summaries, frequency comparisons, compound-heterozygous candidates, and ranked candidate genes when the required evidence is available.
Current Outputs
Per-sample processing and quality-control status.
A searchable cohort variant matrix with sample-level genotypes and carrier summaries.
Pathogenic and Likely Pathogenic findings mapped to the samples that carry them.
Gene-level rare-variant burden results with multiple-testing correction and power context.
Predicted loss-of-function and candidate compound-heterozygous summaries.
A ranked candidate-gene view that keeps its contributing evidence components visible.
Current Scope
The production Study workflow requires a gene panel or explicit gene list. GWAS and polygenic risk scoring are not current Study outputs. SKAT-O is not currently reported; the burden view uses the implemented carrier-based tests described in Results and Interpretation.
Pathway enrichment appears only when pathway definitions and a completed result are available. The current automated workflow does not load pathway definitions by default, so this section can be absent.
Clinical Boundary
Cohort associations and rankings are decision-support evidence, not diagnoses or proof of causality. Study design, ancestry, relatedness, technical batch, phenotype definition, and sample size remain essential to interpretation.
In This Section
Creating a Study
Study setup, sample import methods, comparability, and living-cohort behavior.
Quality Control
Reported metrics, cohort-relative outlier rules, status meanings, and limitations.
Results and Interpretation
Variant matrix, actionable findings, burden, pLoF, frequency, compound-het, and candidate genes.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Creating a Study Study setup, sample import methods, comparability, and living-cohort behavior.](https://folklore.helena.bio/docs/cohort-analysis/creating-a-study)
- [Quality Control Reported metrics, cohort-relative outlier rules, status meanings, and limitations.](https://folklore.helena.bio/docs/cohort-analysis/quality-control)
- [Results and Interpretation Variant matrix, actionable findings, burden, pLoF, frequency, compound-het, and candidate genes.](https://folklore.helena.bio/docs/cohort-analysis/results)

---

## Page: Creating and Maintaining a Study | Folklore Documentation

Source: https://folklore.helena.bio/docs/cohort-analysis/creating-a-study
Canonical: https://folklore.helena.bio/docs/cohort-analysis/creating-a-study
Description: How to create a Folklore Study, choose its clinical scope, add samples from supported sources, and maintain a comparable living cohort.

Documentation / Cohort Analysis / Creating a Study
## Creating and Maintaining a Study
Create a Study when multiple samples need to be interpreted as one defined cohort. The initial setup fixes the Study name, sequencing type, and clinical scope; sample membership can continue to change afterward.
Initial Setup
Enter a clear Study name. When an XLSX manifest is selected, Folklore can derive an initial name from the filename, which you can edit.
Select WES, WGS, or CES to describe the sequencing data in the Study.
Choose the clinical profile and screening panel, or add explicitly scoped genes where supported.
Review the manifest before starting the linked batch. A Study can exist even if subsequent batch ingestion fails, so verify the Study sample list after processing.
A Gene Scope Is Required for Analysis
The production analysis resolves genes from the Study screening panel unless an explicit gene list is supplied. A Study without either cannot start cohort analysis. Select the intended panel during setup and review it whenever the Study question changes.
Ways to Add Samples
Method When to Use It What Happens Next
Study manifest Create a new Study and linked processing batch from the Folklore XLSX manifest. Samples are uploaded and processed through the linked batch workflow.
Upload Files Add VCF files from your computer to an existing Study. Each uploaded sample runs through Variant Analysis and screening with the Study panel.
Import from MEGA Select sample folders or files from configured MEGA storage. Files are downloaded and processed through Variant Analysis and Study screening.
Import Sessions Reuse compatible completed Folklore analyses. Existing classification is retained and the sample is re-screened with the Study panel.
Upload Manifest Register sample metadata from CSV or TSV using sample_id and vcf_filename columns. The manifest maps samples to their expected VCF filenames; it does not itself upload the VCF files.
Sample Metadata
Every sample requires a unique sample ID. Import flows can also record sex, age, and clinical subgroup. The CSV or TSV metadata manifest requires sample_id and vcf_filename ; optional values include sex, age, and subgroup.
Use pseudonymous identifiers and include only metadata necessary for the Study. Clinical subgroup labels should be defined consistently before analysis rather than inferred after inspecting results.
Keep Samples Comparable
Use the same reference genome and compatible contig naming across samples.
Avoid combining materially different capture designs, calling pipelines, or quality filters without accounting for the technical batch.
Check that sample IDs match the intended individuals and are not duplicated.
Define case, control, and clinical-subgroup membership before interpreting cohort differences.
Review per-sample processing and QC status before treating the matrix as complete.
Living-Cohort Behavior
Samples can be added while the Study is active, failed samples can be retried, and samples can be removed from the active cohort. Because carrier counts, allele frequencies, outlier statistics, and gene-level results depend on membership, rerun the analysis after a membership change before using the results clinically.
Removing a sample marks it as removed from the Study workflow. Deleting the entire Study is a separate operation and can be blocked while linked batch processing is still running.
Related Documentation
Study Quality Control
How sample metrics and cohort-relative outlier states are derived.
Results and Interpretation
What the current Study result views contain and how to read them.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Cohort Analysis](https://folklore.helena.bio/docs/cohort-analysis)
- [Study Quality Control How sample metrics and cohort-relative outlier states are derived.](https://folklore.helena.bio/docs/cohort-analysis/quality-control)
- [Results and Interpretation What the current Study result views contain and how to read them.](https://folklore.helena.bio/docs/cohort-analysis/results)

---

## Page: Study Quality Control | Folklore Documentation

Source: https://folklore.helena.bio/docs/cohort-analysis/quality-control
Canonical: https://folklore.helena.bio/docs/cohort-analysis/quality-control
Description: How Folklore reports per-sample cohort QC metrics, applies cohort-relative outlier flags, and distinguishes PASS, OUTLIER, and FAIL states.

Documentation / Cohort Analysis / Quality Control
## Study Quality Control
Folklore computes per-sample summary metrics from each completed Variant Analysis result and compares selected metrics with the other successfully read samples in the same Study. The purpose is to identify technical or compositional differences that require review before cohort-level interpretation.
Reported Metrics
Metric What Is Reported Used for Outlier Flag
Variant count Total classified variant rows for the sample. Yes
Ti/Tv ratio The ratio of single-nucleotide transitions to transversions. Yes
Het/Hom ratio The ratio of heterozygous to homozygous-alternate calls. Yes
Mean depth Mean recorded read depth across variants with a depth value. No; displayed for review
Current Outlier Rule
A successfully read sample is marked OUTLIER when its variant count, Ti/Tv ratio, or Het/Hom ratio is more than 2 sample standard deviations from the Study mean. A deviation in any one of the three metrics is sufficient. Mean depth is reported but is not part of the current automated outlier rule.
QC Status Meanings
PASS
The sample result was readable and none of the three automated cohort-relative metrics exceeded the current outlier boundary.
OUTLIER
The result was readable, but at least one automated metric differed from the Study mean by more than the current boundary. The recorded reason identifies the metric or metrics.
FAIL
The expected classified result was missing or could not be read. This is a processing or data-availability failure, not a statistical outlier.
How to Review an Outlier
Confirm the sample identity, sequencing type, reference build, and expected capture design.
Compare raw coverage and calling QC with the laboratory source data, especially when mean depth or variant count differs.
Check whether ancestry, consanguinity, sex-chromosome content, or a genuine biological feature could explain the difference.
Review whether the Study combines different laboratories, callers, panels, or processing versions.
Resolve failed or unintended samples and rerun the Study before relying on cohort-level results.
Interpretation Limits
These are cohort-relative checks. PASS does not prove that a sample is uncontaminated, correctly labeled, or technically equivalent to every other sample. OUTLIER does not prove an error. In a small or heterogeneous Study, the mean and standard deviation can be unstable or reflect the cohort composition itself.
Principal-component batch or ancestry detection is not part of the current production QC result. The present QC should therefore be combined with laboratory-level QC, provenance review, and an appropriate Study design.
Related Documentation
Creating a Study
How to make sample inputs and Study membership comparable.
Results and Interpretation
How QC context affects cohort-level variant and gene evidence.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Cohort Analysis](https://folklore.helena.bio/docs/cohort-analysis)
- [Creating a Study How to make sample inputs and Study membership comparable.](https://folklore.helena.bio/docs/cohort-analysis/creating-a-study)
- [Results and Interpretation How QC context affects cohort-level variant and gene evidence.](https://folklore.helena.bio/docs/cohort-analysis/results)

---

## Page: Study Results and Interpretation | Folklore Documentation

Source: https://folklore.helena.bio/docs/cohort-analysis/results
Canonical: https://folklore.helena.bio/docs/cohort-analysis/results
Description: How to interpret the current Folklore Study result views, including actionable variants, the cohort matrix, burden, pLoF, frequency, compound-heterozygous candidates, and candidate genes.

Documentation / Cohort Analysis / Results and Interpretation
## Study Results and Interpretation
A completed Study combines individual classifications with cohort-level genotype, carrier, frequency, and gene evidence. Start by confirming sample membership and QC, then move from the variant matrix to the derived gene-level views.
Results Belong to an Analysis Run
Burden, pLoF, compound-heterozygous, pathway, and candidate-gene results are tied to a Study analysis run. If sample membership or the gene scope changes, run the analysis again and interpret the current run rather than combining values from different cohort states.
Current Result Views
Actionable Findings
Pathogenic and Likely Pathogenic variants from the cohort catalog, with the samples and genotypes in which each variant occurs.
Interpretation boundary: Actionable in this view means classification-based prioritization. It does not establish phenotype fit, penetrance, or a final reportable finding.
Variant Matrix
A deduplicated variant catalog linked to sparse sample-level genotypes, cohort carrier and homozygous counts, cohort allele frequency, and classification-discordance flags.
Interpretation boundary: A cohort summary class does not replace the individual sample classification or its evidence trace.
Gene Burden
Carrier-based rare-variant aggregation by gene, compared with an expected population carrier count and accompanied by corrected statistical results and power context.
Interpretation boundary: Results are sensitive to qualifying-variant rules, population comparability, sample size, and technical batch.
Predicted Loss of Function
Gene summaries for frameshift, stop-gained, splice-donor, and splice-acceptor variants, including carriers and available constraint or clinical context.
Interpretation boundary: A predicted loss-of-function consequence is not automatically disease-causing; transcript relevance and gene mechanism require review.
Frequency Comparison
Variants ordered by the absolute difference between Study allele frequency and the available global population frequency.
Interpretation boundary: The displayed comparison is descriptive and should not be read as an ancestry-adjusted association test.
Compound Heterozygotes
Within-gene pairs observed in the same sample for genes in the Study scope.
Interpretation boundary: The pair is a candidate configuration. Phase, variant significance, gene mechanism, and phenotype fit remain separate questions.
Candidate Genes
A ranked synthesis of available Study evidence, with component evidence retained for review.
Interpretation boundary: Ranking supports triage. Exact internal weighting is proprietary and the score is not a calibrated probability of causality.
Cohort Classification Consensus
When the same variant has more than one classification across samples, the cohort catalog uses the most severe observed class in the order Pathogenic, Likely Pathogenic, Uncertain Significance, Likely Benign, and Benign. A discordance flag is stored when more than one non-null class is present.
This is a display and filtering convention, not a reclassification rule. Review the individual analyses, evidence versions, phenotype context, and any manual reclassification before resolving the difference.
Gene Burden Statistics
The current engine collapses qualifying variants into a carrier state for each gene. It reports a two-sided Fisher exact comparison against the expected population carrier count and a CMC carrier-collapsing result. With the current binary-carrier implementation, the CMC p-value is the same as the Fisher p-value.
Folklore applies Benjamini-Hochberg false-discovery-rate correction across tested genes and also reports a Bonferroni-adjusted value. The default significance flag uses an FDR threshold of 0.05. Minimum detectable odds-ratio context is provided to show when the Study is underpowered for modest effects.
SKAT-O is not currently computed or reported. A missing SKAT-O value must not be interpreted as a negative result.
Pathway Results
The interface can display pathway-enrichment results when a completed analysis run contains them. The current automated pipeline does not load pathway definitions by default, so an absent pathway section normally means that no pathway result was generated, not that every pathway tested negative.
Not Current Study Outputs
Genome-wide association study results are not generated by the current production Study pipeline.
Polygenic risk scores are not generated by the current production Study pipeline.
SKAT-O is not a current burden result.
Principal-component ancestry or batch-effect results are not part of the current production QC view.
Interpretation Checklist
Confirm that the intended samples and latest completed run are selected.
Review FAIL and OUTLIER samples before interpreting carrier or frequency values.
Check reference build, pipeline version, capture design, ancestry, relatedness, and case-control definition.
Inspect qualifying variants and individual classifications behind every gene-level signal.
Use corrected statistical values and the reported power context; do not rank genes by raw p-value alone.
Validate clinically important findings with the appropriate laboratory and clinical review process.
Clinical and Statistical Boundary
Enrichment, frequency difference, and candidate ranking identify evidence worth reviewing. They do not prove causality, exclude confounding, or establish a diagnosis. Final interpretation belongs to qualified professionals using the complete clinical and laboratory context.
Related Documentation
Study Quality Control
Understand the sample-level context behind cohort results.
Variant Analysis
Review how individual variant classifications and evidence traces are produced.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Cohort Analysis](https://folklore.helena.bio/docs/cohort-analysis)
- [Study Quality Control Understand the sample-level context behind cohort results.](https://folklore.helena.bio/docs/cohort-analysis/quality-control)
- [Variant Analysis Review how individual variant classifications and evidence traces are produced.](https://folklore.helena.bio/docs/classification)

---

## Page: Structural Variants and CNVs | Folklore Documentation

Source: https://folklore.helena.bio/docs/structural-variants
Canonical: https://folklore.helena.bio/docs/structural-variants
Description: How Folklore normalizes structural-variant calls, assigns an interpretation framework, annotates copy-number evidence, and presents reviewable results.

Documentation / Structural Variants
## Structural Variants and CNVs
Structural-variant analysis is part of Folklore's production Variant Analysis pipeline. It normalizes compatible VCF calls, gathers structural and copy-number evidence, and assigns each supported record to exactly one interpretation framework.
The workflow is caller-neutral: interpretation depends on the normalized event and available evidence, not on the name of the upstream calling tool. Missing fields remain visible as evidence limitations rather than being inferred without support.
Interpretation, Not Variant Calling
Folklore starts from calls already present in the submitted VCF. Read alignment, structural-variant discovery, breakpoint assembly, and upstream caller validation remain the responsibility of the laboratory workflow.
Production Workflow
1
Read the submitted call
Folklore consumes compatible structural-variant records from the uploaded VCF. It does not call structural variants from sequencing reads.
2
Normalize the event
Recognized calls are represented consistently as deletion, duplication, inversion, insertion, breakend, or copy-number variant, with genomic span and copy state when supplied.
3
Annotate clinical context
The pipeline evaluates gene and exon overlap, dosage-sensitive regions, predicted dosage sensitivity, population overlap, and available call-quality evidence.
4
Assign one framework
Constitutional copy-number loss or gain is assigned to the Riggs 2020 framework. Other eligible nuclear events are handled by the nuclear Variant Analysis framework.
5
Present an auditable result
The result preserves the owning framework, classification, contributing evidence, genomic content, population context, and limitations for clinical review.
What the Result Can Show
Event summary
Type, chromosome, start and end coordinates, span, copy state, zygosity, and precision when available.
Classification
A five-tier result and the framework that produced it when the record is eligible for automated interpretation.
Genomic impact
Overlapping genes, exons, dosage-sensitive regions, and per-gene dosage evidence.
Population context
Type-compatible overlap against local structural-variant population references.
Call evidence
Available precision, confidence interval, paired-read, split-read, and related quality fields.
Audit trail
Human-readable evidence criteria and a machine-readable breakdown for supported CNV classifications.
Clinical Boundary
Structural-variant classification is clinical decision support. Breakpoint uncertainty, assay coverage, mosaicism, inheritance, phenotype fit, and orthogonal confirmation can materially change interpretation and require qualified review.
In This Section
Input and Scope
Accepted records, normalization, framework ownership, and unsupported cases.
Interpretation and Limitations
Evidence categories, result provenance, missing-data behavior, and clinical review.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Input and Scope Accepted records, normalization, framework ownership, and unsupported cases.](https://folklore.helena.bio/docs/structural-variants/input-and-scope)
- [Interpretation and Limitations Evidence categories, result provenance, missing-data behavior, and clinical review.](https://folklore.helena.bio/docs/structural-variants/interpretation)

---

## Page: Structural Variant Input and Scope | Folklore Documentation

Source: https://folklore.helena.bio/docs/structural-variants/input-and-scope
Canonical: https://folklore.helena.bio/docs/structural-variants/input-and-scope
Description: The structural-variant records Folklore accepts, how calls are normalized, and which clinical framework owns each supported event.

Documentation / Structural Variants / Input and Scope
## Input and Scope
The current production workflow reads structural-variant and copy-number records from a compatible VCF. It recognizes the normalized event types DEL, DUP, INV, INS, BND, and CNV from standard VCF fields or symbolic alleles.
Sequence-resolved deletions and insertions can also enter structural-variant processing when their reference-versus-alternate length difference is at least 50 base pairs. Smaller sequence changes remain in the small-variant path.
Fields Used When Available
Event type and alternate allele representation.
Start, end, and structural-variant length. Length is represented as a magnitude so deletion sign conventions do not change interpretation.
Copy number, genotype, and derived zygosity.
Breakpoint precision and confidence intervals.
Paired-read, split-read, and other call-support values supplied by the VCF.
Framework Ownership
Every supported record is assigned to one interpretation owner. This prevents the same event from receiving competing automated verdicts from multiple frameworks.
Copy-number loss
When: Deletion or a reported copy state below the reference state
Result: ClinGen/ACMG Riggs 2020 loss framework
Copy-number gain
When: Duplication or a reported copy state above the reference state, after loss has been excluded
Result: ClinGen/ACMG Riggs 2020 gain framework
Other eligible nuclear SV
When: Non-mitochondrial, non-reference event with an annotated gene and no copy-number loss or gain ownership
Result: Nuclear ACMG/AMP Variant Analysis framework
Unsupported record
When: No safe framework owner can be established from the available fields
Result: No automated verdict; the reason remains recorded for review
Caller-Neutral Processing
Folklore does not select interpretation rules by caller brand. Different callers may populate different VCF fields, so missing optional evidence can reduce what the platform can establish, but it does not create a caller-specific classification path.
Current Boundaries
VCF is the production input for this workflow; TSV and BEDPE are not current Folklore analysis inputs.
For a multi-allelic record, structural normalization follows the first alternate allele. Submit unambiguous, normalized records for clinical use.
The analysis uses the first sample represented by the primary single-sample Variant Analysis workflow.
A missing end coordinate can limit interval-based evidence, particularly for insertions and breakends.
A missing gene annotation, a hom-ref genotype, or an unsupported mitochondrial structural event can prevent automated framework assignment.
Repeat expansions and somatic tumor events are outside this germline interpretation workflow.
Continue Reading
Interpretation and Limitations
How structural and copy-number evidence is presented and where clinical review remains essential.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Structural Variants](https://folklore.helena.bio/docs/structural-variants)
- [Interpretation and Limitations How structural and copy-number evidence is presented and where clinical review remains essential.](https://folklore.helena.bio/docs/structural-variants/interpretation)

---

## Page: Structural Variant Interpretation and Limitations | Folklore Documentation

Source: https://folklore.helena.bio/docs/structural-variants/interpretation
Canonical: https://folklore.helena.bio/docs/structural-variants/interpretation
Description: How Folklore presents CNV and structural-variant evidence, classification provenance, missing data, and clinical limitations.

Documentation / Structural Variants / Interpretation
## Interpretation and Limitations
Folklore separates event normalization, evidence annotation, framework ownership, and verdict assembly. This makes it possible to distinguish what the VCF reported, what reference data contributed, and which interpretation framework produced the displayed result.
Evidence Reviewed
Genomic content
Event span, affected genes and exons, and whether breakpoints or the full interval overlap an annotated gene region.
Dosage sensitivity
Curated ClinGen haploinsufficiency and triplosensitivity regions, plus locally available predicted dosage-sensitivity evidence.
Population evidence
Type-compatible structural-variant overlap against local gnomAD SV/CNV references. A true absence is kept distinct from a record that could not be annotated.
Call confidence
Precision status, breakpoint confidence intervals, paired-read and split-read support, and related fields when supplied upstream.
Inheritance and phenotype
Relevant only when the submitted data and current schema provide sufficient evidence. Missing segregation or phenotype specificity is not silently assumed.
Copy-Number Classification
Constitutional copy-number losses and gains are evaluated with the applicable ClinGen/ACMG Riggs 2020 evidence framework. When sufficient evidence is available, the result uses the standard five tiers: Pathogenic, Likely Pathogenic, Variant of Uncertain Significance, Likely Benign, or Benign.
The result view identifies the applied criteria, whether each reviewed criterion contributed, the total used by the framework, and the final tier. Folklore exposes this audit trail without publishing implementation-specific optimization or internal scoring code.
A Structural Type Does Not Guarantee a Riggs Verdict
Riggs 2020 ownership is reserved for constitutional copy-number loss and gain. Other eligible nuclear structural events may enter the nuclear Variant Analysis framework, while records without a safe owner retain no automated verdict and a reviewable reason.
Missing Evidence and Safe Under-Calling
Unavailable evidence does not receive a favorable or pathogenic contribution by default.
Incomplete breakpoint geometry can prevent reliable gene, exon, or region overlap assessment.
Insertions and breakends without an end coordinate have limited interval-based annotation.
The current structural schema cannot establish every inheritance, phenotype-specificity, or transcript-consequence condition described by the clinical standard.
Some intragenic events require exon-level and transcript-level review beyond the automated representation before stronger loss-of-function evidence can be assigned.
Family segregation and de novo status are not automatically inferred for structural variants by this module.
Clinical Review Checklist
Confirm the event and breakpoint uncertainty with the upstream caller and assay quality metrics.
Review affected transcripts, genes, regulatory regions, and the relevance of partial versus complete overlap.
Evaluate inheritance, segregation, phenotype fit, penetrance, and mosaicism outside the automated score where necessary.
Consider orthogonal confirmation before clinical reporting, particularly for imprecise or borderline calls.
Verify the recorded interpretation framework and distinguish an unclassified record from a benign result.
Related Page
Input and Scope
Recognized VCF records, normalization rules, framework ownership, and unsupported cases.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Structural Variants](https://folklore.helena.bio/docs/structural-variants)
- [Input and Scope Recognized VCF records, normalization rules, framework ownership, and unsupported cases.](https://folklore.helena.bio/docs/structural-variants/input-and-scope)

---

## Page: Data and Privacy | Folklore Documentation

Source: https://folklore.helena.bio/docs/data-and-privacy
Canonical: https://folklore.helena.bio/docs/data-and-privacy
Description: How Folklore protects genomic data: EU infrastructure, GDPR compliance, data retention, and zero external calls.

## Data and Privacy
Genomic data is the most sensitive category of personal data under EU law. It is immutable, uniquely identifying, and carries implications for the data subject and their biological relatives. Folklore is designed from the ground up to process this data with clinical-grade security and full GDPR compliance.
The platform operates on dedicated hardware in the European Union, makes zero external API calls during variant processing, and provides transparent data retention with automatic deletion. Every access, modification, and analysis is tracked and auditable.
Key Principles
EU-Only Processing
All genomic data is processed and stored on dedicated hardware in Helsinki, Finland. Data never leaves the European Union at any processing stage. No cloud services with non-EU jurisdiction are used for variant data.
Data Minimization
The platform processes only the genomic data necessary for analysis. VCF files are received in pseudonymized form -- sample identifiers only, no patient names, dates of birth, or national identification numbers.
Zero External Calls
During variant processing, the platform makes zero outbound network calls. All reference databases, annotation tools, and the literature database run locally. No patient data or query parameters are sent to any external service.
Transparent Retention
Uploaded VCF files are deleted after processing completes. Analysis results are retained for the duration specified in the service agreement. All data is deletable on request per GDPR Article 17.
Controller/Processor Separation
The laboratory is the data controller. Helena Bioinformatics acts as data processor under a signed Data Processing Agreement (DPA) that defines responsibilities, retention periods, and breach notification procedures.
Compliance Documents
The following legal documents are available on our website:
Privacy Policy How we collect, use, store, and protect personal and genomic data. Data Processing Agreement (DPA) Standard DPA for laboratory partners, pursuant to GDPR Article 28. Data Protection Impact Assessment (DPIA) Risk assessment for high-risk processing of genetic data, per GDPR Article 35.
In This Section
Infrastructure
Dedicated EU hardware, Helsinki datacenter, no multi-tenant cloud.
GDPR Compliance
Article 9 special category data, controller/processor roles, data subject rights.
Data Retention
What data is kept, for how long, and how deletion works.
No External Calls
Zero outbound network calls during variant processing.

### Links and cited sources

- [Privacy Policy](https://folklore.helena.bio/privacy)
- [Data Processing Agreement (DPA)](https://folklore.helena.bio/dpa)
- [Data Protection Impact Assessment (DPIA)](https://folklore.helena.bio/dpia)
- [Infrastructure Dedicated EU hardware, Helsinki datacenter, no multi-tenant cloud.](https://folklore.helena.bio/docs/data-and-privacy/infrastructure)
- [GDPR Compliance Article 9 special category data, controller/processor roles, data subject rights.](https://folklore.helena.bio/docs/data-and-privacy/gdpr-compliance)
- [Data Retention What data is kept, for how long, and how deletion works.](https://folklore.helena.bio/docs/data-and-privacy/data-retention)
- [No External Calls Zero outbound network calls during variant processing.](https://folklore.helena.bio/docs/data-and-privacy/no-external-calls)

---

## Page: Data Retention | Folklore Documentation

Source: https://folklore.helena.bio/docs/data-and-privacy/data-retention
Canonical: https://folklore.helena.bio/docs/data-and-privacy/data-retention
Description: What data Folklore retains, for how long, and how deletion works.

Documentation / Data and Privacy / Data Retention
## Data Retention
Folklore follows the principle of data minimization: we process only what is necessary and retain data only for as long as required. This page describes what data is kept, for how long, and how deletion works.
Retention by Data Type
Uploaded VCF Files Retained for re-analysis capability
The parsed variant data from the uploaded VCF file is retained on the server to enable re-processing without requiring the laboratory to upload the file again. This supports iterative analysis with updated parameters, new HPO terms, or updated reference databases. Data is deleted when the laboratory deletes the analysis session or upon termination of the service agreement.
Analysis Results Duration specified in service agreement
Classified variants, ACMG evidence, phenotype match scores, screening tiers, and literature evidence are retained as a structured database (DuckDB) for the duration agreed with the laboratory. This allows geneticists to return to previous cases for review without re-processing.
Clinical Profile Data Retained alongside analysis results
Patient phenotype (HPO terms), demographics (age, sex, ethnicity when provided), and clinical context are stored alongside the analysis results. These are deleted when the analysis session is deleted.
Literature Search Results Retained alongside analysis results
Ranked literature publications with relevance scores and evidence categories are stored as part of the session output. Deleted with the session.
Session Metadata Retained for audit purposes
Processing timestamps, pipeline parameters, database versions used, and quality metrics are retained for audit trail compliance. These contain no genomic data.
Account Data Duration of the service agreement
User name, email, organization, and role are retained while the account is active. Deleted within 30 days of account termination or on request.
Usage Data Maximum 12 months
IP addresses, session logs, and page views are retained for security monitoring for up to 12 months, then automatically purged.
Deletion Mechanisms
Automatic Deletion
VCF files are automatically deleted after processing completes. Analysis sessions that exceed the agreed retention period are automatically purged. No manual intervention is required.
On-Demand Deletion
The laboratory (data controller) can request deletion of any analysis session at any time. Helena Bioinformatics will delete all associated data (results, clinical profile, literature evidence, metadata) within 30 days and certify the deletion in writing.
Data Subject Erasure (Article 17)
Data subjects (patients) can exercise their right to erasure through the data controller (laboratory). Helena Bioinformatics assists in fulfilling these requests. Deletion may be subject to legal retention requirements (e.g., clinical record-keeping obligations under national legislation).
Termination Deletion
Upon termination of the service agreement, all personal data is either returned to the data controller or securely deleted within 30 days, at the controller's election. Deletion is certified in writing per the DPA.
What Is NOT Retained
Patient names, dates of birth, or national identification numbers (never received)
Original VCF files after processing (automatically deleted)
Raw sequencing data (FASTQ, BAM -- not accepted by the platform)
Intermediate processing files (temporary and deleted after each pipeline stage)
Audit Trail
Every deletion event is logged in the audit trail: what was deleted, when, by whom (or automatically), and the reason. Audit logs themselves are retained for the minimum period required by applicable legislation and do not contain genomic data.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Data and Privacy](https://folklore.helena.bio/docs/data-and-privacy)

---

## Page: GDPR Compliance | Folklore Documentation

Source: https://folklore.helena.bio/docs/data-and-privacy/gdpr-compliance
Canonical: https://folklore.helena.bio/docs/data-and-privacy/gdpr-compliance
Description: How Folklore complies with GDPR for processing genetic data as special category data.

Documentation / Data and Privacy / GDPR Compliance
## GDPR Compliance
Genetic data is classified as special category data under GDPR Article 9, requiring additional legal safeguards beyond those for ordinary personal data. Folklore is designed to meet these requirements at every level -- legal, organizational, and technical.
Controller and Processor Roles
The distinction between data controller and data processor is fundamental to how genomic data flows through the platform:
Laboratory (Data Controller)
The clinical genetics laboratory determines the purpose and means of processing. The laboratory is responsible for: ensuring a lawful basis for processing the genomic data (typically explicit consent under Article 9(2)(a) or healthcare provision under Article 9(2)(h)), obtaining any necessary patient consents, providing only pseudonymized data (sample IDs, no patient names or identification numbers), and deciding on data retention periods.
Helena Bioinformatics (Data Processor)
Helena Bioinformatics processes genomic data exclusively on documented instructions from the data controller. As processor, we: process data only for the purposes specified in the Data Processing Agreement, implement appropriate technical and organizational security measures, do not engage sub-processors without prior written authorization, assist the controller in responding to data subject requests, and delete or return all data upon termination of the agreement.
Legal Basis for Processing
Genomic Data (Special Category)
Article 9(2)(a): explicit consent of the data subject, or Article 9(2)(h): processing necessary for healthcare provision. The applicable legal basis is determined by the data controller (laboratory), not by Helena Bioinformatics. We process this data under contractual obligation (Article 6(1)(b)) and the DPA.
Account Data
Article 6(1)(b): contractual necessity. Name, email, organization, and role are collected during registration to provide the service.
Usage Data
Article 6(1)(f): legitimate interest. IP addresses, session duration, and pages visited are collected for security monitoring and service improvement.
Data Subject Rights
Data subjects (patients whose genomic data is processed) retain full GDPR rights. Because the laboratory is the data controller, rights requests are typically routed through the laboratory. Helena Bioinformatics assists in fulfilling these requests:
Access (Article 15) Confirmation of processing and a copy of the data being processed. Rectification (Article 16) Correction of inaccurate data. Erasure (Article 17) Deletion of genomic data, subject to legal retention requirements. Restriction (Article 18) Limitation of processing while a dispute is resolved. Portability (Article 20) Export of data in machine-readable format (VCF, PDF). Objection (Article 21) Right to object to specific processing activities.
Data Protection Impact Assessment
A DPIA is mandatory under GDPR Article 35 when processing genetic data on a large scale. Helena Bioinformatics has conducted a comprehensive DPIA covering: the nature and scope of processing, necessity and proportionality assessment, risk identification (unauthorized access, data breach, re-identification), and mitigation measures (encryption, physical isolation, access controls, audit logging, automatic deletion).
The full DPIA summary is available at Data Protection Impact Assessment .
International Transfers
No genomic data or personal data is transferred outside the European Economic Area. All processing occurs on infrastructure located in Finland (EU). In the event that a transfer outside the EEA becomes necessary in the future, Helena Bioinformatics will obtain prior written consent from the data controller and implement appropriate safeguards under GDPR Chapter V (Standard Contractual Clauses or adequacy decision).
Contact
For questions regarding data protection, GDPR compliance, or to exercise data subject rights, contact privacy@helena.bio. The full Privacy Policy and Data Processing Agreement are available on our website.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Data and Privacy](https://folklore.helena.bio/docs/data-and-privacy)
- [Data Protection Impact Assessment](https://folklore.helena.bio/dpia)
- [Privacy Policy](https://folklore.helena.bio/privacy)
- [Data Processing Agreement](https://folklore.helena.bio/dpa)

---

## Page: Infrastructure | Folklore Documentation

Source: https://folklore.helena.bio/docs/data-and-privacy/infrastructure
Canonical: https://folklore.helena.bio/docs/data-and-privacy/infrastructure
Description: Where Folklore processes and stores genomic data: dedicated EU hardware in Helsinki, Finland.

Documentation / Data and Privacy / Infrastructure
## Infrastructure
All genomic data processing occurs on dedicated hardware located in the European Union. The infrastructure is purpose-built for clinical genomics workloads, not shared multi-tenant cloud services.
Overview
Type Dedicated bare-metal server (not shared cloud) Location Helsinki datacenter, Finland (EU) Jurisdiction Finnish and EU data protection law Provider European hosting provider, subject to EU law only Capacity Enterprise-grade server optimized for genomics workloads (multi-core, high memory, NVMe storage)
Why Dedicated Hardware
Multi-tenant cloud providers (AWS, GCP, Azure) share physical infrastructure across customers. Even with logical isolation, genetic data processing on shared hardware introduces risks that dedicated servers eliminate:
Physical isolation
No other customer’s workloads run on the same hardware. There is no risk of side-channel attacks, noisy neighbor performance degradation, or accidental data exposure through shared resources.
Jurisdiction certainty
The server is physically located in Helsinki, Finland. Unlike cloud providers that may move workloads between regions, the physical location of the data is fixed and verifiable.
No US jurisdiction exposure
Major cloud providers are subject to US laws (CLOUD Act, FISA) that can compel disclosure of data stored on their infrastructure regardless of physical location. Our hosting provider is a European company subject to EU law only.
Full administrative control
Helena Bioinformatics has exclusive root-level access to the server. No hosting provider employee has access to the operating system, storage, or network configuration.
Security Measures
TLS 1.3 encryption for all data in transit
AES-256 encryption for data at rest
Network firewall with restrictive inbound/outbound rules
Role-based access control (RBAC) for all platform functions
Comprehensive audit logging of all data access and processing activities
No outbound network access from the variant processing pipeline
Regular security assessments and vulnerability scanning
Automated intrusion detection and alerting
Data Path
When a laboratory uploads a VCF file, the data travels over TLS 1.3 directly to the Helsinki server. The file is parsed, annotated, classified, and scored entirely on this server. Results are stored on the same server. At no point does the data transit through non-EU infrastructure or third-party services.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Data and Privacy](https://folklore.helena.bio/docs/data-and-privacy)

---

## Page: No External Calls | Folklore Documentation

Source: https://folklore.helena.bio/docs/data-and-privacy/no-external-calls
Canonical: https://folklore.helena.bio/docs/data-and-privacy/no-external-calls
Description: How Folklore processes genomic data with zero outbound network calls.

Documentation / Data and Privacy / No External Calls
## No External Calls
During variant processing, Folklore makes zero outbound network calls. No patient data, genomic coordinates, variant identifiers, or query parameters are sent to any external service at any point in the analysis pipeline. This is a fundamental architectural decision, not a configuration option.
What Runs Locally
Every component of the variant analysis pipeline operates on local data stored on the Helsinki server. No external API calls are made during processing:
Ensembl VEP
Variant Effect Predictor runs locally with a local offline cache (Release 113, GRCh38). No calls to the Ensembl REST API. Consequence annotation, transcript selection, and impact classification all run on local data.
gnomAD
Population frequency database (v4.1.0, 759 million variants) is stored locally. Allele frequencies, homozygote counts, and population-specific data are queried from local storage.
ClinVar
Clinical significance database (2025-01, 4.1 million variants) is stored locally. Clinical significance, review status, and star ratings are queried without any connection to NCBI.
dbNSFP
Functional prediction scores (4.9c, 80.6 million sites) -- SIFT, AlphaMissense, MetaSVM, DANN, BayesDel, PhyloP, GERP -- all stored and queried locally.
HPO Ontology
Human Phenotype Ontology with 17,000+ terms and 320,000+ gene-phenotype associations. Phenotype matching and semantic similarity computed locally using the pyhpo library.
ClinGen
Dosage sensitivity data (1,600+ genes) for haploinsufficiency and triplosensitivity stored locally.
SpliceAI Precomputed
Splice impact delta scores for MANE transcripts stored locally. No calls to the Illumina API or SpliceAI web interface.
Literature Database
Local PubMed mirror with 2-3 million genetics-relevant publications. Literature search, relevance scoring, and evidence assessment all run against local data. No calls to NCBI PubMed.
Network Architecture
The variant processing pipeline has no outbound network access by design. Network architecture enforces this at the infrastructure level:
Inbound only
The platform accepts VCF uploads and API requests from authenticated users over TLS 1.3. These are the only inbound connections.
No outbound from processing
The variant analysis, phenotype matching, screening, and literature search services have no outbound network routes. They cannot make HTTP requests, DNS queries, or any other network calls to external services.
Separate update channel
Reference database updates (ClinVar quarterly, PubMed daily) are fetched by a separate background process that does not have access to patient data. The update process downloads public data; it never sends any data outbound.
Why This Matters
No data leakage risk
If the processing pipeline cannot make network calls, it cannot leak data -- even in the event of a software vulnerability or misconfiguration.
No third-party dependency
Processing does not depend on external API availability. The platform operates at full capacity even if every external service is offline.
No third-party data access
No external organization receives genomic data, variant identifiers, or query parameters. There is no risk of external services logging, caching, or retaining patient-related data.
Verifiable by design
The network isolation is enforced at the infrastructure level (firewall rules, Docker network configuration), not at the application level. It can be independently verified through network audit.
AI Clinical Assistant
The AI Clinical Assistant is the one component that may use an external language model API. When this occurs, only anonymized variant data is transmitted -- genomic coordinates and classification results, never patient identifiers, sample IDs, or phenotype data. The AI assistant is an optional feature; all core analysis (classification, phenotype matching, screening, literature search) runs entirely locally without any external calls.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Data and Privacy](https://folklore.helena.bio/docs/data-and-privacy)

---

## Page: Reference Databases | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases
Canonical: https://folklore.helena.bio/docs/databases
Description: Folklore integrates 45 reference databases and curated sources across nuclear germline, phenotype, inheritance, mitochondrial, and structural-variant analysis.

## Reference Databases
Folklore integrates 45 reference databases and curated sources across variant annotation, classification, phenotype matching, screening, inheritance analysis, mitochondrial analysis, and SV/CNV interpretation. These external and curated sources are materialized into internal reference tables and derived views. Production reference data is stored locally on EU-based infrastructure in Helsinki, Finland.
Reference versions are pinned per deployment and recorded through the reference-data manifest. The table below lists selected core sources used by the nuclear germline workflow; it is not the complete inventory of all 45 integrations or of every internal reference table and derived view.
How the database counts are defined
The core Stage 4 enrichment currently loads 16 production classification and reference databases. The platform-wide figure of 45 counts all reference databases and curated sources used across nuclear classification, phenotype, inheritance, screening, mitochondrial, and SV/CNV modules. Smaller numbers in a subsystem, selected-source table, or historical changelog describe that narrower set; they are not alternative totals for the current platform.
Zero External API Calls
Core genomic processing uses locally stored reference data and a local Ensembl VEP cache. Patient variants are not submitted to public genomic databases or third-party annotation APIs during the analysis path.
Database Summary
Database Primary Use ACMG Criteria
gnomAD Population frequencies BA1, BS1, BS2, PM2
ClinVar Clinical significance PS1, PP5, BP6, ClinVar override
dbNSFP Functional predictions PP3, BP4 (BayesDel_noAF)
SpliceAI Splice impact PP3_splice, BP7 guard
gnomAD Constraint Gene-level tolerance PVS1, PP2, BP1
HPO Gene-phenotype mapping PP4
ClinGen Dosage sensitivity BS1, BP2
Orphanet Gene-disease-inheritance BS1 (AD/AR/XLD), PVS1 gate
VCEP Specs Gene-specific thresholds BA1, BS1, PM2, PVS1 gate
ClinGen GDV Gene-disease validity PVS1 constraint gate, disease association gate
DECIPHER G2P Molecular mechanism curation PVS1 GoF/DN guard, disease association gate
GoFCards Gain-of-function curation PVS1 GoF guard
Ensembl VEP Variant effect prediction PVS1, PM1, PM4, BP1, BP3, BP7
Annotation Pipeline
Reference data is loaded into each variant record during Stage 4 of the processing pipeline. After annotation, every variant carries all reference columns directly -- no database lookups are needed during classification or clinical review. The annotation order is:
1
gnomAD v4.1
Population allele frequencies. Positional match on chromosome, position, reference allele, and alternate allele. Loads 6 columns.
2
ClinVar
Clinical significance assertions. Same positional match. Loads 7 columns including review stars and disease associations.
3
dbNSFP 4.9c
Functional predictions from SIFT, AlphaMissense, MetaSVM, DANN, BayesDel, and conservation scores. Loads 9 columns with duplicate variant aggregation.
4
gnomAD Constraint
Gene-level tolerance metrics. Joined on gene symbol. Loads 4 columns: pLI, LOEUF, o/e LoF, and missense Z-score.
5
HPO
Gene-phenotype associations. Joined on gene symbol with deduplication and aggregation. Loads 6 columns.
6
ClinGen
Dosage sensitivity scores. Joined on gene symbol. Loads 2 columns: haploinsufficiency and triplosensitivity.
Ensembl VEP runs as a separate stage (Stage 3) before database annotation, providing consequence predictions and transcript selection that the annotation phases then build upon. SpliceAI scores are accessed from precomputed data during VEP annotation.
In This Section
gnomAD
Population allele frequencies from 807,162 individuals across 8 genetic ancestry groups.
ClinVar
Clinical significance assertions from submitting laboratories worldwide.
dbNSFP
Functional predictions and conservation scores for all possible coding SNVs.
HPO
Gene-phenotype associations from the Human Phenotype Ontology.
ClinGen
Gene dosage sensitivity curation from the Clinical Genome Resource.
Ensembl VEP
Variant Effect Predictor for consequence annotation and transcript selection.
SpliceAI Precomputed
Precomputed splice impact delta scores for all coding variants.
Update Policy
How and when reference databases are updated, validated, and versioned.
SV/CNV Reference Evidence
ClinGen Dosage, dosage regions, and population evidence used by the Riggs workflow.
mtDNA Reference Evidence
Mitochondrial frequencies, haplogroups, predictors, ClinVar assertions, and NUMT flags.
For details on how these databases are combined during ACMG classification, see the Criteria Reference .

### Links and cited sources

- [gnomAD Population allele frequencies from 807,162 individuals across 8 genetic ancestry groups.](https://folklore.helena.bio/docs/databases/gnomad)
- [ClinVar Clinical significance assertions from submitting laboratories worldwide.](https://folklore.helena.bio/docs/databases/clinvar)
- [dbNSFP Functional predictions and conservation scores for all possible coding SNVs.](https://folklore.helena.bio/docs/databases/dbnsfp)
- [HPO Gene-phenotype associations from the Human Phenotype Ontology.](https://folklore.helena.bio/docs/databases/hpo)
- [ClinGen Gene dosage sensitivity curation from the Clinical Genome Resource.](https://folklore.helena.bio/docs/databases/clingen)
- [Ensembl VEP Variant Effect Predictor for consequence annotation and transcript selection.](https://folklore.helena.bio/docs/databases/ensembl-vep)
- [SpliceAI Precomputed Precomputed splice impact delta scores for all coding variants.](https://folklore.helena.bio/docs/databases/spliceai-precomputed)
- [Update Policy How and when reference databases are updated, validated, and versioned.](https://folklore.helena.bio/docs/databases/database-update-policy)
- [SV/CNV Reference Evidence ClinGen Dosage, dosage regions, and population evidence used by the Riggs workflow.](https://folklore.helena.bio/methodology/sv)
- [mtDNA Reference Evidence Mitochondrial frequencies, haplogroups, predictors, ClinVar assertions, and NUMT flags.](https://folklore.helena.bio/methodology/mtdna)
- [Criteria Reference](https://folklore.helena.bio/docs/classification/criteria-reference)

---

## Page: ClinGen Dosage Sensitivity in Folklore | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/clingen
Canonical: https://folklore.helena.bio/docs/databases/clingen
Description: How Folklore uses ClinGen haploinsufficiency and triplosensitivity evidence in nuclear classification, screening, and constitutional SV/CNV interpretation.

Documentation / Reference Databases / ClinGen
## ClinGen
ClinGen provides expert-curated haploinsufficiency and triplosensitivity assessments. Folklore uses these records as gene-level evidence in nuclear classification and screening, and as dosage-sensitivity evidence in the constitutional SV/CNV workflow.
Database Details
Source clinicalgenome.org Producer ClinGen Consortium (NIH-funded)
Role in ACMG Classification
For nuclear SNV and indel classification, selected dosage fields contribute gene-level inheritance and mechanism context:
BS1 Strong Benign
Haploinsufficiency score = 3 is used as a proxy for autosomal dominant inheritance. When triggered, BS1 applies a stricter frequency threshold (AF >= 0.1%) compared to the default recessive threshold (AF >= 5%).
BP2 Supporting Benign
When a variant is a compound heterozygote candidate and ClinGen haploinsufficiency score = 30 (dosage sensitivity unlikely), BP2 applies. This combination suggests the variant is less likely to be pathogenic in a dominant context.
Columns Loaded (2)
ClinGen data is joined on gene symbol. Each variant inherits its gene-level dosage sensitivity scores.
haploinsufficiency_score INTEGER
Haploinsufficiency assessment score. Values 0-3 indicate evidence level for haploinsufficiency. Special values: 30 = autosomal recessive phenotype, 40 = dosage sensitivity unlikely.
triplosensitivity_score INTEGER
Triplosensitivity assessment score. The field is informational for the nuclear SNV/indel classifier and is used as dosage-sensitivity evidence by the dedicated SV/CNV workflow.
Structural variants and CNVs
The SV/CNV module evaluates constitutional deletions and duplications under the Riggs 2020 framework and uses ClinGen dosage regions and gene-level sensitivity evidence. See SV/CNV Methodology .
Haploinsufficiency Score Interpretation
Score Evidence Level Meaning Usage in Classification
3 Sufficient evidence Multiple independent studies demonstrate haploinsufficiency causes disease. Gene is intolerant to loss of one copy. Used as proxy for autosomal dominant inheritance in BS1 threshold selection.
2 Emerging evidence Some evidence suggests haploinsufficiency, but additional studies needed. Informational. Does not affect ACMG criteria thresholds.
1 Little evidence Limited or conflicting evidence for haploinsufficiency. Informational. Does not affect ACMG criteria thresholds.
0 No evidence No published evidence for haploinsufficiency. Does not affect ACMG criteria.
30 Gene associated with autosomal recessive phenotype Disease mechanism requires biallelic variants. AR proxy. Used to calibrate BS1 frequency thresholds and BP2 logic.
40 Dosage sensitivity unlikely Evidence suggests the gene tolerates copy number changes. Used in BP2: dosage sensitivity unlikely supports benign interpretation for compound heterozygotes.
Limitations
ClinGen covers approximately 1,600 genes. Variants in unassessed genes will have NULL dosage scores.
Dosage sensitivity is a gene-level property. It does not distinguish between different variant types or positions within the gene.
The haploinsufficiency score is used as an inheritance proxy, not a direct measure of variant pathogenicity.
Genes with score 0 (no evidence) are not the same as genes with negative evidence -- absence of evidence is not evidence of absence.
Reference
Rehm HL, et al. "ClinGen -- The Clinical Genome Resource." New England Journal of Medicine . 2015;372(23):2235-2242. PMID: 26014595.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)
- [SV/CNV Methodology](https://folklore.helena.bio/methodology/sv)

---

## Page: ClinVar Review Stars and Clinical Significance | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/clinvar
Canonical: https://folklore.helena.bio/docs/databases/clinvar
Description: ClinVar review stars, clinical significance assertions, PS1, PP5, BP6, classification override rules, and pathogenic-variant rescue in Folklore.

Documentation / Reference Databases / ClinVar
## ClinVar Review Stars and Clinical Significance
ClinVar records clinical significance assertions submitted by laboratories, research groups, expert panels, and practice guidelines. Folklore reads the assertion, review status, condition, and variation identifier separately; a label alone does not determine how much evidence the record can carry.
Database Details
Genome Build GRCh38 Source ncbi.nlm.nih.gov/clinvar Producer National Center for Biotechnology Information (NCBI)
Role in ACMG Classification
ClinVar serves two distinct functions in Folklore. First, it provides evidence criteria (PS1, PP5, BP6) based on prior clinical assertions. Second, it provides a classification override mechanism for variants with established clinical significance and no conflicting computational evidence.
PS1 Strong Pathogenic
Same amino acid change reported as Pathogenic or Likely Pathogenic with >= 2 review stars. The higher star threshold reflects greater confidence in the clinical assertion.
PP5 Supporting Pathogenic
Pathogenic or Likely Pathogenic with >= 1 review star but < 2 stars. Lower confidence tier than PS1. Retained for maximum sensitivity despite ClinGen SVI retirement recommendation.
BP6 Supporting Benign
Benign or Likely Benign with >= 1 review star. Retained for maximum sensitivity despite ClinGen SVI retirement recommendation.
Override Classification Priority 3
ClinVar classification is applied directly when no conflicting computational evidence exists and review stars >= 1. ClinVar VUS does not override computational classification. Override is subordinate to BA1 and high-confidence conflict checks.
ClinVar Pathogenic Rescue
During quality filtering (Stage 2), variants that fail quality thresholds are normally excluded from classification. However, variants with ClinVar Pathogenic or Likely Pathogenic status are rescued and retain their quality-pass flag regardless of quality metrics. This ensures that known pathogenic variants in low-coverage regions are never silently discarded.
ClinVar Review Stars Explained
ClinVar assigns review stars based on the level of evidence review and submitter agreement. Folklore uses these stars to calibrate the strength of ClinVar-derived evidence:
Stars Review Status Meaning Criteria
4 Practice guideline Classification from a recognized clinical practice guideline. PS1 eligible
3 Reviewed by expert panel Classified by a ClinGen Expert Panel (VCEP) with full evidence review. PS1 eligible
2 Criteria provided, multiple submitters, no conflicts Two or more submitters agree, using stated criteria, with no conflicting assertions. PS1 eligible
1 Criteria provided, single/conflicting submitters At least one submitter provided criteria, but there may be conflicts or only one submitter. PP5/BP6 eligible
0 No assertion criteria provided Classification submitted without supporting evidence criteria. Not used
Columns Loaded (7)
clinical_significance VARCHAR
Aggregate clinical significance: Pathogenic, Likely pathogenic, Uncertain significance, Likely benign, Benign, or Conflicting classifications.
review_status VARCHAR
Review status text describing the level of review. Maps to review stars.
review_stars INTEGER
Review confidence level from 0 to 4 stars. Determines whether PS1 (>=2 stars) or PP5 (>=1 star) is applied.
clinvar_variation_id INTEGER
Unique ClinVar variation identifier. Links to the ClinVar web entry for full submission details.
clinvar_rsid VARCHAR
dbSNP rsID associated with the ClinVar entry, if available.
disease_name VARCHAR
Condition name(s) associated with the clinical assertion. May contain multiple conditions separated by semicolons.
hgvsp VARCHAR
Protein-level HGVS notation from ClinVar. Used for PS1 matching (same amino acid change as established pathogenic).
Limitations
ClinVar submissions vary in quality. A 1-star submission without criteria may reflect older, less rigorous classification practices.
Conflicting classifications (e.g., one submitter says Pathogenic, another says Benign) are flagged but require manual review to resolve.
ClinVar updates monthly. The deployed version may not include the most recent submissions.
Some ClinVar entries reference GRCh37 coordinates that have been lifted over to GRCh38, which may introduce positional ambiguity for complex variants.
ClinVar VUS does not contribute evidence in either direction and does not override computational classification.
Reference
Landrum MJ, et al. "ClinVar: improvements to accessing data." Nucleic Acids Research . 2020;48(D1):D835-D844. PMID: 31777943.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)

---

## Page: Database Update Policy | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/database-update-policy
Canonical: https://folklore.helena.bio/docs/databases/database-update-policy
Description: How and when reference databases are updated, validated, and versioned in Folklore.

Documentation / Reference Databases / Update Policy
## Database Update Policy
Reference database updates affect variant classification. A variant classified as VUS today may be reclassified as Likely Pathogenic after a ClinVar update adds new clinical evidence. Folklore manages database updates through a controlled process that balances clinical currency with classification stability.
Update Principles
Versioned and Documented
Every reference update is recorded with its source version, load date, row count, provenance, and validation result in the internal reference-data manifest.
Validated Before Deployment
Each database update undergoes regression testing against a reference cohort of known pathogenic and benign variants. Classification changes are reviewed before production deployment.
Atomic Updates
Database updates are applied atomically. All analyses in progress complete with the previous version. New analyses use the updated version. No analysis ever mixes database versions.
Reproducible
Analysis results include the database versions used. Re-analysis of a case with the same database versions will produce identical results.
Update Schedule
Database Frequency Note
gnomAD Major release only Major releases add significant new samples or populations.
ClinVar Quarterly Quarterly cadence balances currency with validation effort.
dbNSFP Major release only Includes predictor algorithm updates.
SpliceAI With Ensembl release Tied to VEP cache version.
gnomAD Constraint With gnomAD Updated alongside population frequency data.
HPO Quarterly Phenotype annotations are actively curated.
ClinGen Quarterly Dosage curation is ongoing.
Ensembl VEP Annual Includes transcript model updates.
Validation Process
Before any database update reaches production, the following validation steps are performed:
1
Record Count Verification
Compare new version record counts against previous version. Unexpected drops in coverage trigger manual review.
2
Schema Compatibility
Verify that all columns and data types expected by the annotation pipeline are present in the new version.
3
Reference Cohort Regression
Run classification on a curated set of variants with known pathogenicity. Compare results against expected classifications.
4
Classification Delta Report
Generate a report of all classification changes between database versions. Review reclassifications for clinical appropriateness.
5
Deployment Approval
Classification delta report is reviewed. If changes are within acceptable limits, the update is approved for production deployment.
Impact on Existing Analyses
Database updates do not retroactively change existing analysis results. Completed analyses retain their original database versions and classifications. To benefit from updated databases, cases can be re-analyzed using the current database versions. The platform tracks which database versions were used for each analysis.
Version History
Database version changes are documented in the Changelog . Each entry includes the database name, old and new versions, number of classification changes in the reference cohort, and the deployment date.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)
- [Changelog](https://folklore.helena.bio/docs/changelog)

---

## Page: dbNSFP | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/dbnsfp
Canonical: https://folklore.helena.bio/docs/databases/dbnsfp
Description: dbNSFP 4.9c functional prediction database -- SIFT, AlphaMissense, MetaSVM, DANN, BayesDel, PhyloP, and GERP scores for missense variant interpretation.

Documentation / Reference Databases / dbNSFP
## dbNSFP
dbNSFP (database of Non-synonymous Functional Predictions) provides precomputed pathogenicity predictions and conservation scores for all possible non-synonymous single nucleotide variants in the human genome. It serves as the source for both BayesDel_noAF (used in ACMG PP3/BP4 criteria) and individual predictor scores used in screening prioritization.
Database Details
Total Fields 434 (9 loaded by Folklore) Genome Build GRCh38 Source sites.google.com/site/jpaborern/dbNSFP Producer Liu X, Li C, Mou C, Dong Y, Tu Y (USF Genomics)
Role in Classification and Screening
dbNSFP data serves two distinct roles in Folklore:
ACMG Classification
BayesDel_noAF (included in dbNSFP) is the primary computational predictor for PP3 and BP4 criteria. ClinGen SVI calibrated thresholds allow PP3 to reach Supporting, Moderate, or Strong evidence strength.
Screening Prioritization
Individual predictor scores (SIFT, AlphaMissense, MetaSVM, DANN) and conservation metrics (PhyloP, GERP) contribute to the weighted deleteriousness component of the screening score.
Duplicate Variant Handling
dbNSFP contains approximately 701,000 duplicate variant entries (0.87% of the dataset) due to multi-transcript annotation. Folklore resolves these by aggregating to the most pathogenic interpretation: MIN(sift_score) since lower SIFT is more pathogenic, MAX for all other scores, with predictions matched to their corresponding extreme scores.
Columns Loaded (9)
Variants are matched by exact positional coordinates. From the 434 available fields, Folklore loads 9 columns covering 4 functional predictors and 2 conservation metrics:
sift_pred VARCHAR
SIFT prediction: "D" (Deleterious) or "T" (Tolerated). Based on sequence homology across related proteins.
sift_score FLOAT
SIFT score (0-1). Lower values indicate higher probability of being deleterious. Threshold: < 0.05 = Deleterious.
alphamissense_pred VARCHAR
AlphaMissense prediction: "P" (Pathogenic), "A" (Ambiguous), or "B" (Benign). DeepMind protein structure-based.
alphamissense_score FLOAT
AlphaMissense score (0-1). Higher values indicate higher probability of pathogenicity.
metasvm_pred VARCHAR
MetaSVM prediction: "D" (Damaging) or "T" (Tolerated). Ensemble of 10 individual predictors combined with SVM.
metasvm_score FLOAT
MetaSVM score (continuous). Positive = damaging, negative = tolerated.
dann_score FLOAT
DANN score (0-1). Deep neural network pathogenicity score. Higher = more likely pathogenic. Can score any SNV.
phylop100way_vertebrate FLOAT
PhyloP conservation score across 100 vertebrate species. Positive = conserved, negative = fast-evolving, zero = neutral.
gerp_rs FLOAT
GERP++ rejected substitution score. Higher = more constrained. Scores > 2 suggest constraint, > 4 strong constraint.
Limitations
dbNSFP covers non-synonymous (missense) single nucleotide variants only. Indels, structural variants, and synonymous variants are not included.
Predictor scores may be NULL for variants not covered by specific prediction algorithms.
Different predictors use different training datasets, so their predictions are not fully independent.
Conservation scores reflect evolutionary constraint at a position, not the impact of a specific amino acid substitution.
BayesDel_noAF deliberately excludes allele frequency to avoid circular reasoning with PM2/BA1/BS1 criteria.
Reference
Liu X, et al. "dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs." Genome Medicine . 2020;12(1):103. PMID: 33261662.
For details on individual predictors, see the Computational Predictors section.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)

---

## Page: Ensembl VEP Consequences and Impact Categories | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/ensembl-vep
Canonical: https://folklore.helena.bio/docs/databases/ensembl-vep
Description: Ensembl VEP Release 113 consequence terms, HIGH, MODERATE, LOW, and MODIFIER impact categories, transcript selection, HGVS output, domains, and ACMG paths.

Documentation / Reference Databases / Ensembl VEP
## Ensembl VEP Consequences and Impact Categories
Ensembl VEP assigns Sequence Ontology consequences, transcript identifiers, HGVS expressions, impact categories, exon positions, and protein domains. Folklore uses those annotations to decide which ACMG criterion paths are eligible; VEP does not assign the final pathogenicity class.
Configuration
Genome Build GRCh38 Cache Local offline cache (no external API calls) Processing Parallelized across chromosomes (up to 48 workers) Source ensembl.org/vep Producer European Molecular Biology Laboratory (EMBL-EBI)
HIGH, MODERATE, LOW, and MODIFIER
VEP assigns each consequence a severity level that determines which ACMG criteria pathways are evaluated:
Impact Example Consequences ACMG Pathway
HIGH frameshift_variant, stop_gained, splice_acceptor_variant, splice_donor_variant PVS1 pathway
MODERATE missense_variant, inframe_insertion, inframe_deletion PP3/BP4 (BayesDel), PM4, BP3
LOW synonymous_variant, splice_region_variant BP7
MODIFIER intron_variant, upstream_gene_variant, downstream_gene_variant Typically no ACMG criteria triggered
Local Execution
VEP runs entirely locally using a pre-downloaded offline cache. No variant data is sent to Ensembl servers. The local FASTA reference file provides sequence context for HGVS notation generation. This ensures both data privacy and processing speed independence from network availability.
Fields Extracted (11)
For each variant, VEP produces annotations across all overlapping transcripts. Folklore selects the most severe consequence transcript and extracts the following fields:
gene_symbol
HGNC gene symbol from the most severe transcript annotation.
gene_id
Ensembl gene identifier (ENSG format).
transcript_id
Ensembl transcript identifier (ENST format) from the most severe consequence.
hgvs_genomic
Genomic HGVS notation (e.g., NC_000001.11:g.12345A>G).
hgvs_cdna
cDNA-level HGVS notation relative to the transcript (e.g., NM_000123.4:c.456A>G).
hgvs_protein
Protein-level HGVS notation (e.g., NP_000114.1:p.Arg152Gly). NULL for non-coding variants.
consequence
Sequence Ontology consequence term(s). Examples: missense_variant, frameshift_variant, splice_donor_variant.
impact
Impact severity: HIGH, MODERATE, LOW, or MODIFIER. Determines which ACMG criteria paths are evaluated.
biotype
Transcript biotype (e.g., protein_coding, nonsense_mediated_decay).
exon_number
Exon number within the transcript, if applicable.
domains
Protein domain annotations (e.g., "Pfam:PF00533,InterPro:IPR011364"). Used for PM1 and PM4 criteria.
ACMG Criteria Dependent on VEP
VEP consequence and domain annotations directly determine which ACMG criteria are evaluated:
PVS1
Impact = HIGH + specific consequence types (frameshift, stop_gained, splice_acceptor, splice_donor)
PM1
Domains field contains Pfam annotation
PM4
In-frame indel consequence + Pfam domain + not repetitive region
PP2
Missense consequence type
BP1
Missense consequence + MODERATE impact
BP3
In-frame indel + repetitive region or no Pfam domain
BP7
Synonymous consequence + not splice region
Limitations
VEP selects one canonical transcript per gene. Variants with different consequences across alternative transcripts may have clinically relevant effects not captured by the primary annotation.
VEP indel representation differs from VCF format. The platform reconciles both formats during annotation matching, but complex multi-allelic sites may require manual review.
Domain annotations depend on Pfam coverage. Novel or uncharacterized protein domains are not represented.
VEP does not predict gain-of-function effects. All consequence annotations reflect loss or disruption of normal function.
Reference
McLaren W, et al. "The Ensembl Variant Effect Predictor." Genome Biology . 2016;17(1):122. PMID: 27268795.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)

---

## Page: gnomAD Allele Frequency, pLI, and LOEUF | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/gnomad
Canonical: https://folklore.helena.bio/docs/databases/gnomad
Description: How Folklore uses gnomAD v4.1 global and ancestry-group frequencies for BA1, BS1, BS2, and PM2, plus pLI and LOEUF for gene-constraint gates.

Documentation / Reference Databases / gnomAD
## gnomAD Allele Frequency, pLI, and LOEUF
Folklore uses two gnomAD v4.1 datasets. Variant-level frequency fields supply BA1, BS1, BS2, and PM2 evidence. Gene-level pLI, LOEUF, and missense constraint help determine whether a loss-of-function or missense criterion path is eligible.
The files are stored locally. No case variant or patient information is sent to gnomAD during processing.
Allele Frequency Evidence
Population allele frequencies from 807,162 individuals across 8 genetic ancestry groups. This is the primary source for population frequency evidence in ACMG classification. A variant observed at high frequency in healthy individuals is unlikely to cause a rare genetic disorder.
Individuals 807,162 across 8 ancestry groups Genome Build GRCh38 Match Key chromosome + position + reference allele + alternate allele Source gnomad.broadinstitute.org Producer Broad Institute of MIT and Harvard
ACMG Criteria Using Variant Frequencies
BA1 Stand-alone Benign
Global allele frequency > 5%. This single criterion classifies the variant as Benign regardless of all other evidence.
BS1 Strong Benign
Frequency higher than expected for the disorder. Inheritance-aware thresholds: >= 0.1% for autosomal dominant (haploinsufficiency score = 3), >= 5% for autosomal recessive.
BS2 Strong Benign
Observed in > 15 homozygous individuals in the healthy population. Indicates the homozygous state is tolerated.
PM2 Moderate Pathogenic
Absent from controls or at extremely low frequency (global AF < 0.01%). Frequency data must be present (non-NULL) to trigger.
Frequency Columns (6)
global_af FLOAT
Global allele frequency across all populations. Primary field for BA1 (>5%), BS1, and PM2 (<0.01%) criteria.
global_ac INTEGER
Global allele count. Number of times the alternate allele was observed across all samples.
global_an INTEGER
Global allele number. Total number of alleles genotyped at this position. Used to assess coverage adequacy.
global_hom INTEGER
Global homozygote count. Number of individuals homozygous for the alternate allele. Used for BS2 (>15 homozygotes).
af_grpmax FLOAT
Maximum allele frequency across ancestry groups. Identifies population-specific enrichment that global AF might mask.
popmax VARCHAR
Population with the highest allele frequency. Reports which ancestry group shows the highest frequency for this variant.
Ancestry Groups
gnomAD v4.1 categorizes individuals into 8 genetic ancestry groups. The popmax field reports which group shows the highest allele frequency for a given variant, which can be clinically relevant for population-specific disease prevalence.
Code Ancestry Group Approximate Samples
AFR African / African American ~30,000
AMR Admixed American / Latino ~16,000
ASJ Ashkenazi Jewish ~5,000
EAS East Asian ~10,000
FIN Finnish ~13,000
MID Middle Eastern ~3,000
NFE Non-Finnish European ~450,000
SAS South Asian ~19,000
Note: Non-Finnish European (NFE) represents the majority of the cohort. Variants enriched in underrepresented populations may have less precise frequency estimates.
gnomAD Gene Constraint
A separate gnomAD dataset provides gene-level constraint metrics derived from the same population data. While variant frequencies tell you how common a specific variant is, constraint metrics tell you how tolerant the gene itself is to different types of mutations. A gene with very few observed loss-of-function variants relative to expectation is likely essential for normal function, and loss-of-function variants in that gene are more likely to be pathogenic.
Match Key gene_symbol Source gnomad.broadinstitute.org/downloads#v4-constraint
ACMG Criteria Using Gene Constraint
PVS1 Very Strong Pathogenic
Loss-of-function variant in a gene intolerant to LoF. Requires pLI > 0.9 OR LOEUF < 0.35. The strongest automated pathogenic criterion.
PP2 Supporting Pathogenic
Missense variant in a gene with constraint against missense variation. Requires pLI > 0.5.
BP1 Supporting Benign
Missense variant in a gene tolerant to loss-of-function (pLI < 0.1). If LoF variants are tolerated, missense variants are even less likely to be pathogenic.
Constraint Columns (4)
pLI FLOAT
Probability of Loss-of-function Intolerance. Ranges 0 to 1. Values > 0.9 indicate the gene is highly intolerant to loss-of-function variants. Used for PVS1 (> 0.9), PP2 (> 0.5), and BP1 (< 0.1).
LOEUF (oe_lof_upper) FLOAT
Loss-of-function Observed/Expected Upper bound Fraction. Lower values indicate stronger constraint. Values < 0.35 trigger PVS1 as an alternative to pLI. The preferred constraint metric in recent literature.
oe_lof FLOAT
Loss-of-function observed/expected ratio. The point estimate of how many LoF variants are observed versus expected. LOEUF is the upper confidence bound of this ratio.
mis_z FLOAT
Missense Z-score. Positive values indicate the gene has fewer missense variants than expected (missense-constrained). Used in the Screening Service for gene relevance scoring.
pLI vs. LOEUF
Both pLI and LOEUF measure gene intolerance to loss-of-function variants, but they use different statistical approaches. pLI is a probability (0 to 1) from a discrete classification model. LOEUF is the upper confidence bound of the observed/expected ratio and provides a continuous measure of constraint. Recent literature favors LOEUF because it better captures the spectrum of constraint rather than forcing genes into discrete categories. Folklore accepts either metric for PVS1 (pLI > 0.9 OR LOEUF < 0.35) to maximize sensitivity.
Limitations
gnomAD excludes individuals with known Mendelian disease diagnoses, but does not screen for carrier status or late-onset conditions.
Coverage varies across the genome. Low-coverage regions may have NULL frequency data, which prevents PM2 from triggering.
Structural variants and complex rearrangements are not represented in the SNV/indel dataset.
Population ancestry group assignment is based on principal component analysis, not self-reported ethnicity.
Rare variants in underrepresented populations may have inflated or absent frequency estimates due to smaller sample sizes.
Gene constraint metrics reflect population-level observations. A gene with low pLI may still harbor pathogenic variants in specific domains or functional regions.
Constraint metrics are gene-wide averages. They do not capture regional variation in constraint within a gene.
References
Chen S, et al. "A genomic mutational constraint map using variation in 76,156 human genomes." Nature . 2024;625:92-100. PMID: 38057664
Karczewski KJ, et al. "The mutational constraint spectrum quantified from variation in 141,456 humans." Nature . 2020;581:434-443. PMID: 32461654

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)
- [PMID: 38057664](https://pubmed.ncbi.nlm.nih.gov/38057664/)
- [PMID: 32461654](https://pubmed.ncbi.nlm.nih.gov/32461654/)

---

## Page: HPO | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/hpo
Canonical: https://folklore.helena.bio/docs/databases/hpo
Description: Human Phenotype Ontology (HPO) gene-phenotype associations used for PP4 criteria and phenotype matching in Folklore.

Documentation / Reference Databases / HPO
## HPO
The Human Phenotype Ontology (HPO) provides a standardized vocabulary of phenotypic abnormalities and their associations with genes and diseases. In Folklore, HPO data is aggregated from 6 curated sources into an enriched gene-phenotype view, covering 5,688 genes with 927K associations. This multi-source approach increases gene coverage by 10% compared to single-source HPO Consortium data and provides richer phenotype profiles for clinical matching.
Database Details
Sources HPO Consortium, Orphanet, DECIPHER G2P, Monarch, ClinVar-MedGen, Manual Curation Producer Helena Bioinformatics (aggregated from multiple public sources)
Role in Classification and Analysis
HPO data supports three analysis functions:
ACMG PP4 Criterion
When patient HPO terms are provided, PP4 triggers if >= 3 patient HPO terms match the gene HPO profile, or >= 2 terms match for a highly specific gene (<= 5 total HPO associations). This provides supporting pathogenic evidence.
Phenotype Matching Service
The dedicated phenotype matching module uses HPO semantic similarity to compute overlap between patient phenotype and gene-associated phenotypes, producing clinical tiers (Tier 1 through Tier 4) for variant prioritization.
Screening Prioritization
When no patient HPO terms are available, the hpo_count field serves as a proxy for clinical relevance. Genes associated with more phenotypes receive higher screening scores, reflecting broader clinical significance.
Data Deduplication
HPO terms are aggregated from 6 sources with source priority ordering (Orphanet > G2P > HPO Consortium > Monarch > ClinVar-MedGen > Manual Curation). The same HPO term from multiple sources is deduplicated per gene, ensuring each unique phenotype is counted once. Low-confidence entries (Orphanet modifying/candidate associations, G2P limited confidence) are excluded from the clinical view. Both hpo_ids and hpo_names are sorted by HPO ID to maintain a reliable 1:1 correspondence between identifiers and names.
Columns Loaded (6)
HPO data is joined on gene symbol. Each variant inherits the complete HPO profile of its associated gene.
hpo_ids VARCHAR
Semicolon-separated HPO term identifiers associated with the gene (e.g., "HP:0001250;HP:0001263;HP:0002069"). Sorted by HPO ID to maintain 1:1 correspondence with hpo_names.
hpo_names VARCHAR
Semicolon-separated HPO term names corresponding to hpo_ids (e.g., "Seizure;Global developmental delay;Dementia"). Sorted to match hpo_ids.
hpo_count INTEGER
Number of unique HPO terms associated with the gene. Used in screening mode as a proxy for clinical breadth when no patient HPO terms are available.
hpo_frequency_data VARCHAR
Frequency of each phenotype in the associated condition, when available. Not all gene-phenotype associations have frequency data.
hpo_disease_ids VARCHAR
OMIM or Orphanet disease identifiers that link the gene to each HPO phenotype.
hpo_gene_id VARCHAR
HPO internal gene identifier used for cross-referencing within the ontology.
Limitations
HPO coverage varies by disease. Well-studied conditions have comprehensive phenotype profiles, while rare diseases may have minimal HPO annotation.
HPO terms are associated at the gene level, not the variant level. Different variants in the same gene may produce different phenotypes.
Frequency data for phenotype-disease associations is incomplete. The absence of frequency data does not mean the phenotype is rare.
HPO is primarily focused on rare diseases. Common complex conditions may have less comprehensive ontology coverage.
Phenotype matching depends on accurate HPO term selection by the clinician. Overly broad or imprecise terms reduce matching specificity.
Reference
Kohler S, et al. "The Human Phenotype Ontology in 2024: phenotypes around the world." Nucleic Acids Research . 2024;52(D1):D1333-D1346. PMID: 37953324.
For details on phenotype-based analysis, see the Phenotype Matching section.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)

---

## Page: SpliceAI Precomputed | Folklore Documentation

Source: https://folklore.helena.bio/docs/databases/spliceai-precomputed
Canonical: https://folklore.helena.bio/docs/databases/spliceai-precomputed
Description: SpliceAI precomputed splice impact scores for all coding variants -- four delta scores predicting splice site gain and loss.

Documentation / Reference Databases / SpliceAI Precomputed
## SpliceAI Precomputed
SpliceAI is a deep learning model developed by Illumina that predicts the impact of genetic variants on mRNA splicing. Folklore uses precomputed SpliceAI scores for all coding variants, eliminating the need for on-the-fly prediction and ensuring consistent, reproducible results.
Database Details
Coverage All coding variants in MANE Select transcripts Genome Build GRCh38 Source github.com/Illumina/SpliceAI Producer Illumina, Inc.
Four Delta Scores
SpliceAI produces four delta scores, each ranging from 0 to 1, representing the change in splice probability caused by the variant:
DS_AG Acceptor Gain
Predicts creation of a new splice acceptor site. A high score indicates the variant may introduce a cryptic acceptor competing with the canonical site.
DS_AL Acceptor Loss
Predicts destruction of an existing splice acceptor site. A high score indicates the canonical acceptor is disrupted.
DS_DG Donor Gain
Predicts creation of a new splice donor site. A high score indicates a cryptic donor may be introduced.
DS_DL Donor Loss
Predicts destruction of an existing splice donor site. A high score indicates the canonical donor is disrupted.
The maximum of the four delta scores (max_score) is used for classification thresholds. This captures the strongest predicted splice effect regardless of mechanism.
Score Interpretation
Max Score Interpretation Clinical Implication
0.0 - 0.1 No predicted splice impact Variant unlikely to affect splicing. BP7 guard satisfied (SpliceAI < 0.1).
0.1 - 0.2 Low predicted impact Some splice effect possible but below PP3_splice threshold.
0.2 - 0.5 Moderate predicted impact PP3_splice triggered (Supporting). ClinGen SVI 2023 threshold for supporting splice evidence.
0.5 - 0.8 High predicted impact Strong prediction of splice disruption. PP3_splice at Supporting strength.
0.8 - 1.0 Very high predicted impact Near-certain splice disruption. PP3_splice at Supporting strength.
Role in ACMG Classification
PP3_splice Supporting Pathogenic
SpliceAI max_score >= 0.2 triggers PP3 at Supporting strength. Applies only when PVS1 does not already apply (ClinGen SVI double-counting guard). Aligned with Walker et al. 2023 ClinGen SVI recommendations.
BP7 guard Benign guard
BP7 (synonymous + no splice impact) requires SpliceAI max_score < 0.1 to confirm the variant does not affect splicing. Without this guard, a synonymous variant near a splice site could be incorrectly classified as benign.
BP4 guard Benign guard
BP4 (computational evidence suggests no impact) additionally requires SpliceAI max_score < 0.1 to ensure no predicted splice disruption before applying benign computational evidence.
PVS1 Double-Counting Guard
When a variant triggers PVS1 (null variant in a LoF-intolerant gene) through a splice consequence (splice_acceptor_variant or splice_donor_variant), SpliceAI PP3_splice is not additionally applied. This prevents counting the same splice disruption evidence twice -- once through the consequence-based PVS1 pathway and again through the SpliceAI prediction pathway. This follows ClinGen SVI 2023 guidelines.
Limitations
SpliceAI predictions are based on primary sequence context. Tissue-specific splicing regulation is not modeled.
Precomputed scores cover coding variants in MANE Select transcripts. Variants in non-MANE transcripts or deep intronic regions may not have scores.
SpliceAI does not predict the functional consequence of aberrant splicing (exon skipping, intron retention, etc.), only the probability that splicing is disrupted.
Scores near the 0.2 threshold should be interpreted with caution. RNA studies can confirm or refute predicted splice effects.
References
Jaganathan K, et al. "Predicting splicing from primary sequence with deep learning." Cell . 2019;176(3):535-548. PMID: 30661751.
Walker LC, et al. "Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing." American Journal of Human Genetics . 2023;110(7):1046-1067. PMID: 37352859.
For more details on SpliceAI interpretation, see the dedicated SpliceAI predictor page.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Reference Databases](https://folklore.helena.bio/docs/databases)
- [SpliceAI predictor](https://folklore.helena.bio/docs/predictors/spliceai)

---

## Page: FAQ | Folklore Documentation

Source: https://folklore.helena.bio/docs/faq
Canonical: https://folklore.helena.bio/docs/faq
Description: Frequently asked questions about Folklore -- data handling, classification methodology, supported formats, and clinical use.

Documentation / FAQ
## Frequently Asked Questions
What file formats does Folklore accept?
Primary variant analysis accepts single-sample VCF 4.1 or 4.2 files in plain-text (.vcf) or bgzipped (.vcf.gz) form. Multi-sample files should be split before upload. Family analysis combines separately processed proband, parent, and sibling samples.
Which genome builds are supported?
GRCh38 (hg38) is the primary build. GRCh37 (hg19) files are accepted and automatically lifted over to GRCh38 using CrossMap. Earlier builds are not supported.
How long does analysis take?
Processing time depends on file size, enabled modules, variant count, and server load. Panel and exome files usually complete faster than whole-genome files. The processing record shows the actual duration for each stage and case.
Where is patient data stored?
All data is processed and stored on EU-based infrastructure in Helsinki, Finland (Hetzner AX162R dedicated server). No patient data is transmitted to external services, cloud providers, or AI APIs.
Is patient data deleted after analysis?
Yes. VCF files and analysis results are retained only for the duration needed for clinical review. Data deletion policies comply with GDPR requirements.
Which ACMG criteria are automated?
Folklore computes criteria that can be derived from the submitted variants, reference databases, phenotype data, and supported family context. Evidence that depends on functional studies, case-control data, or expert assessment remains available for clinical curation. The Criteria Reference documents the status of each criterion.
What computational predictor is used for PP3/BP4?
BayesDel_noAF with ClinGen SVI-calibrated thresholds (Pejaver et al. 2022). The "_noAF" variant excludes allele frequency from its model to avoid circular reasoning with PM2, BA1, and BS1. SpliceAI is used independently for splice impact (PP3_splice).
Does Folklore support structural variants, CNVs, and mitochondrial variants?
Yes. Folklore interprets compatible structural-variant calls already present in the VCF; it does not perform upstream SV calling. Constitutional copy-number loss and gain use the Riggs 2020 framework, while other eligible nuclear structural events can enter the nuclear Variant Analysis framework. Mitochondrial variants use MMDWG 2020. Repeat expansions and somatic tumor variants remain outside these modules.
Does the platform support somatic variant interpretation?
No. Folklore is designed for germline variant interpretation in Mendelian disease contexts. Somatic variant analysis requires tumor-normal paired analysis with different classification frameworks (AMP/ASCO/CAP).
What databases are used for annotation?
The core Stage 4 enrichment uses 16 production classification and reference databases, including gnomAD v4.1, ClinVar 2025-01, dbNSFP 4.9c, SpliceAI, HPO, ClinGen, and Ensembl VEP Release 113. Across nuclear, phenotype, inheritance, mitochondrial, and SV/CNV modules, Folklore integrates 45 total reference databases and curated sources. Smaller counts in subsystem descriptions refer to selected source sets, not the platform-wide total. Production reference data is stored locally.
How does phenotype matching work?
Patient HPO terms are compared against gene-phenotype associations using Lin semantic similarity within the HPO ontology graph. Genes are ranked by the strength of phenotype correlation and assigned clinical priority tiers. See the Phenotype Matching documentation for details.
What does the AI clinical assistant use for its model?
A large language model hosted on dedicated GPU infrastructure within the EU. All inference happens on-premise through a secure internal connection. No patient data is sent to external AI services.
Can the AI assistant modify variant classifications?
No. ACMG classifications are determined by the automated pipeline. The AI assistant can explain why criteria were triggered and discuss the evidence, but it cannot change classifications. Only the reviewing geneticist can override classifications.
Is the clinical interpretation a medical diagnosis?
No. The AI-generated clinical interpretation is a decision support tool. All findings must be independently validated by a qualified clinical geneticist before being used in patient care. Reports include a standard disclaimer stating this requirement.
How often are reference databases updated?
ClinVar is updated quarterly or more frequently for clinically significant changes. gnomAD major releases are adopted within 3 months of publication. All updates undergo validation testing with a reference cohort before deployment. See the Database Update Policy for the complete schedule.
Can I use Folklore for research purposes?
Yes. The platform is suitable for both clinical and research use. For research applications, the same analytical rigor applies but reporting requirements may differ from clinical diagnostic settings.
What happens if ClinVar and the automated classification disagree?
Folklore applies a transparent priority system: BA1 always overrides ClinVar, expert-panel ClinVar assertions can override the automated classification, and conflicting evidence is flagged for manual review. See the ClinVar Integration documentation for the complete decision logic.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)

---

## Page: Trio and Family Variant Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/family-analysis
Canonical: https://folklore.helena.bio/docs/family-analysis
Description: How Folklore processes a complete trio through GLnexus joint genotyping, proband ACMG classification, relationship QC, de novo detection, compound heterozygous phasing, and segregation assessment.

Documentation / Trio and Family Variant Analysis
## Trio and Family Variant Analysis
Family context can add evidence that is unavailable from a proband-only analysis. Folklore uses parental genotypes to identify technically supported de novo candidates, determine whether two variants in a recessive gene are in trans, and measure segregation within the limits of the available pedigree.
The current production workflow supports a complete trio only: proband, mother, and father. Duo, sibling-only, extended-pedigree, and multiple-proband workflows are not part of the present production pipeline.
One Call Set, Two Evidence Layers
The proband classification and the parental genotype evidence are derived from the same normalized joint VCF. The family module records inheritance evidence beside the upstream ACMG classification; it does not silently overwrite or replace that classification.
Required Inputs
Either one compatible joint multi-sample VCF or separate proband, maternal, and paternal gVCFs.
Sample identifiers, sex, family role, and affected status for each member.
An inheritance hypothesis and phenotype specificity when available.
A compatible reference genome and contig naming scheme.
Production Workflow
1
Validate the Complete Trio
The production workflow requires one proband, one mother, and one father with compatible gVCF inputs and recorded pedigree metadata.
2
Prepare the Joint Call Set
Separate member gVCFs are joint-genotyped into one normalized multi-sample VCF. A compatible laboratory-provided joint VCF can be validated and used directly.
3
Classify the Proband
A proband-only VCF is extracted from the joint call set and processed through the standard Variant Analysis workflow. ACMG classification belongs to the proband.
4
Build the Trio Evidence Store
The joint genotypes are combined with the classified proband data. PLINK relationship quality control checks the declared parent-child relationships, parental relatedness, and possible duplicate samples.
5
Compute Inheritance Evidence
The family module evaluates de novo candidates, compound heterozygous phase, segregation scores, and categorical inheritance patterns without replacing the proband ACMG classification.
De Novo Candidates
A candidate de novo variant must be present in the proband and supported by reference genotypes in both parents. Folklore records high- or low-confidence technical states, together with excluded and not-applicable control states. Confidence describes support for de novo origin, not pathogenicity.
The clinical-grade reporting layer applies additional quality, population-frequency, consequence, and splice-signal filters to identify a focused review subset. This reporting layer does not independently change PS2 criterion semantics.
Compound Heterozygous Candidates
Within each gene, the workflow evaluates pairs of heterozygous coding or splice-relevant variants. A pair receives strongest technical support when parental genotypes demonstrate that one variant was inherited from each parent, establishing an in-trans configuration.
An in-trans result establishes candidate phase. It does not by itself establish that both alleles are clinically significant; each variant still requires its own evidence review.
Segregation Limits
A complete trio contributes at most two informative parent-to-proband transmissions. Its maximum raw LOD is approximately 0.602 before phenotype specificity is applied. Depending on the inheritance hypothesis, informative transmissions, and phenotype specificity, a trio result can remain Indeterminate or reach the lower Supporting band.
Moderate, Strong, and Very Strong segregation bands require additional informative relatives in an extended pedigree. Extended-pedigree analysis is outside the current production scope.
Outputs for Clinical Review
A normalized joint multi-sample VCF for the complete trio.
A classified proband dataset produced through the standard ACMG workflow.
Pairwise relationship quality-control results from PLINK Identity-by-Descent analysis.
Per-variant de novo confidence and supporting technical evidence.
Candidate compound heterozygous pairs with parental-origin phase information.
Segregation scores and evidence bands, including explicit trio-only limits.
A separate clinical-grade reporting subset for focused review.
Clinical Boundary
Family Analysis is clinical decision support. It records technically derived inheritance evidence and makes the evidence traceable for review. A qualified genetics professional decides how that evidence affects the final case interpretation.
In This Section
De Novo Variant Analysis
Trio genotype requirements, confidence states, reporting filters, PS2 and PM6 boundaries, and mosaicism limits.
Compound Heterozygous Variants
Candidate pairing, parental-origin phasing, trans and cis states, multiple partners, clinical-grade filtering, and PM3 boundaries.
Segregation Analysis
Per-variant LOD scoring, phenotype multipliers, evidence bands, trio limits, and PP1 or BS4 review boundaries.
Trio Quality Control
PLINK IBD analysis, PI_HAT thresholds, sample-swap and duplicate alerts, related parents, skipped checks, and interpretation limits.
Inheritance Patterns
Categorical pattern summaries, sex-gated X-linked states, unclear results, and ACMG boundaries.
Related Pages
Family and Trio Analysis
Product overview and clinical workflow positioning.
Full Family Analysis Methodology
Detailed methods, thresholds, limitations, tools, and references.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [De Novo Variant Analysis Trio genotype requirements, confidence states, reporting filters, PS2 and PM6 boundaries, and mosaicism limits.](https://folklore.helena.bio/docs/family-analysis/de-novo-variants)
- [Compound Heterozygous Variants Candidate pairing, parental-origin phasing, trans and cis states, multiple partners, clinical-grade filtering, and PM3 boundaries.](https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes)
- [Segregation Analysis Per-variant LOD scoring, phenotype multipliers, evidence bands, trio limits, and PP1 or BS4 review boundaries.](https://folklore.helena.bio/docs/family-analysis/segregation-analysis)
- [Trio Quality Control PLINK IBD analysis, PI_HAT thresholds, sample-swap and duplicate alerts, related parents, skipped checks, and interpretation limits.](https://folklore.helena.bio/docs/family-analysis/trio-quality-control)
- [Inheritance Patterns Categorical pattern summaries, sex-gated X-linked states, unclear results, and ACMG boundaries.](https://folklore.helena.bio/docs/family-analysis/inheritance-patterns)
- [Family and Trio Analysis Product overview and clinical workflow positioning.](https://folklore.helena.bio/platform/family-trio)
- [Full Family Analysis Methodology Detailed methods, thresholds, limitations, tools, and references.](https://folklore.helena.bio/methodology/family-analysis)

---

## Page: Compound Heterozygous Variant Analysis in Trios | Folklore Documentation

Source: https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes
Canonical: https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes
Description: How Folklore identifies compound heterozygous candidate pairs, determines parental origin, distinguishes in-trans from in-cis configurations, and filters pairs for clinical review.

Documentation / Family Analysis / Compound Heterozygous Variants
## Compound Heterozygous Variant Analysis in Trios
A compound heterozygous candidate contains two heterozygous variants in the same gene. In a recessive disorder, the pair is most informative when the variants occur on different parental chromosomes, with one inherited from each parent.
Folklore identifies candidate pairs and records the available phasing evidence. A paired result does not establish that both variants are clinically significant and does not by itself establish a molecular diagnosis.
Current Production Scope
The production workflow phases variants from a complete proband-mother-father trio. Parental genotypes are read from the normalized joint call set, where each parent has an explicit genotype or no-call at every evaluated site.
Candidate Selection
The detector begins with proband variants that are heterozygous, mapped to a gene, and assigned an included molecular consequence. Intronic, upstream, downstream, and synonymous variants are excluded from the initial candidate set.
Missense variants.
Stop-gained and stop-lost variants.
Frameshift variants.
In-frame insertions and deletions.
Splice donor, splice acceptor, and splice-region variants.
Start-lost and initiator-codon variants.
Pair Formation Within Each Gene
Genes with at least two candidate variants are evaluated by pairing each retained variant with every other retained variant in that gene. The detector evaluates each unordered pair once and writes the result back symmetrically to both variants.
To keep pair growth bounded in very large genes, the workflow applies a configurable per-gene candidate limit after consequence filtering and records when that guard is used. Reviewers should consider this limit when a gene contains an unusually large number of potentially relevant variants.
Parental-Origin States
Phase State Confidence Meaning Handling
Trio phased High One variant is inherited from the mother and the other from the father, with an explicit reference genotype on the opposite parental side for each site. The pair is supported as being in trans and is retained for clinical review.
Cis detected Excluded Both variants are inherited from the same parent, or a homozygous-alternate parental state prevents a valid trans interpretation. The pair is not treated as a compound heterozygous candidate.
Cis ambiguous Low The parental genotypes do not distinguish a trans configuration from a same-haplotype configuration. The pair remains visible with unresolved phase.
Parental coverage gap Low At least one parental genotype required for phasing is a no-call. The pair cannot be phased confidently from the available trio data.
Parental genotypes missing Low All four parental genotype observations needed for the pair are no-calls. No parental-origin conclusion is made.
What Counts as an In-Trans Pair
A high-confidence pair requires reciprocal parental transmission. For one variant, the mother is heterozygous and the father is reference. For the partner variant, the father is heterozygous and the mother is reference, or the same pattern occurs in the opposite order.
A homozygous-alternate genotype in either parent excludes a high-confidence trans interpretation. That parent would transmit the alternate allele to every child, which does not represent the required one-allele-from-each-parent pattern.
Multiple Candidate Partners
One variant can form candidate pairs with more than one variant in the same gene. Folklore records a primary partner and separately marks that additional partners exist.
Reviewers should inspect all relevant partners when the multiple-partner flag is present rather than treating the primary partner as the only possible pair.
Clinical-Grade Reporting Subset
The reporting overlay counts a pair only when both variants are high-confidence, have a non-null partner relationship, pass the configured proband filter status, satisfy the population-frequency ceiling, and have HIGH or MODERATE impact or the configured splice-rescue signal.
The configured population ceiling is applied per allele. The reporting overlay narrows the review set; it does not modify the detector result or the stored ACMG classifications of either variant.
PM3 Evidence Boundary
A trio-phased result provides technical support that two variants are in trans. PM3 use still depends on the clinical significance of the partner allele, the gene-disease mechanism, the inheritance model, and the applicable ACMG or ClinGen specifications.
Folklore records the pair, phase source, confidence, and both proband classifications. A qualified genetics professional decides whether the evidence supports PM3 and at what strength.
Clinical Interpretation
Phase and pathogenicity are separate questions. A confirmed in-trans pair can still contain one or two benign variants, and two clinically relevant variants can remain unresolved when parental coverage is insufficient.
Related Documentation
Trio and Family Variant Analysis
Production workflow, inputs, outputs, and present scope.
De Novo Variant Analysis
Candidate detection, confidence states, and PS2 or PM6 review boundaries.
Family Analysis Methodology
Detailed methods, thresholds, tools, and limitations.
ACMG Criteria Reference
Reference for PM3 and the remaining ACMG evidence criteria.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Family Analysis](https://folklore.helena.bio/docs/family-analysis)
- [De Novo Variant Analysis Candidate detection, confidence states, and PS2 or PM6 review boundaries.](https://folklore.helena.bio/docs/family-analysis/de-novo-variants)
- [Family Analysis Methodology Detailed methods, thresholds, tools, and limitations.](https://folklore.helena.bio/methodology/family-analysis)
- [ACMG Criteria Reference Reference for PM3 and the remaining ACMG evidence criteria.](https://folklore.helena.bio/docs/classification/criteria-reference)

---

## Page: De Novo Variant Analysis in Trio Sequencing | Folklore Documentation

Source: https://folklore.helena.bio/docs/family-analysis/de-novo-variants
Canonical: https://folklore.helena.bio/docs/family-analysis/de-novo-variants
Description: How Folklore identifies de novo variant candidates from a jointly genotyped proband-mother-father trio, assigns technical confidence, applies reporting filters, and preserves clinical review of PS2 and PM6 evidence.

Documentation / Family Analysis / De Novo Variants
## De Novo Variant Analysis in Trio Sequencing
A de novo variant is present in the proband and absent from both parents at the same genomic site. In a clinical trio, this pattern can provide inheritance evidence that is unavailable from proband-only sequencing.
Folklore identifies technically supported de novo candidates. De novo origin does not establish pathogenicity, and a confidence state does not by itself determine whether PS2 or PM6 should be applied.
Current Production Scope
De novo analysis requires a complete trio: proband, mother, and father. The three samples are evaluated from one normalized multi-sample call set. When separate member gVCFs are supplied, they are joint-genotyped before inheritance analysis.
Why Joint Genotyping Matters
Independently called files can differ because of caller thresholds, normalization, or representation of the same allele. A missing row in a parent-only variant file is not equivalent to a confirmed reference genotype.
Folklore evaluates one multi-sample VCF for the proband, mother, and father. The proband classification and the parental genotype evidence are derived from that shared call set, reducing ambiguity at candidate sites.
Candidate Detection
The detector begins with variants present in the proband. For each site, it reads the maternal and paternal genotypes from the joint VCF and checks whether both parents are reference at that position.
Genotype state alone is not enough. The detector also evaluates the configured parental depth and genotype-quality requirements before assigning a confidence state.
De Novo Confidence States
State Technical Meaning Clinical Handling
High confidence The candidate is present in the proband, both parents have reference genotypes at the site, and the configured genotype-quality and depth requirements are satisfied. Technically supported de novo origin. Pathogenicity and ACMG criterion use still require clinical review.
Low confidence The inheritance pattern is compatible with de novo origin, but at least one parental genotype is a no-call or does not meet the configured depth and genotype-quality requirements. Retained for review with the reason for reduced confidence.
Excluded The alternate allele is present in at least one parent, so de novo origin is not supported by the trio genotypes. Not reported as a de novo candidate.
Not applicable The required trio or inheritance conditions are not available for a valid de novo assessment. No de novo evidence is assigned.
Explicit Genotypes and No-Calls
The joint call set represents each parent at an evaluated site as reference, heterozygous, homozygous alternate, or no-call. A no-call is not treated as a confirmed reference genotype.
When a parental genotype is a no-call, or a reference call does not meet the configured quality requirements, the candidate cannot receive the high-confidence state. If either parent carries the alternate allele, de novo origin is excluded.
Clinical-Grade Reporting Checks
The reporting layer narrows the technically supported candidate set for clinical review. It does not rerun the detector and does not independently change ACMG criterion semantics.
High-confidence technical support from the trio detector.
Variant-level and proband quality status.
Population-frequency limits used by the reporting profile.
Predicted molecular consequence.
Splice-signal rescue when the configured SpliceAI threshold is met.
PS2 and PM6 Remain Review Decisions
ACMG/AMP uses PS2 and PM6 for de novo evidence under different levels of confirmation. Their application depends on more than the observed trio genotype pattern. The reviewer must consider parentage, phenotype consistency, gene-disease mechanism, penetrance, and the quality of the supporting data.
Folklore records the detector result and its technical basis. A qualified genetics professional decides whether the available evidence satisfies the requirements of a criterion and what strength is justified.
Mosaicism and Other Limits
Low-level parental mosaicism may fall below the detection threshold of the source assay. A variant can therefore appear absent in blood-derived parental data even when mosaicism is present in another tissue or at a lower allele fraction.
Mapping ambiguity, segmental duplication, low-complexity sequence, allelic imbalance, and low coverage can also weaken a de novo inference. The confidence state records the available technical support; it cannot remove these assay-level limits.
What the Reviewer Receives
Each candidate carries its proband and parental genotypes, quality measurements, confidence state, exclusion or downgrade reason, population and consequence context, and reporting status. The original proband ACMG classification remains separately visible.
Related Documentation
Trio and Family Variant Analysis
Production workflow, required inputs, outputs, and current scope.
Family Analysis Methodology
Detailed methods, limitations, reference tools, and evidence boundaries.
ACMG Criteria Reference
Role and review status of PS2, PM6, PP1, and other ACMG criteria.
Compound Heterozygous Variants
Parental-origin phasing and in-trans candidate pairs.
ACMG/AMP framework: Richards S, et al. Genet Med. 2015;17(5):405-424. PMID: 25741868
De novo mutation review: Veltman JA, Brunner HG. Nat Rev Genet. 2012;13(8):565-575. PMID: 22781750

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Family Analysis](https://folklore.helena.bio/docs/family-analysis)
- [Family Analysis Methodology Detailed methods, limitations, reference tools, and evidence boundaries.](https://folklore.helena.bio/methodology/family-analysis)
- [ACMG Criteria Reference Role and review status of PS2, PM6, PP1, and other ACMG criteria.](https://folklore.helena.bio/docs/classification/criteria-reference)
- [Compound Heterozygous Variants Parental-origin phasing and in-trans candidate pairs.](https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)
- [PMID: 22781750](https://pubmed.ncbi.nlm.nih.gov/22781750/)

---

## Page: Inheritance Pattern Classification in Clinical Trios | Folklore Documentation

Source: https://folklore.helena.bio/docs/family-analysis/inheritance-patterns
Canonical: https://folklore.helena.bio/docs/family-analysis/inheritance-patterns
Description: How Folklore derives per-variant inheritance-pattern labels from de novo and compound-heterozygous detector outputs, trio genotypes, chromosome context, and proband sex.

Documentation / Family Analysis / Inheritance Patterns
## Inheritance Pattern Classification in Clinical Trios
Folklore assigns one inheritance-pattern label to each variant present in the proband. The label summarizes existing detector outputs, trio genotypes, chromosome context, and proband sex in a form that can be displayed and filtered consistently.
This step is derivational. It does not replace de novo detection, compound-heterozygous phasing, segregation scoring, ACMG classification, or clinical interpretation.
Authoritative Evidence Remains Separate
The inheritance-pattern column is a summary. The underlying de novo tier, inference basis, compound-heterozygous partner, phase source, confidence state, genotypes, and segregation fields remain the authoritative technical evidence.
How to Read a Pattern
Each proband variant receives at most one summary label. The label is derived consistently from the available family evidence so that results can be grouped and filtered, but it is not a substitute for reviewing the underlying genotypes and detector outputs.
Where more than one family observation could describe the same row, Folklore uses a stable precedence to produce one label. The detailed de novo, compound-heterozygous, segregation, and per-member fields remain available for interpretation.
Pattern Categories
Stored Pattern Technical Basis Clinical Review Boundary
de_novo_high De novo candidate with high technical confidence. Review the underlying parental genotype, depth, quality, and PS2 boundary.
de_novo_low De novo candidate with unresolved or sub-threshold parental evidence. Origin remains insufficiently established for high-confidence use.
compound_het_high A compound-heterozygous partner is recorded with high confidence. Inspect both variants, parental origin, phase source, and PM3 requirements.
compound_het_low A compound-heterozygous partner is recorded with low confidence. Phase remains unresolved or limited by parental genotype data.
x_linked_recessive_male Male proband carries an X variant and the mother is heterozygous. Sex, X-chromosome region, genotype representation, phenotype, and gene mechanism still require review.
x_linked_recessive_female_homozygous Female proband is homozygous on X, with maternal and paternal variant carriage. This rare configuration requires careful pedigree and genotype review.
x_linked_dominant The proband carries an X variant and exactly one parent carries it. The categorical rule does not fully enforce every sex-specific transmission constraint.
mitochondrial The proband and mother carry the mitochondrial variant. The label does not assess heteroplasmy, tissue distribution, or pathogenicity.
autosomal_recessive_homozygous The proband is homozygous alternate on an autosome and both parents are heterozygous. Review gene-disease validity, phenotype fit, and the classifications of the allele.
autosomal_dominant_inherited The proband is heterozygous on an autosome and exactly one parent carries the variant. Affected status, penetrance, age of onset, and phenocopies remain clinically relevant.
unclear_parents_no_call The proband carries the variant and both parental genotypes are no-calls. No parental-origin conclusion can be made from the available calls.
unclear The row matches none of the defined categories. This can include atypical, inconsistent, unresolved, or cis-associated configurations.
Parent-Only Rows
A site present only in a parent has no proband variant index. Such a row is not assigned an inheritance pattern and remains null rather than being forced into an unclear category.
Sex-Gated X-Linked Categories
X-linked categories require proband sex to be recorded as male or female. When sex is absent, Folklore suppresses the X-linked labels and allows the row to fall through to another applicable category or to unclear.
Sex-chromosome labels also account for chromosome context, including the distinction between pseudoautosomal and non-pseudoautosomal regions. Ambiguous or incomplete states are handled conservatively rather than being forced into an X-linked category.
Pattern Versus Pathogenicity
An inheritance label describes how the observed genotypes fit a categorical transmission rule. It does not establish that the variant causes disease, that the gene explains the phenotype, or that an ACMG criterion should be applied.
A benign variant can match a recognizable inheritance pattern. A clinically relevant variant can also remain unclear when parental genotypes, sex, phase, or family structure are insufficient.
Relationship to ACMG Classification
An inheritance pattern does not automatically upgrade, downgrade, or replace the proband ACMG classification.
Family evidence remains available through the dedicated detector and segregation fields. Any criterion application and final classification decision remain separate.
Unavailable and Legacy Results
If a pattern cannot be derived, the family evidence remains available through the underlying detector, genotype, and segregation fields. The absence of a summary label does not erase those results.
Legacy trio analyses can also have a null pattern because they were processed before categorical derivation was introduced. A null value must therefore be read with the algorithm version and evidence summary.
Related Documentation
De Novo Variant Analysis
Parental genotype requirements, confidence states, and PS2 or PM6 boundaries.
Compound Heterozygous Variants
Candidate pairing, phase source, and PM3 boundaries.
Segregation Analysis
LOD scoring, evidence bands, and PP1 or BS4 boundaries.
Trio Quality Control
Cross-sample relationship checks that protect inheritance interpretation.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Family Analysis](https://folklore.helena.bio/docs/family-analysis)
- [De Novo Variant Analysis Parental genotype requirements, confidence states, and PS2 or PM6 boundaries.](https://folklore.helena.bio/docs/family-analysis/de-novo-variants)
- [Compound Heterozygous Variants Candidate pairing, phase source, and PM3 boundaries.](https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes)
- [Segregation Analysis LOD scoring, evidence bands, and PP1 or BS4 boundaries.](https://folklore.helena.bio/docs/family-analysis/segregation-analysis)
- [Trio Quality Control Cross-sample relationship checks that protect inheritance interpretation.](https://folklore.helena.bio/docs/family-analysis/trio-quality-control)

---

## Page: Segregation Analysis and LOD Scores in Clinical Trios | Folklore Documentation

Source: https://folklore.helena.bio/docs/family-analysis/segregation-analysis
Canonical: https://folklore.helena.bio/docs/family-analysis/segregation-analysis
Description: How Folklore computes per-variant segregation LOD scores from trio genotypes, applies phenotype-specificity multipliers, maps evidence bands, and preserves clinical review of PP1 and BS4.

Documentation / Family Analysis / Segregation Analysis
## Segregation Analysis and LOD Scores in Clinical Trios
Segregation analysis asks whether a variant is transmitted through a family in the pattern expected for the stated inheritance hypothesis. Folklore computes this signal separately for each proband variant.
The result is a numerical LOD score, an evidence band, and a consistency flag. These outputs describe the observed trio pattern. They do not independently establish pathogenicity or determine final PP1 or BS4 use.
Inheritance Hypothesis Is Required
When the inheritance hypothesis is missing, unknown, or unsupported, Folklore skips segregation scoring. Proband variants receive a not-applicable state rather than a zero score. This distinguishes an analysis that was not performed from one that was performed and found no supporting segregation.
Per-Variant LOD Calculation
Each informative parent-to-proband transmission contributes log10(2), approximately 0.301, when it matches the selected inheritance model. An uninformative transmission contributes zero. Some model violations can reduce the score.
A complete trio contains at most two parent-to-proband meioses. The maximum raw score produced by two fully informative transmissions is therefore approximately 0.602 before phenotype specificity is applied.
Phenotype Specificity Multiplier
Phenotype State Multiplier Maximum Adjusted Trio LOD
Specific 1.0 Approximately 0.602
Broad 0.5 Approximately 0.301
Unspecified 0.25 Approximately 0.151
The multiplier is applied to the computed LOD score before evidence-band assignment. Under the default unspecified state, even a fully informative trio remains below the Supporting threshold.
Evidence Bands
LOD Range Band Points Interpretation
0.0 to <0.5 Indeterminate 0 Positive or neutral segregation signal below the Supporting threshold.
0.5 to <2.0 Supporting 1 Configured PP1 Supporting-equivalent band.
2.0 to <3.0 Moderate 2 Requires more informative meioses than a complete trio normally provides.
3.0 to <5.0 Strong 4 Extended-pedigree territory under the current scoring model.
>=5.0 Very Strong 8 Requires substantial segregation evidence beyond trio-only analysis.
Supported Inheritance Models
Autosomal recessive compound heterozygous.
Autosomal recessive homozygous.
X-linked recessive.
X-linked dominant.
Mitochondrial inheritance.
Uniparental disomy, using the current recessive-like scoring approximation.
Autosomal dominant de novo cases return zero segregation LOD in the current scorer because de novo evidence is handled through the separate de novo workflow rather than PP1-style familial segregation.
Consistency Flag and Drill-Down Results
The segregation-consistent flag is true for any adjusted LOD above zero, including scores that remain in the Indeterminate band. It therefore means that the observed transmission contributes some positive support, not that PP1 Supporting has been reached.
The drill-down view can include positive but indeterminate variants because the read path selects segregation-consistent rows as well as rows in a non-indeterminate evidence band. Results are ordered by LOD score and genomic position.
PP1 and BS4 Boundary
A numerical band is an input to clinical review. PP1 application depends on the gene-disease relationship, penetrance, phenocopies, pedigree structure, and whether the meioses are genuinely informative for the variant under review.
The current scorer maps negative values back to the Indeterminate band. It does not automatically assign BS4. Evidence against segregation requires separate clinical assessment of phenotype, penetrance, and possible locus or allelic heterogeneity.
Current Pedigree Limit
The implementation currently receives a complete trio. Moderate, Strong, and Very Strong segregation bands require additional informative relatives and are not normally reachable from trio-only data. Extended-pedigree processing remains outside the present production scope.
Related Documentation
Trio and Family Variant Analysis
Production workflow, inputs, outputs, and current scope.
Compound Heterozygous Variants
Parental-origin phasing, trans and cis states, and PM3 boundaries.
ACMG Criteria Reference
Reference for PP1, BS4, and the remaining ACMG criteria.
Family Analysis Methodology
Detailed methods, tools, thresholds, and limitations.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Family Analysis](https://folklore.helena.bio/docs/family-analysis)
- [Compound Heterozygous Variants Parental-origin phasing, trans and cis states, and PM3 boundaries.](https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes)
- [ACMG Criteria Reference Reference for PP1, BS4, and the remaining ACMG criteria.](https://folklore.helena.bio/docs/classification/criteria-reference)
- [Family Analysis Methodology Detailed methods, tools, thresholds, and limitations.](https://folklore.helena.bio/methodology/family-analysis)

---

## Page: Trio Quality Control with PLINK IBD Analysis | Folklore Documentation

Source: https://folklore.helena.bio/docs/family-analysis/trio-quality-control
Canonical: https://folklore.helena.bio/docs/family-analysis/trio-quality-control
Description: How Folklore checks declared trio relationships with PLINK 1.9 IBD analysis, interprets PI_HAT, Z0, Z1, and Z2, and handles sample-swap, duplicate-sample, and consanguinity alerts.

Documentation / Family Analysis / Trio Quality Control
## Trio Quality Control with PLINK IBD Analysis
Trio interpretation depends on the declared proband, mother, and father labels matching the biological relationships represented in the sequencing data. Folklore checks those relationships before relying on the trio for inheritance analysis.
The check estimates pairwise identity by descent with PLINK 1.9. It can flag relationship patterns compatible with a sample swap, duplicate sample, or related parents. It does not establish the cause of an unexpected result.
Why This Check Comes First
De novo detection, parental-origin phasing, and segregation scoring all depend on correctly assigned family roles. A critical identity alert makes those downstream inheritance results unreliable even when the variant-level computations completed successfully.
Input and PLINK Preparation
The QC stage reads the normalized joint multi-sample VCF and maps each VCF sample identifier to proband, mother, or father. Sample identifiers must match the joint VCF columns exactly.
Before pairwise IBD estimation, PLINK converts the VCF to binary format. Half-calls are treated as missing, only strictly biallelic variants are retained, variants below a minor-allele frequency of 0.05 are removed, and sites with more than 5 percent missing genotypes are excluded.
Pairwise IBD Output
Field Meaning
PI_HAT Estimated proportion of alleles shared identical by descent across the pair.
Z0 Estimated probability that the pair shares zero alleles identical by descent.
Z1 Estimated probability that the pair shares one allele identical by descent.
Z2 Estimated probability that the pair shares two alleles identical by descent.
PI_HAT is the primary threshold value used by the current alert rules. Z0, Z1, and Z2 remain visible for review and audit of the pairwise relationship pattern.
Current Alert Rules
Declared Pair Condition Alert Severity Handling
Parent and proband PI_HAT < 0.40 Sample swap Critical QC does not pass. Inheritance results should not drive interpretation until sample identity or pedigree labelling is resolved.
Parent and proband PI_HAT > 0.70 Duplicate sample Warning Possible duplicate sample or unusually high relatedness. The warning requires review but does not by itself set QC passed to false.
Mother and father PI_HAT > 0.125 Consanguinity Warning The declared parents show third-degree-or-closer relatedness under the configured threshold. The family context should be reviewed.
Expected Parent-Child Range
A declared parent-proband pair with PI_HAT from 0.40 through 0.70 does not trigger an alert in the current implementation. The interval is an operational QC range, not a statement that every value inside it represents the same biological relationship with equal certainty.
PI_HAT alone does not reliably distinguish every relationship class. Full siblings can also cluster near 0.50, and unusual family structure or population context can alter the observed pattern. The declared pedigree and the complete IBD profile still require review.
Critical and Warning States
The trio-level passed field is false only when at least one critical alert is present. In the current implementation, a low parent-child PI_HAT creates that critical state.
Duplicate-sample and consanguinity alerts are warnings. They remain visible in the QC record, but a warning alone does not change passed to false. This distinction should not be read as clinical clearance; warning findings still require review.
Skipped Check Versus Passed Check
The QC contract records whether IBD analysis was performed. A legacy or incomplete run without a joint VCF can carry ibd_check_performed=false, reason_skipped=no_joint_vcf, and passed=true.
In that state, passed does not mean that biological relationships were confirmed. It means that no critical PLINK result was produced because the IBD check did not run. The performed flag and skip reason must be read with the passed field.
Execution Failure
A PLINK timeout, malformed output, missing required output column, missing input VCF, or failed subprocess is not converted into a neutral QC result. The trio pipeline fails at the sample_qc stage and records the failure for recovery.
This separates an unavailable computation from an evaluated trio that produced no critical alerts.
Clinical Review Boundary
A QC alert identifies a relationship pattern that needs investigation. It does not prove a sample swap, duplicate, or specific degree of consanguinity. Resolution can require laboratory records, sample identifiers, pedigree confirmation, and repeat or orthogonal testing.
Recorded QC Data
Folklore stores the pairwise IBD values, alerts, pass state, critical-alert flag, PLINK version, and optional skip reason in the trio QC summary. The same summary is written to the versioned trio_qc.json artifact with its generation time and family identifier.
The family workspace displays the pair table in the Members section and places critical sample-integrity alerts above the family result tabs so that the warning remains visible outside the QC table.
Related Documentation
Trio and Family Variant Analysis
Production workflow, required inputs, outputs, and current scope.
De Novo Variant Analysis
How parental genotypes support or weaken de novo origin.
Compound Heterozygous Variants
Parental-origin phasing and trans or cis interpretation.
Segregation Analysis
Per-variant LOD scoring and PP1 or BS4 review boundaries.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Family Analysis](https://folklore.helena.bio/docs/family-analysis)
- [De Novo Variant Analysis How parental genotypes support or weaken de novo origin.](https://folklore.helena.bio/docs/family-analysis/de-novo-variants)
- [Compound Heterozygous Variants Parental-origin phasing and trans or cis interpretation.](https://folklore.helena.bio/docs/family-analysis/compound-heterozygotes)
- [Segregation Analysis Per-variant LOD scoring and PP1 or BS4 review boundaries.](https://folklore.helena.bio/docs/family-analysis/segregation-analysis)

---

## Page: Getting Started | Folklore Documentation

Source: https://folklore.helena.bio/docs/getting-started
Canonical: https://folklore.helena.bio/docs/getting-started
Description: Get started with Folklore -- upload a VCF file, set patient phenotype, and understand your variant analysis results.

Documentation / Getting Started
## Getting Started
Folklore validates and annotates the submitted VCF, applies the relevant germline classification framework, and prepares phenotype, literature, screening, and reporting evidence for clinical review. Family, mitochondrial, cohort, and structural-variant modules extend the core case workflow when their input data is present.
Core Variant Processing
1
VCF Parsing
Your file is read and loaded into the analysis engine.
2
Quality Filtering
Low-quality variants are flagged. ClinVar pathogenic variants are never discarded.
3
Annotation
Variant consequences predicted by Ensembl VEP (coding impact, protein effect, splice region).
4
Reference Database Enrichment
Population frequencies, clinical significance, functional predictions, gene constraint, phenotype associations, and dosage sensitivity loaded from 16 production classification and reference databases. Across all modules, Folklore integrates 45 reference databases and curated sources.
5
ACMG Classification
Applicable evidence criteria are evaluated, Bayesian combining rules are applied, and the complete evidence path is recorded.
6
Export
Results ready for clinical review. Gene-level summaries available immediately.
Processing Time
Runtime depends on variant count, input size, enabled modules, and current server load. Each case records its actual processing duration and stage-level status rather than relying on a fixed turnaround claim.
What the Platform Does Not Do
Folklore does not diagnose. It does not replace the geneticist. It automates the evidence-gathering step of variant interpretation. The geneticist reviews the evidence, applies clinical judgment, considers family history and clinical context, and makes the final clinical decision.
Maximum Sensitivity Approach
No frequency-based or impact-based pre-filtering is applied at any stage. A common variant with a gnomAD allele frequency of 40% still receives a classification -- it will be classified as Benign via BA1, but it is not silently discarded before classification. Nothing is hidden from the reviewing geneticist. The clinician decides clinical relevance based on the complete classification and annotation data.
In This Section
Uploading a VCF File
Supported formats, genome build requirements, and what happens after upload.
Setting HPO Terms
How to provide patient phenotype information and why it matters for prioritization.
Understanding Results
How to read the results interface -- classifications, scores, and annotations explained.
Quality Presets
Three quality filtering levels and the ClinVar rescue mechanism.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Uploading a VCF File Supported formats, genome build requirements, and what happens after upload.](https://folklore.helena.bio/docs/getting-started/uploading-vcf)
- [Setting HPO Terms How to provide patient phenotype information and why it matters for prioritization.](https://folklore.helena.bio/docs/getting-started/setting-hpo-terms)
- [Understanding Results How to read the results interface -- classifications, scores, and annotations explained.](https://folklore.helena.bio/docs/getting-started/understanding-results)
- [Quality Presets Three quality filtering levels and the ClinVar rescue mechanism.](https://folklore.helena.bio/docs/getting-started/quality-presets)

---

## Page: Quality Presets | Folklore Documentation

Source: https://folklore.helena.bio/docs/getting-started/quality-presets
Canonical: https://folklore.helena.bio/docs/getting-started/quality-presets
Description: Three quality filtering levels for variant analysis in Folklore -- Strict, Balanced, and Permissive -- with the ClinVar rescue mechanism.

Documentation / Getting Started / Quality Presets
## Quality Presets
Quality filtering occurs before annotation and classification. Three configurable presets control the stringency of variant filtering based on sequencing quality metrics.
Preset Quality (QUAL) Depth (DP) Genotype Quality (GQ) Recommended Use
Strict >= 30 >= 20 >= 30 High-confidence clinical reporting
Balanced >= 20 >= 15 >= 20 Standard clinical analysis (default)
Permissive >= 10 >= 10 >= 10 Maximum sensitivity / research
When to Use Each Preset
Strict
Use for final clinical reports where only high-confidence variant calls are acceptable. Minimizes false positives at the cost of potentially missing variants in low-coverage regions.
Balanced
Recommended for most clinical analyses. Provides a good balance between sensitivity and specificity. This is the default preset.
Permissive
Use in research settings, for reanalysis of older or lower-quality sequencing data, or in cases where maximum sensitivity is needed even at the cost of more false positives. Review flagged variants carefully.
ClinVar Rescue Mechanism
Regardless of the selected preset, variants with documented ClinVar Pathogenic or Likely Pathogenic significance that fail quality thresholds are not discarded. They are flagged as "rescued" variants and proceed through the full classification pipeline. This prevents clinically significant findings from being silently excluded due to sequencing quality in low-coverage regions.
When reviewing rescued variants, pay careful attention to the quality metrics. Low sequencing coverage means the genotype call itself may be unreliable. Consider confirmation by Sanger sequencing for rescued variants that are clinically actionable.
For the complete quality filtering documentation with technical details, see the Methodology page.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Getting Started](https://folklore.helena.bio/docs/getting-started)
- [Methodology](https://folklore.helena.bio/methodology#quality-filtering)

---

## Page: Setting HPO Terms | Folklore Documentation

Source: https://folklore.helena.bio/docs/getting-started/setting-hpo-terms
Canonical: https://folklore.helena.bio/docs/getting-started/setting-hpo-terms
Description: How to provide patient phenotype information using HPO terms and why it improves variant analysis in Folklore.

Documentation / Getting Started / Setting HPO Terms
## Setting HPO Terms
What Are HPO Terms
The Human Phenotype Ontology (HPO) is a standardized clinical vocabulary for phenotypic abnormalities observed in human disease. It contains over 17,000 terms organized in a hierarchical structure where specific terms (such as "Focal clonic seizure", HP:0002266) are children of broader terms (such as "Seizure", HP:0001250). Using standardized HPO terms enables computational comparison across patients, diseases, and databases.
Why HPO Terms Matter
Providing HPO terms enables three platform features that significantly improve the clinical relevance of results. First, the PP4 ACMG criterion activates when patient phenotype matches a gene with a known disease association, adding supporting pathogenic evidence. Second, the Phenotype Matching service scores every candidate gene against the patient's presentation, producing a 0-100 similarity score and clinical tier assignment. Third, the Screening service uses phenotype correlation to boost prioritization of variants in phenotype-relevant genes.
Selecting HPO Terms
Be specific. "Seizure" (HP:0001250) provides more discrimination than "Abnormality of the nervous system" (HP:0012638). Include findings from all affected organ systems, not just the primary complaint. Negative findings with clinical significance should also be included -- the platform handles negation. Five to fifteen specific terms is optimal for most cases.
The platform provides two input methods: manual HPO term search and selection with real-time autocomplete, and free-text clinical description where the platform automatically extracts HPO terms using ontology matching with negation detection.
HPO Term Selection by Clinical Domain
Neurodevelopmental
Seizure types and onset age, developmental milestones (sitting, walking, speech), brain MRI findings, EEG patterns, behavioral features.
Cardiology
Specific cardiomyopathy type (dilated, hypertrophic, restrictive), arrhythmia pattern, ECG findings, echocardiographic measurements, family history of sudden cardiac death.
Nephrology
Specific renal finding (cysts, proteinuria, hematuria), biopsy findings, extrarenal manifestations.
Metabolic
Specific metabolites elevated or decreased, enzyme activity levels, organ involvement, response to treatment.
Neonatal
Gestational age, birth parameters, feeding difficulties, hypotonia, seizure onset, congenital anomalies, metabolic screening results.
Analysis Without HPO Terms
The platform works fully without HPO terms. Variant classification, annotation, and all non-phenotype-dependent criteria proceed normally. However, phenotype-dependent features -- PP4, phenotype matching, and phenotype-based screening boosts -- will not be available. For clinical cases, providing HPO terms is strongly recommended.
For more on how phenotype information is used in prioritization, see Phenotype Matching .

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Getting Started](https://folklore.helena.bio/docs/getting-started)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)

---

## Page: Understanding Results | Folklore Documentation

Source: https://folklore.helena.bio/docs/getting-started/understanding-results
Canonical: https://folklore.helena.bio/docs/getting-started/understanding-results
Description: How to read and interpret Folklore variant analysis results -- ACMG classifications, annotations, scores, and clinical tiers.

Documentation / Getting Started / Understanding Results
## Understanding Results
Results Organization
Results are organized by gene, not by individual variant. Each gene card shows a summary of the most clinically relevant variant in that gene. Expanding the card reveals all individual variants with their complete annotation data. This gene-centric view helps the geneticist focus on the biological context rather than individual base changes.
Variant Annotations
Each variant displays the following fields. Together, these provide the evidence a geneticist needs to assess clinical relevance.
ACMG Classification
Pathogenic, Likely Pathogenic, VUS, Likely Benign, or Benign. Determined by the Bayesian point system.
ACMG Criteria
Which specific criteria were triggered (e.g., PVS1, PM2, PP3_Strong) and their evidence strength levels.
Confidence Score
A continuous score reflecting the strength of classification evidence. Higher values indicate greater confidence in the classification.
Consequence
The predicted effect on the transcript: frameshift, missense, synonymous, splice donor, stop gained, and others.
Impact
Severity classification from Ensembl VEP: HIGH, MODERATE, LOW, or MODIFIER.
HGVS Notation
Standardized variant description at genomic (g.), coding (c.), and protein (p.) levels.
Genotype
Heterozygous (0/1), homozygous alternate (1/1), or hemizygous. Indicates the number of copies of the variant allele.
Population Frequency
gnomAD global allele frequency and population-specific maximum frequency. Lower frequency generally indicates higher clinical relevance.
ClinVar Significance
What ClinVar reports for this variant, with the review star level indicating assertion quality.
Computational Predictors
BayesDel score (used in classification) plus display predictors: SIFT, AlphaMissense, MetaSVM, DANN, PhyloP, GERP.
SpliceAI Scores
Four delta scores predicting splice impact: acceptor gain, acceptor loss, donor gain, donor loss.
Gene Constraint
pLI, LOEUF, and observed/expected loss-of-function ratio. Indicates how tolerant the gene is to damaging variation.
HPO Associations
Which clinical phenotypes are associated with this gene in the HPO database.
Phenotype Match Score
How well this gene matches the patient presentation (0-100), if HPO terms were provided.
Screening Tier
Tier 1 through 4 priority classification based on multi-dimensional scoring.
Literature Evidence
Relevant publications from the local PubMed database, ranked by clinical relevance.
What Each ACMG Classification Means
Classification Clinical Meaning Action
Pathogenic Variant causes disease. Report to clinician.
Likely Pathogenic Strong evidence variant causes disease (>90% certainty per ACMG). Report to clinician.
VUS Insufficient evidence to classify. May require additional evidence gathering.
Likely Benign Strong evidence variant does not cause disease. Generally not reported.
Benign Variant does not cause disease. Not reported.
Where to Start
The Screening Tier 1 list provides the highest-priority variants based on classification strength, phenotype correlation, population rarity, and gene constraint. Start there. If Tier 1 does not explain the clinical presentation, proceed to Tier 2 and consider additional evidence gathering for VUS candidates with strong phenotype matches.
For details on how screening tiers are assigned, see Screening Tier System . For classification methodology, see the full Methodology page.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Getting Started](https://folklore.helena.bio/docs/getting-started)
- [Screening Tier System](https://folklore.helena.bio/docs/screening/tier-system)
- [Methodology](https://folklore.helena.bio/methodology)

---

## Page: Uploading a VCF File | Folklore Documentation

Source: https://folklore.helena.bio/docs/getting-started/uploading-vcf
Canonical: https://folklore.helena.bio/docs/getting-started/uploading-vcf
Description: Supported VCF formats, genome build requirements, file size limits, and what happens after upload in Folklore.

Documentation / Getting Started / Uploading a VCF File
## Uploading a VCF File
Supported Formats
Folklore accepts VCF files in version 4.1 and 4.2 format, either as plain text (.vcf) or bgzipped (.vcf.gz). Whole genome sequencing files containing approximately 4 million variants (1-2 GB compressed) are fully supported.
Genome Build Requirement
GRCh38 (hg38) is the primary genome build. GRCh37 (hg19) files are also accepted -- the platform performs automatic liftover to GRCh38 during upload, so no manual conversion is required. The liftover step is transparent and reported in the processing summary.
Single Samples and Families
Primary analysis accepts one sample per VCF. Multi-sample VCF files should be split before upload. For family analysis, process the proband, parents, and optional siblings separately, then link the resulting cases in the Family Analysis module for relationship QC, de novo review, compound-heterozygous phasing, and segregation evidence.
See Family Analysis Methodology .
What Happens After Upload
Once uploaded, the file is checked for format and genome-build compatibility. The core path parses and quality-checks the variants, runs local annotation and reference enrichment, applies the relevant classifier, and stores the evidence for clinical review. Variants on chrM and supported structural-variant records are routed to their dedicated interpretation modules.
Expected chromosomes are chr1 through chr22, chrX, chrY, and chrM. Non-standard contigs such as decoys, alternative haplotypes, and unplaced scaffolds are skipped, with a count reported in the processing summary.
Data Handling
The uploaded VCF file is deleted from the server after processing completes. Analysis results are retained for the duration specified in your service agreement. All processing occurs on dedicated EU-based infrastructure in Helsinki, Finland. No variant data is sent to external services during processing. See Data and Privacy for full details.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Getting Started](https://folklore.helena.bio/docs/getting-started)
- [Family Analysis Methodology](https://folklore.helena.bio/methodology/family-analysis)
- [Data and Privacy](https://folklore.helena.bio/docs/data-and-privacy)

---

## Page: Glossary | Folklore Documentation

Source: https://folklore.helena.bio/docs/glossary
Canonical: https://folklore.helena.bio/docs/glossary
Description: Definitions of key terms used throughout the Folklore documentation -- ACMG criteria, genomics concepts, and platform-specific terminology.

Documentation / Glossary
## Glossary
Key terms used throughout the Folklore documentation.
ACMG
American College of Medical Genetics and Genomics. Publisher of the 2015 variant classification guidelines used by Folklore.
ACMG Secondary Findings (SF)
A curated list of 81 genes (v3.2) where pathogenic variants should be reported regardless of the primary testing indication, because early identification can lead to medical interventions.
Allele Frequency (AF)
The proportion of a specific allele in a population. Used for BA1 (>5%), BS1, and PM2 (<0.01%) criteria. Sourced from gnomAD.
AlphaMissense
A protein structure-based pathogenicity predictor from DeepMind. Displayed for clinical reference in Folklore but does not contribute to ACMG classification.
BayesDel_noAF
The ClinGen SVI-calibrated computational meta-predictor used by Folklore for PP3 and BP4 ACMG criteria. The "_noAF" suffix indicates allele frequency is excluded from its model.
Benign
ACMG classification for a variant with a Bayesian total of -6 or below, or a valid stand-alone BA1 result.
BP4
Computational benign evidence. For missense variants, Folklore uses the calibrated BayesDel_noAF benign ranges; supported non-canonical splice variants may use the documented SpliceAI benign path.
ClinGen
Clinical Genome Resource. Provides dosage sensitivity scores and gene-specific ACMG classification specifications (VCEP). Funded by NIH.
ClinVar
NCBI database of clinical significance assertions for genetic variants. Folklore uses ClinVar for PS1, PP5, BP6 criteria and classification override.
Compound Heterozygote
Two different heterozygous variants in the same gene that are established or suspected to be in trans. Their clinical significance depends on the variants, gene mechanism, phase, and phenotype.
Confidence Score
A continuous score (0.0-1.0) reflecting how far a variant is from the classification boundary. Higher scores indicate more certain classifications.
Consequence
The predicted effect of a variant on the gene product, determined by Ensembl VEP. Examples: missense_variant, frameshift_variant, splice_donor_variant.
DANN
Deep Annotation of Noncoding Variants. A deep neural network pathogenicity score applicable to any single nucleotide variant. Displayed for reference.
dbNSFP
Database for Nonsynonymous SNPs Functional Predictions. Version 4.9c provides BayesDel, SIFT, AlphaMissense, MetaSVM, DANN, PhyloP, and GERP scores.
DuckDB
An in-process analytical database engine used by Folklore for variant storage and querying. Each analysis session uses an isolated DuckDB file.
Ensembl VEP
Variant Effect Predictor from EMBL-EBI. Determines variant consequences, transcript selection, and functional annotations. Runs locally with offline cache.
GERP
Genomic Evolutionary Rate Profiling. Measures evolutionary constraint at a genomic position. Scores >4.0 indicate strongly constrained sites.
gnomAD
Genome Aggregation Database. Version 4.1 provides population allele frequencies from 807,162 individuals across 8 genetic ancestry groups.
GRCh38
Genome Reference Consortium Human Build 38 (hg38). The current standard human reference genome used by Folklore.
Haploinsufficiency
A condition where loss of one gene copy (one allele) is sufficient to cause disease. ClinGen curates haploinsufficiency scores (0-3).
HGVS
Human Genome Variation Society nomenclature. Standard notation for describing variants at the genomic (g.), coding DNA (c.), and protein (p.) levels.
HPO
Human Phenotype Ontology. A standardized vocabulary of over 17,000 phenotypic abnormalities used for phenotype matching and PP4 criteria.
Impact
VEP-assigned severity category: HIGH (frameshift, stop gained), MODERATE (missense), LOW (synonymous), MODIFIER (intronic, UTR).
Likely Benign
ACMG classification for a Bayesian total from -1 through -5.
Likely Pathogenic
ACMG classification indicating the variant is probably disease-causing. Bayesian points between 6 and 9.
LOEUF
Loss-of-function Observed/Expected Upper bound Fraction. Lower values indicate stronger gene constraint. Values < 0.35 are considered highly constrained.
MetaSVM
A Support Vector Machine meta-predictor combining 10 individual prediction tools. Displayed for clinical reference.
Pathogenic
ACMG classification for a Bayesian total of 10 or more, subject to the applicable priority and conflict rules.
PhyloP
Phylogenetic P-value. Measures evolutionary conservation across 100 vertebrate species. Scores >2.0 indicate conserved positions.
pLI
Probability of Loss-of-function Intolerance. Ranges 0-1. Values > 0.9 indicate the gene is highly intolerant to loss-of-function variants.
PP3
ACMG pathogenic supporting criterion: computational evidence supports a deleterious effect. In Folklore, triggered by BayesDel_noAF or SpliceAI scores.
PVS1
Loss-of-function evidence in a gene where LoF is an established disease mechanism. Folklore applies transcript, expression, mechanism, disease-association, and positional guards before assigning the available strength.
Screening Tier
Priority ranking assigned by the Screening Service. Tier 1 (immediate review), Tier 2 (moderate priority), Tier 3 (low priority), Tier 4 (very low priority).
SIFT
Sorting Intolerant From Tolerant. Predicts amino acid substitution tolerance based on sequence homology. Scores < 0.05 are Deleterious.
SpliceAI
A deep learning model predicting splice site disruption. Produces four delta scores. Threshold >= 0.2 triggers PP3_splice in Folklore.
VCF
Variant Call Format. The standard file format for genomic variant data. Folklore accepts VCF 4.1 and 4.2 files (plain text or bgzipped).
VCEP
Variant Curation Expert Panel. ClinGen panels that define gene-specific ACMG classification thresholds.
VUS
Variant of Uncertain Significance. ACMG classification indicating insufficient evidence to classify as pathogenic or benign. Bayesian points between 0 and 5.
WES
Whole Exome Sequencing. Captures coding regions of the genome. Typically produces 40,000-60,000 variants per sample.
SV/CNV
Constitutional structural and copy-number variants, including supported deletions and duplications, interpreted in Folklore under the Riggs 2020 framework.
MMDWG
ClinGen Mitochondrial Disease Working Group specification used by the dedicated Folklore mtDNA classifier.
WGS
Whole Genome Sequencing. Captures the entire genome. Typically produces 4-5 million small-variant records per sample, depending on the upstream caller and filtering.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)

---

## Page: Limitations | Folklore Documentation

Source: https://folklore.helena.bio/docs/limitations
Canonical: https://folklore.helena.bio/docs/limitations
Description: Known limitations of the Folklore platform -- variant types, analysis scope, AI boundaries, and database coverage.

Documentation / Limitations
## Limitations
Folklore is designed for germline variant interpretation in Mendelian disease contexts. The following limitations should be considered when using the platform for clinical genomic analysis.
Variant Types
Nuclear SNVs and small insertions/deletions are processed through the ACMG/AMP classifier.
Constitutional deletions and duplications are processed through the dedicated SV/CNV module under the Riggs 2020 scoring framework. The module does not replace upstream structural-variant calling.
Mitochondrial variants use a separate MMDWG 2020 classifier with mtDNA-specific frequency, haplogroup, heteroplasmy, predictor, and NUMT evidence.
Repeat expansions are not detected or interpreted. Disorders caused by tandem-repeat expansion require a dedicated upstream caller and interpretation workflow.
Somatic tumor variants are not supported. Tumor-normal analysis and AMP/ASCO/CAP somatic classification remain outside the product scope.
Genome Build
GRCh38 (hg38) is the primary supported genome build. All reference databases are indexed to GRCh38.
GRCh37 (hg19) VCF files are accepted and automatically lifted over to GRCh38 using CrossMap. Liftover may fail for a small number of variants in regions with structural differences between builds.
Earlier genome builds (hg18 and prior) are not supported.
ACMG Classification
The nuclear classifier computes evidence that can be derived from variant annotations, reference data, phenotype input, and supported family context. Functional assays, case-control evidence, and other expert-curated observations still require reviewer input.
Gene-specific VCEP specifications are applied where available. Genes without a supported VCEP overlay use the documented general thresholds and guards.
Every automated class remains subject to geneticist review. The software records the applied criteria and evidence path but does not replace clinical judgment.
A full trio can establish parental origin for candidate compound-heterozygous pairs. Without informative parental or read-based phasing, the pair remains a candidate rather than a confirmed in-trans result.
Database Coverage
gnomAD v4.1 population frequencies have uneven coverage across ancestry groups. Non-Finnish European samples represent the majority of the cohort. Variants in underrepresented populations may have less precise frequency estimates.
ClinVar assertions vary in quality. Single-submitter assertions with zero review stars carry less weight than expert panel consensus. The platform applies review star filtering but cannot independently validate submitter quality.
dbNSFP functional predictions are limited to missense variants at coding positions. Non-coding variants, UTR variants, and deep intronic variants do not receive BayesDel, SIFT, or AlphaMissense scores.
HPO gene-phenotype associations depend on the completeness of disease gene curation. Recently discovered gene-disease associations may not yet be reflected in the HPO database.
The literature database is a filtered local PubMed snapshot, not a live mirror. Freshness is not guaranteed, so recent, corrected, retracted, or exhaustive literature must be verified against live PubMed and the journal record.
AI Clinical Assistant
The AI assistant generates clinical interpretations that require independent validation by a qualified clinical geneticist. AI-generated content does not constitute a medical diagnosis.
The assistant only has access to data within the current analysis session. It cannot cross-reference findings across different patients or historical cases.
Conversation context is limited to the last 20 messages. Very long diagnostic discussions may lose early context.
SQL generation from natural language may occasionally produce incorrect queries for ambiguous or highly complex questions. The generated SQL is visible to the user for verification.
The literature search operates on the local PubMed mirror. It cannot access preprint servers, full-text content, or databases outside the indexed PubMed corpus.
Input Requirements
Primary variant analysis accepts single-sample VCF files. Family analysis links separately processed proband, parent, and sibling samples into a trio, duo, or extended family configuration.
VCF files must conform to VCF 4.1 or 4.2 specification. Non-standard fields in the INFO or FORMAT columns are ignored.
The platform processes up to approximately 5 million variants per file. Whole genome sequencing files typically contain 4-5 million variants and are fully supported.
Phenotype matching requires HPO terms. Free-text clinical descriptions are supported through automatic HPO term extraction, but manually curated HPO terms produce better matching accuracy.
Intended Use
Folklore is a clinical decision support tool for germline variant interpretation. It is not a diagnostic device. All automated classifications, screening priorities, and AI-generated interpretations must be independently reviewed and validated by qualified clinical professionals before being used in patient care decisions.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)

---

## Page: Literature Evidence | Folklore Documentation

Source: https://folklore.helena.bio/docs/literature
Canonical: https://folklore.helena.bio/docs/literature
Description: How Folklore retrieves, ranks, and presents locally indexed PubMed literature for clinical review.

## Literature Evidence
Folklore searches a local, pre-processed PubMed dataset for publications relevant to a case. The literature stage is integrated with the clinical pipeline and is intended to reduce the candidate set that a geneticist must review, not to replace a systematic literature search.
Results are organized by gene and retain the publication metadata and matching signals that produced their rank. If Phenotype Matching data is available, the user interface also combines clinical priority with literature relevance when ordering gene groups.
Production Workflow
1
Build the case query
The pipeline uses reportable genes and variants from Variant Analysis. Patient HPO terms enrich the query when they are available.
2
Find candidate publications
Folklore searches locally indexed gene mentions and the locally stored publication title and abstract, then deduplicates the candidate set.
3
Assemble review signals
Candidates are associated with available gene mentions, variant notation, MeSH terms, publication metadata, and indicators of possible functional work.
4
Rank for review
A multi-signal relevance model orders publications by phenotype overlap, publication context, gene focus, functional indicators, variant match, and recency.
5
Present the evidence trail
The interface shows the relevance score, its component signals, triage label, abstract, matched entities, and source links.
Local Clinical Search
Case searches run against the locally hosted literature database on EU infrastructure. Patient genes, variants, and phenotype terms are not sent to PubMed or another external search API during clinical retrieval. Database ingestion is a separate maintenance process.
Important Boundary
Relevance scores and evidence labels are search-and-triage aids. They do not establish an ACMG/AMP criterion, validate a functional assay, prove that a reported variant is identical in transcript context, or determine whether a publication is applicable to the patient. The source publication must be reviewed before evidence is used clinically.
In This Section
Literature Evidence
Case inputs, publication discovery, reported fields, and the review workflow.
Relevance Ranking
The clinical signals used to order publications and how to interpret the score.
Evidence Labels
What the Strong, Moderate, Supporting, and Weak triage labels mean.
PubMed Coverage
Current database scope, ingestion filters, freshness, and known gaps.

### Links and cited sources

- [Literature Evidence Case inputs, publication discovery, reported fields, and the review workflow.](https://folklore.helena.bio/docs/literature/literature-evidence)
- [Relevance Ranking The clinical signals used to order publications and how to interpret the score.](https://folklore.helena.bio/docs/literature/relevance-scoring)
- [Evidence Labels What the Strong, Moderate, Supporting, and Weak triage labels mean.](https://folklore.helena.bio/docs/literature/evidence-strength)
- [PubMed Coverage Current database scope, ingestion filters, freshness, and known gaps.](https://folklore.helena.bio/docs/literature/pubmed-coverage)

---

## Page: Literature Evidence Labels | Folklore Documentation

Source: https://folklore.helena.bio/docs/literature/evidence-strength
Canonical: https://folklore.helena.bio/docs/literature/evidence-strength
Description: How to interpret the Strong, Moderate, Supporting, and Weak labels in Folklore literature results.

Documentation / Literature Evidence / Evidence Labels
## Evidence Labels
Strong, Moderate, Supporting, and Weak are Folklore literature-triage labels. They summarize detected search signals so a geneticist can prioritize reading. They are not ACMG/AMP evidence strengths and do not apply a classification criterion.
Label Definitions
Strong
Prioritized for early review because the abstract-level record contains both variant-specific and experimental signals.
Detected signals: Exact variant indicator and functional-data indicator are both present.
Moderate
Potentially useful variant-specific or experimental context is present, but the other signal is absent.
Detected signals: Either the exact variant indicator or the functional-data indicator is present.
Supporting
The paper may help assess gene-phenotype relevance but has no detected exact-variant or functional signal.
Detected signals: A query gene and at least one patient phenotype name match.
Weak
The publication remains in the review set as lower-priority background or a possible extraction miss.
Detected signals: The stronger combinations above are not present.
What the Indicators Establish
Variant indicator
Shows that compatible gene and variant notation were detected in the indexed title or abstract record. It does not establish that the paper studied the same transcript, allele, phenotype, or inheritance context.
Functional indicator
Shows that MeSH metadata or abstract language suggests experimental work. It does not validate the assay or establish a well-established functional effect.
Before Applying Clinical Evidence
Read the full methods and results, not only the PubMed abstract.
Confirm transcript, genomic build, allele, zygosity, inheritance, and phenotype context.
Evaluate assay controls, calibration, biological relevance, and replication before considering functional criteria.
Check publication corrections, expressions of concern, and retraction status at the source.
Apply the current ACMG/AMP and ClinGen specifications independently of the Folklore triage label.
Do Not Translate Labels into ACMG Codes
A Strong literature label is not PS3, a Supporting label is not PP4, and no label is PP5. The automated Variant Analysis classifier and the clinician's evidence review remain separate from literature triage.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Literature Evidence](https://folklore.helena.bio/docs/literature)

---

## Page: Literature Evidence Workflow | Folklore Documentation

Source: https://folklore.helena.bio/docs/literature/literature-evidence
Canonical: https://folklore.helena.bio/docs/literature/literature-evidence
Description: How Folklore constructs a case literature query, retrieves local PubMed records, and presents results for clinical review.

Documentation / Literature Evidence / Workflow
## Literature Evidence Workflow
Literature retrieval normally runs near the end of a case pipeline and is non-blocking: a literature-stage failure does not invalidate the completed Variant Analysis result. When results are available, they are saved with the session and streamed to the interface.
Case Inputs
Genes
Reportable genes and clinically prioritized genes define the primary retrieval scope.
Variants
Protein and cDNA HGVS notation can be compared with variant mentions extracted from titles and abstracts.
Phenotype
Resolved HPO names can enrich ranking when patient phenotype is available; HPO IDs alone are not searched as publication text.
Publication Discovery
Indexed gene mentions
The local gene-mention table provides a structured path to publications in which a candidate gene was identified during ingestion.
Title and abstract fallback
A broader case-insensitive text lookup finds additional title or abstract records. This improves recall but can also introduce incidental or substring matches that require review.
Candidate PMIDs are deduplicated before evidence assembly. The scoring stage evaluates a bounded candidate set and returns a bounded list for the session rather than every matching PubMed record.
What Is Reported
PMID, title, abstract, authors, journal, publication date, DOI, and PubMed Central identifier when present.
Overall literature relevance and the visible phenotype, publication-context, gene-focus, functional, variant, and recency components.
Matched query genes, matched variant notation, and matched phenotype names.
A heuristic functional-data indicator and a Strong, Moderate, Supporting, or Weak triage label.
Links to PubMed and, when a PMC identifier exists, to PubMed Central full text.
How Results Appear
The interface groups publications by matched gene. Within each group, papers are ordered by literature relevance. Gene groups can also incorporate the Phenotype Matching clinical priority, display the associated clinical tier and phenotype rank, and be filtered by gene or tier.
Read the Source
Extraction operates on PubMed metadata, titles, abstracts, publication types, and MeSH descriptors. It does not interpret the full article. Methods, cohort details, transcript context, assay validity, segregation, conflicts, and limitations must be checked in the publication itself.
Continue with Relevance Ranking or review the Evidence Labels .

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Literature Evidence](https://folklore.helena.bio/docs/literature)
- [Relevance Ranking](https://folklore.helena.bio/docs/literature/relevance-scoring)
- [Evidence Labels](https://folklore.helena.bio/docs/literature/evidence-strength)

---

## Page: PubMed Coverage and Freshness | Folklore Documentation

Source: https://folklore.helena.bio/docs/literature/pubmed-coverage
Canonical: https://folklore.helena.bio/docs/literature/pubmed-coverage
Description: The current scope, ingestion filters, freshness limits, and content boundaries of Folklore literature data.

Documentation / Literature Evidence / PubMed Coverage
## PubMed Coverage and Freshness
Folklore uses a filtered local dataset derived from NCBI PubMed baseline and update files. It is not a live mirror of PubMed and does not contain the complete biomedical literature corpus.
Current Operational Snapshot
Source NCBI PubMed XML distribution Indexed publications 1,012,962 Extracted gene mentions 403,803 Extracted variant mentions 102,630 MeSH-derived phenotype records 12,095,118 Configured publication years 1990 onward Snapshot verified 5 August 2026
Counts describe the deployed database at the verification date and can change after a promoted ingestion. Mention counts are extracted records, not unique genes, variants, or diseases.
Freshness Is Not Guaranteed
The service supports PubMed update-file ingestion through a separate staging database and explicit promotion to production. The deployed snapshot is not guaranteed to be current with live PubMed, and this documentation does not promise a daily refresh interval. Search PubMed directly when recent or exhaustive coverage matters.
Ingestion Filters
A publication must meet all configured gates: at least one selected genetics-related MeSH descriptor, at least one accepted publication type, and a publication year within the configured range.
MeSH topics include
Mutation
Genetic Variation
Single-Nucleotide Polymorphism
DNA Sequence Analysis
Genotype
Phenotype
Alleles
Genetic Predisposition to Disease
Accepted types include
Journal Article
Case Reports
Clinical Study
Clinical Trial
Systematic Review
Meta-Analysis
Indexed Content
Publication metadata
PMID, title, abstract, authors, journal, date, publication types, MeSH descriptors, DOI, and PMC identifier when supplied by PubMed.
Gene mentions
Uppercase gene-like tokens from titles and abstracts are screened against a strict human protein-coding gene validation path during ingestion.
Variant mentions
Selected cDNA, protein, and legacy notation patterns are extracted from titles and abstracts and associated with a nearby validated gene when possible.
Phenotype records
The current ingestion path stores PubMed MeSH descriptors as phenotype names. It does not convert those records into HPO or OMIM identifiers.
Not Reliably Covered
Publications added to or corrected in PubMed after the deployed snapshot.
Articles that fail the configured MeSH, publication-type, or date gates, even when clinically relevant.
Preprints and literature sources outside PubMed.
Full article text, supplementary material, tables, figures, and paywalled methods or results.
Variant notation outside the implemented extraction patterns, complex HGVS expressions, and variants whose nearby gene cannot be established.
A guaranteed current retraction or correction status. Verify the live PubMed and journal record before clinical use.
Database Promotion
Baseline and update ingestion write to a separate staging database. A verified staging database can then be promoted to production while preserving the previous clinical-search dataset during ingestion. Promotion is an operational action, not an automatic consequence of receiving a PubMed update file.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Literature Evidence](https://folklore.helena.bio/docs/literature)

---

## Page: Literature Relevance Ranking | Folklore Documentation

Source: https://folklore.helena.bio/docs/literature/relevance-scoring
Canonical: https://folklore.helena.bio/docs/literature/relevance-scoring
Description: The evidence signals Folklore uses to rank locally indexed publications for clinical review.

Documentation / Literature Evidence / Relevance Ranking
## Relevance Ranking
Folklore assigns each candidate publication a normalized relevance score for ordering the review list. The score combines six visible evidence dimensions. It represents search relevance to this case, not study quality, pathogenicity, or ACMG evidence strength.
Signals Used
Phenotype overlap
Compares resolved patient HPO names with structured MeSH-derived terms and words in the title and abstract. Morphological normalization helps with common word forms but is not ontology-based semantic similarity.
Publication context
Uses the indexed publication type to distinguish clinical reports, studies, general journal articles, reviews, and other publication contexts.
Gene focus
Uses the available mention count as a signal for whether a query gene is central to the abstract or merely incidental.
Functional-study indicator
Looks for curated MeSH descriptors and title or abstract language associated with models, assays, expression, splicing, or other experimental work.
Variant match
Compares the query gene and supplied protein or cDNA notation with variant mentions extracted from the publication record.
Recency
Adds publication age as a ranking signal. Older evidence is not treated as invalid; recency only influences ordering.
Phenotype Matching in Literature
The literature ranker uses the human-readable name of each selected HPO term. It can match a MeSH-derived term directly or find morphologically related words in the title and abstract. Multi-word terms require their meaningful word components to be present, but the components do not have to form an exact phrase.
This is intentionally different from the ontology-based semantic similarity used by the dedicated Phenotype Matching service. A missing literature phenotype match does not mean the publication is clinically unrelated.
Variant and Functional Signals
Exact variant indicator
Requires the same gene and compatible supplied HGVS protein or cDNA notation in the extracted publication record. It does not confirm transcript equivalence, allelic phase, zygosity, or patient-level applicability.
Functional-data indicator
Signals that the metadata or abstract appears to describe experimental work. It does not assess assay calibration, biological relevance, reproducibility, or whether PS3/BS3 can be applied.
How to Read the Number
Use a higher relevance score as a prompt to review a paper earlier. Compare the component breakdown and matched entities before deciding why the score is high. Do not compare the number with an ACMG probability, a confidence interval, or a validated evidence-strength calibration.
Result Boundaries
Very low-relevance candidates are omitted from the returned review set.
Only a bounded number of candidate publications are scored and returned for a case.
When no patient HPO terms are present, the phenotype component contributes no signal and scores are not directly comparable with phenotype-enriched searches.
The user interface can combine literature relevance with Phenotype Matching clinical priority when ordering gene groups; publication cards continue to show the literature score itself.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Literature Evidence](https://folklore.helena.bio/docs/literature)

---

## Page: HPO Phenotype Matching for Gene and Variant Prioritization | Helena Bioinformatics

Source: https://folklore.helena.bio/docs/phenotype-matching
Canonical: https://folklore.helena.bio/docs/phenotype-matching
Description: How Folklore uses Human Phenotype Ontology (HPO) semantic matching to rank phenotype-relevant genes and variants without changing their ACMG classification.

## HPO Phenotype Matching for Gene and Variant Prioritization
Folklore compares the patient's selected Human Phenotype Ontology (HPO) terms with the phenotype profile associated with each gene that carries a candidate variant. Ontology relationships allow related concepts to match even when their labels are not identical.
The result is a phenotype match score from 0 to 100 and a phenotype-aware clinical tier. Variant Analysis owns the ACMG classification; Phenotype Matching uses that existing class during prioritization but does not rewrite it. The PP3/BP4 computational evidence path is documented separately under BayesDel_noAF score thresholds . Neither score establishes a diagnosis.
Analysis Flow
1
Prepare the phenotype profile
Select HPO terms directly or review terms proposed from clinical free text. Only confirmed, present findings should be submitted for matching.
2
Identify eligible variants
Folklore evaluates non-reference variants that have a gene-associated HPO profile from Variant Analysis.
3
Compare HPO profiles
Each valid patient term is compared with the HPO terms associated with the variant gene, including related terms in the ontology.
4
Prioritize variants
Phenotype similarity is considered together with ACMG class, functional impact, population frequency, ClinVar evidence, and inheritance context.
5
Review grouped results
Genes are ranked by their strongest result. Expand a gene to inspect its variants and the individual HPO term matches.
HPO Terms Are Required
If no patient HPO terms are provided, Folklore cannot run phenotype matching and returns no phenotype-ranked results. Add or confirm the clinical findings before starting this stage.
What the Results Show
Gene rank
Genes ordered by the best clinical priority result among their variants.
Clinical tier
Tier 1, Tier 2, IF, Tier 3, or Tier 4 for phenotype-aware triage.
Clinical priority score
A tier-coded score used to order results within the same clinical workflow.
HPO term matches
Patient terms and their strongest semantic match in the gene profile.
In This Section
HPO Overview
How standardized phenotype terms enter the analysis.
Semantic Similarity
How patient findings are compared with gene-associated phenotype profiles.
Clinical Tiers
The five phenotype-aware prioritization groups and their safeguards.
Interpreting Scores
How to read phenotype and clinical priority scores.
HPO Term Selection Guide
Practical guidance for preparing a useful phenotype profile.

### Links and cited sources

- [Human Phenotype Ontology (HPO)](https://hpo.jax.org/)
- [BayesDel_noAF score thresholds](https://folklore.helena.bio/docs/predictors/bayesdel)
- [HPO Overview How standardized phenotype terms enter the analysis.](https://folklore.helena.bio/docs/phenotype-matching/hpo-overview)
- [Semantic Similarity How patient findings are compared with gene-associated phenotype profiles.](https://folklore.helena.bio/docs/phenotype-matching/semantic-similarity)
- [Clinical Tiers The five phenotype-aware prioritization groups and their safeguards.](https://folklore.helena.bio/docs/phenotype-matching/clinical-tiers)
- [Interpreting Scores How to read phenotype and clinical priority scores.](https://folklore.helena.bio/docs/phenotype-matching/interpreting-scores)
- [HPO Term Selection Guide Practical guidance for preparing a useful phenotype profile.](https://folklore.helena.bio/docs/phenotype-matching/hpo-term-selection-guide)

---

## Page: Phenotype-Aware Clinical Tiers | Folklore Documentation

Source: https://folklore.helena.bio/docs/phenotype-matching/clinical-tiers
Canonical: https://folklore.helena.bio/docs/phenotype-matching/clinical-tiers
Description: Current Folklore rules for phenotype-aware Tier 1, Tier 2, IF, Tier 3, and Tier 4 prioritization.

Documentation / Phenotype Matching / Clinical Tiers
## Phenotype-Aware Clinical Tiers
The tier answers how strongly a variant should be prioritized for this patient's selected phenotype. It is separate from the variant's ACMG classification and combines phenotype relevance with variant evidence and inheritance context.
Tier Definitions
Tier 1 Actionable Clinical priority: 80.00-99.99
A rare Pathogenic or Likely Pathogenic variant with a phenotype match of at least 50, subject to inheritance safeguards.
Tier 2 Potentially Actionable Clinical priority: 60.00-79.99
A rare HIGH or MODERATE impact VUS with a phenotype match of at least 70 and no strong benign evidence, or a Tier 1 candidate demoted because the observed inheritance is insufficient.
IF Incidental Finding Clinical priority: 40.00-59.99
A rare Pathogenic or Likely Pathogenic variant whose phenotype match is below the Tier 1 threshold. Review separately from referral-phenotype candidates.
Tier 3 Uncertain Clinical priority: 20.00-39.99
A VUS with moderate phenotype and impact support, or a result retained conservatively by a service safeguard.
Tier 4 Unlikely Clinical priority: 0.00-19.99
Benign, Likely Benign, common, or otherwise weakly supported results for the current phenotype.
IF Is Not an ACMG Secondary Findings Determination
The IF group means that a P/LP result does not sufficiently match the referral phenotype. It does not by itself determine reportability, consent scope, or membership in an ACMG Secondary Findings gene list. Apply the appropriate laboratory policy and current professional guidance separately.
Current Assignment Rules
Tier 1 pathway
P/LP, population frequency below 1%, and phenotype match at or above 50. A heterozygous result in an autosomal recessive-only gene is moved to Tier 2 when a second candidate allele is not identified.
Tier 2 pathway
A rare HIGH or MODERATE impact VUS requires phenotype match at or above 70 and must not carry strong benign evidence. Tier 2 is limited to the 15 highest-priority results; overflow is retained in Tier 3 with a recorded reason.
IF pathway
A rare P/LP result below the phenotype threshold is separated from phenotype-relevant Tier 1 results.
Tier 3 pathway
A HIGH or MODERATE impact VUS with phenotype match at or above 30, or another rare VUS with phenotype match at or above 50.
Tier 4 pathway
Benign or Likely Benign classifications, population frequency at or above 1%, reviewed benign ClinVar evidence, and results that do not meet a higher-tier rule.
Additional Safeguards
Inheritance-aware review
Autosomal recessive carrier status can prevent a P/LP result from appearing as immediately actionable. Genes with both dominant and recessive inheritance retain an explanatory note.
Benign evidence
Confirmed benign ClinVar evidence and strong benign ACMG evidence prevent unsupported promotion.
Highly polymorphic regions
HLA-gene results are handled conservatively and do not enter Tier 1 or Tier 2 through this workflow.
Within-tier ordering
Phenotype match, ACMG class, and population rarity order results inside a tier without overriding the tier&apos;s clinical safeguards.
Learn how the two 0-100 values differ in Interpreting Scores .

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)
- [Interpreting Scores](https://folklore.helena.bio/docs/phenotype-matching/interpreting-scores)

---

## Page: HPO Overview | Folklore Documentation

Source: https://folklore.helena.bio/docs/phenotype-matching/hpo-overview
Canonical: https://folklore.helena.bio/docs/phenotype-matching/hpo-overview
Description: How Folklore searches, extracts, and uses Human Phenotype Ontology terms for phenotype matching.

Documentation / Phenotype Matching / HPO Overview
## Human Phenotype Ontology Overview
The Human Phenotype Ontology (HPO) provides standardized identifiers for clinical abnormalities and connects them in a structured hierarchy. Folklore uses these identifiers to compare a patient's observed findings with gene-associated phenotype profiles.
Terms and Relationships
Each HPO concept has an identifier, a preferred name, synonyms, and relationships to broader or more specific concepts. For example, a specific seizure type sits below the broader term "Seizure". These relationships allow clinically related findings to contribute to a match without requiring identical wording.
A specific, observed finding is usually more informative than a broad ancestor term. HPO is a vocabulary for phenotype description, not a diagnosis code and not a substitute for clinical interpretation.
How Terms Enter Folklore
Search by name, synonym, or identifier
The clinical profile search resolves current HPO terms and excludes obsolete entries from suggestions.
Extract findings from free text
Folklore can propose clinical findings from pasted text and map them to HPO. If that path is unavailable, a local ontology-based fallback is used.
Review before analysis
Automatic extraction is an aid, not a final clinical profile. Confirm the concept, specificity, and whether the finding is actually present before accepting it.
Reuse gene associations
Variant Analysis attaches gene-associated HPO profiles to eligible variants. Phenotype Matching compares the patient profile against those annotations.
Negated Findings Need Manual Review
Do not add an absent or explicitly negated feature as a present HPO term. Automatic extraction can miss contextual negation, especially in complex clinical text. Folklore's phenotype score currently evaluates selected positive findings; it does not use absent findings as exclusion evidence.
Data Boundaries
Ontology coverage
Matching is limited to HPO concepts and gene-phenotype associations available to the analysis environment.
Annotation maturity
Recently described genes and phenotypes may have incomplete profiles and can be under-ranked.
Unresolved identifiers
HPO identifiers that cannot be resolved are omitted from the similarity calculation.
Clinical context
Age, onset, severity, and inheritance remain important even when they are not fully represented by the selected terms.
Reference
Kohler S, et al. "The Human Phenotype Ontology in 2024: phenotypes around the world." Nucleic Acids Research . 2024;52(D1):D1333-D1346. PMID: 37953324.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)

---

## Page: HPO Term Selection Guide | Folklore Documentation

Source: https://folklore.helena.bio/docs/phenotype-matching/hpo-term-selection-guide
Canonical: https://folklore.helena.bio/docs/phenotype-matching/hpo-term-selection-guide
Description: Practical guidance for selecting, reviewing, and refining HPO terms before phenotype matching.

Documentation / Phenotype Matching / HPO Term Selection Guide
## HPO Term Selection Guide
A focused, accurate patient phenotype is more useful than a long list of uncertain or repetitive terms. Select findings that are present, clinically meaningful, and as specific as the available evidence supports.
Selection Principles
Record observed findings
Use findings documented in the patient rather than suspected diagnoses or features copied from a candidate syndrome.
Choose the most specific supported term
Prefer a specific seizure type, structural finding, or biochemical abnormality when the record supports it. Do not infer detail that was not observed.
Cover the whole presentation
Include relevant findings across organ systems, not only the referral complaint. The combination is often more discriminating than any single feature.
Avoid parent-child duplication
If a specific child term fully represents a finding, usually omit its broad ancestor so one feature is not counted twice.
Keep absent findings separate
Do not select a negated feature as present. Negative phenotypes are not currently used as exclusion evidence in the similarity score.
Useful Clinical Detail
Onset and course
Use onset- or progression-specific HPO concepts when those concepts exist and are documented.
Morphology and anatomy
Prefer the precise structural or imaging finding over a broad organ-system abnormality.
Laboratory phenotype
A specific metabolite, enzyme, hematologic, or biochemical abnormality may be more informative than a general symptom.
Neurodevelopment
Capture the affected domain, seizure type, tone, behavior, imaging, and developmental trajectory where observed.
Multisystem findings
Include secondary organ-system features that help distinguish syndromic from isolated disease.
Age-appropriate absence
Do not add a feature that cannot yet be assessed because of the patient&apos;s age as though it were present or absent.
Using Folklore's Input Tools
HPO Search
Search by clinical wording, synonym, or HPO identifier. Review the preferred name and choose the most specific current concept supported by the record.
Free-Text Extraction
Paste relevant clinical text to obtain proposed terms, then review every proposal. Confirm the mapping, remove duplicates and generic findings, and check negation manually before analysis.
Pre-Run Checklist
1
Every selected term describes a finding that is present in this patient.
2
The term is as specific as the clinical evidence allows.
3
Broad ancestors and duplicate synonyms have been removed.
4
Important findings outside the primary organ system are included.
5
Automatically extracted terms and possible negation have been reviewed.
Refine and Re-Run
Phenotype matching can be repeated after the clinical profile changes. Re-run when new findings are confirmed or when a corrected term set materially changes the patient phenotype used for prioritization.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)

---

## Page: Interpreting Phenotype Matching Scores | Folklore Documentation

Source: https://folklore.helena.bio/docs/phenotype-matching/interpreting-scores
Canonical: https://folklore.helena.bio/docs/phenotype-matching/interpreting-scores
Description: How to distinguish Folklore phenotype match scores, individual HPO similarities, and tier-coded clinical priority scores.

Documentation / Phenotype Matching / Interpreting Scores
## Interpreting Phenotype Matching Scores
Phenotype Matching produces related but different measurements. Read the score label and clinical tier together; a value without its context can be misleading.
Three Result Levels
Individual HPO similarity (0-1)
For each patient term, the closest term in the gene profile and its Lin similarity. Folklore counts an individual match as significant only when it is greater than 0.5.
Phenotype match score (0-100)
The average best match across valid patient HPO terms, scaled to 100. It summarizes semantic overlap only.
Clinical priority score (0-99.99)
A tier-coded ordering value. Its band identifies Tier 1, Tier 2, IF, Tier 3, or Tier 4; evidence then orders results within that band.
A Match Score Is Not a Probability
A phenotype score of 80 does not mean an 80% probability that the variant causes disease. It reflects ontology-based similarity between the selected patient findings and the available gene annotations.
Recommended Review Order
1. Confirm the tier
Start with the rule-based clinical group and read any demotion or inheritance note.
2. Inspect the variant evidence
Review ACMG classification, criteria, population frequency, ClinVar evidence, impact, genotype, and inheritance.
3. Inspect individual HPO matches
Identify which findings drive the score and whether they are specific, independent, and clinically accurate.
4. Check annotation coverage
Consider whether an emerging gene, atypical presentation, or incomplete phenotype profile could distort the ranking.
Factors That Change the Score
Specificity
Specific terms usually discriminate better than broad system-level findings.
Additional patient terms
Every valid selected term contributes to the directional average; an unmatched term can lower it.
Redundant terms
Selecting both a broad parent and a specific child can over-weight one feature.
Gene annotation coverage
Sparse or outdated gene-phenotype profiles can understate a real clinical relationship.
Atypical presentation
A patient outside the established disease spectrum may receive a lower score.
Invalid identifiers
Terms that cannot be resolved in the loaded ontology are ignored.
Clinical Boundary
Use phenotype matching to focus review, not to exclude a gene or make a diagnosis by itself. Final interpretation remains with a qualified professional and should incorporate the full patient, family, laboratory, and literature context.
Improve the input profile with the HPO Term Selection Guide .

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)
- [HPO Term Selection Guide](https://folklore.helena.bio/docs/phenotype-matching/hpo-term-selection-guide)

---

## Page: HPO Semantic Similarity | Folklore Documentation

Source: https://folklore.helena.bio/docs/phenotype-matching/semantic-similarity
Canonical: https://folklore.helena.bio/docs/phenotype-matching/semantic-similarity
Description: How Folklore compares patient HPO terms with gene-associated phenotype profiles and produces a directional 0-100 match score.

Documentation / Phenotype Matching / Semantic Similarity
## HPO Semantic Similarity
Folklore compares HPO concepts through their position in the ontology rather than by label alone. It uses the established Lin similarity measure with information content derived from OMIM disease annotations.
From Individual Terms to One Score
1
Resolve the selected patient HPO identifiers and the HPO profile associated with the variant gene.
2
For each valid patient term, compare it with every valid term in the gene profile.
3
Keep the strongest semantic match for that patient term.
4
Average those patient-term matches and convert the result to a 0-100 phenotype match score.
Why Specific Terms Matter
Information content reflects how informative a term is in disease annotations. A broad finding shared by many disorders contributes less discrimination than a specific finding associated with a smaller set of disorders. Exact term matches receive the highest individual similarity, while closely related terms can still score strongly through a shared informative ancestor.
The Comparison Is Directional
The final score asks how well the gene profile explains each selected patient term. Extra phenotypes in the gene profile are not penalized. By contrast, an additional patient term with no close match can lower the average. This is why confirmed, relevant term selection matters.
How to Read the Output
Phenotype match score
A 0-100 semantic similarity summary for the patient-to-gene comparison.
Individual matches
The strongest gene-profile match for every valid patient HPO term.
Matched-term count
The number of individual term similarities above the service threshold.
Unresolved terms
Identifiers that are not present in the loaded ontology do not contribute to the average.
Limitations
Absent or negated findings are not used as negative evidence in the similarity score.
Redundant parent and child terms can over-represent one clinical feature because each selected patient term contributes to the average.
A low score can reflect incomplete gene annotation or an atypical presentation, not only biological irrelevance.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)

---

## Page: Computational Predictors | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors
Canonical: https://folklore.helena.bio/docs/predictors
Description: How Folklore uses computational predictors for variant classification -- BayesDel for ACMG PP3/BP4, SpliceAI for splice impact, and displayed reference predictors.

Documentation / Computational Predictors
## Computational Predictors
Role in Classification
In ACMG variant classification, two criteria depend on computational predictions: PP3 (computational evidence supports a deleterious effect) and BP4 (computational evidence suggests no impact). These are evidence criteria -- they contribute to the overall classification but do not determine it on their own.
Folklore uses two computational tools that directly influence ACMG classification: BayesDel_noAF for missense variant pathogenicity assessment, and SpliceAI for splice impact prediction. Both were selected based on ClinGen Sequence Variant Interpretation (SVI) Working Group recommendations and are calibrated against clinical truth sets.
Classification Tools vs. Displayed Predictors
Not all computational predictions shown in the results interface contribute to ACMG classification. The platform distinguishes between tools that drive classification decisions and predictors displayed for clinical reference.
Tool Role ACMG Criteria What It Measures
BayesDel_noAF Classification PP3, BP4 Missense pathogenicity (calibrated meta-predictor)
SpliceAI Classification PP3_splice, BP4/BP7 guard Splice site creation or disruption
SIFT Displayed -- Amino acid substitution tolerance
AlphaMissense Displayed -- Protein structure-based pathogenicity
MetaSVM Displayed -- Ensemble of multiple prediction methods
DANN Displayed -- Deep neural network pathogenicity
PhyloP Displayed -- Evolutionary conservation (100 vertebrates)
GERP Displayed -- Evolutionary constraint at genomic position
Why BayesDel_noAF
The ClinGen SVI Working Group evaluated multiple computational tools against clinical truth sets and published calibrated evidence strength thresholds for BayesDel_noAF (Pejaver et al. 2022). This tool was selected for three reasons: it has ClinGen-calibrated thresholds that map directly to ACMG evidence strength levels (Supporting, Moderate, Strong); it explicitly excludes allele frequency from its model, avoiding circular reasoning with the frequency-based criteria PM2, BA1, and BS1; and it is precomputed in the dbNSFP database, requiring no external API calls during processing.
Unlike approaches that rely on a single damaging/benign threshold, BayesDel_noAF with ClinGen SVI calibration provides evidence strength modulation -- the same tool can contribute Supporting, Moderate, or Strong evidence depending on the score magnitude. This reflects the clinical reality that a variant with a very high pathogenicity score provides stronger evidence than one just above the threshold.
Why Displayed Predictors Are Shown
SIFT, AlphaMissense, MetaSVM, DANN, PhyloP, and GERP do not contribute to ACMG criteria, but they are shown in the results interface because experienced geneticists use them as additional clinical context. A geneticist reviewing a VUS may find it informative that AlphaMissense predicts the variant as pathogenic based on protein structure, even though this does not change the formal ACMG classification. These predictions help clinicians form their independent assessment alongside the automated classification.
Important
Computational predictions are supporting evidence. They contribute PP3 or BP4 criteria to the ACMG framework, but they do not determine classification on their own. A variant is never classified as Pathogenic based solely on computational evidence, and a variant is never classified as Benign based solely on the absence of computational predictions.
In This Section
SpliceAI
Deep learning splice impact prediction. Directly contributes PP3_splice to ACMG classification.
SIFT
Amino acid substitution tolerance based on sequence homology. Displayed for reference.
AlphaMissense
DeepMind protein structure-based pathogenicity prediction. Displayed for reference.
MetaSVM
Ensemble meta-predictor combining multiple pathogenicity tools. Displayed for reference.
DANN
Deep neural network pathogenicity score for any single nucleotide variant. Displayed for reference.
Conservation Scores
PhyloP and GERP evolutionary constraint metrics. Displayed for reference.
BayesDel
BayesDel_noAF thresholds for PP3 and BP4, evidence-strength modulation, and double-counting guards.
Reference: Pejaver V, et al. Am J Hum Genet. 2022;109(12):2163-2177. PMID: 36413997

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [SpliceAI Deep learning splice impact prediction. Directly contributes PP3_splice to ACMG classification.](https://folklore.helena.bio/docs/predictors/spliceai)
- [SIFT Amino acid substitution tolerance based on sequence homology. Displayed for reference.](https://folklore.helena.bio/docs/predictors/sift)
- [AlphaMissense DeepMind protein structure-based pathogenicity prediction. Displayed for reference.](https://folklore.helena.bio/docs/predictors/alphamissense)
- [MetaSVM Ensemble meta-predictor combining multiple pathogenicity tools. Displayed for reference.](https://folklore.helena.bio/docs/predictors/metasvm)
- [DANN Deep neural network pathogenicity score for any single nucleotide variant. Displayed for reference.](https://folklore.helena.bio/docs/predictors/dann)
- [Conservation Scores PhyloP and GERP evolutionary constraint metrics. Displayed for reference.](https://folklore.helena.bio/docs/predictors/conservation-scores)
- [BayesDel BayesDel_noAF thresholds for PP3 and BP4, evidence-strength modulation, and double-counting guards.](https://folklore.helena.bio/docs/predictors/bayesdel)
- [PMID: 36413997](https://pubmed.ncbi.nlm.nih.gov/36413997/)

---

## Page: AlphaMissense Score Interpretation for Missense Variants | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors/alphamissense
Canonical: https://folklore.helena.bio/docs/predictors/alphamissense
Description: Interpret AlphaMissense scores and Pathogenic, Ambiguous, and Benign labels. Folklore displays the result as context but does not use it for ACMG PP3 or BP4.

Documentation / Computational Predictors / AlphaMissense
## AlphaMissense Score Interpretation for Missense Variants
AlphaMissense estimates the effect of missense substitutions from protein sequence and structural context. Folklore displays the prediction for review; it does not use AlphaMissense to assign PP3 or BP4.
What the Score Represents
AlphaMissense is a Google DeepMind model for missense substitutions. It combines protein sequence and structural context to estimate whether an amino acid change is likely to impair protein function.
The model does not score splice variants, truncating variants, in-frame indels, or non-coding variants. Folklore uses BayesDel_noAF and SpliceAI for the formal PP3 and BP4 evidence paths.
How to Read the Score
AlphaMissense scores range from 0 to 1, with higher scores indicating a greater likelihood of pathogenicity. Variants are classified into three categories:
Category Label in Results Meaning
Pathogenic P The variant is predicted to be disease-causing based on protein structural analysis
Ambiguous A Insufficient confidence for a clear prediction in either direction
Benign B The variant is predicted to be tolerated by the protein structure
Strengths and Limitations
AlphaMissense's primary strength is that it incorporates protein three-dimensional structural information. While sequence-based tools like SIFT can only assess conservation at a position, AlphaMissense understands whether the amino acid sits in the protein core, on the surface, at an interaction interface, or near a catalytic site. It was trained on human population data and primate conservation data, giving it a human-specific perspective on pathogenicity.
The main limitation is scope: AlphaMissense only predicts impact for missense variants. It does not assess in-frame indels, splice variants, nonsense variants, or any non-coding variants. Its "Ambiguous" category covers a meaningful fraction of all possible missense variants where the model lacks confidence.
Role in Folklore
AlphaMissense predictions are displayed in the variant detail view as additional clinical context. They do not contribute to PP3 or BP4 ACMG criteria. The formal classification uses BayesDel_noAF with ClinGen SVI calibrated thresholds. See BayesDel thresholds for details.
Reference: Cheng J, et al. Science. 2023;381(6664):eadg7492. PMID: 37733863

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [BayesDel_noAF and SpliceAI](https://folklore.helena.bio/docs/predictors/bayesdel)
- [PMID: 37733863](https://pubmed.ncbi.nlm.nih.gov/37733863/)

---

## Page: BayesDel_noAF Score Thresholds for ACMG PP3 and BP4 | Helena Bioinformatics

Source: https://folklore.helena.bio/docs/predictors/bayesdel
Canonical: https://folklore.helena.bio/docs/predictors/bayesdel
Description: ClinGen SVI-calibrated BayesDel_noAF score thresholds for ACMG PP3 and BP4, including Bayesian points, indeterminate scores, double-counting safeguards, and ACMG/AMP version 4 beta status.

Documentation / Computational Predictors / BayesDel
## BayesDel_noAF Score Thresholds for ACMG PP3 and BP4
BayesDel_noAF with ClinGen SVI Calibration
Folklore uses BayesDel_noAF as the single computational tool for determining PP3 (computational evidence supports a deleterious effect) and BP4 (computational evidence suggests no impact). BayesDel is a meta-predictor that integrates deleteriousness scores from multiple tools into a single calibrated score. The "_noAF" variant explicitly excludes allele frequency from its model, which is critical: since the ACMG framework already has frequency-based criteria (PM2, BA1, BS1), using a predictor that includes frequency would double-count the same evidence.
The evidence strength thresholds were calibrated by the ClinGen Sequence Variant Interpretation (SVI) Working Group against clinical truth sets (Pejaver et al. 2022). These thresholds map BayesDel_noAF score ranges directly to ACMG evidence strength levels, enabling evidence strength modulation -- a feature not possible with simpler binary (damaging/benign) approaches.
BayesDel contributes computational evidence inside Variant Analysis. It does not compare the patient's HPO profile or rank genes by phenotype; that separate step is documented under HPO Phenotype Matching .
ACMG/AMP Version 4 Status
ClinGen describes the ACMG/AMP/CAP/ClinGen version 4 guidelines as beta and the associated standards as forthcoming. Its public ACMGv4 data profile is marked Draft, remains under active development, and may change significantly.
Folklore does not claim conformance with a final version 4 specification. The current implementation maps calibrated PP3/BP4 evidence strengths to Bayesian points. The v4 draft also represents evidence with numeric scores; this is a structural correspondence, not evidence that Folklore implements the final standard.
Two Independent Paths
PP3 and BP4 are evaluated through two independent paths that capture different biological mechanisms. These paths do not overlap: a given variant is assessed through the missense path, the splice path, or both, depending on the available data.
Path Tool Applies To Evidence
Missense BayesDel_noAF Variants with BayesDel score available PP3 (Strong / Moderate / Supporting) or BP4 (Moderate / Supporting)
Splice SpliceAI Variants with SpliceAI score available PP3_splice (Supporting)
PP3 Pathogenic Evidence (BayesDel)
The ClinGen SVI calibration provides three evidence strength levels for pathogenic computational evidence. Higher BayesDel scores produce stronger evidence. This graduated approach reflects the clinical reality that a variant with a very high pathogenicity score provides more compelling evidence than one near the minimum threshold.
BayesDel_noAF Score Evidence Strength Criteria Label Bayesian Points
>= 0.518 Strong PP3_Strong +4
0.290 -- 0.517 Moderate PP3_Moderate +2
0.130 -- 0.289 Supporting PP3 +1
PP3 Splice Evidence (SpliceAI)
Independently of the BayesDel missense assessment, SpliceAI provides splice-specific evidence. When the maximum SpliceAI delta score is 0.2 or above, PP3_splice is triggered as Supporting pathogenic evidence (+1 Bayesian point). This threshold follows ClinGen SVI 2023 recommendations (Walker et al., 2023).
PP3_splice is excluded when PVS1 (loss-of-function) applies to the same variant. This prevents double-counting splice disruption that already contributed to the Very Strong PVS1 criterion. See SpliceAI for full details.
BP4 Benign Evidence (BayesDel)
Low BayesDel scores provide evidence that a variant is computationally predicted to be benign. Two evidence strength levels are calibrated:
BayesDel_noAF Score Evidence Strength Criteria Label Bayesian Points
<= -0.361 Moderate BP4_Moderate -2
-0.360 -- -0.181 Supporting BP4 -1
BP4 at any level requires that SpliceAI max score is below 0.1 or absent. A variant cannot receive computational benign evidence if there is any predicted splice impact.
PM1 + PP3 Point-Sum Cap
The ClinGen SVI Working Group recommends that the combined evidence from PM1 (variant in a functional domain) and PP3 (computational prediction) should not exceed Strong equivalent (4 Bayesian points). This prevents over-counting evidence when a variant is both in a known functional domain and computationally predicted to be damaging -- since these two observations are not fully independent.
In practice, PM1 contributes Moderate evidence (2 points). When PM1 is triggered and BayesDel reaches the Strong threshold (>= 0.518), PP3 is automatically downgraded from Strong (4 points) to Moderate (2 points), keeping the combined total at 4 points. When PM1 is not triggered, PP3_Strong is applied at full strength.
Scenario PM1 PP3 Combined Points
No functional domain, BayesDel >= 0.518 -- (0 pts) PP3_Strong (4 pts) 4
Pfam domain, BayesDel >= 0.518 PM1 (2 pts) PP3_Moderate (2 pts, capped) 4
Pfam domain, BayesDel 0.290-0.517 PM1 (2 pts) PP3_Moderate (2 pts) 4
Pfam domain, BayesDel 0.130-0.289 PM1 (2 pts) PP3 (1 pt) 3
Indeterminate Zone
BayesDel_noAF scores between -0.180 and 0.129 fall in the indeterminate zone -- neither PP3 nor BP4 is applied. This is intentional: variants in this range do not have sufficient computational signal to contribute evidence in either direction. Approximately 20-30% of rare missense variants fall in this zone (Stenton et al. 2024), which prevents computational evidence from being over-applied.
Expected Evidence Yield
For a typical case with approximately 75 rare missense variants in disease genes, the expected distribution is approximately: 1 variant receiving PP3_Strong, 3-5 receiving PP3_Moderate, 3-5 receiving PP3_Supporting, 41-49 receiving BP4, and 17-19 in the indeterminate zone (Stenton et al. 2024). PP3_Strong is rare enough to avoid excessive reclassification of VUS variants.
Why Not a Multi-Predictor Consensus
Some variant classification systems use a weighted consensus across multiple individual predictors (SIFT, PolyPhen, CADD, etc.) to determine PP3/BP4. Folklore uses BayesDel_noAF as a single calibrated tool instead, for three reasons: it has ClinGen SVI-calibrated thresholds directly mapping to ACMG evidence strength levels; it excludes allele frequency, avoiding circular reasoning with PM2/BA1/BS1; and it provides evidence strength modulation (Supporting, Moderate, Strong) which a binary consensus approach cannot. The individual predictors (SIFT, AlphaMissense, MetaSVM, DANN, PhyloP, GERP) remain displayed in the results for additional clinical context.
ClinGen SVI calibration: Pejaver V, et al. Am J Hum Genet. 2022;109(12):2163-2177. PMID: 36413997
Evidence yield: Stenton SL, et al. Genet Med. 2024;26(11):101213. PMID: 39030733
ClinGen SVI splice: Walker LC, et al. Am J Hum Genet. 2023;110(7):1046-1067. PMID: 37352859
BayesDel original: Feng BJ. Hum Mutat. 2017;38(3):243-251.
ACMG/AMP version 4 status: ClinGen at ACMG 2026 .
ACMGv4 machine-readable profile: ClinGen draft profile .

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [HPO Phenotype Matching](https://folklore.helena.bio/docs/phenotype-matching)
- [ACMG/AMP/CAP/ClinGen version 4 guidelines as beta](https://www.clinicalgenome.org/about/events/clingen-at-acmg-2026/)
- [public ACMGv4 data profile](https://dataexchange.clinicalgenome.org/clinvar-gks/drafts/acmgv4-variant-pathogenicity/)
- [SpliceAI](https://folklore.helena.bio/docs/predictors/spliceai)
- [PMID: 36413997](https://pubmed.ncbi.nlm.nih.gov/36413997/)
- [PMID: 39030733](https://pubmed.ncbi.nlm.nih.gov/39030733/)
- [PMID: 37352859](https://pubmed.ncbi.nlm.nih.gov/37352859/)

---

## Page: PhyloP and GERP Conservation Score Interpretation | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors/conservation-scores
Canonical: https://folklore.helena.bio/docs/predictors/conservation-scores
Description: Interpret PhyloP and GERP scores, including PhyloP above 2 and GERP above 4. Folklore displays both as context but does not use them directly for PP3 or BP4.

Documentation / Computational Predictors / Conservation Scores
## PhyloP and GERP Conservation Score Interpretation
PhyloP and GERP measure evolutionary constraint at a genomic position. Higher positive scores indicate fewer observed substitutions than expected. Folklore displays both scores for review but does not use either score directly for PP3 or BP4.
What Conservation Can Tell You
Conservation identifies positions that have changed less often than expected across species. It does not determine whether a particular substitution is damaging. A conserved position can tolerate one amino acid change and reject another.
PhyloP measures deviation from a neutral substitution rate. GERP reports rejected substitutions. The two metrics use different calculations and should be read independently.
PhyloP (100-Way Vertebrate)
PhyloP measures evolutionary conservation across 100 vertebrate species using phylogenetic models. It compares the observed substitution rate at each position against the neutral expectation derived from the phylogenetic tree.
Score Range Interpretation
> 2.0 Conserved. Fewer substitutions than expected -- position is under purifying selection.
0.0 -- 2.0 Weakly conserved or neutral. Limited evolutionary signal.
<= 0.0 Fast-evolving. More substitutions than expected -- position may be under positive or relaxed selection.
GERP++ (Genomic Evolutionary Rate Profiling)
GERP++ measures evolutionary constraint by comparing observed substitutions against the expected number at each position across a multi-species alignment. The difference between expected and observed substitutions is the "rejected substitution" score -- higher values indicate that more mutations have been eliminated by natural selection.
Score Range Interpretation
> 4.0 Strongly constrained. The position has rejected a large number of substitutions across evolution.
2.0 -- 4.0 Moderately constrained. Some evidence of purifying selection.
<= 0.0 Unconstrained. No evidence of evolutionary pressure to maintain this position.
Why Conservation Scores Are Context, Not Evidence
Conservation tells you that a position is important, but it does not tell you whether a specific change at that position is damaging. A highly conserved position could tolerate certain amino acid substitutions (for example, leucine to isoleucine at a hydrophobic core position) while being completely intolerant of others (for example, leucine to proline, which would break the helix). This is why conservation scores are displayed as additional context rather than used as direct ACMG evidence.
Role in Folklore
PhyloP and GERP scores are displayed in the variant detail view as additional clinical context. They do not contribute to PP3 or BP4 ACMG criteria. The formal classification uses BayesDel_noAF (which itself incorporates conservation signals internally) with ClinGen SVI calibrated thresholds. See BayesDel thresholds for details.
PhyloP reference: Pollard KS, et al. Genome Res. 2010;20(1):110-121. PMID: 19858363
GERP reference: Davydov EV, et al. PLoS Comput Biol. 2010;6(12):e1001025. PMID: 21152010

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [BayesDel thresholds](https://folklore.helena.bio/docs/predictors/bayesdel)
- [PMID: 19858363](https://pubmed.ncbi.nlm.nih.gov/19858363/)
- [PMID: 21152010](https://pubmed.ncbi.nlm.nih.gov/21152010/)

---

## Page: DANN Score Interpretation for Genetic Variants | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors/dann
Canonical: https://folklore.helena.bio/docs/predictors/dann
Description: Interpret DANN scores from 0 to 1, including the 0.95 damaging threshold and the broad ambiguous range. Folklore displays DANN without using it for ACMG PP3 or BP4.

Documentation / Computational Predictors / DANN
## DANN Score Interpretation for Genetic Variants
DANN scores range from 0 to 1. Scores at or above 0.95 are labelled damaging, while 0.5 to 0.95 remains a broad ambiguous interval. Folklore displays DANN for review but does not use it to assign PP3 or BP4.
What the Score Represents
DANN is a neural-network score derived from genomic annotations. It can score single nucleotide variants in coding and non-coding regions, including positions where missense-specific predictors do not apply.
The 0.95 Threshold
DANN scores range from 0 to 1, with higher scores indicating a greater likelihood of pathogenicity.
Score Range Interpretation
>= 0.95 Predicted damaging with high confidence
0.5 -- 0.95 Ambiguous range -- insufficient confidence for a clear prediction
< 0.5 Predicted benign
Strengths and Limitations
DANN's primary strength is breadth: it can score any single nucleotide variant in the genome, not just missense variants in coding regions. This makes it useful as a reference for intronic, synonymous, and UTR variants where protein-specific tools like SIFT or AlphaMissense are not applicable.
The main limitation is its wide ambiguous range (0.5 to 0.95), which means many variants receive scores that are neither clearly damaging nor clearly benign. The binary threshold approach may also miss nuanced pathogenicity signals.
Role in Folklore
DANN scores are displayed in the variant detail view as additional clinical context. They do not contribute to PP3 or BP4 ACMG criteria. The formal classification uses BayesDel_noAF with ClinGen SVI calibrated thresholds. See BayesDel thresholds for details.
Reference: Quang D, Chen Y, Xie X. Bioinformatics. 2015;31(5):761-763. PMID: 25338716

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [BayesDel thresholds](https://folklore.helena.bio/docs/predictors/bayesdel)
- [PMID: 25338716](https://pubmed.ncbi.nlm.nih.gov/25338716/)

---

## Page: MetaSVM | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors/metasvm
Canonical: https://folklore.helena.bio/docs/predictors/metasvm
Description: MetaSVM predictor in Folklore -- ensemble meta-predictor combining multiple pathogenicity tools. Displayed for clinical reference.

Documentation / Computational Predictors / MetaSVM
## MetaSVM
Displayed for clinical reference. Does not contribute to ACMG classification.
What MetaSVM Is
MetaSVM is a meta-predictor that combines scores from 10 individual pathogenicity prediction tools using a Support Vector Machine (SVM) classifier. Rather than relying on any single tool's assessment, it aggregates signals from multiple methods -- each capturing different aspects of variant impact -- into a single consensus score. This ensemble approach generally produces more reliable predictions than any individual component tool.
Score Interpretation
MetaSVM produces a continuous score. Positive values indicate a predicted damaging effect, negative values indicate a predicted tolerated effect. The further from zero, the higher the confidence.
Score Prediction Label in Results
Positive The variant is predicted to be damaging D (Deleterious)
Negative The variant is predicted to be tolerated T (Tolerated)
Strengths and Limitations
As an ensemble method, MetaSVM reduces the bias of individual predictors and generally achieves higher accuracy than any single component tool. It captures multiple dimensions of variant impact simultaneously -- conservation, protein structure, physicochemical properties, and more.
The main limitation is that because MetaSVM combines the same underlying tools used by other predictors, its errors are correlated with them. It is also limited to missense variants and cannot assess splice, nonsense, or non-coding variants.
Role in Folklore
MetaSVM predictions are displayed in the variant detail view as additional clinical context. They do not contribute to PP3 or BP4 ACMG criteria. The formal classification uses BayesDel_noAF with ClinGen SVI calibrated thresholds. See BayesDel thresholds for details.
Reference: Dong C, et al. Hum Mol Genet. 2015;24(8):2125-2137. PMID: 25552646

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [BayesDel thresholds](https://folklore.helena.bio/docs/predictors/bayesdel)
- [PMID: 25552646](https://pubmed.ncbi.nlm.nih.gov/25552646/)

---

## Page: SIFT Score Interpretation: Deleterious vs Tolerated | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors/sift
Canonical: https://folklore.helena.bio/docs/predictors/sift
Description: Interpret SIFT scores below and above 0.05, understand the sequence-conservation model, and see why Folklore displays SIFT without using it for ACMG PP3 or BP4.

Documentation / Computational Predictors / SIFT
## SIFT Score Interpretation: Deleterious vs Tolerated
A SIFT score below 0.05 is Deleterious; a score of 0.05 or above is Tolerated. Folklore displays the score for review but does not use SIFT to assign PP3 or BP4.
What the Score Represents
SIFT estimates whether a missense substitution is tolerated at its protein position. The model compares homologous sequences and treats substitutions at conserved positions as more likely to impair function.
The 0.05 Threshold
SIFT scores range from 0 to 1. Unlike most pathogenicity predictors, lower scores indicate greater predicted damage.
Score Range Prediction Label in Results
< 0.05 The amino acid change is predicted to affect protein function D (Deleterious)
>= 0.05 The amino acid change is predicted to be tolerated T (Tolerated)
Strengths and Limitations
SIFT is one of the oldest and most widely cited computational predictors in genetics, with a straightforward biological rationale: positions conserved across evolution are likely functionally important. It is applicable to any missense variant in any protein with sufficient homologous sequences.
However, SIFT is based on sequence conservation alone and does not consider protein three-dimensional structure, post-translational modifications, or protein-protein interactions. It may miss gain-of-function variants (where the new amino acid has an active harmful effect rather than a loss of the original function) and is less effective for positions with low sequence conservation across species.
Role in Folklore
SIFT predictions are displayed in the variant detail view as additional clinical context. They do not contribute to PP3 or BP4 ACMG criteria. The formal classification uses BayesDel_noAF with ClinGen SVI calibrated thresholds. See BayesDel thresholds for details.
Reference: Ng PC, Henikoff S. Nucleic Acids Res. 2003;31(13):3812-3814. PMID: 12824425

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [BayesDel thresholds](https://folklore.helena.bio/docs/predictors/bayesdel)
- [PMID: 12824425](https://pubmed.ncbi.nlm.nih.gov/12824425/)

---

## Page: SpliceAI Score Interpretation and ACMG Thresholds | Folklore Documentation

Source: https://folklore.helena.bio/docs/predictors/spliceai
Canonical: https://folklore.helena.bio/docs/predictors/spliceai
Description: Interpret SpliceAI delta scores and the 0.1, 0.2, 0.5, and 0.8 thresholds used for PP3_splice, BP4, BP7, and PVS1 double-counting control in Folklore.

Documentation / Computational Predictors / SpliceAI
## SpliceAI Score Interpretation and ACMG Thresholds
SpliceAI reports donor and acceptor gain or loss. Folklore reads the maximum delta score through explicit thresholds: values below 0.1 support benign guards, values from 0.2 enter the PP3_splice path, and PVS1 blocks duplicate computational evidence for the same loss-of-function mechanism.
What the Four Scores Represent
SpliceAI is a deep learning model developed by Illumina that predicts whether a genetic variant will disrupt normal mRNA splicing. Splicing is the process by which the cell removes non-coding sections (introns) from the pre-mRNA and joins the coding sections (exons) to produce the final messenger RNA. Variants that disrupt this process can lead to abnormal proteins or complete loss of protein production, even if they do not directly change the amino acid sequence.
The model was published in Cell (Jaganathan et al., 2019) and is widely used in clinical genetics laboratories. It is one of two computational tools in Folklore that directly influences ACMG classification.
Four Delta Scores
SpliceAI produces four scores, each representing a different type of splice disruption. Each score ranges from 0 (no impact) to 1 (certain disruption). The maximum of the four scores is used for classification thresholds.
Score Name Meaning
DS_AG Acceptor Gain The variant creates a new splice acceptor site where one did not exist
DS_AL Acceptor Loss The variant destroys an existing splice acceptor site
DS_DG Donor Gain The variant creates a new splice donor site where one did not exist
DS_DL Donor Loss The variant destroys an existing splice donor site
Clinical Score Thresholds
Max Score Range Interpretation
0.0 -- 0.1 No predicted splice impact. The variant is unlikely to affect splicing.
0.1 -- 0.2 Low predicted impact. Some splice effect possible but below the clinical evidence threshold.
0.2 -- 0.5 Moderate predicted impact. Supporting evidence for spliceogenicity per ClinGen SVI 2023.
0.5 -- 0.8 High predicted impact. Strong prediction of splice disruption.
0.8 -- 1.0 Very high predicted impact. Near-certain splice disruption.
How Folklore Uses SpliceAI
SpliceAI scores are used in three distinct ways within the classification engine. Each role has a specific threshold and clinical rationale.
Condition ACMG Effect Rationale
Max score >= 0.2 PP3_splice (Supporting pathogenic) Supporting evidence for splice impact, applied as an independent path separate from the BayesDel missense assessment. This threshold follows ClinGen SVI 2023 recommendations.
Max score < 0.1 Required for BP4 A variant cannot receive computational benign evidence (BP4) if SpliceAI predicts any splice impact. This prevents benign classification when splice disruption is possible.
Max score <= 0.1 Required for BP7 Synonymous variants are only classified as Likely Benign (via BP7) when SpliceAI confirms no splice impact. Synonymous variants near splice junctions can be pathogenic through aberrant splicing.
PVS1 Double-Counting Guard
When PVS1 (loss-of-function) is triggered for a variant, PP3_splice is not applied. This prevents counting the same biological mechanism -- splice disruption leading to loss of function -- as both PVS1 and PP3 evidence. A splice donor variant that causes a frameshift already receives Very Strong evidence through PVS1; adding PP3_splice on top would double-count the same observation. This follows ClinGen SVI 2023 recommendations (Walker et al., 2023).
Score Source
SpliceAI scores are precomputed from Ensembl MANE transcript predictions (Release 113). They are not computed at runtime. This ensures reproducibility -- the same variant always receives the same SpliceAI scores regardless of when the analysis is run.
Limitations
SpliceAI predictions are computational. RNA splicing studies (RT-PCR, minigene assays) remain the gold standard for confirming splice-altering effects. Scores are computed on MANE Select transcripts only, so non-MANE transcript-specific splicing effects may be missed. Deep intronic variants beyond the SpliceAI prediction window (typically +/- 50bp from precomputed scores) may not be captured. The model was trained on known splice sites, so novel splice mechanisms not represented in the training data may not be predicted.
Clinical Relevance
Approximately 10-15% of pathogenic variants cause disease through aberrant splicing. SpliceAI provides a fast, reproducible screen for splice impact that can guide whether RNA studies are warranted. When reviewing variants with high SpliceAI scores, consider recommending RT-PCR or minigene assay confirmation.
Reference: Jaganathan K, et al. Cell. 2019;176(3):535-548. PMID: 30661751
ClinGen SVI: Walker LC, et al. Am J Hum Genet. 2023;110(7):1046-1067. PMID: 37352859

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Computational Predictors](https://folklore.helena.bio/docs/predictors)
- [PMID: 30661751](https://pubmed.ncbi.nlm.nih.gov/30661751/)
- [PMID: 37352859](https://pubmed.ncbi.nlm.nih.gov/37352859/)

---

## Page: References | Folklore Documentation

Source: https://folklore.helena.bio/docs/references
Canonical: https://folklore.helena.bio/docs/references
Description: Scientific publications and standards cited in the Folklore documentation.

Documentation / References
## References
Scientific publications and standards cited throughout the Folklore documentation.
ACMG Classification
Richards S, Aziz N, Bale S, et al. " Standards and guidelines for the interpretation of sequence variants: a joint consensus recommendation of the American College of Medical Genetics and Genomics and the Association for Molecular Pathology. " Genet Med . 2015 ; 17(5):405-424 . PMID: 25741868
Tavtigian SV, Greenblatt MS, Harrison SM, et al. " Modeling the ACMG/AMP variant classification guidelines as a Bayesian classification framework. " Genet Med . 2018 ; 20(9):1054-1060 . PMID: 29300386
Tavtigian SV, Harrison SM, Boucher KM, Biesecker LG. " Fitting a naturally scaled point system to the ACMG/AMP variant classification guidelines. " Hum Mutat . 2020 ; 41(11):1734-1737 . PMID: 32720330
Computational Predictors
Pejaver V, Byrne AB, Feng BJ, et al. " Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3/BP4 criteria. " Am J Hum Genet . 2022 ; 109(12):2163-2177 . PMID: 36413997
Jaganathan K, Kyriazopoulou Panagiotopoulou S, McRae JF, et al. " Predicting splicing from primary sequence with deep learning. " Cell . 2019 ; 176(3):535-548 . PMID: 30661751
Walker LC, Hoya M, Wiggins GAR, et al. " Using the ClinGen/ACMG/AMP framework to assess splicing impact of sequence variants. " Hum Mutat . 2023 ; 44:1-12 . PMID: 36864581
Cheng J, Novati G, Pan J, et al. " Accurate proteome-wide missense variant effect prediction with AlphaMissense. " Science . 2023 ; 381(6664):eadg7492 . PMID: 37733863
Quang D, Chen Y, Xie X. " DANN: a deep learning approach for annotating the pathogenicity of genetic variants. " Bioinformatics . 2015 ; 31(5):761-763 . PMID: 25338716
Population Databases
Chen S, Francioli LC, Goodrich JK, et al. " A genomic mutational constraint map using variation in 76,156 human genomes. " Nature . 2024 ; 625:92-100 . PMID: 38057664
Karczewski KJ, Francioli LC, Tiao G, et al. " The mutational constraint spectrum quantified from variation in 141,456 humans. " Nature . 2020 ; 581:434-443 . PMID: 32461654
Landrum MJ, Lee JM, Benson M, et al. " ClinVar: improving access to variant interpretations and supporting evidence. " Nucleic Acids Res . 2018 ; 46(D1):D1062-D1067 . PMID: 29165669
Annotation Tools
McLaren W, Gil L, Hunt SE, et al. " The Ensembl Variant Effect Predictor. " Genome Biol . 2016 ; 17:122 . PMID: 27268795
Liu X, Li C, Mou C, et al. " dbNSFP v4: a comprehensive database of transcript-specific functional predictions and annotations for human nonsynonymous and splice-site SNVs. " Genome Med . 2020 ; 12:103 . PMID: 33261662
Phenotype Ontology
Kohler S, Gargano M, Matentzoglu N, et al. " The Human Phenotype Ontology in 2024: phenotypes around the world. " Nucleic Acids Res . 2024 ; 52(D1):D1333-D1346 . PMID: 37953324
Lin D. " An information-theoretic definition of similarity. " Proc 15th Int Conf Machine Learning . 1998 ; pp. 296-304 .
Clinical Genetics Resources
Rehm HL, Berg JS, Brooks LD, et al. " ClinGen -- the Clinical Genome Resource. " N Engl J Med . 2015 ; 372:2235-2242 . PMID: 26014595
Stenton SL, Kremer LS, Gusic M, et al. " Systematic application of computational variant interpretation tools for germline variant classification. " Am J Hum Genet . 2024 ; 111:1-15 . PMID: 38552641

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [PMID: 25741868](https://pubmed.ncbi.nlm.nih.gov/25741868/)
- [PMID: 29300386](https://pubmed.ncbi.nlm.nih.gov/29300386/)
- [PMID: 32720330](https://pubmed.ncbi.nlm.nih.gov/32720330/)
- [PMID: 36413997](https://pubmed.ncbi.nlm.nih.gov/36413997/)
- [PMID: 30661751](https://pubmed.ncbi.nlm.nih.gov/30661751/)
- [PMID: 36864581](https://pubmed.ncbi.nlm.nih.gov/36864581/)
- [PMID: 37733863](https://pubmed.ncbi.nlm.nih.gov/37733863/)
- [PMID: 25338716](https://pubmed.ncbi.nlm.nih.gov/25338716/)
- [PMID: 38057664](https://pubmed.ncbi.nlm.nih.gov/38057664/)
- [PMID: 32461654](https://pubmed.ncbi.nlm.nih.gov/32461654/)
- [PMID: 29165669](https://pubmed.ncbi.nlm.nih.gov/29165669/)
- [PMID: 27268795](https://pubmed.ncbi.nlm.nih.gov/27268795/)
- [PMID: 33261662](https://pubmed.ncbi.nlm.nih.gov/33261662/)
- [PMID: 37953324](https://pubmed.ncbi.nlm.nih.gov/37953324/)
- [PMID: 26014595](https://pubmed.ncbi.nlm.nih.gov/26014595/)
- [PMID: 38552641](https://pubmed.ncbi.nlm.nih.gov/38552641/)

---

## Page: Screening | Folklore Documentation

Source: https://folklore.helena.bio/docs/screening
Canonical: https://folklore.helena.bio/docs/screening
Description: How Folklore prioritizes variants after ACMG classification using multi-dimensional scoring, clinical context, and tiered ranking.

Documentation / Screening
## Screening
After Variant Analysis assigns each variant an ACMG category, a clinician still needs a focused review queue. Screening applies a multi-dimensional prioritization layer to Pathogenic, Likely Pathogenic, and optionally VUS findings within the selected gene scope.
Screening orders the available candidates by configured clinical context and retains a score breakdown and justification for review. The number of high-priority results depends on the case, selected mode, available phenotype data, and configured filters.
How Screening Works
1
Resolve the Gene Scope
Selected panels and custom genes define which genes are screened. Without either, Folklore uses its built-in fallback set rather than loading every gene.
2
Calculate Component Scores
Each variant is evaluated across seven dimensions: gene constraint, computational deleteriousness, phenotype relevance, dosage sensitivity, consequence severity, compound heterozygote potential, and age-appropriate gene prioritization.
3
Apply Clinical Boosts
Patient-specific context (ACMG class, phenotype match tier, ethnicity, family history, sex-linked inheritance, consanguinity, pregnancy status) adds additional priority boosts.
4
Assign Tiers
Variants are ranked by their final score and explicit priority rules, then assigned to one of four review tiers. Tier 1 is the highest-priority review queue.
Key Design Principles
The seven component scores, selected clinical boosts, final score, tier, and human-readable justification are retained for review. Some context contributions affect the total without appearing as separate output columns.
Pathogenic and Likely Pathogenic classifications contribute a strong prioritization boost, but final tier assignment remains inheritance-aware. A heterozygous finding in an autosomal-recessive gene can be demoted when no compound-heterozygous partner is supported.
Scoring adapts to clinical context. A neonatal screening case uses different weights than an adult proactive screening case. Diagnostic mode with HPO terms emphasizes phenotype matching.
Component and final scores are bounded to 0.0-1.0. They are prioritization values, not probabilities of pathogenicity or clinical actionability.
Screening runs after classification and records the applied mode, component scores, boosts, and final tier for audit and review.
Screening vs. Classification
Classification assigns the evidence-based ACMG class. Screening is a separate prioritization layer that orders findings for review using the selected mode and patient context. A screening tier does not modify or override the underlying ACMG classification.
In This Section
Scoring Components
Seven dimensions of variant scoring: constraint, deleteriousness, phenotype, dosage, consequence, compound het, and age relevance.
Tier System
Four-tier priority ranking with clinical actionability labels and boost mechanisms.
Screening Modes
Diagnostic, neonatal, pediatric, proactive adult, carrier, and pharmacogenomics modes.
Age-Aware Prioritization
How patient age influences scoring weights and gene relevance from neonatal through elderly.
Gene Panels and Custom Genes
How panels define the genes in scope and contribute curated disease, age, priority, and ClinGen context.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Scoring Components Seven dimensions of variant scoring: constraint, deleteriousness, phenotype, dosage, consequence, compound het, and age relevance.](https://folklore.helena.bio/docs/screening/scoring-components)
- [Tier System Four-tier priority ranking with clinical actionability labels and boost mechanisms.](https://folklore.helena.bio/docs/screening/tier-system)
- [Screening Modes Diagnostic, neonatal, pediatric, proactive adult, carrier, and pharmacogenomics modes.](https://folklore.helena.bio/docs/screening/screening-modes)
- [Age-Aware Prioritization How patient age influences scoring weights and gene relevance from neonatal through elderly.](https://folklore.helena.bio/docs/screening/age-aware-prioritization)
- [Gene Panels and Custom Genes How panels define the genes in scope and contribute curated disease, age, priority, and ClinGen context.](https://folklore.helena.bio/docs/screening/gene-panels)

---

## Page: Age-Aware Screening Prioritization | Folklore Documentation

Source: https://folklore.helena.bio/docs/screening/age-aware-prioritization
Canonical: https://folklore.helena.bio/docs/screening/age-aware-prioritization
Description: How Folklore maps precise patient age to screening context and combines panel age annotations with curated fallback categories.

Documentation / Screening / Age-Aware Prioritization
## Age-Aware Screening Prioritization
Age relevance helps order findings according to when the associated gene-disease context is most useful for review. It does not decide whether a condition is penetrant, reportable, or actionable for an individual patient.
Age Groups
Group Operational Range Screening Context
Neonatal 0-28 days Early-onset and time-sensitive neonatal context.
Infant 29-365 days, or age_years below 1 Accepts neonatal and pediatric panel relevance.
Child 1 to under 12 years Pediatric panel relevance and childhood-onset fallback context.
Adolescent 12 to under 18 years Accepts both pediatric and adult panel relevance.
Adult 18 to under 65 years Adult panel relevance and adult-onset fallback context.
Elderly 65 years and older Adult panel relevance with a distinct elderly component profile.
Day Precision Protects the Neonatal Boundary
When age in days is available, 0-28 days maps to Neonatal and 29-365 days maps to Infant. Year-based age is used afterward. If neither value is usable, the service falls back to Adult.
Panel Context Takes Precedence
A selected panel entry can specify neonatal, pediatric, adult, or all-age relevance. Matching entries receive stronger age context; non-matching panel genes remain in scope with a conservative value rather than being discarded. All-age entries retain a minimum relevance across age groups. Panel priority and available ClinGen gene-disease validity metadata also contribute to this context.
Fallback Categories
When no panel metadata is available, Folklore uses curated early-onset, childhood, metabolic, adult cancer, cardiac, carrier, and secondary-findings categories. The current built-in secondary-findings subset contains 66 unique gene symbols. It is not the complete 81-gene ACMG SF v3.2 list despite the legacy internal label.
The exact internal values assigned to these categories are proprietary. The category and resulting age-relevance component remain visible for interpretation.
Interpretation Limits
A lower age-relevance value does not mean that the variant is benign or clinically irrelevant.
Age relevance does not model penetrance, age of onset, competing risk, or individual treatment eligibility.
Panel curation can change over time. Review the panel and analysis provenance used for the result.
Diagnostic cases with patient HPO terms use the Diagnostic profile, where phenotype replaces age relevance as the dominant contextual signal.
Related Documentation
Gene Panels and Custom Genes
How curated age annotations and panel metadata enter Screening.
Screening Modes
How phenotype and age select the active prioritization profile.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Screening](https://folklore.helena.bio/docs/screening)
- [Gene Panels and Custom Genes How curated age annotations and panel metadata enter Screening.](https://folklore.helena.bio/docs/screening/gene-panels)
- [Screening Modes How phenotype and age select the active prioritization profile.](https://folklore.helena.bio/docs/screening/screening-modes)

---

## Page: Gene Panels and Custom Genes | Folklore Documentation

Source: https://folklore.helena.bio/docs/screening/gene-panels
Canonical: https://folklore.helena.bio/docs/screening/gene-panels
Description: How Folklore gene panels define screening scope, carry curated disease and age context, and combine with custom genes.

Documentation / Screening / Gene Panels and Custom Genes
## Gene Panels and Custom Genes
A gene panel is both a scope boundary and a source of curated clinical context. Screening loads variants only from the resolved gene set, then uses panel metadata alongside the standard component scores and patient context.
Panels Restrict the Search Space
Selecting a panel does not merely boost its genes after a genome-wide screen. The selected panel genes define which classified variants are loaded into Screening. Add custom genes when the clinical question extends beyond the selected panels.
How the Gene Set Is Resolved
Panels only
The union of genes in the selected panels is screened.
Custom genes only
Only the supplied custom genes are screened.
Panels and custom genes
The two sets are combined. If the same symbol occurs in both, the custom-gene metadata takes precedence for that run.
No panel or custom gene
Folklore uses a built-in fallback assembled from its current secondary-findings, pediatric, adult-onset, and carrier lists. It does not screen every annotated gene.
Panel Gene Metadata
Field Use in Screening
Gene symbol Defines the gene in scope. Managed panel entries are validated against the deployed HGNC symbol index and aliases are normalized to an approved symbol.
Disease name and display label Provides human-readable panel context in screening results and reports.
Age relevance Records neonatal, pediatric, adult, or all-age relevance for age-aware prioritization.
Priority score Provides a bounded panel-specific prioritization value from 0.0 to 1.0.
ClinGen status Modulates panel context when a gene-disease validity label is available. The label remains curation metadata, not a variant classification.
Notes Stores additional administrative or curation context for managed panel entries.
Built-in and Organization Panels
Built-in panels are shared platform resources. Organization panels are visible within their organization. Organization administrators can create panels and maintain their gene entries; built-in panel changes require platform administration. Panels can be deactivated without immediately deleting their definition.
Available built-in panels can change as curation is updated. Review the selected panel name, active status, gene count, and gene list for the analysis being interpreted rather than relying on a remembered panel composition.
Custom Genes
Custom genes are supplied with an individual analysis or batch and do not require creation of a reusable organization panel. They can carry the same prioritization and display context used by panel genes. Use approved gene symbols and record why each gene was added to the clinical scope.
Interpretation Boundaries
Absence from the selected gene scope is not evidence that a gene or variant is clinically irrelevant.
Panel priority and ClinGen metadata influence prioritization; they do not replace variant-level ACMG evidence.
An all-age label means the panel entry remains relevant across age groups, not that every associated condition is equally actionable at every age.
Changing panel membership or metadata can change the returned result set and scores. Preserve the analysis provenance used for clinical review.
Related Documentation
Age-Aware Prioritization
How panel age annotations interact with patient age.
Scoring Components
The seven component scores calculated for each in-scope variant.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Screening](https://folklore.helena.bio/docs/screening)
- [Age-Aware Prioritization How panel age annotations interact with patient age.](https://folklore.helena.bio/docs/screening/age-aware-prioritization)
- [Scoring Components The seven component scores calculated for each in-scope variant.](https://folklore.helena.bio/docs/screening/scoring-components)

---

## Page: Screening Scoring Components | Folklore Documentation

Source: https://folklore.helena.bio/docs/screening/scoring-components
Canonical: https://folklore.helena.bio/docs/screening/scoring-components
Description: The clinical evidence dimensions used to prioritize in-scope variants in Folklore Screening, with interpretation and missing-data boundaries.

Documentation / Screening / Scoring Components
## Screening Scoring Components
Every in-scope variant is evaluated across seven evidence dimensions. The component values are combined according to the active clinical profile, then selected patient-context signals are added before the final review tier is assigned.
Folklore shows the component values so reviewers can understand the direction of prioritization. The exact internal weighting and calibration are proprietary. Component and total values are prioritization scores, not probabilities of pathogenicity or clinical benefit.
Evidence Dimensions
Component What It Measures Primary Inputs
Gene Constraint How intolerant the gene is to loss-of-function or missense variation, interpreted in the context of the variant consequence. pLI, LOEUF, missense constraint
Deleteriousness Whether multiple in-silico signals support a damaging protein or splicing effect. BayesDel_noAF, SpliceAI, AlphaMissense, DANN, SIFT, MetaSVM, PhyloP, GERP
Phenotype Relevance Patient-to-gene HPO overlap in diagnostic cases, or a conservative gene-disease association signal when no patient HPO terms are supplied. Patient HPO terms and gene HPO annotations
Dosage Sensitivity Whether the predicted consequence is compatible with available gene haploinsufficiency evidence. ClinGen dosage curation and consequence
Consequence The predicted severity and coding relevance of the annotated consequence. VEP consequence and transcript biotype
Compound-Heterozygous Potential Whether a qualifying heterozygous coding or splice-relevant variant has a same-gene partner or an upstream candidate flag. Genotype, consequence, same-gene variants
Age Relevance How the gene fits the patient age and the selected panel annotation. Patient age, panel age relevance, curated fallback categories
Deleteriousness Is an Ensemble Signal
The current engine uses eight protein, splice, meta-prediction, and conservation inputs. BayesDel_noAF is the primary calibrated signal, with other predictors providing complementary coverage. This ensemble contributes to screening priority and must not be counted again as independent ACMG evidence without reviewing the Variant Analysis evidence trace.
Phenotype and No-Phenotype Behavior
When patient HPO terms are present, the phenotype component measures direct overlap with gene HPO annotations and the Diagnostic profile gives phenotype greater influence. Without patient HPO terms, the component uses a deliberately capped gene-disease association signal. Non-coding consequences are discounted so a heavily annotated gene does not dominate solely because it has many known phenotypes.
Compound-Heterozygous Boundary
The component identifies candidate pairs among qualifying heterozygous coding or splice-relevant variants in the same gene. It does not prove that the variants are in trans, establish biallelic pathogenicity, or replace Family Analysis phasing.
Missing Data
Missing annotations use conservative defaults. Missing BayesDel is handled by redistributing its contribution across the remaining predictor slots; other absent predictor values do not receive the same per-field redistribution. Scores from variants with different annotation completeness should therefore be compared with care.
Clinical Boundary
A high component value explains why a variant moved upward in the review queue. It does not change the ACMG class, confirm a molecular diagnosis, or determine whether a finding is reportable.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Screening](https://folklore.helena.bio/docs/screening)

---

## Page: Screening Modes | Folklore Documentation

Source: https://folklore.helena.bio/docs/screening/screening-modes
Canonical: https://folklore.helena.bio/docs/screening/screening-modes
Description: The six recorded Folklore screening strategies and how phenotype and age select the active prioritization profile.

Documentation / Screening / Screening Modes
## Screening Modes
Folklore records one of six screening strategies for provenance and interpretation. The active component profile is selected first by phenotype availability and otherwise by the patient's age group. The exact internal weights are proprietary; the clinical emphasis and current limitations are documented here.
Available Modes
Diagnostic
Used when the case has patient HPO terms or explicitly requests diagnostic screening. Phenotype relevance becomes the dominant clinical component and age relevance is not used in the component profile.
Neonatal Screening
Records a newborn screening strategy. Without patient HPO terms, the active profile is still selected from the precise patient age group.
Pediatric Screening
Records child or adolescent screening. Infant, child, and adolescent ages share the current pediatric component profile while retaining different age-relevance matching.
Proactive Adult
The default requested strategy for screening without a specific phenotype. Adult and elderly ages use distinct active profiles.
Carrier Screening
Records a reproductive carrier-screening intent. It currently uses the same age-derived component profile as other no-phenotype strategies rather than a dedicated carrier weight profile.
Pharmacogenomics
Records a pharmacogenomics intent. It currently uses the age-derived component profile rather than a dedicated pharmacogenomics weight profile.
Automatic Profile Selection
Patient HPO terms activate the Diagnostic profile regardless of the requested strategy. Without HPO terms, Folklore uses a neonatal, pediatric, adult, or elderly profile derived from age. When no usable age is supplied, Adult is the fallback age group. Proactive Adult is the default recorded strategy.
What Changes Between Profiles
Diagnostic screening emphasizes patient phenotype. No-phenotype profiles balance gene constraint, predicted effect, dosage sensitivity, consequence, candidate biallelic context, gene-disease burden, and age relevance. Age relevance has progressively greater influence in the adult and elderly profiles, while neonatal and pediatric profiles retain broader early-onset context.
Mode Labels Do Not Guarantee Dedicated Content
Carrier Screening and Pharmacogenomics are accepted and stored strategies, but they do not yet select dedicated component profiles. The result should therefore be interpreted from its actual component values, gene scope, and provenance rather than from the mode label alone.
Related Documentation
Scoring Components
The evidence dimensions combined by the active profile.
Age-Aware Prioritization
Precise age groups and panel age-relevance behavior.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Screening](https://folklore.helena.bio/docs/screening)
- [Scoring Components The evidence dimensions combined by the active profile.](https://folklore.helena.bio/docs/screening/scoring-components)
- [Age-Aware Prioritization Precise age groups and panel age-relevance behavior.](https://folklore.helena.bio/docs/screening/age-aware-prioritization)

---

## Page: Screening Tier System | Folklore Documentation

Source: https://folklore.helena.bio/docs/screening/tier-system
Canonical: https://folklore.helena.bio/docs/screening/tier-system
Description: How Folklore assigns final screening priority tiers, applies context-aware promotion, and keeps tier separate from ACMG classification and clinical actionability.

Documentation / Screening / Tier System
## Screening Tier System
Screening tiers organize returned variants into review queues after component scoring and patient-context prioritization. The returned tier uses the final bounded score together with explicit classification- and phenotype-priority rules.
Tier Definitions
Tier 1 High Priority
The first review queue. It includes findings promoted by strong classification or phenotype context and findings whose final score reaches the highest returned band.
Tier 2 Moderate Priority
The second review queue, including moderate phenotype support and selected inheritance-aware carrier demotions.
Tier 3 Lower Priority
Findings that pass the configured minimum score but do not meet the higher-priority rules.
Tier 4 Below Default Filter
Findings below the default minimum score. They are normally omitted and require both a lower minimum-score setting and explicit Tier 4 inclusion.
What Can Raise Priority
Pathogenic and Likely Pathogenic classifications receive direct priority treatment.
A VUS carrying stored PVS1 evidence receives a dedicated prioritization treatment but remains a VUS.
High clinical phenotype tiers can promote a finding independently of the numeric score band.
Patient context can include ancestry, family history, sex-linked context, consanguinity, parental-sample context, pregnancy or family-planning context, and selected panel metadata.
The final score is capped at 1.0 before the returned tier is assigned.
Autosomal-Recessive Carrier Guard
A heterozygous finding in a gene recorded as autosomal recessive is returned in Tier 2 instead of Tier 1 when no compound-heterozygous candidate is supported. The current guard does not apply when the stored evidence includes PVS1. This is a review-priority safeguard, not a carrier interpretation.
Result Limits
Tier 1 is capped at 20 returned variants by default and can be configured up to 100. Findings beyond the configured Tier 1 cap are not automatically reassigned to Tier 2. The default minimum total score is 0.20, and Tier 4 is excluded by default.
Clinical Actionability Labels
Folklore also stores a separate actionability label: Immediate, Monitoring, Future, or Research. The current Immediate rule is limited to Pathogenic or Likely Pathogenic findings in the service's built-in 66-gene secondary-findings subset. Other P/LP findings are labeled Monitoring. Score bands can also place non-P/LP findings into Monitoring or Future.
These labels support display and triage. They do not determine diagnosis, reporting, surveillance, or patient management.
Tier Is Not Classification
A VUS can be Tier 1, and a clinically important P/LP finding can be subject to inheritance-aware prioritization. The ACMG class and its evidence trace remain authoritative and are never replaced by the screening tier.

### Links and cited sources

- [Documentation](https://folklore.helena.bio/docs)
- [Screening](https://folklore.helena.bio/docs/screening)


## Page: Phenotype Matching in Genomic Variant Interpretation

Source: https://folklore.helena.bio/phenotype-matching
Canonical: https://folklore.helena.bio/phenotype-matching
Description: A scientific guide to HPO phenotype matching, ontology-aware semantic similarity and phenotype-driven gene and variant prioritization in genomic review.

# Phenotype matching in genomic variant interpretation

How HPO and ontology-aware semantic similarity connect clinical findings with gene-associated phenotypes and create a reviewable order of genomic evidence.

Helena Bioinformatics. Reviewed 8 August 2026. The page includes a 22 minute self-hosted scientific briefing, optional English captions and an edited text transcript covering all 23 slides.

## Clinical context

Sequencing can produce millions of observations. Technical filtering reduces that list, but it does not answer which candidates fit the person being evaluated. Phenotype matching compares a reviewed patient profile with phenotype knowledge associated with candidate genes. The comparison uses HPO concepts and their relationships, so related findings can retain useful similarity when records use synonyms or different levels of specificity.

## HPO as structured clinical data

The Human Phenotype Ontology assigns stable identifiers to phenotypic abnormalities and places them in a directed hierarchy. A clinical note can use familiar language while the coded finding provides a computable representation. Selected terms still need qualified review. Age, disease stage, negated observations, missing findings and false precision can alter the available signal.

## Ontology-aware semantic similarity

Semantic similarity compares concepts through their position in the ontology and through their information content. Broad, common findings contribute less discrimination than rare, specific shared concepts. The system evaluates patient terms against a candidate gene profile and considers the complete profile rather than a count of matching words. The comparison does not depend on a language model judgment that two phrases sound similar.

## Folklore workflow

Folklore stores reviewed HPO terms with the analysis session. Candidate genes carry phenotype profiles derived from curated knowledge. Gene-level relevance is presented beside variant classification, consequence, frequency, inheritance and case evidence. Molecular evidence remains variant-specific. Reviewers can move from compact gene summaries to the variants, sources and term relationships behind a candidate.

## Interpretation boundary

A phenotype match score is evidence. It is not a diagnosis. It can focus attention and organise a large result set. It cannot prove causality, repair an incomplete phenotype profile or replace assessment of zygosity, inheritance, penetrance, segregation and molecular mechanism. A common benign variant does not become pathogenic because its gene matches the phenotype. A qualified clinical geneticist interprets the complete record.

## Failure modes

An incomplete profile can under-rank a relevant gene. Gene and phenotype knowledge changes, so a configured reference release can contain annotation gaps. Broad findings carry limited discriminatory value. Phenocopies can produce similar patterns through different mechanisms. Incorrect HPO terms can mislead the comparison. Absence of a strong match is not proof against causality.

## Edited transcript coverage

The visible page contains an edited 23-part transcript in initial server-rendered HTML. Its sections cover the interpretation bottleneck, phenotype coding, HPO structure, exact matching failures, semantic similarity, input quality, patient and candidate profiles, set-to-set comparison, whole-genome scale, integrated priority, clinical tiers, case context, a fictional case without patient data, failure modes, human review and primary sources. The video includes the complete narration with optional English captions that can be turned on or off in the player.

## Primary sources

1. Gargano MA et al. The Human Phenotype Ontology in 2024. Nucleic Acids Research. 2024;52:D1333 to D1346. doi:10.1093/nar/gkad1005.
2. Köhler S et al. Clinical diagnostics in human genetics with semantic similarity searches in ontologies. American Journal of Human Genetics. 2009;85:457 to 464. doi:10.1016/j.ajhg.2009.09.003.
3. Jacobsen JOB et al. Phenotype-driven approaches to enhance variant prioritization and diagnosis of rare disease. Human Mutation. 2022;43:1071 to 1081. doi:10.1002/humu.24380.
4. Human Phenotype Ontology official resource: https://hpo.jax.org/.

## Related canonical resources

Phenotype matching platform: https://folklore.helena.bio/platform/phenotype-matching
Technical documentation: https://folklore.helena.bio/docs/phenotype-matching
Semantic similarity documentation: https://folklore.helena.bio/docs/phenotype-matching/semantic-similarity


## Page: Computational Predictors in ACMG Variant Classification

Source: https://folklore.helena.bio/computational-predictors
Canonical: https://folklore.helena.bio/computational-predictors
Description: A clinical guide to computational predictors in ACMG variant classification, including BayesDel_noAF, SpliceAI, PP3, BP4, calibration, double-counting safeguards and interpretation limits.

# Computational predictors in ACMG variant classification

Computational predictors estimate properties such as missense deleteriousness, splice disruption and evolutionary constraint. They do not classify a variant on their own. Their formal contribution depends on scope, calibrated score intervals and independence from evidence already counted elsewhere.

Helena Bioinformatics. Reviewed 10 August 2026. The page includes a 24 minute self-hosted scientific briefing, optional English captions and an edited transcript covering all 20 slides.

## Clinical context

ACMG and AMP classification weighs evidence from population data, observations in affected individuals and families, functional studies, phenotype, allelic data and computational predictions. PP3 records computational support for a deleterious effect. BP4 records computational support for no impact. A predictor result is one evidence stream inside that larger record.

## Three admission checks

Scope: confirm that the predictor applies to the variant type, mechanism and clinically relevant transcript.

Calibration: use the exact tool and score interval that were validated for a defined evidence strength. A developer label such as damaging is not enough.

Independence: do not count the same information twice. Shared training signals, allele frequency, protein domains and predicted loss of function can create dependence between criteria.

## Formal missense path: BayesDel_noAF

Folklore uses BayesDel_noAF for formal missense computational evidence. The noAF model excludes allele frequency, reducing circularity with population criteria. The current pathogenic intervals are 0.130 to 0.289 for PP3 Supporting, 0.290 to 0.517 for PP3 Moderate and 0.518 or higher for PP3 Strong. The benign intervals are minus 0.360 to minus 0.181 for BP4 Supporting and minus 0.361 or lower for BP4 Moderate. Scores from minus 0.180 through 0.129 are indeterminate.

The missense path includes relevance and disease-association guards. When PM1 and PP3 both apply, their combined contribution is capped at a Strong equivalent. Benign missense evidence also requires no concerning SpliceAI signal.

## Formal splice path: SpliceAI

SpliceAI reports acceptor gain, acceptor loss, donor gain and donor loss delta scores. Folklore uses the maximum score. The current production intervals are 0.2 to below 0.5 for PP3_splice Supporting, 0.5 to below 0.8 for PP3_splice Moderate and 0.8 or higher for PP3_splice Strong. The higher levels remain subject to consequence, transcript, gene-mechanism and classification safeguards.

PP3_splice is suppressed when PVS1 already captures the same predicted loss-of-function mechanism. Pure missense variants with BayesDel data but no VEP splice consequence do not receive a second computational splice criterion. Missense and splice predictions are not added as independent votes when they describe the same variant consequence. A SpliceAI score at or below 0.1 can support defined benign splice guards, but eligibility depends on consequence and criterion-specific rules.

## Predictors displayed as clinical context

SIFT estimates whether an amino acid substitution is tolerated from sequence homology. AlphaMissense adds a protein model perspective. MetaSVM combines multiple annotations through a support vector machine. DANN scores single nucleotide variants with a neural network. PhyloP and GERP describe evolutionary conservation and constraint. Folklore displays these outputs for qualified review, but they do not add separate ACMG votes in the current formal classification path.

Agreement among displayed predictors does not multiply evidence. Disagreement can identify a difference in tool scope, transcript or biological mechanism that merits review. A missing score means that computational evidence is unavailable, not that the variant is benign.

## Worked example

For a synthetic missense variant with BayesDel_noAF 0.35 and SpliceAI 0.02, the missense score enters the Moderate pathogenic interval and contributes PP3_Moderate. Concordant AlphaMissense or SIFT outputs remain contextual and do not add more PP3 criteria. Classification then returns to population frequency, gene-disease validity, phenotype, segregation, functional data and other independent evidence.

## Interpretation boundary

Computational predictors can focus attention and contribute calibrated evidence. They do not establish a gene-disease relationship, prove causality for the patient's phenotype or replace functional and segregation evidence. Even a highly accurate model can fail for a specific transcript, mechanism or gene. A qualified clinical geneticist interprets the complete evidence record.

## Primary sources

1. Richards S et al. Standards and guidelines for the interpretation of sequence variants. Genetics in Medicine. 2015;17:405 to 424. PMID:25741868.
2. Pejaver V et al. Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3 and BP4 criteria. American Journal of Human Genetics. 2022;109:2163 to 2177. PMID:36413997.
3. Walker LC et al. Using the ACMG and AMP framework to capture evidence related to predicted and observed impact on splicing. American Journal of Human Genetics. 2023;110:1046 to 1067. PMID:37352859.
4. Jaganathan K et al. Predicting splicing from primary sequence with deep learning. Cell. 2019;176:535 to 548. PMID:30661751.

## Related canonical resources

Predictor documentation: https://folklore.helena.bio/docs/predictors
BayesDel_noAF documentation: https://folklore.helena.bio/docs/predictors/bayesdel
SpliceAI documentation: https://folklore.helena.bio/docs/predictors/spliceai
Variant classification platform: https://folklore.helena.bio/platform/variant-classification


## Page: Folklore Clinical Variant Interpretation MCP Connector

Source: https://folklore.helena.bio/docs/folklore-connector
Canonical: https://folklore.helena.bio/docs/folklore-connector
Description: Use Folklore Clinical Variant Interpretation MCP from an MCP-compatible AI client.

# Folklore Clinical Variant Interpretation MCP Connector

Folklore Clinical Variant Interpretation MCP is the official public, read-only MCP server from Helena Bioinformatics. It lets an MCP-compatible AI client look up one supported GRCh38 germline variant and return structured Folklore evidence, automated ACMG/AMP decision support, transparent provenance and related scientific literature. It does not provide patient diagnoses or treatment recommendations, and its results require review by a qualified genetics professional.

## Connection details

The connector uses Model Context Protocol over Streamable HTTP. Its public endpoint is https://api.helena.bio/folklore/v1/mcp. The available tools are search_variant_evidence, search_variant_literature and get_publication_details. Authentication is not required, and all operations are read-only.

## Request contract

The tool accepts GRCh38 chromosome coordinates, genomic or coding HGVS, protein HGVS, SPDI and rsID. The current public scope covers nuclear germline SNVs and simple insertions or deletions shorter than 50 base pairs. GRCh37, mitochondrial, structural and somatic variants return an unsupported result.

The request contains only the assembly and variant expression. Patient, phenotype, family, segregation and case information are outside the public contract.

## Returned evidence

A resolved result contains the GRCh38 identity, current annotation, automated ACMG/AMP decision-support classification, applied evidence, ClinVar assertions, population frequency and computational predictions when available. It also records source provenance, data versions, limitations and a link to the matching public Folklore record.

The response status is resolved, ambiguous, not found, invalid, unsupported or temporarily unavailable.

## Clinical use

Folklore provides variant-level decision support for qualified genetics professionals. Its automated classification does not establish a patient diagnosis or determine treatment. A qualified professional must review the evidence and confirm clinically significant findings before they enter a clinical report or patient-care decision.

## Related canonical resources

MCP clients and public listings: https://folklore.helena.bio/integrations
All-version public adapter DOI: https://doi.org/10.5281/zenodo.21922951
Version 1.2.2 public adapter DOI: https://doi.org/10.5281/zenodo.21922952
Platform limitations: https://folklore.helena.bio/docs/limitations
Classification methodology: https://folklore.helena.bio/methodology
Privacy Policy: https://www.helena.bio/privacy
Terms of Service: https://www.helena.bio/terms


## Page: Folklore Clinical Variant Interpretation MCP Integrations

Source: https://folklore.helena.bio/integrations
Canonical: https://folklore.helena.bio/integrations
Description: Connect MCP-compatible AI clients to Folklore Clinical Variant Interpretation MCP and find verified public listings.

# Folklore Clinical Variant Interpretation MCP Integrations

Folklore Clinical Variant Interpretation MCP is the official public, read-only MCP server from Helena Bioinformatics. It interprets supported GRCh38 germline variants using structured Folklore evidence, automated ACMG/AMP variant-level decision support, transparent provenance and related scientific literature. The server does not provide patient diagnoses or treatment recommendations. Its results are intended for review by qualified genetics professionals. MCP-compatible clients connect through Streamable HTTP at https://api.helena.bio/folklore/v1/mcp. Authentication is not required. The tools are search_variant_evidence, search_variant_literature and get_publication_details. The Official MCP Registry identity is io.github.helena-bioinformatics/folklore.

## Direct client connections

ChatGPT can use Folklore as a custom MCP app. Claude can use it as a custom connector. Codex, Cursor, VS Code, Windsurf, Perplexity, Grok and Gemini CLI can add the same remote endpoint through their MCP or connector settings. Select HTTP transport when a client asks for the transport type.

A tested custom connection proves compatibility for that account or workspace. It does not prove public provider-directory approval.

## Verified public listings

The Official MCP Registry is the canonical machine-readable publication source. Folklore Clinical Variant Interpretation MCP version 1.2.2 is published there. Public listings were also verified on Glama, MCPBeat and mcpservers.org on 13 August 2026.

MCP.so and the awesome-mcp-servers list were under review on that date. ChatGPT and Claude directory submissions were complete, but public approval had not been independently verified. Smithery publication was deferred. A submitted or deferred integration must not be described as an active listing.

## Request and result boundary

The tool accepts one supported GRCh38 nuclear germline SNV or simple indel smaller than 50 base pairs. It accepts coordinates, selected genomic, coding and protein HGVS, SPDI and rsID. Ambiguous expressions return candidates for the user to choose.

Folklore returns variant-level evidence, provenance, automated ACMG/AMP decision support and explicit limitations. It accepts no patient, phenotype, family, segregation or case context. Results require review by a qualified genetics professional and must not be presented as a diagnosis, treatment recommendation or standalone clinical report.

## Canonical resources

Technical connector guide: https://folklore.helena.bio/docs/folklore-connector
MCP Server Card: https://folklore.helena.bio/.well-known/mcp/server-card.json
Public integration repository: https://github.com/helena-bioinformatics/folklore-mcp
All-version public adapter DOI: https://doi.org/10.5281/zenodo.21922951
Version 1.2.2 public adapter DOI: https://doi.org/10.5281/zenodo.21922952
Official MCP Registry: https://registry.modelcontextprotocol.io/v0/servers?search=io.github.helena-bioinformatics%2Ffolklore
