Scientific briefing
Computational predictors in ACMG variant classification
How calibrated missense and splice scores enter PP3 and BP4, why several displayed tools do not become extra votes, and where clinical review remains essential.
Helena Bioinformatics | Reviewed 10 August 2026 | 24 minute video
24:28 / Scientific briefing with optional English captions
Read the edited transcriptClinical context
A predictor estimates one property. Classification weighs the complete evidence record.
Computational predictors can estimate missense deleteriousness, splice disruption or evolutionary constraint. ACMG classification asks a wider question: whether the total evidence supports pathogenicity or benignity for a specific disease and inheritance model.
The score can contribute through PP3 or BP4 only when its scope, calibrated interval and relationship to other evidence are explicit. Popularity, agreement among several correlated tools or an extreme-looking output is not enough.
Evidence admission
Three checks come before the score
01
Scope
Confirm the variant type, biological mechanism and clinically relevant transcript supported by the tool.
02
Calibration
Use a predefined model and a validated interval that maps to a stated evidence strength.
03
Independence
Cap or omit evidence that repeats frequency, domain, conservation or loss-of-function information already counted.
Formal computational evidence
Missense and splice predictions follow separate paths
Missense
BayesDel_noAF
The noAF model excludes allele frequency, reducing overlap with population criteria.
- PP3 Supporting
- 0.130 to 0.289
- PP3 Moderate
- 0.290 to 0.517
- PP3 Strong
- 0.518 or higher
- BP4 Supporting
- minus 0.360 to minus 0.181
- BP4 Moderate
- minus 0.361 or lower
The indeterminate interval is minus 0.180 through 0.129. Relevance, disease-association, splice and PM1 dependency guards still apply.
Splice
SpliceAI
The maximum acceptor gain, acceptor loss, donor gain or donor loss delta score enters the splice path.
- PP3_splice Supporting
- 0.2 to below 0.5
- PP3_splice Moderate
- 0.5 to below 0.8
- PP3_splice Strong
- 0.8 or higher
- Benign splice guard
- 0.1 or lower
PVS1 blocks duplicate evidence for the same loss-of-function mechanism. Consequence, transcript, disease-association and missense mutual-exclusion safeguards determine eligibility.
Current implementation note
The video introduces the 0.2 SpliceAI entry threshold. The current production classifier also modulates eligible splice predictions at 0.5 and 0.8. The intervals and safeguards shown on this page reflect the live implementation reviewed on 10 August 2026.
Displayed clinical context
Visible does not mean counted
Folklore displays additional predictors so the geneticist can inspect conservation, sequence tolerance, protein-model context and broader deleteriousness signals. They do not add separate ACMG votes in the current formal classification path.
SIFT
Amino acid substitution tolerance from sequence homology.
AlphaMissense
Missense pathogenicity context from protein sequence and structure-aware modelling.
MetaSVM
A meta-predictor that combines multiple variant annotations.
DANN
A neural-network score for single nucleotide variants.
PhyloP
Position-level acceleration or conservation relative to neutral evolution.
GERP
Evolutionary constraint from rejected substitutions.
Worked example
One calibrated contribution, then back to the case
A synthetic missense variant has BayesDel_noAF 0.35 and SpliceAI 0.02. AlphaMissense reports likely pathogenic and SIFT reports deleterious. BayesDel 0.35 falls in the Moderate pathogenic interval, so the formal computational contribution is PP3_Moderate. The displayed predictors do not add more PP3 evidence.
The classification then returns to population frequency, gene-disease validity, phenotype, segregation, functional data, inheritance and other independent evidence. The computational result has narrowed the question without deciding the case.
Interpretation boundary
A predictor can focus attention and contribute calibrated evidence. It cannot establish the gene-disease relationship, prove that the molecular effect explains the patient or replace functional and segregation evidence. A qualified clinical geneticist interprets the complete record.
Common failure modes
Where predictor evidence can mislead
- Wrong scope. The model does not support the variant type, mechanism or clinically relevant transcript.
- Uncalibrated threshold. A developer label or visually extreme score is treated as evidence strength without validation.
- Correlated votes. Several tools repeat overlapping training data, conservation features or component scores.
- Mechanism counted twice. A prediction repeats frequency, domain or loss-of-function evidence already represented by another criterion.
- Missing interpreted as benign. An unavailable score is treated as evidence of no impact.
Primary scientific sources
- 1. Richards S et al. Standards and guidelines for the interpretation of sequence variants. Genetics in Medicine. 2015;17:405 to 424. PMID:25741868
- 2. Pejaver V et al. Calibration of computational tools for missense variant pathogenicity classification and ClinGen recommendations for PP3 and BP4 criteria. American Journal of Human Genetics. 2022;109:2163 to 2177. PMID:36413997
- 3. Walker LC et al. Using the ACMG and AMP framework to capture evidence related to predicted and observed impact on splicing. American Journal of Human Genetics. 2023;110:1046 to 1067. PMID:37352859
- 4. Jaganathan K et al. Predicting splicing from primary sequence with deep learning. Cell. 2019;176:535 to 548. PMID:30661751
Accessible text version
Edited video transcript
Repeated spoken phrases and verbal corrections have been removed for reading. The sequence follows all 20 slides. The SpliceAI section includes the current production strength modulation described above. The video includes optional English captions generated directly from the supplied SRT timings.
- 01
Computational predictors
Computational predictors are among the most searched topics in variant interpretation. Names such as BayesDel, AlphaMissense, MetaSVM, DANN, SIFT and SpliceAI lead to a practical question: what should any score change in an ACMG classification? This briefing separates predictors that provide formal computational evidence from predictors shown as clinical context. It also separates missense and splice mechanisms and explains calibration, evidence strength and double counting. The goal is a disciplined workflow: define the biological question, choose the appropriate predictor, apply a validated threshold, assign only the permitted evidence strength and combine the result with independent evidence.
- 02
Evidence, not a verdict
A computational predictor does not classify a variant. It estimates a property that may be relevant to classification. SIFT estimates whether an amino acid substitution is tolerated at a conserved protein position. SpliceAI estimates whether sequence context may create or disrupt splice sites. BayesDel combines several signals into a missense pathogenicity score. ACMG classification asks whether the total evidence supports pathogenicity or benignity for a particular disease and inheritance model. That record can include population frequency, functional studies, segregation, phenotype, allelic data and computational evidence.
- 03
The ACMG evidence frame
The ACMG and AMP framework organises evidence by direction and strength. Pathogenic evidence codes begin with P and benign evidence codes begin with B. Computational evidence enters through PP3 on the pathogenic side and BP4 on the benign side. Quantitative calibration maps defined score intervals to defined evidence strengths. A predictor is not admitted because it is popular or technically sophisticated. It is admitted because a specified range has been evaluated against an appropriate truth set under a documented method.
- 04
One stream in the complete record
Population evidence asks whether an allele frequency is compatible with the disease mechanism. Functional and case evidence asks what has been observed in experiments, families and affected individuals. Computational evidence infers biological impact from sequence patterns and annotations. It is fast, reproducible and available for many variants, including variants that lack a patient series. Its limitation is that it remains an inference rather than a direct observation of causality. A rare variant with a high predictor score may therefore remain a VUS.
- 05
PP3 and BP4
PP3 records computational support for a deleterious effect. BP4 records computational support for no impact. Running several tools and taking a majority vote can repeat correlated information because predictors often share conservation features, training variants and annotations. A defensible process chooses one predefined tool for a formal evidence path, maps its score to a calibrated strength and leaves an indeterminate interval where neither criterion applies. A missing damaging prediction is not automatically benign evidence.
- 06
Scope, calibration and independence
Three rules make computational evidence defensible. Choose the tool before seeing the answer. Use validated score intervals rather than developer labels alone. Protect independence from evidence already counted elsewhere. A predictor that includes allele frequency can overlap with population criteria. A splice prediction can describe the same mechanism already captured by PVS1. These checks determine whether a score can enter an auditable clinical workflow or should remain contextual information.
- 07
Two formal paths in Folklore
Folklore uses BayesDel_noAF for the formal missense path and SpliceAI for the formal splice path. BayesDel_noAF excludes allele frequency and reduces circularity with PM2, BA1 and BS1. SpliceAI evaluates local sequence context for donor or acceptor gain and loss. A variant may raise a missense question, a splice question or both. The scores are not averaged. Each path has its own thresholds, consequences and safeguards.
- 08
Why BayesDel_noAF
BayesDel is a meta-predictor that combines information from multiple annotations and prediction methods. The noAF model removes allele frequency, which matters because frequency is already assessed through independent ACMG population criteria. ClinGen SVI calibration also gives the score a graded interpretation. A score just above a pathogenic boundary does not carry the same weight as an extreme score, and an indeterminate region receives no computational criterion.
- 09
BayesDel_noAF thresholds
The current Folklore intervals are 0.130 to 0.289 for PP3 Supporting, 0.290 to 0.517 for PP3 Moderate and 0.518 or higher for PP3 Strong. Scores from minus 0.360 to minus 0.181 support BP4, while minus 0.361 or lower supports Moderate benign evidence. Scores from minus 0.180 through 0.129 are indeterminate. The missense path also applies relevance and disease-association guards. PM1 and PP3 are capped at a combined Strong equivalent, and benign missense evidence requires no concerning splice signal.
- 10
SpliceAI thresholds and current safeguards
SpliceAI produces acceptor gain, acceptor loss, donor gain and donor loss delta scores, and Folklore reads their maximum. The current production intervals are 0.2 to below 0.5 for PP3_splice Supporting, 0.5 to below 0.8 for Moderate and 0.8 or higher for Strong. These strengths remain subject to consequence, disease-association and transcript safeguards. If PVS1 already captures the same loss-of-function mechanism, the splice criterion is suppressed. A score at or below 0.1 participates only in defined benign guards whose eligibility depends on the variant consequence.
- 11
Do not count one signal twice
ACMG codes are combined as evidence contributions, so dependence matters. Two predictors can share training sets or component annotations. A splice model and a loss-of-function rule can describe the same transcript disruption. A predictor containing frequency can repeat rarity evidence. More displayed signals do not necessarily mean more independent information. Calibration controls the strength of one tool, while dependency guards control its relationship with the rest of the record.
- 12
Formal evidence and displayed context
Several predictor scores can appear together on a result screen, but not all of them contribute to automated classification. BayesDel_noAF and SpliceAI connect to explicit formal evidence paths. SIFT, AlphaMissense, MetaSVM, DANN, PhyloP and GERP are displayed for clinical context. This separation lets a reviewer inspect useful signals without turning every output into another ACMG vote.
- 13
SIFT, PhyloP and GERP
SIFT evaluates a missense substitution using homologous protein sequences. PhyloP measures whether a genomic position evolves more slowly or quickly than expected under a neutral model. GERP estimates evolutionary constraint from substitutions rejected over time. Conservation can identify an important position, but it does not prove that every possible substitution at that position is damaging. These scores remain contextual in Folklore.
- 14
AlphaMissense, MetaSVM and DANN
AlphaMissense combines protein sequence modelling with structural context. MetaSVM combines multiple annotations through a support vector machine. DANN uses a neural network over genomic annotations and can score single nucleotide variants beyond coding missense positions. Their architectures and scopes differ, but difference does not guarantee statistical independence. Folklore displays them for review while the formal missense path uses BayesDel_noAF.
- 15
Agreement, disagreement and missing scores
Agreement can provide useful context, but it does not multiply formal evidence. If BayesDel crosses a calibrated threshold and other displayed tools point in the same direction, PP3 is still recorded once through the predefined path. Disagreement may reflect different transcripts, variant scopes or mechanisms and should trigger inspection. A missing output means that evidence is unavailable. It does not mean that the predictor returned a benign result.
- 16
Synthetic worked example
Consider a synthetic missense variant with BayesDel_noAF 0.35 and SpliceAI 0.02. AlphaMissense predicts likely pathogenic and SIFT predicts deleterious. The missense mechanism is relevant, and BayesDel 0.35 falls in the Moderate pathogenic interval, so the formal computational contribution is PP3_Moderate. AlphaMissense and SIFT do not add more PP3 points. The reviewer returns to population frequency, gene-disease validity, phenotype, segregation, functional data and other independent evidence.
- 17
Review checklist
Before a predictor changes classification, confirm scope, calibration and independence. Scope asks whether the tool supports the variant type, mechanism and clinically relevant transcript. Calibration asks whether the exact score interval maps to a documented evidence strength. Independence asks whether the same biological or statistical signal has already entered the record. A score that fails scope is not applicable. A score that lacks calibration may remain contextual. Dependent evidence should be capped or omitted.
- 18
Prediction narrows attention
Computational predictors can highlight a residue, suggest a splice mechanism, identify a constrained position and focus deeper review. They do not establish the disease relationship of a gene, prove that a molecular effect explains the patient or replace functional and segregation data. Even an accurate model can fail for a specific transcript, mechanism or gene. The final interpretation returns to the complete observed evidence and the clinical context.
- 19
Primary sources
Richards and colleagues defined the ACMG and AMP sequence variant interpretation framework. Pejaver and colleagues provided quantitative calibration recommendations for computational missense predictors. Walker and colleagues specified how predicted and observed splice effects should enter the ACMG framework, including protection against double counting. Jaganathan and colleagues described SpliceAI. Folklore documentation records the current operational thresholds and safeguards used by the platform.
- 20
Continue with the documentation
The Folklore documentation provides a page for each predictor. The BayesDel page records missense score intervals, evidence strengths and dependency guards. The SpliceAI page records the four delta scores, current production intervals and PVS1 safeguards. Separate pages cover SIFT, AlphaMissense, MetaSVM, DANN, PhyloP and GERP. Read the model, threshold, mechanism and ACMG role together rather than reading a prediction in isolation.