Documentation / Reference Databases / SpliceAI Precomputed

SpliceAI precomputed dataset: fields, coverage, and provenance

Folklore reads a managed GRCh38 SpliceAI asset and exposes its donor and acceptor delta-score (DS) and delta-position (DP) fields with source, file-set, and version provenance. Coverage is limited to records represented in that asset. For score meaning and classifier thresholds, use the dedicated predictor guide.

Review one public GRCh38 variant against Folklore's current evidence snapshot. Public search accepts a variant expression only. Do not enter patient or case information.

Search a public variant

Database Details

CoverageManaged precomputed SNV and indel records available to the GRCh38 pipeline
Genome BuildGRCh38
ProvenanceSource files and asset version recorded with the managed pipeline snapshot
ProducerIllumina, Inc.

Four Delta Scores

SpliceAI produces four model delta scores, each ranging from 0 to 1, representing a predicted change in splice-site usage. The values are not calibrated probabilities that a splice event occurs:

DS_AGAcceptor Gain

The model predicts increased acceptor usage at a nearby position relative to the reference sequence.

DS_ALAcceptor Loss

The model predicts reduced usage of an existing acceptor relative to the reference sequence.

DS_DGDonor Gain

The model predicts increased donor usage at a nearby position relative to the reference sequence.

DS_DLDonor Loss

The model predicts reduced usage of an existing donor relative to the reference sequence.

The managed record also exposes the maximum of the four delta scores (max_score). The dedicated predictor guide explains how Folklore interprets that model output; this dataset page documents the asset and fields.

Delta Positions

The paired DP_AG, DP_AL, DP_DG, and DP_DL values report the predicted position of the corresponding splice change relative to the variant. A delta position helps a reviewer locate the predicted event. It does not identify an experimentally observed transcript or determine the functional consequence of that transcript.

Limitations

SpliceAI predictions are based on primary sequence context. Tissue-specific splicing regulation is not modeled.

A managed precomputed dataset has a defined genome build, file set, and coverage. A missing record means the prediction is unavailable in that dataset, not that the variant has no splice effect.

SpliceAI does not establish the transcript-level or functional consequence of a predicted splice change. Its scores are model outputs, not calibrated probabilities of disruption.

The asset does not establish an observed splice effect or its clinical relevance. Score interpretation, RNA evidence, and the full evidence context require qualified review.

References

Jaganathan K, et al. "Predicting splicing from primary sequence with deep learning." Cell. 2019;176(3):535-548. PMID: 30661751.

Walker LC, et al. "Using the ACMG/AMP framework to capture evidence related to predicted and observed impact on splicing." American Journal of Human Genetics. 2023;110(7):1046-1067. PMID: 37352859.

Continue with the SpliceAI predictor guide or review how computational evidence fits the ACMG/AMP framework.