Helena
Back to Studies

Deterministic Classification Performance Study

Cohort 4: Deterministic Classification Performance Study

Cohort 4 is the largest Folklore study to date: 100 of 100 real-world cases completed under locked classifier version v3.39.1. The primary analysis covers 367 evaluable outcomes from 395 reported outcomes, followed by session-level forensic review of 19 priority variants across 16 cases. The study separates deterministic classifier defects, reference-data gaps, manual evidence limitations and legitimate inter-laboratory differences, then converts confirmed findings into a defined correction and full-cohort regression programme.

Public scientific report

Cite this study

The archived report, persistent DOI and full metadata record are available through OSF.

Suggested citation

Mitev V. Cohort 4: Deterministic Classification Performance Study. Helena Bioinformatics public scientific report, 1.0, August 2026. doi:10.17605/OSF.IO/NPHW8.

Study specification

DocumentHEL-RWP-C4-2026-001-PUBLIC v1.0
Study typeLarge-scale fixed-version real-world performance study under HEL-RWP-FRAMEWORK-v1.0
PeriodCohort closed July 2026; forensic review completed August 2026
Classifier versionFolklore v3.39.1 at study closure
Execution100/100 cases completed; no failed session
Primary denominator367 evaluable outcomes from 395 reported outcomes (92.9%)
Forensic review set19 priority variants across 16 cases
Comparator roleCellGenetics peer comparison, evaluated against multi-source evidence rather than treated as single-source ground truth
PrivacyAggregate results only; no patient, case, session, batch or variant identifiers are published
L1

Layer 1 - Cohort execution and evaluability

Completed cases100/100 (100%)

All planned real-world cases completed under the locked study version; no session failed.

Evaluable outcomes367/395 (92.9%)

The denominator includes outcomes for which a platform-versus-comparator classification assessment was methodologically supportable.

Not evaluable28 outcomes

Reported separately and excluded from concordance fractions rather than silently treated as disagreement.

L2

Layer 2 - Deterministic classification performance

FULL concordance223/367 (60.8%; Wilson 95% CI 55.6-65.6%)

Exact five-tier ACMG/AMP agreement.

FULL + CLINICAL246/367 (67.0%; Wilson 95% CI 62.1-71.6%)

Exact or clinically equivalent adjacent-category agreement.

PARTIAL outcomes113/367 (30.8%)

Adjacent-boundary methodological differences, concentrated at Pathogenic/Likely Pathogenic/VUS and VUS/Likely Benign boundaries.

DISCORDANT outcomes8/367 (2.2%)

A small, reviewable set concentrated in risk-allele and mechanism-specific boundaries.

Forensic deep dive19 variants across 16 cases

Each priority difference was traced from session DuckDB evidence through deterministic ACMG criteria and available reference data.

L3

Layer 3 - Findings and correction programme

Confirmed deterministic defects4 mechanisms plus 1 family-architecture finding

Disease-aware BS1 selection, dual-mechanism/PVS1 handling, BP4_Moderate strength handling, carrier/class separation, and a separate family-level architecture correction.

Reference-data gapsConfirmed as a distinct source category

Some differences depend on evidence absent from the current automated reference layer; these are assigned to source-ingestion work, not mislabelled as combining-logic defects.

Manual evidence limitationsExpected scope boundary

Functional, segregation, de-novo and case-specific literature evidence cannot always be derived from variant input alone and is not counted as a classifier defect.

Correction and regressionDefined and cohort-wide

Confirmed defects will be corrected, relevant reference sources will be added, and all 100 cases will be reprocessed in shadow mode against the locked baseline.

Study conclusionSuccessful

The cohort met its scientific purpose: it measured performance at scale, identified bounded defects and converted them into a verifiable improvement programme.

Principal findings and action programme

The forensic review assigns material differences to the source that can actually change them. This preserves a clean distinction between software defects, missing reference evidence, non-automatable clinical evidence and legitimate peer-method differences.

Deterministic classifier defects

Confirmed and bounded

Four deterministic mechanisms require correction: disease-aware BS1 source selection, dual-mechanism PVS1 eligibility, BP4_Moderate strength handling, and separation of carrier context from variant class. A separate family-architecture finding is tracked with the same regression programme.

Action: Implement the corrections and prove each one with targeted regression fixtures plus a 100-case shadow re-run.

Reference-data coverage gaps

Source-ingestion programme

Some comparator evidence was not available in the automated reference layer at study time. The gap is attributable to missing source coverage rather than faulty ACMG combining logic.

Action: Add the authoritative sources to the reference-data lifecycle, preserve provenance and versioning, then re-evaluate affected outcomes.

Manual clinical evidence

Recognised automation boundary

Segregation, de-novo status, functional assays and case-specific literature interpretation may require human evidence not present in a VCF or reference snapshot.

Action: Keep these outcomes in specialist review and do not count the absence of non-automatable evidence as a classifier defect.

Inter-laboratory classification differences

Expected and reviewable

Qualified laboratories can reach adjacent ACMG/AMP categories through different admissible evidence sets and strength assignments. CellGenetics manual calls are therefore evaluated, not accepted as ground truth by default.

Action: Retain both-defensible differences as methodological variation; correct Folklore only where the evidence trace demonstrates a platform defect or data gap.