Deterministic Classification Performance Study
Cohort 4: Deterministic Classification Performance Study
Cohort 4 is the largest Folklore study to date: 100 of 100 real-world cases completed under locked classifier version v3.39.1. The primary analysis covers 367 evaluable outcomes from 395 reported outcomes, followed by session-level forensic review of 19 priority variants across 16 cases. The study separates deterministic classifier defects, reference-data gaps, manual evidence limitations and legitimate inter-laboratory differences, then converts confirmed findings into a defined correction and full-cohort regression programme.
Public scientific report
Cite this study
The archived report, persistent DOI and full metadata record are available through OSF.
Suggested citation
Mitev V. Cohort 4: Deterministic Classification Performance Study. Helena Bioinformatics public scientific report, 1.0, August 2026. doi:10.17605/OSF.IO/NPHW8.Study specification
Layer 1 - Cohort execution and evaluability
All planned real-world cases completed under the locked study version; no session failed.
The denominator includes outcomes for which a platform-versus-comparator classification assessment was methodologically supportable.
Reported separately and excluded from concordance fractions rather than silently treated as disagreement.
Layer 2 - Deterministic classification performance
Exact five-tier ACMG/AMP agreement.
Exact or clinically equivalent adjacent-category agreement.
Adjacent-boundary methodological differences, concentrated at Pathogenic/Likely Pathogenic/VUS and VUS/Likely Benign boundaries.
A small, reviewable set concentrated in risk-allele and mechanism-specific boundaries.
Each priority difference was traced from session DuckDB evidence through deterministic ACMG criteria and available reference data.
Layer 3 - Findings and correction programme
Disease-aware BS1 selection, dual-mechanism/PVS1 handling, BP4_Moderate strength handling, carrier/class separation, and a separate family-level architecture correction.
Some differences depend on evidence absent from the current automated reference layer; these are assigned to source-ingestion work, not mislabelled as combining-logic defects.
Functional, segregation, de-novo and case-specific literature evidence cannot always be derived from variant input alone and is not counted as a classifier defect.
Confirmed defects will be corrected, relevant reference sources will be added, and all 100 cases will be reprocessed in shadow mode against the locked baseline.
The cohort met its scientific purpose: it measured performance at scale, identified bounded defects and converted them into a verifiable improvement programme.
Principal findings and action programme
The forensic review assigns material differences to the source that can actually change them. This preserves a clean distinction between software defects, missing reference evidence, non-automatable clinical evidence and legitimate peer-method differences.
Deterministic classifier defects
Confirmed and boundedFour deterministic mechanisms require correction: disease-aware BS1 source selection, dual-mechanism PVS1 eligibility, BP4_Moderate strength handling, and separation of carrier context from variant class. A separate family-architecture finding is tracked with the same regression programme.
Action: Implement the corrections and prove each one with targeted regression fixtures plus a 100-case shadow re-run.
Reference-data coverage gaps
Source-ingestion programmeSome comparator evidence was not available in the automated reference layer at study time. The gap is attributable to missing source coverage rather than faulty ACMG combining logic.
Action: Add the authoritative sources to the reference-data lifecycle, preserve provenance and versioning, then re-evaluate affected outcomes.
Manual clinical evidence
Recognised automation boundarySegregation, de-novo status, functional assays and case-specific literature interpretation may require human evidence not present in a VCF or reference snapshot.
Action: Keep these outcomes in specialist review and do not count the absence of non-automatable evidence as a classifier defect.
Inter-laboratory classification differences
Expected and reviewableQualified laboratories can reach adjacent ACMG/AMP categories through different admissible evidence sets and strength assignments. CellGenetics manual calls are therefore evaluated, not accepted as ground truth by default.
Action: Retain both-defensible differences as methodological variation; correct Folklore only where the evidence trace demonstrates a platform defect or data gap.