Helena

Documentation / Cohort Analysis

Cohort Analysis (Studies)

A Study is Folklore's workspace for analyzing a group of samples together. It preserves the individual Variant Analysis result for each sample while adding cohort-level quality, carrier, frequency, and gene evidence.

Studies are living cohorts. Samples can be added, removed, or retried, and the cohort matrix and downstream results should be regenerated when membership changes.

Individual Classification Remains Authoritative

Cohort Analysis does not replace or silently rewrite a sample's Variant Analysis result. The cohort layer aggregates classifications and genotypes for comparison. Where classifications differ across samples, Folklore records the most clinically severe observed class for cohort display and marks the variant as discordant for review.

Required Study Context

A defined sequencing type: WES, WGS, or CES.

A screening panel or explicit gene list for the analysis in scope.

Unique sample identifiers and compatible VCF or completed-analysis inputs.

Comparable reference build, calling approach, and analysis provenance across samples.

Production Workflow

1

Create the Study

Name the Study, select the sequencing type, and define the clinical profile or screening panel that determines the genes in scope.

2

Add and Process Samples

Upload VCF files, import from connected storage, reuse completed analyses, or start from a batch manifest. Each sample is classified through Variant Analysis before cohort aggregation unless a compatible completed analysis is reused.

3

Run Cohort Quality Control

Folklore reports per-sample variant count, Ti/Tv ratio, heterozygous-to-homozygous ratio, and mean depth, then flags cohort-relative outliers.

4

Build the Cohort Matrix

A deduplicated variant catalog and sparse sample-genotype matrix provide carrier counts, homozygous counts, cohort allele frequency, and classification-discordance flags.

5

Generate Study Results

The current workflow produces actionable findings, burden results, pLoF summaries, frequency comparisons, compound-heterozygous candidates, and ranked candidate genes when the required evidence is available.

Current Outputs

Per-sample processing and quality-control status.

A searchable cohort variant matrix with sample-level genotypes and carrier summaries.

Pathogenic and Likely Pathogenic findings mapped to the samples that carry them.

Gene-level rare-variant burden results with multiple-testing correction and power context.

Predicted loss-of-function and candidate compound-heterozygous summaries.

A ranked candidate-gene view that keeps its contributing evidence components visible.

Current Scope

The production Study workflow requires a gene panel or explicit gene list. GWAS and polygenic risk scoring are not current Study outputs. SKAT-O is not currently reported; the burden view uses the implemented carrier-based tests described in Results and Interpretation.

Pathway enrichment appears only when pathway definitions and a completed result are available. The current automated workflow does not load pathway definitions by default, so this section can be absent.

Clinical Boundary

Cohort associations and rankings are decision-support evidence, not diagnoses or proof of causality. Study design, ancestry, relatedness, technical batch, phenotype definition, and sample size remain essential to interpretation.

In This Section