Skip to contents

Internal function that runs Gene Set Variation Analysis (gsva) on expression data using the cached CTD chemical gene sets. Unlike ORA/GSEA/CAMERA — which return a single p-value per chemical (group-level inference) — GSVA produces per-sample enrichment scores: one row per chemical, one column per sample. These scores are suitable for downstream clustering, association tests against phenotypes, survival analysis, or heatmap visualization.

A SummarizedExperiment is handed to GSVA unchanged rather than reduced to its assay, because GSVA is itself SE-in/SE-out: the scores then come back in a container that still carries colData, which is what makes per-sample scores interpretable.

Usage

.run_gsva(
  expr,
  id_type = NULL,
  cache_dir,
  interaction_types = NULL,
  assay = NULL,
  ...
)

Arguments

expr

Numeric expression matrix (genes x samples) or a SummarizedExperiment. rownames(expr) must be Entrez IDs or HGNC symbols matching the cached CTD gene sets.

id_type

Either "entrez", "symbol", or NULL for auto-detection from rownames(expr).

cache_dir

Directory holding the cached CTD .rda files.

interaction_types

Character vector of CTD InteractionActions values to retain when building gene sets, or NULL for all.

assay

Assay name or index to use when expr is a SummarizedExperiment; NULL (default) takes the first. Ignored for a matrix.

...

Forwarded to gsvaParam (e.g. kcdf, minSize, maxSize, tau, maxDiff).

Value

GSVA enrichment scores with CTD chemical IDs in rows and samples in columns, in the same container as expr: a numeric matrix for a matrix input, a SummarizedExperiment (with colData preserved) for a SummarizedExperiment input.