ctdR is an R package that identifies chemicals significantly associated with a set of genes using data from the Comparative Toxicogenomics Database (CTD).
📖 Documentation, tutorials & function reference: luigicorsaro.com/ctdR — full package website with articles and examples.
⭐ Find ctdR useful? Star it on GitHub — it takes one second and helps other researchers discover the package.
✅ Bioconductor status: accepted into Bioconductor, currently available in Bioconductor devel (3.24). From the next Bioconductor release it will install from the standard repository (instructions below).
Features
Four enrichment methods through a unified enrichment_CTD() interface, selected via the method argument:
-
ORA: Over-Representation Analysis (hypergeometric test). Input: gene list. Backend:
stats::phyper, computed directly. -
GSEA — Gene Set Enrichment Analysis (rank-based, permutational). Input: ranked gene list. Backend:
fgsea::fgsea. -
CAMERA: competitive gene-set test with inter-gene correlation correction. Input: expression matrix or
SummarizedExperiment, plus design + contrast. Backend:limma::camera. -
GSVA: Gene Set Variation Analysis (per-sample scoring). Input: expression matrix or
SummarizedExperiment. Output: chemical × sample scores, returned in whichever of the two you supplied. Backend:GSVA::gsva.
Additional features:
-
Provenance tracking:
ctd_provenance()reports which CTD release an analysis ran on, read from theReport createdline of the file’s own header. CTD re-releases continuously and does not version its filenames, so this is what lets a result name the data behind it. Attached to every result, and readable from the cache before any analysis -
Size thresholds chosen for CTD, not inherited:
minGSSizedefaults to 2 because the median chemical has 4 target genes, and there is no upper limit because no CTD gene set is large enough to be untestable. Chemicals the filter does leave out are reported rather than dropped in silence - Automatic caching — CTD data is parsed once and cached locally for fast repeated analyses
- Human-only filtering — automatically restricts interactions to Homo sapiens (OrganismID 9606)
-
Auto-detected gene identifiers: Entrez vs HGNC SYMBOL detected from
rownames()of the input matrix orSummarizedExperiment(CAMERA / GSVA); override viaid_type -
Visualization —
plot_CTD()auto-dispatches to bar/dot plot (ORA/GSEA/CAMERA) or heatmap (GSVA)
Data Licensing Disclaimer
This package does NOT bundle, redistribute, or embed any data from the Comparative Toxicogenomics Database.
CTD data are created and maintained by NC State University and are subject to specific licensing terms and conditions. Users are solely responsible for:
- Downloading the required data files directly from https://ctdbase.org
- Reading and accepting the CTD Terms of Service before using the data
- Complying with all applicable CTD data licensing requirements, including proper citation of CTD in any publications or derived works
By using ctdR with CTD data, you acknowledge that you have read, understood, and agreed to the CTD Terms of Service.
Installation
Status: ctdR has been accepted into Bioconductor (package page). For now it is available only in Bioconductor devel (3.24): install it from the devel repository as shown below. After the next Bioconductor release it will be installable from the standard (release) repository.
From Bioconductor devel (current installation method)
Start R (version “4.6”) and enter:
if (!require("BiocManager", quietly = TRUE))
install.packages("BiocManager")
## The following initializes the development version of Bioconductor
BiocManager::install(version = "devel")
BiocManager::install("ctdR")Bioconductor dependencies are installed automatically.
From Bioconductor release (after the next Bioconductor release)
Once ctdR is part of a Bioconductor release, the standard installation is enough, with no need to switch to devel:
if (!require("BiocManager", quietly = TRUE))
install.packages("BiocManager")
BiocManager::install("ctdR")Quick Start
Step 1 — Download CTD data (once)
Download CTD_chem_gene_ixns.csv.gz from the CTD website:
https://ctdbase.org/reports/CTD_chem_gene_ixns.csv.gz
Decompress the file:
Step 2 — Import into ctdR (once)
library(ctdR)
import_CTD("~/Downloads/CTD_chem_gene_ixns.csv")
#> Reading CTD chemical-gene interactions from: ~/Downloads/CTD_chem_gene_ixns.csv
#> Filtered to 1234567 human interactions
#> Mapping genes for 12345 chemicals (this may take a while)...
#> CTD data cached successfully in: /Users/you/Library/Caches/ctdRThe data is now cached locally. You only need to do this once (or again when you download a newer CTD release).
import_CTD() also accepts a URL in place of a local path; remote files are downloaded and cached via BiocFileCache, with a one-time reminder pointing to the CTD data-licensing terms. The package assumes no default URL, so you always supply the source yourself.
Step 3 — Run enrichment analysis
The first argument of enrichment_CTD() is polymorphic and named x: a data.frame for ORA / GSEA, and for CAMERA / GSVA either a numeric matrix or a SummarizedExperiment.
ORA / GSEA — gene-list paradigm
# Your gene list: a data frame with an EntrezID column and a numeric one
genes <- data.frame(
EntrezID = c("7124", "3569", "7157", "672", "1956"),
pvalue = c(0.001, 0.003, 0.01, 0.02, 0.05)
)
# Over-Representation Analysis (default)
ora_results <- enrichment_CTD(genes, method = "ORA")
head(ora_results)
# Better, when you have the whole differential expression table: pass it
# with a threshold and say which p-value to judge on. The genes below the
# threshold are tested against every gene in the table as background, so
# the list and the background cannot disagree. Left to itself the
# background falls back to all CTD genes, which is wider than anything
# your experiment could detect and makes p-values too small.
ora_results <- enrichment_CTD(de_table, method = "ORA",
alpha = 0.05, alpha_column = "padj")
# Gene Set Enrichment Analysis (uses the second column as ranking)
gsea_results <- enrichment_CTD(genes, method = "GSEA")
head(gsea_results)
# Customize multiple testing correction (default is "BH")
ora_bonf <- enrichment_CTD(genes, method = "ORA", pAdjustMethod = "bonferroni")CAMERA — multi-sample paradigm with correlation correction
# Suppose `se` is a SummarizedExperiment whose assay holds normalised
# expression (genes x samples) with Entrez IDs (or HGNC SYMBOLs) as
# rownames, and whose colData carries the experimental group.
# A plain numeric matrix works too; the SummarizedExperiment keeps the
# sample annotation attached to the data.
design <- model.matrix(~ se$group)
camera_results <- enrichment_CTD(
se,
method = "CAMERA",
design = design,
contrast = 2 # last column of `design` = treat vs ctrl
)
head(camera_results)GSVA — per-sample scoring
# SummarizedExperiment in, SummarizedExperiment out: the scores arrive
# with the sample annotation still attached. Pass a matrix instead and
# you get a plain matrix back.
gsva_scores <- enrichment_CTD(se, method = "GSVA")
dim(gsva_scores)
SummarizedExperiment::colData(gsva_scores)Utilities and interoperability
# Read a CTD file into a DataFrame via the Bioconductor import() convention
library(BiocIO)
ctd <- import(CTDFile("~/Downloads/CTD_chem_gene_ixns.csv"))
# Retrieve processed tables from the cache
chem <- ctd_cache("chemicals")
# Export the CTD gene sets for a third-party engine (e.g. EnrichmentBrowser)
gene_sets <- as_genesets_CTD("entrez")In ctdR, CTD always means the Comparative Toxicogenomics Database. It is unrelated to the CRAN package CTD (“Connect The Dots”) and to the ctd (CellTypeDataset) object used by EWCE.
End-to-end runnable example
For a production-shaped reference pipeline — full GSE311566 Female PBMCs dataset (Dex vs DMSO), a-priori power analysis, declared significance thresholds, limma differential expression, and all four enrichment methods with BH-adjusted alpha cutoffs — see:
inst/scripts/example_gse311566_full_pipeline.R
Run it from the package root after installing ctdR and populating the CTD cache (import_CTD()):
Outputs (DE table, per-method significant chemicals, plots, power report, sessionInfo) land in ./example_outputs/. This script is the example referenced by the companion paper; the vignette covers the same workflow on a bundled subset (~34 KB, no alpha cutoffs) for didactic purposes and BiocCheck-friendly build times.
Input Format
The first argument x of enrichment_CTD() depends on the method:
ORA / GSEA — data frame
| Column | Description |
|---|---|
EntrezID |
Character or numeric Entrez gene IDs |
| (a numeric column) | A value per gene. GSEA ranks on the second column, or on stat when present. ORA ignores it unless you pass alpha, which thresholds the column named by alpha_column. |
CAMERA / GSVA: numeric matrix or SummarizedExperiment
A genes × samples numeric matrix with rownames(x) set to either Entrez IDs or HGNC SYMBOLs, or a SummarizedExperiment wrapping one. Identifier type is auto-detected from rownames(x); override with the id_type argument if needed.
Prefer the SummarizedExperiment when you have sample annotation. It keeps the assay and the annotation in one object, so subsetting or reordering samples moves both together instead of leaving you to keep a separate group vector in step with the columns. Use assay to pick which assay to read when the object carries more than one; the default is the first.
Output
ORA results (data.frame)
| Column | Description |
|---|---|
ChemicalID |
CTD chemical identifier (e.g. "D000082") |
ChemicalName |
Human-readable chemical name |
GeneRatio |
Proportion of input genes in the chemical’s set |
BgRatio |
Background ratio |
pvalue |
Raw p-value |
padj |
Adjusted p-value (BH method) |
foldEnrichment |
GeneRatio / BgRatio |
geneID |
Enriched gene symbols |
Count |
Number of overlapping genes |
GSEA results (data.frame)
| Column | Description |
|---|---|
ChemicalID |
CTD chemical identifier |
ChemicalName |
Human-readable chemical name |
pval |
Raw p-value |
padj |
Adjusted p-value |
ES |
Enrichment score |
NES |
Normalized enrichment score |
size |
Size of the gene set |
leadingEdge |
Leading-edge gene subset |
foldEnrichment |
abs(ES) / mean(ES) |
Enriched_GENE |
Comma-separated enriched gene symbols |
CAMERA results (data.frame)
| Column | Description |
|---|---|
ChemicalID |
CTD chemical identifier |
ChemicalName |
Human-readable chemical name |
NGenes |
Number of gene-set genes matched in rownames(x)
|
Direction |
"Up" or "Down" — direction of mean effect |
Correlation |
Estimated (or pre-specified) inter-gene correlation |
pvalue |
Raw p-value from limma::camera()
|
padj |
Adjusted p-value |
GSVA results (matrix or SummarizedExperiment)
CTD chemical IDs in rows and samples in columns, containing per-sample enrichment scores (typically in [-1, 1]). The container follows the input: a matrix in returns a numeric matrix, a SummarizedExperiment in returns a SummarizedExperiment whose assay holds the scores and whose colData is carried over, so the annotation needed to interpret the scores stays next to them. Unlike the other methods, GSVA does not return p-values — the scores are descriptive features for downstream analyses (clustering, association with outcomes, heatmap visualization).
Dependencies
Mirrors the Imports field of DESCRIPTION.
Bioconductor
- fgsea — fast GSEA implementation
-
limma — CAMERA backend (
limma::camera()) - GSVA — per-sample gene-set scoring
- SummarizedExperiment: standard container accepted by CAMERA / GSVA
-
BiocIO:
import()convention behindCTDFile - BiocFileCache: local cache of the processed CTD tables
-
S4Vectors:
DataFramereturned byimport() - AnnotationDbi — gene ID mapping
- org.Hs.eg.db — human gene annotation
Continuous Integration & Security
ctdR uses GitHub Actions for continuous integration and automated security auditing:
| Workflow | Purpose |
|---|---|
| R-CMD-check | Package build & check on macOS, Ubuntu, Windows |
| test-coverage | Code coverage via covr + Codecov |
| security-scan | Automated cybersecurity pipeline (see below) |
The security-scan workflow runs on every push/PR and weekly, and includes four jobs:
-
Dependency vulnerability audit — scans all installed R packages against the Sonatype OSS Index via
oysteR, flagging packages with known CVEs. -
Static code analysis — runs
lintron the entire package source to detect code quality and potential security issues. - Dependency review (PR only) — uses GitHub’s dependency-review-action to flag new dependencies with high-severity vulnerabilities before merging.
- Secret & credential scan — uses TruffleHog to detect accidentally committed secrets or API keys in the repository history.
Contributing
Contributions are welcome! Please open an issue or submit a pull request.
Citation
If you use ctdR in your research, please cite:
Corsaro L (2026). ctdR: Enrichment Analysis of Chemical-Gene Interactions
from the Comparative Toxicogenomics Database. R package version 0.99.2.
doi: 10.5281/zenodo.19344200. https://github.com/drake69/ctdR
Additionally, please cite CTD as required by their Terms of Service:
Davis AP, Wiegers TC, Johnson RJ, Sciaky D, Wiegers J, Mattingly CJ. Comparative Toxicogenomics Database (CTD): update 2023. Nucleic Acids Research. 2023;51(D1):D1257-D1262.
⭐ If you find ctdR useful, please star the repo
A star helps other researchers discover the package and lets us see who’s using it. Click the ⭐ at the top of the GitHub page — it takes one second and means a lot.
License
Apache License 2.0 - see LICENSE for details.