Skip to contents

ctdR is an R package that identifies chemicals significantly associated with a set of genes using data from the Comparative Toxicogenomics Database (CTD).

📖 Documentation, tutorials & function reference: luigicorsaro.com/ctdR — full package website with articles and examples.

⭐ Find ctdR useful? Star it on GitHub — it takes one second and helps other researchers discover the package.

✅ Bioconductor status: accepted into Bioconductor, currently available in Bioconductor devel (3.24). From the next Bioconductor release it will install from the standard repository (instructions below).

Features

Four enrichment methods through a unified enrichment_CTD() interface, selected via the method argument:

  • ORA: Over-Representation Analysis (hypergeometric test). Input: gene list. Backend: stats::phyper, computed directly.
  • GSEA — Gene Set Enrichment Analysis (rank-based, permutational). Input: ranked gene list. Backend: fgsea::fgsea.
  • CAMERA: competitive gene-set test with inter-gene correlation correction. Input: expression matrix or SummarizedExperiment, plus design + contrast. Backend: limma::camera.
  • GSVA: Gene Set Variation Analysis (per-sample scoring). Input: expression matrix or SummarizedExperiment. Output: chemical × sample scores, returned in whichever of the two you supplied. Backend: GSVA::gsva.

Additional features:

  • Provenance tracking: ctd_provenance() reports which CTD release an analysis ran on, read from the Report created line of the file’s own header. CTD re-releases continuously and does not version its filenames, so this is what lets a result name the data behind it. Attached to every result, and readable from the cache before any analysis
  • Size thresholds chosen for CTD, not inherited: minGSSize defaults to 2 because the median chemical has 4 target genes, and there is no upper limit because no CTD gene set is large enough to be untestable. Chemicals the filter does leave out are reported rather than dropped in silence
  • Automatic caching — CTD data is parsed once and cached locally for fast repeated analyses
  • Human-only filtering — automatically restricts interactions to Homo sapiens (OrganismID 9606)
  • Auto-detected gene identifiers: Entrez vs HGNC SYMBOL detected from rownames() of the input matrix or SummarizedExperiment (CAMERA / GSVA); override via id_type
  • Visualization — plot_CTD() auto-dispatches to bar/dot plot (ORA/GSEA/CAMERA) or heatmap (GSVA)

Data Licensing Disclaimer

This package does NOT bundle, redistribute, or embed any data from the Comparative Toxicogenomics Database.

CTD data are created and maintained by NC State University and are subject to specific licensing terms and conditions. Users are solely responsible for:

  1. Downloading the required data files directly from https://ctdbase.org
  2. Reading and accepting the CTD Terms of Service before using the data
  3. Complying with all applicable CTD data licensing requirements, including proper citation of CTD in any publications or derived works

By using ctdR with CTD data, you acknowledge that you have read, understood, and agreed to the CTD Terms of Service.

Installation

Status: ctdR has been accepted into Bioconductor (package page). For now it is available only in Bioconductor devel (3.24): install it from the devel repository as shown below. After the next Bioconductor release it will be installable from the standard (release) repository.

From Bioconductor devel (current installation method)

Start R (version “4.6”) and enter:

if (!require("BiocManager", quietly = TRUE))
    install.packages("BiocManager")

## The following initializes the development version of Bioconductor
BiocManager::install(version = "devel")
BiocManager::install("ctdR")

Bioconductor dependencies are installed automatically.

From Bioconductor release (after the next Bioconductor release)

Once ctdR is part of a Bioconductor release, the standard installation is enough, with no need to switch to devel:

if (!require("BiocManager", quietly = TRUE))
    install.packages("BiocManager")

BiocManager::install("ctdR")

From GitHub (latest development version)

# install.packages("devtools")
devtools::install_github("drake69/ctdR")

Quick Start

Step 1 — Download CTD data (once)

Download CTD_chem_gene_ixns.csv.gz from the CTD website:

https://ctdbase.org/reports/CTD_chem_gene_ixns.csv.gz

Decompress the file:

gunzip CTD_chem_gene_ixns.csv.gz

Step 2 — Import into ctdR (once)

library(ctdR)

import_CTD("~/Downloads/CTD_chem_gene_ixns.csv")
#> Reading CTD chemical-gene interactions from: ~/Downloads/CTD_chem_gene_ixns.csv
#> Filtered to 1234567 human interactions
#> Mapping genes for 12345 chemicals (this may take a while)...
#> CTD data cached successfully in: /Users/you/Library/Caches/ctdR

The data is now cached locally. You only need to do this once (or again when you download a newer CTD release).

import_CTD() also accepts a URL in place of a local path; remote files are downloaded and cached via BiocFileCache, with a one-time reminder pointing to the CTD data-licensing terms. The package assumes no default URL, so you always supply the source yourself.

Step 3 — Run enrichment analysis

The first argument of enrichment_CTD() is polymorphic and named x: a data.frame for ORA / GSEA, and for CAMERA / GSVA either a numeric matrix or a SummarizedExperiment.

ORA / GSEA — gene-list paradigm

# Your gene list: a data frame with an EntrezID column and a numeric one
genes <- data.frame(
  EntrezID = c("7124", "3569", "7157", "672", "1956"),
  pvalue   = c(0.001, 0.003, 0.01, 0.02, 0.05)
)

# Over-Representation Analysis (default)
ora_results <- enrichment_CTD(genes, method = "ORA")
head(ora_results)

# Better, when you have the whole differential expression table: pass it
# with a threshold and say which p-value to judge on. The genes below the
# threshold are tested against every gene in the table as background, so
# the list and the background cannot disagree. Left to itself the
# background falls back to all CTD genes, which is wider than anything
# your experiment could detect and makes p-values too small.
ora_results <- enrichment_CTD(de_table, method = "ORA",
                              alpha = 0.05, alpha_column = "padj")

# Gene Set Enrichment Analysis (uses the second column as ranking)
gsea_results <- enrichment_CTD(genes, method = "GSEA")
head(gsea_results)

# Customize multiple testing correction (default is "BH")
ora_bonf <- enrichment_CTD(genes, method = "ORA", pAdjustMethod = "bonferroni")

CAMERA — multi-sample paradigm with correlation correction

# Suppose `se` is a SummarizedExperiment whose assay holds normalised
# expression (genes x samples) with Entrez IDs (or HGNC SYMBOLs) as
# rownames, and whose colData carries the experimental group.
# A plain numeric matrix works too; the SummarizedExperiment keeps the
# sample annotation attached to the data.
design <- model.matrix(~ se$group)

camera_results <- enrichment_CTD(
  se,
  method   = "CAMERA",
  design   = design,
  contrast = 2  # last column of `design` = treat vs ctrl
)
head(camera_results)

GSVA — per-sample scoring

# SummarizedExperiment in, SummarizedExperiment out: the scores arrive
# with the sample annotation still attached. Pass a matrix instead and
# you get a plain matrix back.
gsva_scores <- enrichment_CTD(se, method = "GSVA")
dim(gsva_scores)
SummarizedExperiment::colData(gsva_scores)

Step 4 — Visualize results

# ORA / GSEA / CAMERA: bar or dot plot of top chemicals
plot_CTD(ora_results,    type = "bar")
plot_CTD(gsea_results,   type = "dot", n = 10)
plot_CTD(camera_results, type = "bar")  # colour encodes Direction

# GSVA: heatmap of the top-variance chemicals across samples
plot_CTD(gsva_scores)

Utilities and interoperability

# Read a CTD file into a DataFrame via the Bioconductor import() convention
library(BiocIO)
ctd <- import(CTDFile("~/Downloads/CTD_chem_gene_ixns.csv"))

# Retrieve processed tables from the cache
chem <- ctd_cache("chemicals")

# Export the CTD gene sets for a third-party engine (e.g. EnrichmentBrowser)
gene_sets <- as_genesets_CTD("entrez")

In ctdR, CTD always means the Comparative Toxicogenomics Database. It is unrelated to the CRAN package CTD (“Connect The Dots”) and to the ctd (CellTypeDataset) object used by EWCE.

End-to-end runnable example

For a production-shaped reference pipeline — full GSE311566 Female PBMCs dataset (Dex vs DMSO), a-priori power analysis, declared significance thresholds, limma differential expression, and all four enrichment methods with BH-adjusted alpha cutoffs — see:

inst/scripts/example_gse311566_full_pipeline.R

Run it from the package root after installing ctdR and populating the CTD cache (import_CTD()):

Rscript inst/scripts/example_gse311566_full_pipeline.R

Outputs (DE table, per-method significant chemicals, plots, power report, sessionInfo) land in ./example_outputs/. This script is the example referenced by the companion paper; the vignette covers the same workflow on a bundled subset (~34 KB, no alpha cutoffs) for didactic purposes and BiocCheck-friendly build times.

Input Format

The first argument x of enrichment_CTD() depends on the method:

ORA / GSEA — data frame

Column Description
EntrezID Character or numeric Entrez gene IDs
(a numeric column) A value per gene. GSEA ranks on the second column, or on stat when present. ORA ignores it unless you pass alpha, which thresholds the column named by alpha_column.

CAMERA / GSVA: numeric matrix or SummarizedExperiment

A genes × samples numeric matrix with rownames(x) set to either Entrez IDs or HGNC SYMBOLs, or a SummarizedExperiment wrapping one. Identifier type is auto-detected from rownames(x); override with the id_type argument if needed.

Prefer the SummarizedExperiment when you have sample annotation. It keeps the assay and the annotation in one object, so subsetting or reordering samples moves both together instead of leaving you to keep a separate group vector in step with the columns. Use assay to pick which assay to read when the object carries more than one; the default is the first.

Output

ORA results (data.frame)

Column Description
ChemicalID CTD chemical identifier (e.g. "D000082")
ChemicalName Human-readable chemical name
GeneRatio Proportion of input genes in the chemical’s set
BgRatio Background ratio
pvalue Raw p-value
padj Adjusted p-value (BH method)
foldEnrichment GeneRatio / BgRatio
geneID Enriched gene symbols
Count Number of overlapping genes

GSEA results (data.frame)

Column Description
ChemicalID CTD chemical identifier
ChemicalName Human-readable chemical name
pval Raw p-value
padj Adjusted p-value
ES Enrichment score
NES Normalized enrichment score
size Size of the gene set
leadingEdge Leading-edge gene subset
foldEnrichment abs(ES) / mean(ES)
Enriched_GENE Comma-separated enriched gene symbols

CAMERA results (data.frame)

Column Description
ChemicalID CTD chemical identifier
ChemicalName Human-readable chemical name
NGenes Number of gene-set genes matched in rownames(x)
Direction "Up" or "Down" — direction of mean effect
Correlation Estimated (or pre-specified) inter-gene correlation
pvalue Raw p-value from limma::camera()
padj Adjusted p-value

GSVA results (matrix or SummarizedExperiment)

CTD chemical IDs in rows and samples in columns, containing per-sample enrichment scores (typically in [-1, 1]). The container follows the input: a matrix in returns a numeric matrix, a SummarizedExperiment in returns a SummarizedExperiment whose assay holds the scores and whose colData is carried over, so the annotation needed to interpret the scores stays next to them. Unlike the other methods, GSVA does not return p-values — the scores are descriptive features for downstream analyses (clustering, association with outcomes, heatmap visualization).

Dependencies

Mirrors the Imports field of DESCRIPTION.

CRAN

  • ggplot2 — publication-quality plots
  • readr — fast CSV reading

Bioconductor

Continuous Integration & Security

ctdR uses GitHub Actions for continuous integration and automated security auditing:

Workflow Purpose
R-CMD-check Package build & check on macOS, Ubuntu, Windows
test-coverage Code coverage via covr + Codecov
security-scan Automated cybersecurity pipeline (see below)

The security-scan workflow runs on every push/PR and weekly, and includes four jobs:

  1. Dependency vulnerability audit — scans all installed R packages against the Sonatype OSS Index via oysteR, flagging packages with known CVEs.
  2. Static code analysis — runs lintr on the entire package source to detect code quality and potential security issues.
  3. Dependency review (PR only) — uses GitHub’s dependency-review-action to flag new dependencies with high-severity vulnerabilities before merging.
  4. Secret & credential scan — uses TruffleHog to detect accidentally committed secrets or API keys in the repository history.

Contributing

Contributions are welcome! Please open an issue or submit a pull request.

Citation

If you use ctdR in your research, please cite:

Corsaro L (2026). ctdR: Enrichment Analysis of Chemical-Gene Interactions
from the Comparative Toxicogenomics Database. R package version 0.99.2.
doi: 10.5281/zenodo.19344200. https://github.com/drake69/ctdR

Additionally, please cite CTD as required by their Terms of Service:

Davis AP, Wiegers TC, Johnson RJ, Sciaky D, Wiegers J, Mattingly CJ. Comparative Toxicogenomics Database (CTD): update 2023. Nucleic Acids Research. 2023;51(D1):D1257-D1262.

⭐ If you find ctdR useful, please star the repo

A star helps other researchers discover the package and lets us see who’s using it. Click the ⭐ at the top of the GitHub page — it takes one second and means a lot.

License

Apache License 2.0 - see LICENSE for details.