Skip to contents

Returns the cached CTD chemical gene sets as a named list, in the shape expected by third-party Bioconductor enrichment engines — for example the gs argument of EnrichmentBrowser::sbea(), or the input to GSEABase::GeneSetCollection(). This lets you run the CTD gene sets through existing infrastructure instead of, or alongside, enrichment_CTD.

Usage

as_genesets_CTD(
  id_type = c("entrez", "symbol"),
  interaction_types = NULL,
  cache_dir = .ctd_cache_dir()
)

Arguments

id_type

Either "entrez" (default) or "symbol": the identifier space of the returned genes, to match your expression data or downstream engine.

interaction_types

Optional character vector of CTD interaction actions (e.g. "increases^expression") to retain. NULL (default) keeps all interactions.

cache_dir

Directory holding the cached CTD files. Defaults to the ctdR user cache populated by import_CTD.

Value

A named list: names are CTD ChemicalIDs, elements are character vectors of gene identifiers (Entrez IDs or HGNC symbols).

Details

ctdR builds chemical-centric gene sets: each CTD chemical becomes a gene set of its interacting genes. enrichment_CTD consumes these internally; this helper exposes the same gene sets so they integrate with the wider ecosystem. Call import_CTD once first to populate the cache.

The returned representation is deliberately a plain named list (not a bespoke class): it is accepted directly by EnrichmentBrowser::sbea() and is a one-line conversion away from a GeneSetCollection. ctdR does not take a dependency on those engines — you supply them.

See also

import_CTD to populate the cache; enrichment_CTD for the built-in enrichment interface.

Examples

# Examples write to a temporary cache, so running them cannot
# disturb CTD data you have already imported. Set the same option
# yourself to keep an analysis isolated from your main cache.
options(ctdR.cache = tempfile())

sample_file <- system.file(
    "extdata", "CTD_chem_gene_ixns_sample.csv",
    package = "ctdR"
)
import_CTD(sample_file)
#> Reading CTD chemical-gene interactions from: /home/runner/work/_temp/Library/ctdR/extdata/CTD_chem_gene_ixns_sample.csv
#> Filtered to 86 human interactions
#> Mapping genes for 10 chemicals...
#> Warning: 10 ChemicalID(s) appear with more than one ChemicalName in the CTD file; only the first name per ID is retained. Affected IDs: D000082, D001564, D002104, D003907, D004958 ... (and 5 more)
#> CTD data cached successfully in: /tmp/Rtmpt9pODL/file1b13732f5ff8
#>   10 chemicals | 17 unique genes | 0 s
#>   CTD release: Mon Jan 01 00:00:00 EST 2024
gene_sets <- as_genesets_CTD("entrez")
length(gene_sets)
#> [1] 10
gene_sets[[1]]
#>  [1] "7124" "3569" "7157" "672"  "1956" "4609" "3845" "207"  "5290" "3553"
#> [11] "7422" "6774"

if (FALSE) { # \dontrun{
# Hand the CTD gene sets to an existing engine (EnrichmentBrowser):
library(EnrichmentBrowser)
res <- sbea(method = "camera", se = my_se, gs = as_genesets_CTD("entrez"))
gsRanking(res)
} # }