Skip to contents

Returns the record of which CTD release produced a result: the Report created date CTD stamps into its own file header, where the file came from, when it was imported, and how much of it was kept.

An analysis is only reproducible if the version of the data behind it can be named. CTD is re-released continuously and its downloads are not versioned in the filename, so the release date inside the header is the only thing that identifies which snapshot a result came from. This carries it from the file all the way to the object you report on.

Usage

ctd_provenance(x)

# S3 method for class 'ctd_provenance'
print(x, ...)

Arguments

x

An object returned by enrichment_CTD, or the DataFrame returned by importing a CTDFile. Omit it to read the record of the data currently cached, which answers "which release am I about to analyse" before any analysis has been run.

...

Ignored, present for compatibility with the generic.

Value

An object of class ctd_provenance: a list with report_created (the CTD release string, NA if the file carried none), source, accessed, n_chemicals, n_interactions and ctdR_version. Returns NULL, with a warning, when x carries no record.

Details

Where the record is stored depends on what the object is, because the methods do not all return the same container. Objects that provide a metadata() slot keep it there, which covers the SummarizedExperiment returned by GSVA and the DataFrame returned by importing a CTDFile. The data frames returned by ORA, GSEA and CAMERA, and a plain score matrix, have no such slot, so there it rides on an attribute. Use this accessor rather than reaching for either directly, so that code keeps working whichever method produced the object.

The record survives subsetting, ordering, head() and the common dplyr verbs. It does not survive merge() or subset(), which drop attributes; retrieve it before those if you need it afterwards.

See also

import_CTD, which reads the record, and enrichment_CTD, which attaches it to its results.

Examples

# Examples write to a temporary cache, so running them cannot
# disturb CTD data you have already imported. Set the same option
# yourself to keep an analysis isolated from your main cache.
options(ctdR.cache = tempfile())

sample_file <- system.file(
    "extdata", "CTD_chem_gene_ixns_sample.csv",
    package = "ctdR"
)
import_CTD(sample_file)
#> Reading CTD chemical-gene interactions from: /home/runner/work/_temp/Library/ctdR/extdata/CTD_chem_gene_ixns_sample.csv
#> Filtered to 86 human interactions
#> Mapping genes for 10 chemicals...
#> Warning: 10 ChemicalID(s) appear with more than one ChemicalName in the CTD file; only the first name per ID is retained. Affected IDs: D000082, D001564, D002104, D003907, D004958 ... (and 5 more)
#> CTD data cached successfully in: /tmp/Rtmpt9pODL/file1b1341a7f1ac
#>   10 chemicals | 17 unique genes | 0 s
#>   CTD release: Mon Jan 01 00:00:00 EST 2024

genes <- data.frame(
    EntrezID = c("7124", "3569", "7157", "672", "1956"),
    pvalue = c(0.001, 0.003, 0.01, 0.02, 0.05)
)
# Which release is cached, before running anything:
ctd_provenance()
#> CTD provenance
#>   Report created: Mon Jan 01 00:00:00 EST 2024
#>   Source:         /home/runner/work/_temp/Library/ctdR/extdata/CTD_chem_gene_ixns_sample.csv
#>   Imported:       2026-10-07 06:48:38 UTC
#>   Retained:       10 chemicals, 86 chemical-gene pairs
#>   ctdR version:   0.99.11

res <- enrichment_CTD(genes, method = "ORA")
#> background: all 17 genes in the CTD sets, because no 'universe' was given. If your experiment could only detect some of them, pass those as 'universe': a background wider than what was measurable makes p-values too small.
ctd_provenance(res)
#> CTD provenance
#>   Report created: Mon Jan 01 00:00:00 EST 2024
#>   Source:         /home/runner/work/_temp/Library/ctdR/extdata/CTD_chem_gene_ixns_sample.csv
#>   Imported:       2026-10-07 06:48:38 UTC
#>   Retained:       10 chemicals, 86 chemical-gene pairs
#>   ctdR version:   0.99.11