Changelog
Source:NEWS.md
Changes in version 0.99.11
Documentation
- The installation instructions in the vignette and the README now reflect that ctdR has been accepted into Bioconductor. They said the package was still under review and pointed to GitHub as the only source. The package is currently available in Bioconductor devel (3.24) and will install from the standard repository after the next Bioconductor release; GitHub remains available for the latest development version.
Changes in version 0.99.10
Infrastructure
- The package source no longer carries the pkgdown site. The site configuration, its assets, the article written only for the website and the workflow that builds it now live on the
docs-sitebranch, which is where the published site is built from. Bioconductor requires them to be kept out of the package, and they were only ever needed to produce the website, never to build, check or install the package. The site itself is unchanged and stays at https://drake69.github.io/ctdR/.
Changes in version 0.99.9
Documentation
Bug fix. The README’s first code example did not run. It built its gene list with a column named
entrez_ids, while the package needsEntrezID, so copying it producedmapIds must have at least one key to match against: an error from AnnotationDbi, naming neither the column at fault nor the function that wanted it. The example and the input schema table are corrected, andenrichment_CTD()now checks for the column itself and says which columns it did find.The README covers what this release changed: the size thresholds, the provenance record, and the
alphaform of ORA, which is the one worth reaching for when the whole differential expression table is at hand. Its list of Bioconductor dependencies was also missing four of them.The vignette has a References section. Every method the package calls is published, and the Bioconductor submission guidance asks for the formal citations; there were none. It lists the CTD release paper, the GSEA method and the fgsea implementation separately, CAMERA and limma, GSVA, EnrichmentBrowser for the interoperability section, the independent-filtering paper the size-threshold reasoning rests on, and the GEO accession behind the worked example.
inst/CITATIONreads the version from the DESCRIPTION instead of naming it. It had saidR package version 0.99.2through eight releases, because nothing made it move. The Zenodo concept DOI is unchanged and verified to resolve to the current release.Releases are no longer cut automatically on a version bump. The workflow now runs only on demand. Firing on every push to main meant a bump travelling with other work cut a release before the work around it was ready, which is how 0.99.9 was tagged while its citation file still named 0.99.2. Cutting a release is a decision.
Significant user-visible changes
ORA no longer goes through
clusterProfiler::enricher(). The hypergeometric test is computed directly withstats::phyper(), andclusterProfilerhas been removed fromImports. The p-values are unchanged:scripts/ora_equivalence_check.Rin the revisions repository compares the two implementations across 24 configurations and finds them identical. The reason is cost, not correctness.clusterProfileraccounted for 59 of the package’s 175 hard dependencies for this single call, and its visualization layer, which ctdR never used, put Pandoc, cairo, fontconfig, freetype2, libuv and glpk among the system requirements of every installation. The dependency closure drops from 175 to 116 packages and the system requirements from 14 to 7.-
maxGSSizeno longer defaults to 500. There is no upper limit. The lower threshold was corrected in this same release because it had been inherited from a tool tuned for KEGG and GO; the upper one had been inherited from exactly the same place and kept without question.The criterion is the one used for the lower threshold, read from the other end. A set of M genes cannot, even when every one of the m input genes falls inside it, produce a p-value below
choose(M, m) / choose(N, m), roughly(M/N)^m. That floor rises with M, so a large enough set is untestable. Whether CTD holds any such set is a measurement, not an opinion: with N = 28,571 and an input list of 169, the largest still-testable set is about 26,600 genes, 93% of the universe, while the largest chemical in CTD has 16,536. No chemical is untestable from above.A cap at 500 excluded 265 chemicals, among them benzo(a)pyrene, valproic acid, sodium arsenite, bisphenol A, aflatoxin B1 and particulate matter: the canonical compounds of toxicology, whose sets are large because the literature on them is. This is where CTD parts company with GO, in which a large term is one that has stopped meaning anything. Removing the cap costs 3% more tests, 8,235 against 7,970.
In the bundled RNA-seq example the cap excluded dexamethasone, the treatment the experiment applied, which without it ranks third of 8,193 chemicals at an adjusted p-value of 0.0002.
maxGSSizeremains settable for anyone who wants a cap of their own. The default
minGSSizefor ORA is now 2, chosen for CTD instead of inherited from a general-purpose tool. The previous value, 10, came fromenricher()’s own default, which suits KEGG and GO (median set sizes 72 and 11) but not CTD, where the median chemical has 4 target genes: it silently excluded 7,831 of 11,067 chemicals from testing altogether. Coverage goes from 26.8% to 72.0% of chemicals. One-gene sets remain excluded on purpose: their hypergeometric p-value equals the ratio of input genes to background whichever gene they contain, so they measure membership rather than enrichment.-
Bug fix. Running an example, knitting the vignette or running
R CMD checkoverwrote whatever CTD data the user had imported, replacing it with the ten-chemical sample those examples run on. The examples calledimport_CTD()against the real user cache undertools::R_user_dir(); only the test suite knew to redirect it. They now write to a temporary cache, as does the vignette. Writing outside the session temporary directory is also against both CRAN and Bioconductor policy, so this was a defect on two counts.The examples set
options(ctdR.cache = tempfile())visibly rather than hiding it, because the same option is how a user isolates one analysis from their main cache. The test suite now redirects the cache from a singlesetup.Ras well: five of its six importing files were writing to the real one, so running the tests carried the same cost. Tests that need a cache of their own restore the suite’s, where they previously cleared the option, which sent every later test in the run back to the user’s cache. -
New
ctd_provenance()returns the record of which CTD release an analysis ran on: theReport createddate CTD stamps into its own file header, where the file came from, when it was imported, how many chemicals and chemical-gene pairs were retained, and the ctdR version. CTD re-releases continuously and does not version its download filenames, so that header date is the only thing identifying which snapshot a result came from.import_CTD()now reads it, caches it and prints it, andenrichment_CTD()attaches it to every result.Where it is stored follows the container:
metadata()on objects that have the slot, which covers theSummarizedExperimentGSVA returns and theDataFramefrom importing aCTDFile; an attribute on the data frames from ORA, GSEA and CAMERA. Use the accessor rather than either directly. The record survives subsetting, ordering,head()and the common dplyr verbs; it does not survivemerge()orsubset(), which drop attributes, and the accessor says so rather than returning an empty answer. -
import_CTD()no longer assumes the CTD header is 27 lines long. A CTD download has no header row: the field names sit inside the commented preamble. The names are now located by finding the commented line that lists at least three known CTD field names, taking the last such line, and the file is read withcomment = "#"as suggested in review. The hard-codedskip, the row dropped afterwards to compensate, and the patch that stripped"# "from the first column name are all gone.This matters beyond tidiness. A fixed line count fails silently: insert one comment line upstream and every column shifts, with the analysis proceeding on misaligned data. Matching field names fails loudly, and it adds no assumption the package was not already making, since those names are referenced throughout.
The bundled sample file now mirrors the structure of a real CTD download, header included. It previously carried an uncommented, duplicated header row, shaped so the old hard-coded skip would work, which meant tests and examples never exercised the format users actually have. Its preamble is deliberately a different length from a real download’s, so that nothing can come to depend on the count again.
Bug fix. The ORA
universeargument was unusable in the default identifier mode, and failed silently. Withgene_id_type = "symbol", the gene sets are keyed by HGNC symbol and the input gene list is converted for that reason, but the universe was passed through as given. A universe of Entrez IDs therefore intersected the background at nothing, the size filter then removed every gene set, and the call returned zero rows instead of an error. The vignette’s own example of restricting the background to expressed genes shipped in that state. The universe is now converted alongside the input; identifiers that are not Entrez IDs, and Entrez IDs that do not map, are kept as they are, so a universe of symbols or a mixture of the two works too.plot_CTD()draws each method on its own measure of effect: fold enrichment for ORA, the normalized enrichment score for GSEA, with the axis labelled accordingly. It previously drewFoldEnrichmentfor both, which is why removing the fabricated GSEA fold broke plotting until this release. A result frame missing the column it needs now says which one, instead of failing insidedata.frame()with a row-count mismatch.-
Breaking change.
FoldEnrichmentis gone from GSEA results. It was computed asabs(ES) / mean(ES), where the divisor is the mean enrichment score across whichever chemicals happened to be tested in the same run. That made it a property of the run rather than of the chemical: the same chemical scored against a different collection got a different value, with nothing about the chemical having changed. It also duplicated, badly, a quantity fgsea already computes properly:NormalizedEnrichmentScore, the NES, which scales the score for gene set size and is what the field compares. Sharing a name with ORA’s fold enrichment, which is a genuine observed-over-expected ratio, invited a cross-method comparison that never meant anything.The shared output schema is unaffected. It has always been the five leading columns, with method-specific extras differing by method, so GSEA was never obliged to carry a column ORA has.
The example script checks the cache the package actually uses. It was rebuilding the path with
rappdirs::user_cache_dir("ctdR"), the location ctdR kept its cache in before moving toBiocFileCache, so it inspected a directory the package no longer writes to: it passed where old files happened to remain and refused to run on a clean machine with a perfectly good cache. It now asks the package, throughctd_cache()andctd_provenance(), and reports the CTD release it found.New tests cover both cache states, empty and populated, on the bundled ten-chemical sample. They pin what the package says when nothing has been imported and what it returns when something has, which is the contract the example script branches on.
-
New
alphaandalpha_columnarguments for ORA. Handenrichment_CTD()the whole differential-expression table and say which p-value column to judge on, instead of filtering first and then describing the background separately:enrichment_CTD(de, method = "ORA", alpha = 0.05, alpha_column = "padj")The genes under the threshold become the list to test and every row becomes the background, so the two are derived from one object and cannot disagree. Filtering first and passing
universeasks the caller to reconnect two things that were together a moment earlier, and that reconnection is where the background goes wrong.alpha_columntakes a name or an index and defaults to the second column. Naming it matters on a real table, where the second column is usually a fold change:limma::topTable()putslogFCthere. Which p-value to judge on is the researcher’s decision, and the function reports the column it used along with how many genes passed.alphaanduniverseare mutually exclusive: withalphathe background is already decided, so passing both is an error rather than a precedence rule applied in silence. Passing
universeto"GSEA","CAMERA"or"GSVA"now warns instead of being dropped without comment. Only ORA needs the argument, because only ORA takes an input that does not record what was measurable: GSEA ranks the whole list supplied, and CAMERA and GSVA intersect the gene sets withrownames(x).universeis now a named argument ofenrichment_CTD()rather than something passed through.... It appears in the help page and in autocompletion, and a misspelling raises an error instead of being swallowed silently by...and running the analysis on the wrong background.-
ORA now says which background it used when none was given. The default is every gene in the CTD gene sets, which is a fallback rather than a recommendation: the package cannot know what a given platform measured. A background wider than what the experiment could detect makes p-values too small, because genes that could never have been selected still count in it. The error is anti-conservative, so it was worth a message rather than a footnote.
On the RNA-seq analysis bundled with the package, the default background returns 32 chemicals at FDR < 0.05 and the correct one, the genes that entered the differential test, returns 19. Thirteen of the thirty-two come from the background alone. The bundled example script now passes
universeand explains why, and the vignette carries the comparison. The vignette explains how to read the size of a chemical’s gene set, which in CTD also reflects how much the chemical has been studied. The natural suspicion, that large sets are padded with genes responding to everything, does not hold when measured: genes in sets above 500 appear in a median of 47 chemicals, against 232 for genes in sets of 4 or fewer. Small sets are the ones built from the usual suspects. The consequence is about what a result means, and holds for any enrichment analysis run against a curated database: the question answered is not whether a chemical is associated with a gene list, but whether it is associated as far as the published literature records. A chemical absent from the output is uninvolved only as far as anyone currently knows.
The bundled example script prints its per-method summary one method at a time, with fixed-width columns, ten chemicals rather than five, and the overlap, set size and fold enrichment beside the p-values. It previously printed one wide data frame, which R wrapped into three detached blocks, so reading off which chemical ranked third meant counting rows across all three. The expected-hit check likewise says in words when a chemical was never tested, rather than printing a bare
NAthat reads as an absent result.The ORA
universeargument now accepts any vector of gene identifiers, not only a character one. A DE table read back withread.delim()gives integer Entrez IDs, so the most ordinary use of the argument,universe = de$EntrezIDto restrict the background to measured genes, used to fail while the same column passed as the input gene list worked. Both are now coerced. The previous backend accepted a non-character universe and then silently ignored it, computing against a background the caller had not asked for.ORA now reports what its size filter removed. A chemical excluded for having too few or too many target genes is absent from the results, not present with an unremarkable p-value, and the two cases used to be indistinguishable.
ora()emits a message giving how many chemicals went untested and on which side of the thresholds they fell.-
Breaking change. ORA results now have 10 columns instead of 13.
ChemicalID,ChemicalName,Method,PValue,PValueAdjusted,GeneRatio,BackgroundRatio,EnrichedGenes,CountandFoldEnrichmentare unchanged. Three columns are gone, none of which carried information the remaining ones do not:-
QValueheld Storey’s q-value from theqvaluepackage. On result sets of the size a CTD analysis produces it was identical toPValueAdjusted, because the q-value estimator falls back to Benjamini-Hochberg when it cannot estimate the proportion of true nulls, so the two columns held the same numbers. Reproducing it would mean taking the dependency back for a duplicate. UsePValueAdjustedfor false-discovery control. -
RichFactorandzScorecame from the previous backend and were passed through undocumented: neither appeared in the output schema described in the vignette.
This also fixes a defect. The previous output carried two columns named
FoldEnrichment, one from the backend and one from ctdR’s own rename. Every ORA call raised a duplicated-column warning frommerge(), andresults$FoldEnrichmentreturned whichever of the two came first. There is now one. -
Internal
The internal
ora()engine takesuniverse,minGSSizeandmaxGSSizeas explicit arguments rather than forwarding an opaque...to another package, so an unrecognised argument now raises an error instead of being silently discarded..parse_ratio()has been removed. Fold enrichment is computed from the counts directly instead of being parsed back out of the"n/d"strings.New
tools/check_internal_params.R, wired into CI, fails the build when a documented function has an argument without its@param.R CMD checkskips that cross-check for topics marked\keyword{internal}, so such an argument used to ship undocumented with the check still reporting Status OK. Six internal topics that were already in that state have been documented.The ORA test suite checks
ora()againststats::phyper()and closed-form values rather than against another implementation. After the migration the hypergeometric distribution is the reference; an equivalence test againstclusterProfilerwould have pinned ctdR’s correctness to a package it no longer depends on.
Changes in version 0.99.8
New features
enrichment_CTD()now accepts aSummarizedExperimentfor the matrix-based methods ("CAMERA"and"GSVA"), in addition to a plain expression matrix. The assay and the sample annotation stay in one object, so subsetting or reordering samples cannot silently desynchronise them from the group labels used to build the design matrix."GSVA"returns the container it was given: a matrix in returns a numeric matrix, aSummarizedExperimentin returns aSummarizedExperimentwhose assay holds the scores and whosecolDatais carried over from the input. Per-sample scores therefore arrive with the sample annotation needed to interpret them.New
assayargument onenrichment_CTD()selects which assay to use when the input carries more than one, by name or by index. Defaults to the first.plot_CTD()accepts GSVA scores wrapped in aSummarizedExperiment.
Changes
The bundled
inst/extdata/GSE311566_subset.rdsexample is now stored as aSummarizedExperimentinstead of alist(expr, coldata). Code reading it must useassay(se)andse$groupin place of$exprand$coldata$group.metadata()records the GEO source, the subsetting applied, and the assay units.The vignette and the RNA-seq workflow tutorial build a
SummarizedExperimentand carry it through the analysis.SummarizedExperimentwas added toImports. It was already a hard transitive dependency throughGSVA, so the installation footprint is unchanged.
Changes in version 0.99.7
New features
New
CTDFileclass (aBiocIO::BiocFilesubclass) with animport()method on theBiocIOgeneric.import(CTDFile(path_or_url))reads a CTD chemical-gene interactions file into aS4Vectors::DataFrameof validated human interactions, following the Bioconductor import/export convention.New
as_genesets_CTD()exports the CTD chemical gene sets as a named list keyed byChemicalID, ready for third-party enrichment engines such asEnrichmentBrowser::sbea(). Supportsid_typeandinteraction_types.New
ctd_cache()retrieves the processed cached tables ("chemicals","interactions") without loading the cache files by hand.import_CTD()(andimport(CTDFile)) now accept a local path or a URL. Remote URLs are downloaded and cached viaBiocFileCache; no default URL is assumed, and a one-time CTD data-licensing reminder is shown on the first remote fetch.
Changes
The processed-data cache moved from a
rappdirsdirectory of.rdafiles to aBiocFileCachestore undertools::R_user_dir("ctdR", "cache").rappdirsis no longer a dependency.pAdjustMethodnow accepts any value instats::p.adjust.methods(previously restricted toBH,bonferroni,fdr,none); its documentation referencesstats::p.adjust.
Documentation
- The vignette gains an “Interoperability with existing Bioconductor infrastructure” section (reading via
CTDFile/import(), and handing gene sets toEnrichmentBrowser::sbea()) plus a note disambiguating the “CTD” acronym.
Changes in version 0.99.6
New features
enrichment_CTD()gains aninteraction_typesargument: a character vector of CTDInteractionActionsvalues (e.g."increases^expression","decreases^expression") that filters each chemical’s gene set at enrichment time. Gene sets are rebuilt on the fly from the newctd_interactions.rdacache without re-importing;NULL(default) retains full backward compatibility. Supported by all four methods (ORA, GSEA, CAMERA, GSVA). Requiresimport_CTD()to be re-run once to generatectd_interactions.rda.enrichment_CTD()gains agene_id_typeargument ("symbol"or"entrez") controlling whether theEnrichedGenesoutput column reports HGNC symbols (with Entrez ID fallback for unmapped genes) or raw Entrez IDs. Default"symbol"is backward compatible.enrichment_CTD(method = "ORA")now forwardsuniverse,minGSSize, andmaxGSSizetoclusterProfiler::enricher()via.... Settinguniverse = de$EntrezIDrestricts the background to measured genes, avoiding inflated fold-enrichment estimates. GSEA and GSVA already exposedminSize/maxSize; CAMERA...was already forwarded.import_CTD()now cachesctd_interactions.rda, a long-format table of(ChemicalID, EntrezID, InteractionActions)triples used by the newinteraction_typesfilter.import_CTD()reports elapsed time, chemical count, and unique gene count on completion.import_CTD()detects and warns when the sameChemicalIDappears with multipleChemicalNamevalues (CTD data quality issue; first name retained). Reports an informational message when the same name is shared by multipleChemicalIDs (legitimate parent/derivative pairs).
Bug fixes
Fixed
EnrichedGenescolumn in GSEA output: was incorrectly set to data frame row indices instead of gene identifiers. Gene labels are now mapped viaAnnotationDbi::mapIds()in.run_gsea().clusterProfiler::enricher()messages (“No gene can be mapped”, “Expected input gene ID”, etc.) are now suppressed viasuppressMessages().
Documentation
New pkgdown article
vignettes/articles/tutorial_rnaseq_workflow.Rmd: a complete RNA-seq → chemical enrichment workflow using the full GSE311566 dataset (downloaded from GEO at runtime),limmaDE, all four methods with recommended parameters, direction-aware GSEA, and a per-method Dexamethasone ranking recap.Vignette gains a “Gene set size filters and background universe” section documenting
universe,minGSSize/maxGSSize(ORA),minSize/maxSize(GSEA, GSVA), and CAMERA’s implicit minimum of 2 genes.Added a new
interaction_typesparameter description to the vignette explaining the CTDverb^nounvocabulary and direction-aware analysis.Added a full end-to-end pipeline example at
inst/scripts/example_gse311566_full_pipeline.R. The script downloads the complete GSE311566 Female PBMCs normalised-count matrix (Dex vs DMSO), computes an a-priori power analysis with declared alpha thresholds, runslimma-based differential expression, and exercises all four enrichment methods (ORA, GSEA, CAMERA, GSVA) with BH-adjusted significance cutoffs. It refuses to fall back to the bundled toy CTD sample; the user must populate the CTD cache with the real chemical-gene interactions file first. Complements the vignette (which uses the bundled subset without alpha cutoffs) by providing the production-shaped example linked from the README and the companion paper.
Changes in version 0.99.5
Documentation
- Added an end-to-end real-data example to the vignette using a small subset of GEO series GSE311566 (human PBMCs, dexamethasone vs. vehicle, female donors). The example walks through loading the bundled subset, a deliberately minimal base-R differential expression with
t.test+p.adjust, and the four ctdR methods (ORA, GSEA, CAMERA, GSVA) on the resulting DE. - Bundled
inst/extdata/GSE311566_subset.rds(~34 KB) containing log2-normalised counts for 1,500 top-variance genes plus the 17 genes referenced by the toy CTD sample, across 7 samples (4 DMSO + 3 Dex). - Added a reproducible provenance script at
inst/scripts/make_gse311566_subset.Rand a per-file documentation README atinst/extdata/README.md. - The
enrichment_CTD()@examplesblock no longer relies on\donttest{}: CAMERA and GSVA examples now run directly on the bundled subset, satisfying the BiocCheck recommendation against\dontrun{}/\donttest{}in man pages.
Testing
- New
tests/testthat/test-e2e-gse311566.Rruns the full data-to-enrichment pipeline on the bundled GSE311566 subset and asserts that Dexamethasone (D003907) ranks in the top 3 by GSEA p-value and in the top 6 by CAMERA p-value, plus structural checks on the GSVA output. Guards against silent regressions in ID mapping, output schema, or sort order that the vignette and man-page examples would only catch as “still runs”.
Changes in version 0.99.4
Breaking changes
These changes are made now, while the package is pre-1.0 and not yet accepted into Bioconductor, so that the public column schema is stable before any external code depends on it.
Input column name (ORA / GSEA)
- The input data frame must now provide an
EntrezIDcolumn (wasentrez_ids). The numeric value column can still be named freely.
Unified output schema (ORA, GSEA, CAMERA)
All three data-frame-returning methods now share the same leading columns, in this order:
ChemicalID, ChemicalName, Method, PValue, PValueAdjusted, ...
The new Method column carries the method label ("ORA", "GSEA", "CAMERA"), making cross-method rbind / dplyr::bind_rows straightforward.
Rows are sorted by PValueAdjusted ascending in all three methods (previously ORA results were unsorted).
Column renames (full table)
| Method | Old column | New column |
|---|---|---|
| GSEA | pval |
PValue |
| GSEA | ES |
EnrichmentScore |
| GSEA | NES |
NormalizedEnrichmentScore |
| GSEA | size |
GeneSetSize |
| GSEA | leadingEdge |
LeadingEdge |
| GSEA | Enriched_GENE |
EnrichedGenes |
| ORA | pvalue |
PValue |
| ORA | padj |
PValueAdjusted |
| ORA | BgRatio |
BackgroundRatio |
| ORA | qvalue |
QValue |
| ORA | geneID |
EnrichedGenes |
| ORA | foldEnrichment |
FoldEnrichment |
| ORA | ID |
ChemicalID |
| ORA | Description |
(dropped — was always == ID) |
| CAMERA | PValue |
(kept as) PValue
|
| CAMERA | NGenes |
GeneSetSize |
| CAMERA | FDR |
(dropped — PValueAdjusted recomputed per pAdjustMethod) |
| All | (new) | Method |
Cross-method semantic alignment: - GeneSetSize replaces both GSEA’s size and CAMERA’s NGenes - EnrichedGenes replaces both ORA’s geneID and GSEA’s Enriched_GENE - PValue / PValueAdjusted are spelled the same across all methods
Internal refactor
-
enrichment_CTD()shrank from 101 lines to ~30 by delegating argument validation to.validate_enrichment_args()and adding.run_ora()/.run_gsea()runners mirroring the existing.run_camera()/.run_gsva()shape. -
.format_enrichment_result()is the single source of truth for the engine -> canonical-schema mapping: each runner passes arename = c(old = "New")anddrop = c(...)and the formatter handles padj recomputation, metadata merge, column ordering, sort and row-name reset. -
gsea()andora()engines now return their underlying tool’s native column casing (pval/pvalue,ES,NES, …); the only semantic rename they apply themselves is lifting the primary-key column (pathwayfor fgsea,IDfor clusterProfiler) toChemicalID, so internal callers never see the misleading generic name. -
gsea()signature dropped the unusedchemicalsandpAdjustMethodarguments. Itsentrez_idsparameter was renamed togene_table(still internal-only). -
.run_camera()shrank to 48 lines (was 57) and.run_ora()/.run_gsea()are now under 30 lines each. BiocCheck’s “function length > 50” NOTE is satisfied for the enrichment-table pipeline.
Changes in version 0.99.2
New features
-
enrichment_CTD()now supports four enrichment methods through a unified interface, selectable via themethodargument:-
"ORA"(default) — Over-Representation Analysis viaclusterProfiler::enricher()(unchanged). -
"GSEA"— rank-based Gene Set Enrichment Analysis viafgsea::fgsea()(unchanged). -
"CAMERA"— competitive gene-set test accounting for inter-gene correlation, vialimma::camera(). Input: a numeric expression matrix (genes x samples) plus a design matrix and a contrast. -
"GSVA"— per-sample Gene Set Variation Analysis viaGSVA::gsva(), returning a chemical x sample score matrix.
-
- New first argument
x(polymorphic): a data frame for ORA/GSEA, a numeric matrix for CAMERA/GSVA. Auto-detects identifier type (Entrez vs HGNC SYMBOL) fromrownames(x); an explicitid_typeoverride is also accepted. -
plot_CTD()now dispatches on input class and method-specific columns: bar/dot plots of fold enrichment for ORA/GSEA, bar/dot plots of-log10(padj)coloured by direction of enrichment for CAMERA, and a sample-level heatmap of the top-variance chemicals for GSVA.
Deprecation
- The first argument of
enrichment_CTD()was renamedentrez_ids->x. Calls using the old name still work but emit a deprecation warning and will be removed in a future release.
Changes in version 0.99.1
Bioconductor reviewer feedback
- Removed
renvfrom the package. Dependencies are managed viaDESCRIPTIONand installed byr-lib/actions/setup-r-dependencies(pak-based) in CI.
Bug fixes
-
import_CTD()now storesChemicalName_GeneSymbols$geneas a character vector instead of a factor, fixing “universe must be a character vector” inclusterProfiler::enricher(). - Inlined
parse_ratio()as a private helper becauseDOSE::parse_ratiois no longer exported.DOSEis dropped fromImports. - Aligned
man/gsea.Rdparameter name withR/gsea.R(ChemicalName_GeneEntrezIds), removing a codoc-mismatch WARNING. - Declared
plot_CTD()ggplot2 NSE column references inutils::globalVariables(), removing “no visible binding” NOTEs.
Infrastructure
- Pinned
trufflesecurity/trufflehogto a concrete version; the bare@v3tag does not exist in the action repo. - Made
oysteRaudit fail-soft when OSS Index credentials are missing (still fails the build on real vulnerabilities). - Removed the
dependency-reviewjob: GitHub’s Dependency Graph does not support RDESCRIPTIONfiles.
Changes in version 0.99.0
Improvements
- Bumped version to 0.99.0 for Bioconductor submission.
- Fixed R CMD check to pass with 0 errors, 0 warnings, 0 notes.
- Fixed CI workflows for macOS, Ubuntu, and Windows.
- Added Codecov integration for coverage reporting.
- Added GitHub Pages site with usage examples.
- Fixed broken README badges.
Changes in version 0.1.2
New features
- Added
import_CTD()function to import and cache CTD chemical-gene interaction data from user-downloaded files. - Added Over-Representation Analysis (ORA) method via
clusterProfiler::enricher. -
enrichment_CTD()now supports two methods:"ORA"(default) and"GSEA". - Clear error message when CTD data has not been imported, with instructions on how to download and import the required file.
Improvements
- Separated data import from analysis — users call
import_CTD()once, thenenrichment_CTD()as many times as needed. - Added fold enrichment calculation to both ORA and GSEA results.
-
gsea()now receiveschemicalsas an explicit parameter instead of relying on the parent environment. - Removed deprecated
npermparameter fromfgsea::fgseacall. - Removed leftover
browser()calls.
Documentation
- Added comprehensive roxygen documentation for all exported and internal functions, including parameter descriptions, return value tables, and examples.
- Added package-level help page (
?ctdR) with quick start guide. - Added data licensing disclaimer throughout documentation and DESCRIPTION.
- Added README.md with badges, installation instructions, and usage examples.
- Added vignette with complete workflow.
Changes in version 0.1.1
- Added GSEA analysis via
fgsea::fgsea. - Initial caching mechanism using
rappdirs.