
Find multicellular programmes with DIALOGUE
dialogue_sc.RdDIALOGUE looks for programmes of cell-type-specific genes whose activity covaries across the samples that several cell types share. Where co-expression modules ask what varies together within a cell type, this asks what varies together between them, sample by sample.
Usage
dialogue_sc(
object,
cell_type_col,
sample_col,
features,
quality_col = NULL,
gene_ids = NULL,
pmd_params = params_dialogue_pmd(),
hlm_params = params_dialogue_hlm(),
refine_params = params_dialogue_refine(),
.verbose = TRUE
)Arguments
- object
SingleCells,SingleCellsSubsetorMetaCellsclass.- cell_type_col
String. Column in the obs table holding the cell type labels.
- sample_col
String. Column in the obs table holding the sample labels. The random intercept in stage two is over these. For
MetaCellsthe meta cells must have been built within samples, otherwise the level is not well-defined.- features
Named list of numeric matrices, one per cell type. Names must match the levels in
cell_type_col, row names must cover that cell type's cells, and each needs at least two columns. Rows are matched by name, not by position.- quality_col
Optional string. Column in the obs table to use as the cell quality covariate. If
NULL, defaults to the z-scored log library size.- gene_ids
Optional character. Genes to consider when building signatures. If
NULL, usesget_hvg()on the object.- pmd_params
List, see
params_dialogue_pmd().- hlm_params
List, see
params_dialogue_hlm().- refine_params
List, see
params_dialogue_refine().- .verbose
Boolean or integer. Verbosity.
Details
The algorithm works in three stages. Each cell type's features are collapsed to one row per sample and put through a sparse multi-CCA, giving every programme a weight vector per cell type plus a provisional gene signature. Then, for every ordered pair of cell types and every candidate gene, a mixed model asks whether a cell's own programme score tracks the partner's expression of that gene in the same sample. Finally the partners are meta-analysed and the scores refit onto the surviving genes by non-negative least squares.
features is mandatory and there is no default. DIALOGUE does not compute
it: it is whatever low-dimensional description of each cell type you trust,
and the method is only as good as that choice. Do not hand it a slice of a
global PCA. Those components mostly carry between-cell-type identity, which
is near-constant inside one cell type and gets dropped by the ANOVA filter,
leaving the decomposition to work off whatever is left. Run a PCA per cell
type instead, which for SingleCells means SingleCellsSubset() followed by
calculate_pca_sc() on each subset.
The method is unforgiving about study design, and the failure modes are
errors rather than bad answers. It needs at least two cell types, at least
five samples present in every cell type, and enough cells per sample per
cell type to clear abn_c in params_dialogue_pmd(). A handful of samples
with thousands of cells each is the regime it was built for; many samples
with a dozen cells each is not.
Examples
# a planted multicellular programme recovered across three cell types
data <- generate_dialogue_test_data()
dir <- tempfile("bixverse_dlg")
dir.create(dir)
object <- load_r_data(
SingleCells(dir_data = dir),
counts = data$counts,
obs = data$obs,
var = data$var,
sc_qc_param = params_sc_min_quality(
min_unique_genes = 10L,
min_lib_size = 50L,
min_cells = 10L
),
.verbose = FALSE
)
res <- dialogue_sc(
object,
cell_type_col = "cell_grp",
sample_col = "sample_id",
features = data$features,
gene_ids = data$var$gene_id,
pmd_params = params_dialogue_pmd(k = 2L, n_permutations = 20L),
.verbose = FALSE
)
res
#> DialogueResult (multicellular programmes)
#> Source class: SingleCells
#> Cell types: 3 (cell_type_1, cell_type_2, cell_type_3)
#> Shared samples: 14
#> Programmes: 2
#> mcp_01 - worst pair p: 0.9586 | spans 2 cell type(s)
#> mcp_02 - worst pair p: 0.05929 | spans 3 cell type(s)
#> Genes with a verdict: 17
#> Signature genes: 6 permissive | 9 strict
unlink(dir, recursive = TRUE, force = TRUE)