
Pre-scan multiple 10x CellRanger h5 files for multi-sample loading
prescan_tenx_h5_files.RdWalks each input file, reads the per-file feature ids, restricts to the
chosen feature_type (V3 only; ignored for V2), builds the intersection
or union universe of gene ids, and returns the file tasks expected by
load_multi_tenx_h5().
V2 and V3 files can be mixed in a single batch. Gene matching is done on
the feature id (ensembl-style) since gene symbols collide.
Usage
prescan_tenx_h5_files(
h5_paths,
feature_type = "Gene Expression",
gene_universe = c("intersection", "union"),
.verbose = TRUE
)Arguments
- h5_paths
Character vector of file paths to 10x h5 files. If names are provided, these will be used as experimental identifiers.
- feature_type
String. Modality to keep across all files (V3 only; ignored for V2). Defaults to
"Gene Expression".- gene_universe
One of
"intersection"or"union".- .verbose
Boolean. Controls verbosity.
Value
A list with:
universe - Character vector of gene ids in the universe.
universe_size - Length of the universe.
file_tasks - Named list of per-file task structures, each containing
exp_id,h5_path,version,no_cells,no_genes,feature_typeandgene_local_to_universe(integer vector,NAfor features outside the universe / non-target modality, 0-indexed).
Examples
# gene universe across two 10x h5 files
data <- generate_single_cell_test_data(
syn_data_params = params_sc_synthetic_data(n_cells = 200L, n_genes = 40L)
)
features <- data.table::data.table(
id = data$var$gene_id,
name = data$var$ensembl_id,
feature_type = "Gene Expression"
)
files <- c(a = tempfile(fileext = ".h5"), b = tempfile(fileext = ".h5"))
for (f in files) {
write_tenx_h5_sc(f, data$counts, data$obs$cell_id, features)
}
scan_res <- prescan_tenx_h5_files(h5_paths = files, .verbose = FALSE)
scan_res$universe_size
#> [1] 40
unlink(files)