anthropics/knowledge-work-plugins

scvi-tools

Deep learning for single-cell analysis using scvi-tools.

Ver código-fonte
Documento original do Skill

Renderizado do repositório de origem, preservando títulos, exemplos, código, tabelas, links e imagens.

scvi-tools Deep Learning Skill

This skill provides guidance for deep learning-based single-cell analysis using scvi-tools, the leading framework for probabilistic models in single-cell genomics.

How to Use This Skill

  1. Identify the appropriate workflow from the model/workflow tables below
  2. Read the corresponding reference file for detailed steps and code
  3. Use scripts in scripts/ to avoid rewriting common code
  4. For installation or GPU issues, consult references/environment_setup.md
  5. For debugging, consult references/troubleshooting.md

When to Use This Skill

  • When scvi-tools, scVI, scANVI, or related models are mentioned
  • When deep learning-based batch correction or integration is needed
  • When working with multi-modal data (CITE-seq, multiome)
  • When reference mapping or label transfer is required
  • When analyzing ATAC-seq or spatial transcriptomics data
  • When learning latent representations of single-cell data

Model Selection Guide

Data TypeModelPrimary Use Case
scRNA-seqscVIUnsupervised integration, DE, imputation
scRNA-seq + labelsscANVILabel transfer, semi-supervised integration
CITE-seq (RNA+protein)totalVIMulti-modal integration, protein denoising
scATAC-seqPeakVIChromatin accessibility analysis
Multiome (RNA+ATAC)MultiVIJoint modality analysis
Spatial + scRNA referenceDestVICell type deconvolution
RNA velocityveloVITranscriptional dynamics
Cross-technologysysVISystem-level batch correction

Workflow Reference Files

WorkflowReference FileDescription
Environment Setupreferences/environment_setup.mdInstallation, GPU, version info
Data Preparationreferences/data_preparation.mdFormatting data for any model
scRNA Integrationreferences/scrna_integration.mdscVI/scANVI batch correction
ATAC-seq Analysisreferences/atac_peakvi.mdPeakVI for accessibility
CITE-seq Analysisreferences/citeseq_totalvi.mdtotalVI for protein+RNA
Multiome Analysisreferences/multiome_multivi.mdMultiVI for RNA+ATAC
Spatial Deconvolutionreferences/spatial_deconvolution.mdDestVI spatial analysis
Label Transferreferences/label_transfer.mdscANVI reference mapping
scArches Mappingreferences/scarches_mapping.mdQuery-to-reference mapping
Batch Correctionreferences/batch_correction_sysvi.mdAdvanced batch methods
RNA Velocityreferences/rna_velocity_velovi.mdveloVI dynamics
Troubleshootingreferences/troubleshooting.mdCommon issues and solutions

CLI Scripts

Modular scripts for common workflows. Chain together or modify as needed.

Pipeline Scripts

ScriptPurposeUsage
prepare_data.pyQC, filter, HVG selectionpython scripts/prepare_data.py raw.h5ad prepared.h5ad --batch-key batch
train_model.pyTrain any scvi-tools modelpython scripts/train_model.py prepared.h5ad results/ --model scvi
cluster_embed.pyNeighbors, UMAP, Leidenpython scripts/cluster_embed.py adata.h5ad results/
differential_expression.pyDE analysispython scripts/differential_expression.py model/ adata.h5ad de.csv --groupby leiden
transfer_labels.pyLabel transfer with scANVIpython scripts/transfer_labels.py ref_model/ query.h5ad results/
integrate_datasets.pyMulti-dataset integrationpython scripts/integrate_datasets.py results/ data1.h5ad data2.h5ad
validate_adata.pyCheck data compatibilitypython scripts/validate_adata.py data.h5ad --batch-key batch

Example Workflow

bash
# 1. Validate input data
python scripts/validate_adata.py raw.h5ad --batch-key batch --suggest

# 2. Prepare data (QC, HVG selection)
python scripts/prepare_data.py raw.h5ad prepared.h5ad --batch-key batch --n-hvgs 2000

# 3. Train model
python scripts/train_model.py prepared.h5ad results/ --model scvi --batch-key batch

# 4. Cluster and visualize
python scripts/cluster_embed.py results/adata_trained.h5ad results/ --resolution 0.8

# 5. Differential expression
python scripts/differential_expression.py results/model results/adata_clustered.h5ad results/de.csv --groupby leiden

Python Utilities

The scripts/model_utils.py provides importable functions for custom workflows:

FunctionPurpose
prepare_adata()Data preparation (QC, HVG, layer setup)
train_scvi()Train scVI or scANVI
evaluate_integration()Compute integration metrics
get_marker_genes()Extract DE markers
save_results()Save model, data, plots
auto_select_model()Suggest best model
quick_clustering()Neighbors + UMAP + Leiden

Critical Requirements

  1. Raw counts required: scvi-tools models require integer count data
python
   adata.layers["counts"] = adata.X.copy()  # Before normalization
   scvi.model.SCVI.setup_anndata(adata, layer="counts")
  1. HVG selection: Use 2000-4000 highly variable genes
python
   sc.pp.highly_variable_genes(adata, n_top_genes=2000, batch_key="batch", layer="counts", flavor="seurat_v3")
   adata = adata[:, adata.var['highly_variable']].copy()
  1. Batch information: Specify batch_key for integration
python
   scvi.model.SCVI.setup_anndata(adata, layer="counts", batch_key="batch")

Quick Decision Tree

Need to integrate scRNA-seq data?
├── Have cell type labels? → scANVI (references/label_transfer.md)
└── No labels? → scVI (references/scrna_integration.md)

Have multi-modal data?
├── CITE-seq (RNA + protein)? → totalVI (references/citeseq_totalvi.md)
├── Multiome (RNA + ATAC)? → MultiVI (references/multiome_multivi.md)
└── scATAC-seq only? → PeakVI (references/atac_peakvi.md)

Have spatial data?
└── Need cell type deconvolution? → DestVI (references/spatial_deconvolution.md)

Have pre-trained reference model?
└── Map query to reference? → scArches (references/scarches_mapping.md)

Need RNA velocity?
└── veloVI (references/rna_velocity_velovi.md)

Strong cross-technology batch effects?
└── sysVI (references/batch_correction_sysvi.md)

Key Resources

do mesmo repositório

Mais Skills

Todos os Skills
anthropics
Oficial

documentation

Write and maintain technical documentation. Trigger with "write docs for", "document this", "create a README", "write a runbook", "onboarding guide", or when the user needs help with any form of technical writing — API docs, architecture docs, or operational runbooks.

instalações
3
GitHub Stars
25,3 mil
Atualizado
21 de set.
anthropics
Oficial

data-visualization

Create effective data visualizations with Python (matplotlib, seaborn, plotly). Use when building charts, choosing the right chart type for a dataset, creating publication-quality figures, or applying design principles like accessibility and color theory.

instalações
2
GitHub Stars
25,3 mil
Atualizado
21 de set.
anthropics
Oficial

debug

Structured debugging session — reproduce, isolate, diagnose, and fix. Trigger with an error message or stack trace, "this works in staging but not prod", "something broke after the deploy", or when behavior diverges from expected and the cause isn't obvious.

instalações
2
GitHub Stars
25,3 mil
Atualizado
21 de set.
anthropics
Oficial

research-synthesis

Synthesize user research into themes, insights, and recommendations. Use when you have interview transcripts, survey results, usability test notes, support tickets, or NPS responses that need to be distilled into patterns, user segments, and prioritized next steps.

instalações
2
GitHub Stars
25,3 mil
Atualizado
21 de set.