
A starter stack for single-cell analysis
Single-cell data is deceptively simple — a big matrix of cells by genes — and surprisingly subtle to analyze well. Here is a dependable stack for getting from raw counts to biological insight.
Pick an ecosystem
You have two excellent options, and most labs standardize on one:
- Python / scverse: AnnData as the data structure, Scanpy for the standard workflow, and scvi-tools for probabilistic deep-learning models.
- R / Bioconductor: Seurat for an all-in-one toolkit, backed by the broader Bioconductor ecosystem.
A typical Scanpy workflow
1import scanpy as sc2adata = sc.read_h5ad("pbmc.h5ad")3sc.pp.filter_cells(adata, min_genes=200)4sc.pp.normalize_total(adata, target_sum=1e4)5sc.pp.log1p(adata)6sc.pp.highly_variable_genes(adata, n_top_genes=2000)7sc.pp.pca(adata)8sc.pp.neighbors(adata)9sc.tl.leiden(adata)10sc.pl.umap(adata, color="leiden")
When to reach for deep learning
Batch effects and multi-dataset integration are where probabilistic models earn their keep. scVI learns a shared latent space across batches; scANVI adds semi-supervised label transfer. These integrate cleanly with the AnnData objects Scanpy produces.
Don't collect data you could download
Before you sequence anything, check CZ CELLxGENE Census — tens of millions of harmonized cells you can query and pull straight into AnnData. Reference atlases make annotation faster and your conclusions stronger.
Tools mentioned



Single-Cell & Spatial
scvi-tools
Probabilistic, deep-learning models for single-cell omics


Programming the wet lab
Liquid handlers you script in Python, hardware-agnostic robot frameworks and modern lab notebooks are closing the loop between code and the bench.

Open structure prediction after AlphaFold 3
AlphaFold 3 restricts its weights to non-commercial use. Here is the open ecosystem that grew up around it — Boltz, Protenix and OpenFold3.