From Subjective to Principled Single-Cell Data Analysis: Evaluating Annotation Behavior, Optimizing Integration, and Benchmarking Visualization Pipelines

Zhiqian Zhai
Ph.D., 2026
LI, JINGYI
Over recent years, single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics have transformed transcriptomic research by enabling high-resolution gene expression profiling and making it possible to generate atlas-scale datasets efficiently and cost-effectively. Yet rigorous scRNA-seq analysis remains challenging because key tasks, including cell-type annotation, data integration, and visualization, are often affected by substantial methodological and practical limitations. Cell-type annotation relies on a flexible multi-step pipeline whose functions and parameter choices are often shaped by analysts’ judgment and expertise, introducing subjectivity that can reduce reproducibility. Integration across batches is essential for large-scale and multi-condition studies, but it requires balancing the removal of between-batch variation with the preservation of true cell identity. Existing integration workflows still lack principled annotation-free strategies for feature selection and hyperparameter tuning. Visualization of atlas-scale data is further complicated by computational demands and by the strong influence of preprocessing steps such as normalization and integration, which remain insufficiently evaluated in existing benchmark studies. This dissertation addresses these challenges through three complementary projects on subjectivity in annotation, principled integration optimization, and systematic benchmarking of large-scale visualization pipelines.
My first project examines how graduate students enrolled in an advanced bioinformatics course at UCLA performed cell-type annotation. We also collected information on their pipeline-tuning practices, annotation results, and educational and research backgrounds. Our study revealed that participants performed well in identifying major cell types but often struggled with closely related cell subtypes. Participants who achieved higher cell-type annotation accuracy tended to incorporate data-quality control into analysis pipelines or have prior publications with single-cell analysis. Subjective choices of parameter values influence clustering results, and tuning beyond default settings typically improves clustering accuracy. However, because cell-type label assignment is driven largely by prespecified marker genes and subjective judgments of their expression, parameter tuning has limited impact on improving annotation accuracy. We also identified a confirmation bias: prior expectations about cell types influence subjective annotation decisions. For comparison, we evaluated an AI agent, Biomni, for automated cell-type annotation and found its performance worse than 70–90% of participants. Our findings underscore the importance of transparent reporting of analysis pipelines, including parameter choice justifications and prior expectations about cell types.
My second project presents IntegrateRigor, a data-driven, method-agnostic framework that performs batch-stable gene selection and optimizes integration hyperparameters across state-of-the-art methods, without relying on cell identity annotations. IntegrateRigor selects genes whose expression patterns are stable across batches using a gene-wise likelihood-based batch stability score, excluding batch-sensitive genes that can bias batch alignment during integration. IntegrateRigor then identifies optimal integration results across methods and hyperparameters by defining a dataset-level integration score that balances between-batch variation removal against cell identity preservation. In a colorectal cancer single-cell and spatial transcriptomics dataset, IntegrateRigor revealed previously uncharacterized cancer-immune interface states in the tumor microenvironment that were masked by both underintegration under default settings and over-integration in previous literature. Across diverse single-cell and spatial transcriptomics datasets, IntegrateRigor consistently improves integration performance by balancing over-integration and under-integration. By transforming integration from a heuristic step into a statistically principled, dataset-adaptive procedure, IntegrateRigor improves the reproducibility and discovery power of large-scale single-cell analyses.
My third project proposes a comprehensive benchmark study evaluating 180 method pipelines that combine state-of-the-art normalization, integration, and visualization methods. We assessed performance on five semi-synthetic atlas-scale scRNA-seq datasets that mimic real data and include ground truths. Each pipeline was evaluated using ten quantitative metrics that assess visualization accuracy and scalability. This benchmark provided systematic, evidence-based guidance for selecting appropriate methods for visualizing atlas-scale single-cell data.
2026