> For the complete documentation index, see [llms.txt](https://academy.dnanexus.com/llms.txt). Markdown versions of documentation pages are available by appending `.md` to page URLs; this page is available as [Markdown](https://academy.dnanexus.com/public-datasets-on-the-dnanexus-platform/single-cell/cellxgene-annotate.md).

# CELLxGENE Annotate

## Necessary Disclaimers and Legal

The user is responsible for reviewing and complying with the license requirements of the software, and data referenced in this documentation.

Users are responsible for the costs associated with analyzing data, software and its storage in their project spaces.&#x20;

Instance type availability and pricing are subject to the contract between the user or the user’s organization and DNAnexus.

## Citations for CELLxGENE Annotate

CELLxGENE Annotate is a cell annotation tool developed by the Chan Zuckerberg Initiative (CZI) in collaboration with the open-source community. Detailed documentation of the CELLxGENE Annotate application can be found [here](https://cellxgene.cziscience.com/).

## Overview of the CELLxGENE Annotate

The CELLxGENE Annotate interface is divided into 3 sections&#x20;

* The left-hand sidebar contains categorical metadata and fields curated for each dataset.
* The Center panel contains an embedding plot where each dot represents a cell.&#x20;
* The right-hand sidebar contains information about genes and gene sets.&#x20;

<img src="/files/ovPqnQj0VCR1cIlDqPyj" alt="" height="311" width="546">

Data Requirements: H5AD File Format

* The input data must be provided as an H5AD file (.h5ad). Please refer to [this document ](https://cellxgene.cziscience.com/docs/05__Annotate%20and%20Analyze%20Your%20Data/5_1__Getting%20Started:%20Install,%20Launch,%20Quick%20Start)for input requirements.

CELLxGENE Annotation Applet can be the final step in a three-stage single-cell RNA-seq analysis pipeline.&#x20;

* Raw sequencing data is first processed through nf-core/scrnaseq for quality control and alignment, then passed to nf-core/scdownstream for integration, clustering, and automated cell type annotation.&#x20;
* The resulting .h5ad file is then loaded into the CELLxGENE Annotation Applet for interactive exploration and manual refinement of cell metadata.

## Using CELLxGENE Annotate Applet on DNAnexus

### Copying the applet and example dataset into a project

To explore the demo analyses, copy the applet and example dataset from the public project into your own project space:

1. Create a project for your cellxgene annotation analyses, billed to your own organization. Tutorials on project setup can be found [here](https://academy.dnanexus.com/overview-of-the-platform/setting-up-a-project).
2. Navigate to the Resources tab, locate the project titled "Public Datasets AWS US (East)," and select the "cellxgene-annotation" folder.
3. Select the applet: cellxgene\_v1.0.3 and example datasets (.h5ad).

<img src="/files/FTNJT0V1IorsapMyv37Q" alt="" height="120" width="624">

4. Click "Copy" in the top-right menu and select the project created in Step 1.
5. Return to your project space to begin exploring the applet.

## Example dataset for testing

We provided three examples for testing this applet (examples folder) and one gene set example (GENE\_SET\_FATTY\_ACID\_OXIDATION.tsv)

1. pbmc3k.h5ad: a small PBMC dataset from CELLxGENE Github: <https://github.com/chanzuckerberg/cellxgene/blob/main/example-dataset/pbmc3k.h5ad>&#x20;
2. A-multi-tissue-single-cell-tumor-microenvironment-atlas.h5ad

* Source: <https://cellxgene.cziscience.com/collections/3f7c572c-cd73-4b51-a313-207c7f20f188>
* 391,963 cells across multiple tissue types

3. GSE174609\_scdownstream\_output.h5ad

* an .h5ad file generated from the nf-core/scdownstream pipeline<br>

Recommended Instance Types for Testing

| Dataset                                                      | Instance                     |
| ------------------------------------------------------------ | ---------------------------- |
| pbmc3k.h5ad (suggested test dataset)                         | mem1\_ssd1\_v2\_x8 (default) |
| A-multi-tissue-single-cell-tumor-microenvironment-atlas.h5ad | mem2\_ssd1\_v2\_x16          |
| GSE174609\_scdownstream\_output.h5ad                         | mem2\_ssd1\_v2\_x16          |

## Quick start

1. Click the applet to run the analysis.

<img src="/files/kXdq59sfkfRCMca4vf5O" alt="" height="308" width="507">

2. In the input box, choose the .h5ad file. For the demo, we used the pbmc3k.h5ad (2638 cells).

<img src="/files/CwroK2MmUEiYNjkNajn2" alt="" height="249" width="624">

3. Choose an instance type. The default instance is mem1\_ssd1\_v2\_x8.&#x20;

<img src="/files/7rDkjy0nTATnzEOReRNc" alt="" height="355" width="624">

4. Click Start to launch the analysis, then Open URL to access the interface.

<img src="/files/tCfrgYvNObglLKuSHWkX" alt="" height="95" width="536">

5. Name your user-generated data directory.&#x20;

<img src="/files/06sF76QEXMNeo8GU5mNI" alt="" height="246" width="435">

Note: This name will be used to save your outputs (gene sets and annotations) to the destination selected when launching the job. Only alphanumeric characters and underscores are accepted

\ <br>

<img src="/files/xQopenZqCTJv4EIY1eme" alt="" height="340" width="624">

Figure 2: CELLxGENE Annotate interface of example data.&#x20;

## Example use cases

For additional analyses, refer to the: [CELLxGENE Annotate Document](https://cellxgene.cziscience.com/docs/04__Analyze%20Public%20Data/4_1__Hosted%20Tutorials)&#x20;

There are seven examples for testing applets.

### Explore cell type composition of a dataset

1. In the left panel, locate the louvain metadata field
2. Click the droplet icon next to louvain to color the embedding by cell type
3. Click the > arrow to expand louvain and view the cell count for each population
4. Click the Toggle color legend on the top to show name of cell type
5. Click on individual cell type values (e.g. CD4 T cells) to highlight that population on the embedding

<img src="/files/oFuRHWDloJAmn0cCTMcl" alt="" height="324" width="595">

<br>

6\. Color by numerical metadata fields (n\_genes, n\_counts, percent\_mito) to explore data quality&#x20;

<img src="/files/6rQTxiKNMcbV34IjjBhC" alt="" height="317" width="595">

<br>

### Finding Cells Where a Gene Is Expressed

The right panel provides gene search functionality to visualize expression levels across the embedding. The applet supports gene search by Ensembl ID or gene symbol depending on .h5ad input file

* In the Quick Gene Search box, type CD19 and press Enter to add the gene.
* Click the droplet icon next to CD19 to color the embedding by its expression.<br>

<img src="/files/6r3rNve9SqgLMygifQMg" alt="" height="333" width="624">

### Identifying marker genes between two cell populations

1. In the left panel, expand louvain and select CD4 T cells
2. Click the "1" button in the toolbar to assign this population as group 1. Then, the toolbar shows "1: 1144 cells"
3. In the left panel, deselect CD4 T cells and select CD8 T cells
4. Click the "2" button in the toolbar to assign this population as group 2. Then, the toolbar shows "2: 316 cells"
5. Click the venn diagram icon in the toolbar to run Find Marker Genes
6. In the right panel, expand the results for group 1 and group 2

<img src="/files/LdyHXojBKX88eZjXaRDV" alt="" height="315" width="590">

<br>

7. Then, you can rename the Pop1 and Pop2

<img src="/files/wSLmfh1YgcnwBtKshJua" alt="" height="205" width="459">

8. Expand the CD4 T high to view the ranked marker gene list for CD4 T cells

<img src="/files/K5MgmfFkRMidBpZcUdca" alt="" height="309" width="580">

### Import gene set

1. In the right panel, click "Create new" next to Gene Sets
2. Paste a comma-separated list of pathway genes

ACADM,ACADS,ACADVL,ADIPOR1,ADIPOR2,ALOX12,BDH2,CPT1A,CPT1B,ECH1,ECHS1,HACL1,HADHB,HAO1,HAO2,PPARA,PPARD,PPARGC1A

3. Name the gene set (e.g. fatty\_acid\_oxidation)
4. Click Create a gene set

<img src="/files/3AuXrLtJZjWr8tNt2iZ6" alt="" height="307" width="442">

5. The dataset will show only genes available in the data

<img src="/files/QOLEepYVotdXvH1atUNi" alt="" height="512" width="408">

6. Click the droplet icon next to the gene set name to color the UMAP&#x20;

<img src="/files/OdzdRePLcKu3XX0UIoQX" alt="" height="300" width="560">

<br>

### Selecting and Subsetting Cells

Cells can be selected using the lasso tool, categorical filters, or gene expression thresholds. Selections can be combined for more precise subsetting.

Example using categorical filters and lasso selection:

1. In the left panel, expand louvain and select CD8 T cells
2. In the right panel, search and add gene NKG7
3. Click the droplet icon next to NKG7 to color the embedding by expression

<img src="/files/LSoMAIS2y6CiKijxO0UF" alt="" height="310" width="581">

For lasso selection:

1. Click the lasso tool icon in the toolbar to activate it
2. Click and drag to draw a freehand region around the target cluster on the UMAP<br>

<img src="/files/0XQrTwKKM9dGba7KB4fr" alt="" height="331" width="624">

3. Click to subset icon to subset your selected cells

<img src="/files/bf0Qi4HxR7LauKJSpVvU" alt="" height="333" width="624">

4. Click deselect to undo

<img src="/files/Pko1W8SnIHzjlokduC4T" alt="" height="335" width="624">

### Manually annotate a cell cluster using lasso

1. In the left panel, click the droplet icon next to louvain to color the embedding by existing cell type labels
2. Click "Create new category" at the top of the left panel
3. Enter the category name
4. In the dropdown, select louvain to duplicate existing labels as a starting point
5. Click "Create new category"&#x20;
6. Activate the lasso tool in the toolbar
7. Draw a lasso region around a cluster of interest on the UMAP&#x20;
8. Click the + button next to the test category in the left panel
9. Enter add labels to the selected cell. For example, name lasso

<img src="/files/urq9QGZg3I2osQwtUJY7" alt="" height="239" width="624">

## Technical Consideration

### Each Job Runs One Dataset

Each job on DNAnexus launches a single instance associated with one dataset. Users provide a single .h5ad file as input, and only that dataset is available throughout the session.

To analyze or explore a different dataset, a new job must be created with a different input file.

A running job can be accessed simultaneously by multiple users via a shared URL. However, all users connected to the same job will interact with the same fixed dataset for the duration of that session.

### Terminate instance

After finishing your session, close the browser tab and terminate the run in the Monitor panel to avoid unnecessary compute costs

### Feature Name Column (feature\_name\_col)

By default, CELLxGENE Annotate uses the existing index of the var table for gene searches in the right panel. If the dataset is indexed by Ensembl IDs (e.g. ENSG00000141510) (for example: A-multi-tissue-single-cell-tumor-microenvironment-atlas.h5ad) but you prefer to search by gene symbols (e.g. TP53), use the optional feature\_name\_col input to specify the column containing gene symbols. The script will reindex the dataset accordingly before launching CELLxGENE Annotate.&#x20;

Note: Check the AnnData object to confirm which column contains gene symbols before specifying this parameter.

<img src="/files/ZPya1FzP5JQHMn41QTiR" alt="" height="324" width="624">

### Output File Management on DNAnexus

Generated File Types

Each working session produces two output files:

* {name}-cell-labels-{########}.csv — cell annotation labels
* {name}-gene-sets-{########}.csv — gene sets and DEG results

Why Multiple Files Appear: CELLxGENE Annotate checks for file changes every minute using MD5 hash comparison. If the content has changed, it uploads a new copy to DNAnexus. Since DNAnexus allows multiple files with the same name, each upload is saved as a separate version — resulting in multiple files per session.

File content grows incrementally:

* cell-labels gains a new column each time a new annotation category is added
* gene-sets gains a new row each time a gene set is added or DEG is run

Session ID

Each session is assigned a unique 8-character suffix (########). Restarting CELLxGENE Annotate starts a new session with a new suffix and a fresh set of files.

Which File to Use: use the file with the latest timestamp within a session. It is the most complete and should be treated as the final output.

## Limitations

### Large File Handling

Large .h5ad files require an appropriate instance type to avoid out-of-memory errors. This is especially relevant when subsetting large cell populations or running DEG analysis. Finding a suitable configuration may require multiple attempts.

### Output Synchronization

The applet uploads output to the project every 60 seconds. This may result in duplicate file names in the output folder. Therefore, always use the file with the most recent timestamp as the final output. Please see “Output File Management on DNAnexus” for explanation.
