跳至章节信息跳至正文
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Doublet 检测

🧠 关键要点
⚙️ 环境设置
步骤
yml
  1. 安装 conda:

    • 在创建环境之前,请确保 conda 已安装在你的系统中。

  2. 保存 yml 内容:

    • 将 yml 选项卡中的内容保存为文件 environment.yml。

  3. 创建环境:

    • 打开终端或命令提示符。

    • 运行以下命令:

      conda env create -f environment.yml
  4. 激活环境:

    • 创建好环境后,使用以下命令激活它:

      conda activate <environment_name>
    • 请将 <environment_name> 替换为 environment.yml 文件中指定的环境名称。该名称在 yml 文件中如下所示:

      name: <environment_name>
  5. 验证安装:

    • 通过运行以下命令,检查环境是否创建成功:

      conda env list
🗄️ 获取数据和笔记本

本书使用 lamindb 存储、共享和加载数据集与笔记本,托管实例为 theislab/sc-best-practices。感谢提供免费托管的 Lamin Labs。

  1. 安装 lamindb

    • 安装 lamindb Python 软件包:

    pip install lamindb
  2. 可选择创建 Lamin 账户

  3. 验证你的设置

    • 下面用 Python API 检查连接;命令行方式可使用 lamin connect:

    import lamindb as ln
    
    ln.Artifact.connect("theislab/sc-best-practices").df()

    运行后应显示最多 100 条已保存的数据集记录。

  4. 访问数据集(Artifact)

    import lamindb as ln
    af = ln.Artifact.connect("theislab/sc-best-practices").get(key="key_of_dataset", is_latest=True)
    obj = af.load()

    加载后的对象可直接在内存中分析。如需指定数据版本,可调整 lamindb.Artifact.connect("theislab/sc-best-practices").get("SOMEIDXXXX") 中的 ID。

  5. 访问笔记本(Transform)

    lamin load <notebook url>

    该命令将笔记本下载到当前工作目录。与 Artifacts 类似,可通过指定 ID 获取旧版本。

研究动机

在 质量控制 章节中,我们根据异常高的计数(Count)及样本内分布过滤了部分疑似低质量观测。本章进一步利用抗体衍生标签(Antibody-Derived Tags, ADTs)识别不同细胞类型共占液滴形成的异型双细胞(heterotypic Doublet)。其依据是正常情况下相互排斥的细胞类型标记同时出现 Sun et al., 2021。

环境设置

import warnings

import scanpy as sc

warnings.filterwarnings("ignore")
sc.settings.verbosity = 0
sc.set_figure_params(
    dpi=80,
    facecolor="white",
    frameon=False,
)

import lamindb as ln

ln.track()
→ connected lamindb: theislab/sc-best-practices
→ found notebook doublet_detection.ipynb, making new version -- anticipating changes
→ created Transform('xo1WtNUmrJCK0007', key='doublet_detection.ipynb'), re-started Run('33z2HhX1JxiWbGxS') at 2026-07-29 11:07:33 UTC
→ notebook imports: lamindb-core==2.3.1 muon==0.1.9 scanpy==1.12.3
• recommendation: to identify the notebook across renames, pass the uid: ln.track("xo1WtNUmrJCK")

加载数据

加载前一章保存的 MuData 对象,参见 归一化(normalization):

af = ln.Artifact.connect("theislab/sc-best-practices").get(
    key="surface-protein/cite_normalization.h5mu", is_latest=True
)
mdata = af.load()
mdata
Loading...

根据细胞类型标记识别 Doublet

先比较通常互斥的表面标记,例如 T 细胞标记 CD3 与 B 细胞标记 CD19。两者同时显示较强信号时,可提示 T/B 细胞组成的 双细胞(Doublet)。不过,背景结合、细胞间蛋白转移或异常生物学状态也可能产生非典型共表达,因此不能仅凭两个标记确证 Doublet。

同样,可比较 T 细胞标记 CD3 与单核细胞标记 CD14,寻找疑似 T 细胞–单核细胞 Doublet。

sc.pl.scatter(mdata["prot"], x="CD3", y="CD19-1", color="log1p_total_counts")
<Figure size 361x320 with 2 Axes>

散点图左下方的观测在两个标记上信号都低;左上和右下主要为单一标记阳性,右上则为两个标记信号均较高的观测。

本例将右上方的双阳性观测作为疑似 Doublet,进一步设置阈值过滤。

接着绘制 CD3 与 CD14,检查另一组互斥标记。

sc.pl.scatter(mdata["prot"], x="CD3", y="CD14-1", color="log1p_total_counts")
<Figure size 361x320 with 2 Axes>

根据本例图形,将 CD3 与 CD19-1 均大于 25,或 CD3 与 CD14-1 均大于 25 的观测标记为疑似 Doublet。这里的 25 是前一章 背景去噪与缩放(denoised and scaled by background, dsb) 处理后的数值阈值,不是原始 Count,也不是可直接迁移到其他数据或归一化方法的通用阈值。

genes2filter = ["CD3", "CD19-1", "CD14-1"]
temp = mdata["prot"][:, genes2filter].X.T.tolist()
mdata["prot"].obs["doublets_markers"] = [
    (temp[0][i] > 25 and temp[1][i] > 25) or (temp[0][i] > 25 and temp[2][i] > 25)
    for i in range(mdata.shape[0])
]
mdata["prot"].obs["doublets_markers"] = (
    mdata["prot"].obs["doublets_markers"].astype(str)
)

由于混合了多个细胞的信号,Doublet 常具有较高的原始总 Count。下面比较标记组与未标记组的 log1p 总 Count,作为辅助检查;分布差异本身不能证明每个标签都正确。

sc.pl.violin(mdata["prot"], keys="log1p_total_counts", groupby="doublets_markers")
<Figure size 372.24x320 with 1 Axes>

去除满足上述双阳性规则的观测,并同步保留 MuData 中相应的细胞。

mdata = mdata[mdata["prot"].obs["doublets_markers"] == "False"].copy()
mdata
Loading...

原教程报告按该规则去除了 608 个疑似 Doublet。

本章根据 ADT 中通常互斥的标记过滤异型 Doublet;这一方法难以识别同类型细胞组成的同型双细胞(homotypic doublet),也会受抗体面板限制。还可结合单细胞 RNA 测序(Single-Cell RNA Sequencing, scRNA-seq)方法,参见 Doublet 检测。

af_doublet_detection = ln.Artifact.from_mudata(
    mdata,
    key="surface-protein/cite_doublet_detection.h5mu",
    description="CITE-seq data after doublet detection",
)
af_doublet_detection.save()
输出
→ returning artifact with same hash: Artifact(uid='I1sQCEySclEA4f2N0005', key='surface-protein/cite_doublet_detection.h5mu', description='CITE-seq data after doublet detection', suffix='.h5mu', kind='dataset', otype='MuData', size=1452761836, hash='FhXN0hS1nI7WfLBybIFE9g', n_files=None, n_observations=105907, branch_id=1, created_on_id=1, space_id=1, storage_id=1, run_id=112, schema_id=None, created_by_id=7, created_at=2026-07-29 11:07:15 UTC, is_locked=False, version_tag=None, is_latest=True); to track this artifact as an input, use: ln.Artifact.get()
... uploading tu0VvP9QqQazgAiZ0000.h5mu: 100.0%
• replacing the existing cache path /var/cache/user/marchena/.cache/lamindb/lamin-eu-central-1/VPwcjx3CDAa2/surface-protein/cite_doublet_detection.h5mu
Artifact(uid='I1sQCEySclEA4f2N0005', key='surface-protein/cite_doublet_detection.h5mu', description='CITE-seq data after doublet detection', suffix='.h5mu', kind='dataset', otype='MuData', size=1452761836, hash='FhXN0hS1nI7WfLBybIFE9g', n_files=None, n_observations=105907, branch_id=1, created_on_id=1, space_id=1, storage_id=1, run_id=112, schema_id=None, created_by_id=7, created_at=2026-07-29 11:07:15 UTC, is_locked=False, version_tag=None, is_latest=True)
ln.finish()
输出
• please hit CTRL + s to save the notebook in your editor ... ✓
→ finished Run('33z2HhX1JxiWbGxS') after 27s at 2026-07-29 11:08:00 UTC
→ go to: https://lamin.ai/theislab/sc-best-practices/transform/xo1WtNUmrJCK0007
→ to update your notebook from the CLI, run: lamin save /groups/nils/members/javier/single-cell-best-practices/jupyter-book/surface_protein/doublet_detection.ipynb

贡献者

我们衷心感谢以下人员的贡献:

作者

  • Javier Marchena-Hurtado

  • Daniel Strobl

  • Ciro Ramírez-Suástegui

审阅者

  • Lukas Heumos

  • Anna Schaar

References
  1. Sun, B., Bugarin-Estrada, E., Overend, L. E., Walker, C. E., Tucci, F. A., & Bashford-Rogers, R. J. M. (2021). Double-jeopardy: scRNA-seq doublet/multiplet detection using multi-omic profiling. Cell Reports Methods, 1(1), 100008. 10.1016/j.crmeth.2021.100008