⚙️ 环境设置
安装 conda:
在创建环境之前,请确保 conda 已安装在你的系统中。
保存 yml 内容:
将 yml 选项卡中的内容保存为文件
environment.yml。
创建环境:
打开终端或命令提示符。
运行以下命令:
conda env create -f environment.yml
激活环境:
创建好环境后,使用以下命令激活它:
conda activate <environment_name>请将
<environment_name>替换为environment.yml文件中指定的环境名称。该名称在 yml 文件中如下所示:name: <environment_name>
验证安装:
通过运行以下命令,检查环境是否创建成功:
conda env list
name: surface-protein
channels:
- conda-forge
dependencies:
- python=3.13
- scanpy=1.12
- muon=0.1.9
- python-igraph=1.0.0
- ipykernel=7.2.0
- pip==26.0.1
- pip:
- lamindb==2.3.1
- harmonypy==0.0.9
- ipywidgets==8.1.8
🗄️ 获取数据和笔记本
本书使用 lamindb 存储、共享和加载数据集与笔记本,托管实例为 theislab
安装 lamindb
安装 lamindb Python 软件包:
pip install lamindb可选择创建 Lamin 账户
按照 说明注册并登录
验证你的设置
下面用 Python API 检查连接;命令行方式可使用
lamin connect:
import lamindb as ln ln.Artifact.connect("theislab/sc-best-practices").df()运行后应显示最多 100 条已保存的数据集记录。
访问数据集(Artifact)
加载一个 Artifact 及其对应的对象:
import lamindb as ln af = ln.Artifact.connect("theislab/sc-best-practices").get(key="key_of_dataset", is_latest=True) obj = af.load()加载后的对象可直接在内存中分析。如需指定数据版本,可调整
lamindb.Artifact.connect("theislab/sc-best-practices").get("SOMEIDXXXX")中的 ID。访问笔记本(Transform)
加载笔记本:
lamin load <notebook url>该命令将笔记本下载到当前工作目录。与
Artifacts类似,可通过指定 ID 获取旧版本。
研究动机¶
在 质量控制 章节中,我们根据异常高的计数(Count)及样本内分布过滤了部分疑似低质量观测。本章进一步利用抗体衍生标签(Antibody-Derived Tags, ADTs)识别不同细胞类型共占液滴形成的异型双细胞(heterotypic Doublet)。其依据是正常情况下相互排斥的细胞类型标记同时出现 Sun et al., 2021。
环境设置¶
import warnings
import scanpy as sc
warnings.filterwarnings("ignore")
sc.settings.verbosity = 0
sc.set_figure_params(
dpi=80,
facecolor="white",
frameon=False,
)
import lamindb as ln
ln.track()→ connected lamindb: theislab/sc-best-practices
→ found notebook doublet_detection.ipynb, making new version -- anticipating changes
→ created Transform('xo1WtNUmrJCK0007', key='doublet_detection.ipynb'), re-started Run('33z2HhX1JxiWbGxS') at 2026-07-29 11:07:33 UTC
→ notebook imports: lamindb-core==2.3.1 muon==0.1.9 scanpy==1.12.3
• recommendation: to identify the notebook across renames, pass the uid: ln.track("xo1WtNUmrJCK")
加载数据¶
加载前一章保存的 MuData 对象,参见 归一化(normalization):
af = ln.Artifact.connect("theislab/sc-best-practices").get(
key="surface-protein/cite_normalization.h5mu", is_latest=True
)
mdata = af.load()
mdata根据细胞类型标记识别 Doublet¶
先比较通常互斥的表面标记,例如 T 细胞标记 CD3 与 B 细胞标记 CD19。两者同时显示较强信号时,可提示 T/B 细胞组成的 双细胞(Doublet)。不过,背景结合、细胞间蛋白转移或异常生物学状态也可能产生非典型共表达,因此不能仅凭两个标记确证 Doublet。
同样,可比较 T 细胞标记 CD3 与单核细胞标记 CD14,寻找疑似 T 细胞–单核细胞 Doublet。
sc.pl.scatter(mdata["prot"], x="CD3", y="CD19-1", color="log1p_total_counts")
散点图左下方的观测在两个标记上信号都低;左上和右下主要为单一标记阳性,右上则为两个标记信号均较高的观测。
本例将右上方的双阳性观测作为疑似 Doublet,进一步设置阈值过滤。
接着绘制 CD3 与 CD14,检查另一组互斥标记。
sc.pl.scatter(mdata["prot"], x="CD3", y="CD14-1", color="log1p_total_counts")
根据本例图形,将 CD3 与 CD19-1 均大于 25,或 CD3 与 CD14-1 均大于 25 的观测标记为疑似 Doublet。这里的 25 是前一章 背景去噪与缩放(denoised and scaled by background, dsb) 处理后的数值阈值,不是原始 Count,也不是可直接迁移到其他数据或归一化方法的通用阈值。
genes2filter = ["CD3", "CD19-1", "CD14-1"]
temp = mdata["prot"][:, genes2filter].X.T.tolist()mdata["prot"].obs["doublets_markers"] = [
(temp[0][i] > 25 and temp[1][i] > 25) or (temp[0][i] > 25 and temp[2][i] > 25)
for i in range(mdata.shape[0])
]
mdata["prot"].obs["doublets_markers"] = (
mdata["prot"].obs["doublets_markers"].astype(str)
)由于混合了多个细胞的信号,Doublet 常具有较高的原始总 Count。下面比较标记组与未标记组的 log1p 总 Count,作为辅助检查;分布差异本身不能证明每个标签都正确。
sc.pl.violin(mdata["prot"], keys="log1p_total_counts", groupby="doublets_markers")
去除满足上述双阳性规则的观测,并同步保留 MuData 中相应的细胞。
mdata = mdata[mdata["prot"].obs["doublets_markers"] == "False"].copy()
mdata原教程报告按该规则去除了 608 个疑似 Doublet。
本章根据 ADT 中通常互斥的标记过滤异型 Doublet;这一方法难以识别同类型细胞组成的同型双细胞(homotypic doublet),也会受抗体面板限制。还可结合单细胞 RNA 测序(Single-Cell RNA Sequencing, scRNA-seq)方法,参见 Doublet 检测。
af_doublet_detection = ln.Artifact.from_mudata(
mdata,
key="surface-protein/cite_doublet_detection.h5mu",
description="CITE-seq data after doublet detection",
)
af_doublet_detection.save()输出
→ returning artifact with same hash: Artifact(uid='I1sQCEySclEA4f2N0005', key='surface-protein/cite_doublet_detection.h5mu', description='CITE-seq data after doublet detection', suffix='.h5mu', kind='dataset', otype='MuData', size=1452761836, hash='FhXN0hS1nI7WfLBybIFE9g', n_files=None, n_observations=105907, branch_id=1, created_on_id=1, space_id=1, storage_id=1, run_id=112, schema_id=None, created_by_id=7, created_at=2026-07-29 11:07:15 UTC, is_locked=False, version_tag=None, is_latest=True); to track this artifact as an input, use: ln.Artifact.get()
... uploading tu0VvP9QqQazgAiZ0000.h5mu: 100.0%
• replacing the existing cache path /var/cache/user/marchena/.cache/lamindb/lamin-eu-central-1/VPwcjx3CDAa2/surface-protein/cite_doublet_detection.h5mu
Artifact(uid='I1sQCEySclEA4f2N0005', key='surface-protein/cite_doublet_detection.h5mu', description='CITE-seq data after doublet detection', suffix='.h5mu', kind='dataset', otype='MuData', size=1452761836, hash='FhXN0hS1nI7WfLBybIFE9g', n_files=None, n_observations=105907, branch_id=1, created_on_id=1, space_id=1, storage_id=1, run_id=112, schema_id=None, created_by_id=7, created_at=2026-07-29 11:07:15 UTC, is_locked=False, version_tag=None, is_latest=True)ln.finish()输出
• please hit CTRL + s to save the notebook in your editor ... ✓
→ finished Run('33z2HhX1JxiWbGxS') after 27s at 2026-07-29 11:08:00 UTC
→ go to: https://lamin.ai/theislab/sc-best-practices/transform/xo1WtNUmrJCK0007
→ to update your notebook from the CLI, run: lamin save /groups/nils/members/javier/single-cell-best-practices/jupyter-book/surface_protein/doublet_detection.ipynb
贡献者¶
我们衷心感谢以下人员的贡献:
作者¶
Javier Marchena-Hurtado
Daniel Strobl
Ciro Ramírez-Suástegui
审阅者¶
Lukas Heumos
Anna Schaar
- Sun, B., Bugarin-Estrada, E., Overend, L. E., Walker, C. E., Tucci, F. A., & Bashford-Rogers, R. J. M. (2021). Double-jeopardy: scRNA-seq doublet/multiplet detection using multi-omic profiling. Cell Reports Methods, 1(1), 100008. 10.1016/j.crmeth.2021.100008