# pyg_lib **Repository Path**: onescience-ai/pyg_lib ## Basic Information - **Project Name**: pyg_lib - **Description**: 本项目为基于flagos统一中间层实现的pyg_lib库 - **Primary Language**: Unknown - **License**: BSD-3-Clause - **Default Branch**: master - **Homepage**: None - **GVP Project**: No ## Statistics - **Stars**: 0 - **Forks**: 0 - **Created**: 2026-09-15 - **Last Updated**: 2026-09-21 ## Categories & Tags **Categories**: Uncategorized **Tags**: None ## README # pyg_lib 这是面向 FlagOS 的 `pyg-lib 0.9.0` FlagTree Triton 重构版本,本项目保留 `pyg_lib` 公开导入名、Python API、`torch.ops.pyg` schema 和自定义类接口。 加速设备上的 Scatter/Segment/Gather、Sampled 运算、分组矩阵乘、CSR Softmax、 图随机游走、点云邻域、Spline、Graclus 和 HashMap 使用 FlagTree Triton kernel; CPU 图采样、METIS 图划分和其他 CPU 算法保留 C++ 主机实现。 文档与项目链接: - [pyg-lib 上游项目](https://github.com/pyg-team/pyg-lib) - [PyG 文档](https://pytorch-geometric.readthedocs.io) - [英文说明](README.en.md) ## 功能概览 本库为 PyG 提供底层图计算、采样和点云算子。主要功能包括: - Scatter:`sum`、`mean`、`mul`、`min`、`max` - COO/CSR Segment 和 Gather - Sampled:`sampled_add`、`sampled_sub`、`sampled_mul`、`sampled_div` - `grouped_matmul`、`segment_matmul` - CSR Softmax 前向和反向 - `grid_cluster`、`nearest`、`knn`、`radius`、`fps` - B-Spline basis、weighting 及其梯度 - Random Walk、Graclus、设备 HashMap - CPU 同构/异构邻居采样、子图提取 - CPU METIS 图划分 离散索引、CSR 指针和 HashMap key 不参与求导。浮点输入和权重按照各算子的接口 支持一阶 Autograd;`fused_scatter_reduce` 当前仅支持前向。 ## 环境要求 - Python >= 3.10 - 已安装与目标 FlagOS 设备匹配的 PyTorch >= 2.5 - 已安装 FlagTree 扩展 Triton;不要用普通 PyPI Triton 覆盖它 - CMake >= 3.19 - 支持 C++20 的 C/C++ 编译器 - setuptools、wheel;Ninja 可选 - 运行 Python 测试需要 pytest ## 默认安装 本项目会编译 CPU 主机扩展。使用 Git 新克隆的仓库时,先初始化两个 CPU 构建依赖然后再安装: ```bash cd /path_to_dir/pyg-lib git submodule update --init --recursive FLAGOS_BACKEND=hygon MAX_JOBS=4 python -m pip install . --no-deps --no-build-isolation ``` 开发和调试推荐 editable 安装: ```bash FLAGOS_BACKEND=hygon MAX_JOBS=4 python -m pip install -e . --no-deps --no-build-isolation ``` 构建 wheel: ```bash FLAGOS_BACKEND=hygon MAX_JOBS=4 python -m pip wheel . --no-deps --no-build-isolation --wheel-dir dist/hygon python -m pip install --force-reinstall --no-deps dist/hygon/pyg_lib-0.9.0+triton.hygon-*.whl ``` 验证安装来源、版本和后端: ```bash python -c "import pyg_lib; print(pyg_lib.__version__); print(pyg_lib.backend_name()); print(pyg_lib.__file__)" ``` 首次使用新的 Triton kernel 配置时会进行 JIT 编译,因此首次调用可能较慢。 ### 按硬件后端安装 源码安装或构建 wheel 时,通过 `FLAGOS_BACKEND` 选择目标硬件: ```bash # NVIDIA(默认值;不设置变量时使用) FLAGOS_BACKEND=nvidia python -m pip install . --no-deps --no-build-isolation # Hygon/DCU FLAGOS_BACKEND=hygon python -m pip install . --no-deps --no-build-isolation # MThreads/MUSA FLAGOS_BACKEND=mthreads python -m pip install . --no-deps --no-build-isolation ``` 可选值只有 `nvidia`、`hygon` 和 `mthreads`,默认是 `nvidia`。 变量只在安装或构建阶段读取。生成的 wheel 带有后端版本标记,例如 `0.9.0+triton.hygon`;安装完成后修改环境变量不会动态切换后端。 ## 第三方源码依赖 ### parallel-hashmap `third_party/parallel-hashmap` 是 header-only C++ 库,用于保留的 CPU HashMap、 邻居采样、分布式 relabel/merge 和部分 CPU matmul 实现。 它不是 Python 包,`pip` 不会从 PyPI 自动安装一个名为 parallel-hashmap 的依赖: - 安装预构建 wheel 时不需要该目录;代码已经编译进 `libpyg.so`。 - 从本次交付的完整仓库安装时,直接使用仓库中已有的头文件。 - 从 Git 新克隆的仓库安装时,必须先运行上述 submodule 初始化命令。 - 从本项目生成的 sdist 安装时,头文件已经打入源码压缩包,无须再下载。 如果删除该目录,已经安装的 wheel 仍可运行,但以后从仓库重新编译会失败。 ### METIS 和 GKlib METIS 是 CPU 图划分库,实现 `pyg_lib.partition.metis`;GKlib 是 METIS 自带的 基础工具子模块。本项目在非 Windows 平台默认从 `third_party/METIS` 编译 64 位索引版本,并静态链接进 `libpyg.so`。 因此 METIS 同样不是由 `pip` 另行安装的 Python 依赖: - wheel 使用者不需要安装系统 METIS,也不需要保留 METIS 源码; - 源码构建需要仓库或 sdist 中完整的 METIS/GKlib; - 删除它不会立即破坏已有 wheel,但会使源码重建失败,并且无法重新生成 `pyg_lib.partition.metis` 的主机实现; - METIS 图划分始终在 CPU 执行,不是 Triton 设备 kernel。 ### 已退出的 CUDA 第三方依赖 原 CUDA 版本使用的 CUTLASS、cuCollections 和 CCCL 已经退出 CMake、源码及 wheel 构建依赖。改造仓库中同名空目录没有运行作用,可以不保留。 ## 基本用法 ### Scatter 和 Segment ```python import torch import pyg_lib src = torch.tensor([[1., 2.], [3., 4.], [5., 6.]], device="cuda") index = torch.tensor([0, 0, 1], device="cuda") out = pyg_lib.ops.scatter( src, index, dim=0, dim_size=2, reduce="sum", ) # tensor([[4., 6.], # [5., 6.]], device="cuda:0") ptr = torch.tensor([0, 2, 3], device="cuda") out_csr = pyg_lib.ops.segment_csr(src, ptr, reduce="mean") ``` ### Sampled 运算与 Autograd ```python import torch import pyg_lib left = torch.randn(8, 16, device="cuda", requires_grad=True) right = torch.randn(5, 16, device="cuda", requires_grad=True) left_index = torch.tensor([0, 0, 3, 7], device="cuda") right_index = torch.tensor([1, 1, 2, 4], device="cuda") out = pyg_lib.ops.sampled_mul( left, right, left_index, right_index, ) out.sum().backward() ``` 重复索引的梯度会累加到对应输入位置。 ### 分段矩阵乘 ```python import torch import pyg_lib x = torch.randn(8, 16, device="cuda", requires_grad=True) ptr = torch.tensor([0, 5, 8], device="cuda") weight = torch.randn(2, 16, 32, device="cuda", requires_grad=True) bias = torch.randn(2, 32, device="cuda", requires_grad=True) out = pyg_lib.ops.segment_matmul(x, ptr, weight, bias) out.mean().backward() ``` ### 点云邻域 ```python import torch import pyg_lib x = torch.rand(128, 3, device="cuda") y = torch.rand(32, 3, device="cuda") knn_edge_index = pyg_lib.ops.knn(x, y, k=8) radius_edge_index = pyg_lib.ops.radius( x, y, r=0.25, max_num_neighbors=16, ) ``` 返回张量形状为 `[2, E]`:第一行是 query 索引,第二行是 reference 索引。 ### Spline ```python import torch import pyg_lib pseudo = torch.rand(64, 2, device="cuda", requires_grad=True) kernel_size = torch.tensor([5, 5], device="cuda") is_open = torch.tensor([1, 1], dtype=torch.uint8, device="cuda") basis, weight_index = pyg_lib.ops.spline_basis( pseudo, kernel_size, is_open, degree=1, ) ``` 输入和执行都在 CPU。不同 METIS 版本或随机状态可能给出编号不同但同样合法的划分, 测试时应检查划分合法性和质量,不应固定要求某一组标签编号。 ## 正确性测试 完整 Python 测试: ```bash cd /path_to_dir/pyg-lib python -m pytest -q -rs test/ ``` ## 性能测试 运行全部迁移算子族的预热微基准: ```bash cd /path_to_dir/pyg-lib/benchmark python flag_tree.py --device 0 --warmups 5 --repeats 20 --inner-loops 20 python scale.py --device 0 --warmups 5 --repeats 20 --inner-loops 50 ```