恒美微站
首页
关于我们
建站服务
主题模板
案例展示
资讯中心
联系我们
提交pr后写代码审查稿的skill
首页
资讯中心
/
提交pr后写代码审查稿的skill
提交pr后写代码审查稿的skill
发布时间:2026/9/29 4:38:43
提交pr后写代码审查稿的skillskill原则skill正文examples 好例子可以一起提交给agentskill原则这个 Skill 的核心原则很简单不要从 diff 出发解释 PR而要从系统真实的运行链路出发理解 PR。先搞清楚任务是什么、系统怎么运行、问题断在哪里再解释为什么这些修改能把链路接回来最后用测试证明它真的有效。归根结底只有五步确定任务 → 恢复链路 → 找到断点 → 解释修改 → 证据闭环目标不是写得多而是做到具体、聚焦、低认知跳跃、可验证。skill正文在粘贴这个skill的时候发现一件非常神奇的事。因为skill内部也放了代码块所以我无法用代码块把这个skill正常显示。最后我使用了四个而不是三个来包裹没想到也能让代码块成立。外面的四个和内部的三个不一样所以就没有冲突--- name: tech-change-walkthrough description: Create a clear Markdown walkthrough for a completed technical change or PR. Reconstruct the real execution path, locate the failure points, explain why each change is correct, and connect the changes to concrete validation evidence. --- # Tech Change Walkthrough ## 1. Goal Produce a code-review walkthrough that lets a reader who did not work on the task understand: **what the task was → how the system runs → where it was wrong → why these changes fix it → how we know they work** The explanation should be **specific, focused, and easy to follow without assuming prior knowledge of the task**. --- ## 2. Method ### Step 1 — Establish the task Before writing, determine: - the task and acceptance criteria; - the PR and final commit; - the actual change scope; - what has been tested. If the core task or acceptance criteria remain materially ambiguous after inspecting the available evidence, ask the user rather than inventing an interpretation. ### Step 2 — Reconstruct the real path Start from the real entry point: a CLI command, request, API call, service startup, test, or other actual trigger. Trace it to the observable result. For each important step, identify: - the relevant file / function; - what that step does; - what data or decisions it receives and passes onward. Also establish the minimum project map needed to show **where the changed code sits in the larger system**. ### Step 3 — Find the breakpoints Locate where the old behavior diverges from the correct system behavior. Pay special attention to handoff points: Does an upstream component already know the correct fact, while the downstream code ignores it, loses it, or guesses it again? Do not force every bug into this pattern; describe the actual root cause when it is different. ### Step 4 — Group the diff into logical changes Do not explain the PR in Git diff order. Group edits by the problem they solve. For every logical change, answer: 1. **Role** — What does this code do in the execution path? 2. **Expected truth** — What information or contract should it rely on, and where does that truth come from? 3. **Old behavior** — What exactly was wrong? 4. **New behavior** — How does the change restore the correct relationship? 5. **Proof** — Who consumes the result, and what test or observable behavior shows that it works? Use real paths, functions, parameters, shapes, or values when they materially improve understanding. ### Step 5 — Close the loop with evidence Show why the logical changes together are sufficient to complete the end-to-end path. Then connect the important claims to evidence: - targeted tests; - integration or end-to-end runs; - real hardware runs; - accuracy or performance metrics; - acceptance thresholds; - CI results. Clearly distinguish **tested facts** from **reasoning based on code paths that has not been directly tested**. --- ## 3. Output Structure Use this structure unless the task clearly benefits from a small adjustment. markdown # Change — Code Review Walkthrough PR: ... Final commit: ... Scope: ... ## One-sentence summary What the PR accomplishes. **The core idea behind the solution.** ## From entry point to result Minimal execution flow showing how the feature actually runs. ### Where this change sits in the project Minimal project/module map needed to locate the changed code. ### Execution-path table | Step | File / function | What happens here | Old breakpoint | Change | |---|---|---|---|---| ## Core changes | # | Location | Before | After | What it fixes | |---|---|---|---|---| ### 1. Logical change - **Role:** ... - **Expected truth:** ... - **Before:** ... - **After:** ... - **Why this is correct:** ... - **How it is verified:** ... ### 2. ... ## Tests and results Tests, real runs, metrics, thresholds, and any untested areas. ## Reviewer questions Only include questions that are genuinely useful for this change. ## Final takeaway **One concise principle that connects the major changes.** --- ## 4. Writing Rules - Explain the **final causal model**, not the chronological debugging history. - Give the reader only the project context required to understand this change. - Do not assume the reader already understands the task, module, function, or parameter; add the smallest missing explanation where it is needed. - Prefer tables for execution chains, before/after comparisons, and test results. - Prefer small, relevant code snippets over large diffs. - Use diagrams only when they make flow or location easier to understand. - Do not repeat the same explanation in the flow diagram, project map, table, and change sections. - Never present an inferred compatibility claim as if it were directly tested. Before finishing, verify that a new reviewer can answer five questions without guessing: **What are we changing? How does this path actually run? Where does it break? Why do these changes fix it? What evidence proves it?** Use examples as a reference for depth and clarity, not as a fill-in-the-blank template.examples 好例子可以一起提交给agent# Voxtral TTS 适配 Ascend NPU —— 代码审查讲解稿大白话版traedeepseek 本文由 **Trae DeepSeek** 编写。 对应 PR[sgl-project/sglang-omni#1725](https://github.com/sgl-project/sglang-omni/pull/1725) 最终提交0e4c2f21 范围1 个提交、5 个文件4 个生产文件 1 个测试文件 --- ## 一句话总结 这个 PR 让 **Voxtral TTS 能在华为昇腾 NPU 上跑起来**单卡 多卡都能跑同时不破坏原来的 CUDA 功能。核心思路只有一句 **把以前写死的硬件假设改成问框架要正确答案。** 总共只改了 4 个生产文件52 行没有重写任何 NPU 专用代码。 --- ## 第一件事先看懂从运行命令到结果的完整链路 **这是理解后面 4 处修改的前提**。如果你只知道改了哪 4 处但不知道每一步是谁在决定、数据怎么传就不知道为什么非要改这 4 处。下面的图和表就是答案。 ### 完整链路总览 mermaid flowchart LR A[① 运行命令br/docker run --device/dev/davinci0,1br/python -m sglang_omni.cli servebr/--config voxtral_tp2.yaml] B[② 配置解析cli/serve.pybr/config.py 声明 3 个 stagebr/preprocessing → tts_generation → vocoderbr/tts_generation: gpu [0,1] tp_size 2br/vocoder: gpu 0] C[③ 平台识别核心物理设备从哪来br/sglang_omni/platforms/__init__.pybr/current_platform 检测当前环境] C --|Ascend: torch.npu.is_available()| C1[NPUOmniPlatformbr/get_device(id) → npu:id] C --|NVIDIA: is_cuda()| C2[CUDAOmniPlatformbr/get_device(id) → cuda:id] B -- D[④ mp_runner 给每个 TP 进程发身份br/gpu_id / tp_rank / tp_size / nccl_portbr/每个 rank 一个独立进程] D -- E[⑤ Stage 工厂 stages.py改动①②br/device current_platform.get_device(gpu_id)br/builder(tp_rank, tp_size, nccl_port)] E -- F[⑥ engine_builder.py改动②br/adjust_overrides → tp_size 进 ServerArgsbr/infra_kwargs → tp_rank/nccl_port] F -- G[⑦ engine_factory.build共享未改br/create_sglang_infrastructurebr/把 TP 身份送进 SGLang runtime] G -- H[⑧ SGLang runtime每 rank 一个进程br/QKVParallelLinear 把权重切成br/本卡分片如 4 卡每卡 8 个 head] H -- I[⑨ sglang_model.py改动③br/attention 按本卡 head 数算 q_size/kv_size] I -- J[⑩ audio_tokenizer.py改动④br/reshape 输出不假设内存连续] J -- K[结果有效 WAVbr/TP1 / TP2 均 1088/1088 达标] classDef fix fill:#c8e6c9,color:#1b5e20; classDef dev fill:#e1f5fe,color:#01579b; classDef shared fill:#fff3e0,color:#e65100; class E,F,I,J fix; class C,C1,C2 dev; class A,B,D,G,H,K shared; ### 设备到底是怎么被选出来的从最底层看 一句话**进程启动那一刻框架问一次这台机器是什么环境答案全局共用**。Voxtral 只需要听答案不需要自己再猜。 mermaid flowchart TD A[进程启动] -- B{sglang.srt.platforms.current_platformbr/检测当前环境只检测一次} B --|torch.npu.is_available()| C[NPUOmniPlatformbr/device_type npu] B --|is_cuda()| D[CUDAOmniPlatformbr/device_type cuda] B --|其他| E[ROCm / CPU / XPU / MUSA] C -- F[current_platform.get_device(0) → npu:0] D -- G[current_platform.get_device(0) → cuda:0] F -- H[stage 工厂把 device 交给引擎 / 声码器] G -- H style C fill:#c8e6c9,color:#1b5e20 style F fill:#c8e6c9,color:#1b5e20 style D fill:#bbdefb,color:#0d47a1 style G fill:#bbdefb,color:#0d47a1 注意docker 启动时把**物理卡** /dev/davinci0,1 映射成了容器内的**逻辑设备号** npu:0,npu:1。所以流程里所有 gpu: 0 / 1 都是逻辑卡号真正落到哪块物理卡由 docker --device 决定。 --- ### 项目地图sglang-omni 大致是什么架构/我们大概在哪里修改 一句话**多模态推理框架**核心思想是把一条多模态流水线声明式地拆成多个 stage阶段由 yaml 声明、由框架拉起进程。有的 stage 是 SGLang 引擎跑大模型有的是普通处理。 | 顶层目录 | 负责什么 | |---|---| | sglang_omni/cli/ | 命令行入口serve / config / check-gpu | | sglang_omni/config/ | 配置解析 / 合并 / 校验yaml 点号覆盖 | | sglang_omni/models/模型/ | **每个模型一个插件包我们的地盘** | | sglang_omni/pipeline/ | 流水线骨架mp_runner 起进程、coordinator、tp_control | | sglang_omni/scheduling/ | 引擎工厂 engine_factory共享不针对某个模型 | | sglang_omni/comm/ relay/ | stage 之间传数据共享内存 / NCCL / CUDA IPC | | sglang_omni/platforms/ | 硬件抽象层npu / cuda / rocm…→ current_platform | | sglang_omni/serve/ | HTTP / OpenAI 兼容 API 服务 | | sglang_omni/proto/ | 消息 / 请求协议 | **Voxtral 插件站在框架流程里的哪个位置** 把框架整条流程 Voxtral 插件的位置画出来框架负责调度骨架插件只提供器官 用户发 HTTP 请求/v1/audio/speechOpenAI 兼容 └─ serve/HTTP 服务收请求、返回音频 └─ pipeline/MultiProcessPipelineRunner → mp_runner 按 yaml 拉起 3 个进程 ├─ 进程1: preprocessing 工厂来自 models/voxtral_tts/pipeline/stages.py ├─ 进程2: tts_generation 工厂→builder→engine_factory→SGLang 引擎跑 Voxtral 模型 └─ 进程3: vocoder 工厂→加载 audio_tokenizer.py声学 token → 波形 └─ 结果走 HTTP 返回给用户 关键认知**进程管理、设备检测、引擎装配、HTTP 全是框架pipeline / scheduling / platforms / serve干的**Voxtral 插件在整个流程里只负责三件事—— ① 每个 stage 用什么工厂搭stages.py ② 模型 forward 怎么算sglang_model.py被塞进 SGLang 引擎里跑 ③ 声学 token 怎么变回波形audio_tokenizer.py。 所以 4 处改动全部落在 models/voxtral_tts/ 这个插件包里框架一行没动。 **每个模型插件 同一个骨架** 所有模型都长一样一个骨架、多个器官所以一旦你认识一个模型的布局就能快速认领任意模型 models/voxtral_tts/ ├─ config.py 声明流水线、阶段、工厂、TP 设置yaml 用到的类 ├─ pipeline/ │ ├─ stages.py 阶段工厂把进程变成执行器改动①、②在这 │ └─ engine_builder.py 引擎装配改动② ├─ sglang_model.py SGLang 里的模型定义 / attention改动③ ├─ audio_tokenizer.py 声码器改动④ └─ model_runner.py / request_builders.py / io.py ... **问题定位路径** - 根据代码架构找到 models/voxtral_tts/,理解这个项目的骨架。 - 根据关键词cuda/nccl/view → 顺着命令链路cli → config → mp_runner → stages → builder → sglang_model → audio_tokenizer验证断点 → 定位到改动点。 ### 链路表格每一步谁在决定、传什么 | 环节 | 谁在决定 / 传什么 | 改前的断点为什么跑不起来 | 本次改动 | |---|---|---|---| | ① 运行命令 | 用户docker 映射物理卡 serve 命令 | — | 无 | | ② 配置解析 | cli/serve.py 读 yaml声明 stage、gpu、tp_size | — | 无 | | ③ 平台识别 | platforms/__init__.py 检测 npu/cuda | — | 无框架已有#1306 已合入 | | ④ TP 身份下发 | mp_runner.py 给每个 rank 发 gpu_id/tp_rank/tp_size/nccl_port | — | 无框架已有 | | ⑤ Stage 工厂 | stages.py 决定 device、把 TP 信息往下传 | **vocoder 写死 cuda:{gpu_id}TP 参数不收** | **改动① ②** | | ⑥ Engine builder | engine_builder.py 把 TP 参数送进 SGLang | **builder 不接收 TP 参数信息在此中断** | **改动②** | | ⑦ 引擎装配 | 共享 engine_factory.build 创建 SGLang runtime | — | 无复用已有钩子 | | ⑧ 权重切分 | QKVParallelLinear 按 TP 切成每卡分片 | — | 无 | | ⑨ Attention | sglang_model.py 算 q_size/kv_size | **每卡还按全局 head 数算切错** | **改动③** | | ⑩ 声码器 | audio_tokenizer.py 把输出还原成 shape | **.view() 要求内存连续NPU 上会炸** | **改动④** | **为什么必须改这 4 处**上层框架③④已经把设备是什么、我是第几个 rank、我有几张卡都算好了Voxtral 的代码只要**老老实实消费这些事实**就行。但旧代码在这 4 个交接点上**自己另做了一套假设**写死 cuda、丢掉 TP 参数、用全局尺寸解释局部数据、要求内存连续于是在 NPU 上——要么拿不到正确设备要么多卡切错要么非连续张量报错。每一处修改都是在把交接点重新接回框架已有的事实。 --- ## 4 处修改 | # | 文件sglang_omni/models/voxtral_tts/ 下 | 以前写死 | 现在问框架 | 解决什么问题 | |---|------|------------|--------------|-------------| | 1 | pipeline/stages.py | 设备名写死 cuda:{gpu_id} | current_platform.get_device(gpu_id) | NPU 拿不到正确设备 | | 2 | pipeline/engine_builder.py | builder 不收多卡参数信息断开 | 收下并传给 SGLang runtime | 模型太大不能多卡分片 | | 3 | sglang_model.py | 每张卡都按全局 head 数算 | 按本卡分到的 head 数算 | 开多卡时 attention 切错 | | 4 | audio_tokenizer.py | .view() 要求内存连续 | .reshape() 只要求形状对 | NPU 上非连续张量报错 | --- ### ① stages.py设备名不再自己拼问平台要 **这个函数是干嘛的**create_generation_executor 和 create_vocoder_executor 是流水线的**开工函数工厂**。框架把某个 stage 的进程拉起来后就会调这里的工厂去搭执行器generation 工厂负责搭 SGLang 推理引擎tts_generation 阶段vocoder 工厂负责加载声码器把音频 token 还原成波形。它俩都接收框架分配的 gpu_id用来决定用哪张卡。 - **以前**vocoder 直接写死 device fcuda:{gpu_id}。 - **现在**改成 device str(current_platform.get_device(gpu_id))。 **改前 / 改后对比** python # 改前vocoder不管什么机器都当 CUDA if gpu_id is not None: device fcuda:{gpu_id} # 改后vocoder问平台 —— CUDA 机器得 cuda:3NPU 机器得 npu:3 if gpu_id is not None: device str(current_platform.get_device(gpu_id)) 效果**不用写任何 if/else 判断**。generation 那边用了同一句设备解析并且额外把 TP 参数传给 builder见 ②。 对应链路位置第 ⑤ 步。框架在第 ③ 步已经把这台机器是 npu 还是 cuda测出来了这里只是**第一次真正去用它**。 **单卡到底改了哪几处拆开看更准确** | 改动 | 是否影响单卡 | 原因 | |---|---|---| | ① 设备名 | ✅ 单卡**关键** | 这是 Voxtral 插件里**唯一**写死设备类型的地方。框架的 current_platform、mp_runner、engine_factory 早就在 NPU 上跑通了插件只是薄薄一层改掉这唯一的硬件假设单卡就搭上了框架的便车 | | ④ view→reshape | ✅ 单卡也要一行加固 | 与多卡无关是NPU 上张量可能不连续的健壮性修复CUDA 上行为完全不变 | | ② TP 参数下传 | ❌ 纯多卡 | TP1 时 tp_size1、nccl_portNone存了也没用 | | ③ 按本卡 head 数算 | ❌ 纯多卡 | TP1 时 tp_size1 → num_heads_per_tp num_heads结果和原来一模一样等于没改 | **改完 device 往哪传上下关系** - **上游谁调用它**每个 stage 是一个独立进程进程里由 stage_workers.py 按 yaml 的 factory 路径导入工厂函数 → 把 gpu_id、TP 身份等参数按签名过滤后 factory(**args) → **返回值就是本 stage 的执行器**进程再拿它跑收请求 → 执行 → 转发的循环。 - **下游device 落点** - generation 工厂device → VoxtralTtsEngineBuilder().build(model_path, devicedevice, ...) → 决定 **SGLang 引擎绑在哪张卡**返回引擎执行器。 - vocoder 工厂device → _load_audio_tokenizer(checkpoint_dir, {}, device) → 决定 **声码器加载并跑在哪张卡**包成 _VoxtralTTSVocoder(...).build_scheduler(...) 返回执行器。 - **跨 stage 传的不是 device 而是数据**preprocessing文本→token→ tts_generationtoken→声学token→ vocoder声学token→波形→ HTTP 返回。所以改 device 的效果是**让本 stage 的执行器落在正确的卡上**而不是往下一个 stage 传参数。 - **怎么验证**改前在 NPU 上跑会报张量在 npu:0 但模型绑在 cuda:0的设备不匹配错误改后两边都在 npu 上就不再报错。 **师兄追问只要改 device 就能单卡跑为什么这么简单** 因为NPU 支持这个重活**框架早就做完了**platforms/current_platform、mp_runner、engine_factory 都已在 NPU 上跑通Voxtral 插件只是薄薄一层它单卡时唯一的硬件假设就是写死的 cuda: 设备名。 严谨拆开看**单卡 ①关键 ④一行加固**②③ 是纯多卡改动TP1 时 tp_size1算出来的值跟原来一模一样数学上等价于没改。所以不是 Voxtral 简单而是**框架把路铺好了① 只是插件层的最后一颗钉子**。 ### ② engine_builder.py把多卡身份传进引擎 **这个函数是干嘛的**VoxtralTtsEngineBuilder 是 Voxtral 的**引擎装配器**继承共享基类 TtsEngineBuilder。它负责把 Voxtral 的模型配置一步步装成可用的 SGLang 引擎加载权重、搭 runtime。共享引擎工厂 engine_factory.build() 会按固定顺序调用基类定义的两个钩子adjust_overrides()往 SGLang ServerArgs 里写参数和 infra_kwargs()往基础设施创建传参。 - **以前**builder 的 __init__ 不接收任何 TP 参数上层调度第 ④ 步 mp_runner已经分好的卡rank/size/端口到 builder 这儿就**断了**。 - **现在**收下 tp_rank / tp_size / nccl_port通过框架**已有的两个钩子** adjust_overrides() 和 infra_kwargs() 传给 SGLang runtime。 **改前 / 改后对比** python # 改前不接收、不保存、也不下传任何 TP 参数 class VoxtralTtsEngineBuilder(TtsEngineBuilder): def __init__(self) - None: self.decrypted_config_file None self.voice_embeddings {} # 改后收下参数并实现基类已有的两个钩子往下传 class VoxtralTtsEngineBuilder(TtsEngineBuilder): def __init__(self, *, tp_rank0, tp_size1, nccl_portNone): ... self.tp_rank tp_rank self.tp_size tp_size self.nccl_port nccl_port def adjust_overrides(self, overrides): # 钩子1tp_size 进 ServerArgs overrides[tp_size] self.tp_size def infra_kwargs(self): # 钩子2rank / 端口 进基础设施 return {tp_rank: self.tp_rank, nccl_port: self.nccl_port} 效果模型太大单卡装不下时能分到多张卡TP上跑。 对应链路位置第 ⑤→⑥→⑦ 步。adjust_overrides 把 tp_size 写进 ServerArgsinfra_kwargs 把 tp_rank/nccl_port 传进 create_sglang_infrastructure——**接回共享的 engine_factory不是自己造新的启动路径**。 ### ③ sglang_model.pyattention 按本卡分到的头数算 **这个函数是干嘛的**VoxtralSGLangAttention 是 Voxtral 模型在 SGLang 里的**自注意力层**每张卡上都会构建一份。它的 qkv_projQKVParallelLinear按 TP 把权重切成每卡分片__init__ 里算出的 q_size/kv_size 决定 forward 时怎么把切分后的 QKV 张量拆成 Q、K、V并喂给 RadixAttention。 - **以前**每张卡都用**全局** head 数算切分尺寸 q_size num_heads * head_dim。 - **现在**用 get_attention_tp_size() 拿到卡数算每卡**本地**的头数。 **改前 / 改后对比**以 n_heads32, n_kv_heads8, TP4 为例 python # 改前每卡都按“全局 head 数”算 —— 本卡实际只有 8 个 Q、2 个 KV self.q_size self.num_heads * self.head_dim # 32*128 ✗ self.kv_size self.num_kv_heads * self.head_dim # 8*128 ✗ self.attn RadixAttention(self.num_heads, ..., num_kv_headsself.num_kv_heads, ...) # 改后按“本卡本地 head 数”算 tp_size get_attention_tp_size() # 从 runtime 拿到卡数如 4 self.num_heads_per_tp self.num_heads // tp_size # 8 self.num_kv_heads_per_tp max(1, self.num_kv_heads // tp_size) # 2 self.q_size self.num_heads_per_tp * self.head_dim # 8*128 ✓ self.kv_size self.num_kv_heads_per_tp * self.head_dim # 2*128 ✓ self.attn RadixAttention(self.num_heads_per_tp, ..., num_kv_headsself.num_kv_heads_per_tp, ...) 为什么多卡时 QKVParallelLinear第 ⑧ 步把 32 个 head 切成每卡一份但旧代码还按 32 个算切分尺寸**就切错了**。单卡时全局本地所以以前没事一开多卡就崩。这个改法和项目里 ming_tts、qwen3_omni 用的是一模一样的套路。 对应链路位置第 ⑨ 步。**输入已经是本卡分片清单也必须按本卡算**——这是用全局尺寸解释局部数据的直接逆操作。 ### ④ audio_tokenizer.pyview → reshape **这个函数是干嘛的**Attention.forward 是声码器audio tokenizer里的**自注意力前向计算**。它做完自注意力后要把输出张量 reshape 成 (batch, seqlen, 本卡head数×head_dim) 的形状再交给输出投影 wo最终还原成音频波形。view 和 reshape 目标形状完全一样区别只在内存布局的假设。 - **以前**output.view(...) 要求张量内存连续不连续就报错。 - **现在**output.reshape(...) 只要求形状对不连续时自动重排。 **改前 / 改后对比** python # 改前view 要求内存连续NPU 上非连续输出直接报错 output output.view(bsz, seqlen, self.n_local_heads * self.args.head_dim) # 改后reshape 只要求形状对必要时自动重排 output output.reshape(bsz, seqlen, self.n_local_heads * self.args.head_dim) 效果**数学结果完全一样**只是删掉了内存必须连续这个业务上不需要的假设NPU 上 attention 输出经常不连续。 对应链路位置第 ⑩ 步声码器拿 attention 输出还原 shape 的地方。 --- ## 测试与结果 - 新增 **4 组定向单测**25 个全过tests/unit_test/voxtral_tts/test_pipeline.py分别锁住上面 4 个边界。 - 昇腾 NPU 真机跑了 **1088 条 SeedTTS 数据**TP1 和 TP2 都通过WER 1.17% / 1.37%达标。 | Generation TP | 完成样本 | Corpus WER | p95 WER | WER50% | 结果 | |---:|---:|---:|---:|---:|---| | 1 | 1088/1088 | 1.172% | 9.09% | 0 | 通过 | | 2 | 1088/1088 | 1.373% | 10.00% | 0 | 通过 | 验收门槛corpus WER ≤ 1.50%、p95 ≤ 11.36%、WER50% 样本数为 0。 --- ## 一个简单的类比讲的时候用这个就够了 把模型想成**换了个厨房** - 以前菜谱写死用燃气灶cuda → 现在**每次做饭前问厨房管理员current_platform今天用哪套灶具**。 - 以前默认一个人做菜单卡 → 现在支持多人分工做菜多卡但必须把**有几个人、我是第几个、集合地点**传清楚这就是 TP 参数。 - 以前要求食材必须按固定队形摆view 要求连续 → 现在只要求总数够就行队形你自己调整reshape。 --- ## 审查时最可能被问的 3 个问题 1. **为什么不直接 cuda → npu** 因为那是把一种写死换成另一种写死还会弄坏 CUDA。设备类型本来就该由 current_platform 管。 2. **参数叫 nccl_portNPU 也用** 这是共享接口的**历史命名**只表示多进程会合的端口不是强制用 NVIDIA 的 NCCL。 3. **CUDA 会不会被改坏** 设计上不会——CUDA 也走同一个 current_platformTP1 默认路径和数值都没变。但这台机器没有 NVIDIA 卡CUDA 真机测试要等 CI 补。 --- ## 最后收束一句话记住这次改动 **设备交给平台层多卡身份交给现有 builder/runtime本卡只解释自己分到的数据张量改形状不假设内存连续。** 这四句话正好对应四个生产修改点也解释了为什么补丁小、为什么能同时保留 CUDA 语义、为什么真实昇腾 TP1/2 都能跑完 1088 条数据。