Skip to content

refactor(runtime)!: adopt the modern Infini stack - #506

Open
voltjia wants to merge 58 commits into
mainfrom
refactor/adopt-modern-infini-stack
Open

voltjia wants to merge 58 commits into
mainfrom
refactor/adopt-modern-infini-stack

Conversation

@voltjia

@voltjia voltjia commented Jul 21, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

  • Rebase and preserve the InfiniCore-to-InfiniLM runtime migration on InfiniLM main at 80bb09ecebc9aabf198b9b866a89456bca1df946.
  • Synchronize the effective InfiniCore changes that landed after the migration source diverged, while retaining InfiniLM's stronger graph ownership, cancellation, and allocation-lifetime behavior.
  • Replace legacy/deprecated operator paths with current InfiniRT, InfiniOps, and InfiniCCL APIs.
  • Enable the currently validated NVIDIA and Moore execution paths, including explicit FlashAttention, paged segmented graph replay, tensor parallelism, and real-weight model coverage.
  • Fix Moore communicator initialization so short-lived worker threads select the native device without constructing and tearing down thread-local InfiniCore runtimes.

Current head: 5de9a0283ec66ed30006372c535a75983407cdfb.

The branch contains 43 commits on top of current main and remains mergeable.

Related to InfiniTensor/InfiniCore#1373.

Migration and Runtime Changes

  • InfiniLM owns the migrated runtime context, tensors, graph integration, operator adapters, distributed wrappers, and Python bindings.
  • Segmented graph replay keeps capture-unsafe MhaKVCache work in host segments while the remaining operators run in device-graph segments; TP paged decode no longer falls back to a wholly eager engine.
  • Static graph cache metadata stays device-resident and observes in-place replay updates.
  • The deprecated causal-softmax backend is removed from execution; the adapter composes current triangular-mask and softmax operations.
  • Linear bias is implemented as Gemm with beta=0 plus broadcast Add, including correct row/column-parallel placement and pre-transposed weights.
  • The dense factory enables validated Baichuan, ChatGLM, FM9G, GLM4, InternLM3, Llama, MiniCPM/MiniCPM4, Qwen2, and Qwen3 families.
  • Moore accepts its native device name and selects a dedicated 23-operator InfiniOps manifest; NVIDIA retains the 24-operator manifest including sampling.
  • Communicator-init threads call InfiniRT's native backend/device selection directly before infinicclCommInitRank. This avoids heap corruption caused by destroying a short-lived InfiniCore Runtime in each worker thread.

Current Upstream Stack

#506 uses AllGather, Send, and Recv, so #57/#58/#59 remain required. #69 is an independent master-based prerequisite; validation combined its MARCH_TYPE=310 behavior with the API stack.

Newly Validated Moore Capability

InfiniOps #819/#962 provide the missing paged prefill/decode attention closure. InfiniCCL #69 passes the actual MUSA architecture into MCCL so S5000's existing BF16 collective support is visible.

The selected formal matrix passed 13/13 commands:

IDs Workloads Result
D01-D03 9g-8B explicit FlashAttention + graph, batches 1/4/16 PASS
D05 9g-8B paged attention + graph, batch 32 PASS
D14 Qwen3-32B paged FlashAttention + graph, TP4 PASS
D15 Llama-3.2-3B paged FlashAttention + graph PASS
D17 Baichuan2-7B paged FlashAttention + graph, TP2 PASS
D18 ChatGLM3-6B paged FlashAttention + graph PASS
D19 InternLM3-8B paged FlashAttention + graph PASS
D20, D22 MiniCPM4-8B generation and benchmark graph paths PASS
D21 GLM-4-9B paged FlashAttention + graph PASS
D26 MiniCPM4-8B Eagle speculative decoding PASS

Every row reported segmented graph execution with host_segments > 0; none used whole-engine eager fallback. D14 and D17 additionally prove TP4/TP2 communicator setup and BF16 AllReduce. D14 completed in 191.391s and D17 in 74.904s after the two single-purpose fixes.

The formal operator smoke also passed D64/D128 paged prefill/decode coverage: 4 passed, 8 deselected.

All selected Moore commands are greedy/default sampling. Non-greedy sampling remains gated because the Moore manifest does not yet include a supported top_k_top_p_sampling_from_logits implementation.

Preserved NVIDIA Validation

  • Real-weight two-token smokes passed for Llama-3.2-3B, FM9G 9g-8B, Baichuan2-7B, ChatGLM3-6B, InternLM3-8B, GLM-4-9B, MiniCPM4-8B, Qwen2-compatible FM9G-70B TP8, and Qwen3-0.6B.
  • Explicit FlashAttention passed eager and segmented graph paths for Qwen3 and Llama, including Qwen3-32B BF16 TP4.
  • Qwen3 TP2 paged segmented graph passed batch sizes 1 and 16 without CUDA_LAUNCH_BLOCKING.
  • Bias, pre-transposition, TP1/TP2, and paged FlashAttention combinations passed the existing focused matrix.

Gates Intentionally Retained

  • FlashInfer attention: the selected linked provider implements sampling, not an attention backend.
  • Compressed-tensors W8A8, AWQ/GPTQ, MXFP4/Quark, and INT8 KV cache: InfiniLM does not yet have complete execution chains for these formats.
  • GPT-2: the required NVIDIA LayerNorm provider is unavailable in the selected closure.
  • Mistral: the current implementation does not consume sliding_window, so a short prompt is not sufficient semantic validation.
  • Mamba, Qwen3-Next, and Qwen3.5: required causal-convolution/selective-scan/gated-delta provider chains are incomplete.
  • Complete MoE and multimodal model-specific paths remain out of scope.
  • Moore non-greedy sampling remains unsupported as described above.

Verification

  • Static contract suite: 83/83 passed.
  • Build-script unit suite: 20/20 passed.
  • Current GitHub Check Format and Ruff jobs: passed.
  • git diff --check: passed.
  • Independent Moore InfiniLM extension build: passed in 118.662s.
  • Two-device and four-device BF16 AllReduce probes: passed with correct numeric outputs.
  • InfiniCCL [DEV] 修改 lanch_server, 支持 vllm benchmark 测试补全文本 #69: explicit mixed-architecture and native-detection builds passed; CTest 2/2 in both configurations; default examples build passed; two-device BF16 AllReduce passed.
  • The InfiniOps pre-rebase hardware-tested commits map one-for-one to the current #819/#962 commits under git range-diff.

Type of Change

  • refactor
  • fix
  • test
  • docs
  • build / CI
  • breaking change

Landing Order

  1. InfiniOps #819.
  2. InfiniOps #962, then retarget it from #819 to master.
  3. InfiniCCL [DEV] 修改 lanch_server, 支持 vllm benchmark 测试补全文本 #69 can land independently.
  4. InfiniCCL 支持海光运行 #57 -> [BUG] test_ppl.py run fail #58 -> Feature/use logsoft max in ppl #59; rebase the stack onto the [DEV] 修改 lanch_server, 支持 vllm benchmark 测试补全文本 #69-updated master before updating the final component pin.
  5. InfiniCore #1406 updates its InfiniCCL and InfiniOps component pins to the merged heads.
  6. This PR.

This PR remains draft until the upstream component PRs and final pins land, but it is ready for code review against the dependency order above.

@voltjia
voltjia force-pushed the refactor/adopt-modern-infini-stack branch 3 times, most recently from 7c36e2f to 077867b Compare August 13, 2026 15:26
@voltjia
voltjia force-pushed the refactor/adopt-modern-infini-stack branch 2 times, most recently from 03d7533 to dcfbebc Compare August 31, 2026 07:07
@voltjia
voltjia force-pushed the refactor/adopt-modern-infini-stack branch from 2a44095 to 5de9a02 Compare September 2, 2026 14:17
Comment thread csrc/config/config_factory.cpp Outdated
throw std::invalid_argument("infinilm::config::ConfigFactory::createConfig: Unsupported model config type: " + model_type);
}

static const std::unordered_set<std::string> kModernModelTypes{

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

hmm

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里的第二份模型白名单确实多余,已删除。配置创建以 get_model_config_map() 注册表为准;这次恢复的是注册机制,不会把未编译、未注册的模型声明为可运行。

本轮统一修正与验证记录:#592

const std::string quant_method = quantization_config.value("quant_method", "");

// Determine the quantization scheme from the JSON config
if (quant_method == "compressed-tensors") {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

为啥在这儿就都不支持了,我感觉是不是到了算子调用再拦住比较好。不然回头把东西补回来的时候又要一串一串改。

以及之前的全量测试好像确实忘记加量化相关的东西了

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已恢复 compressed-tensors、AWQ、GPTQ、Quark/MXFP4 的配置解析和参数布局,不再在 QuantConfig 中统一拒绝。未接通的量化执行/权重处理仍在实际使用入口明确报错;新增原生配置测试覆盖各量化类型,不能据此宣称量化推理已恢复。

本轮统一修正与验证记录:#592

Comment thread csrc/config/config_factory.cpp Outdated
throw std::invalid_argument("infinilm::config::ConfigFactory::createConfig: Unsupported model config type: " + model_type);
}

static const std::unordered_set<std::string> kModernModelTypes{

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

所以现在是要求显示列举支持的模型类型了么?本来应该是不需要的

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

前面有 it = config_map.find(model_type);

如果it有值,就说明支持这个model_type。 可以不用枚举

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

不需要维护第二份模型列表。已删除 kModernModelTypes,保留已有 config_map.find(model_type) 的注册表检查,并用自定义注册模型的原生测试确认无需修改工厂即可扩展。

本轮统一修正与验证记录:#592

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同上,这样以后恢复起来岂不是很费劲。由缺失算子支持造成的问题还是建议直接暴露在算子层,而不是直接从基建里把痕迹都移除了

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已恢复 KV dtype 的配置记录和 INT8/FP16/BF16 scheme 映射,缺失能力在实际执行时报告。INT8 KV 的量化/反量化和 attention provider 仍是未完成的执行能力,配置可解析并不代表 INT8 KV 推理通过。

本轮统一修正与验证记录:#592

Comment thread csrc/engine/compiler/paged_compiler.hpp Outdated
};

std::optional<CompiledResult> compiled_short_decode_b1_;
std::optional<CompiledResult> compiled_baichuan_prefill_b1_s10_;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这都是啥呀···怎么还有这种东西

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已删除 compiled_baichuan_prefill_b1_s10_ 和按模型 profile 构建的 short-decode 分支。通用 PagedCompiler 只保留按 batch/输入兼容性选择的 decode graph。

本轮统一修正与验证记录:#592

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

怎么多了这么多东西,

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这些新增内容中,Baichuan 固定长度 prefill、ChatGLM/InternLM 专用 short-decode profile 属于超出迁移范围的特化,已全部移除。保留通用 decode 捕获、输入校验和 replay 存储更新;NVIDIA 双卡 TP 图模式已验证。

本轮统一修正与验证记录:#592

Comment thread csrc/engine/compiler/paged_compiler.hpp Outdated
};

std::optional<CompiledResult> compiled_short_decode_b1_;
std::optional<CompiledResult> compiled_baichuan_prefill_b1_s10_;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

反正这个通用文件不应该出现baichuan独占的内容

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同意,已移除通用头文件和实现中的全部 Baichuan 专属字段、环境变量、profile 与固定输入判定,使用通用 decode graph/eager 路径。

本轮统一修正与验证记录:#592

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

超能力!

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已移除这份会话产生的过时规划文档;仓库边界和构建步骤以 README/CONTRIBUTING 为准。

本轮统一修正与验证记录:#592

Comment thread csrc/cache/kv_cache.cpp
size_t num_blocks_per_layer = config.num_blocks();
size_t block_size = config.block_size();

infinicore::Shape kv_shape;

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里相当于paged attention也使用了对flash attention友好的kv cache排布,paged attention算子是否有相关修改?

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里不该修改的。
可能是之前的InfiniLM中cache创建,attn计算,cache更新,这三个行为的组织方式不好。
ai看着容易混淆了。

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

后续可以尝试调整到一起去。

static,paged, flash三个attn计算,在自己命名空间里,都各自对应了一份cache创建,attn计算,cache更新的函数。

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

是一起修改的:cache 使用 [K/V, blocks, block_size, heads, dim];paged_caching_infiniops.cc 调 ReshapeAndCacheFlash;paged_attention.cc 的 decode 转到 MhaKVCache;prefill 转到 MhaVarlen,三者消费同一布局。这里保留统一布局,若只恢复 cache 的旧排布,会与新的写入和读取接口不匹配。本轮 NVIDIA paged 两请求及 flash TP2 graph 推理已验证。

本轮统一修正与验证记录:#592

@@ -97,8 +97,8 @@ std::tuple<infinicore::Tensor, infinicore::Tensor> FlashAttentionImpl::do_kv_cac
auto k_cache_layer = kv_cache->narrow({{0, 0, 1}})->squeeze(0);
auto v_cache_layer = kv_cache->narrow({{0, 1, 1}})->squeeze(0);
infinicore::op::paged_caching_(
k_cache_layer->permute({0, 2, 1, 3}), // permute to BHSD for paged_caching_

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里是修改了paged caching算子所用的kv cache排布么?

@pengcheng888 pengcheng888 Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

那这样修改后flash attention还能说话么

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

是的,paged_caching_ 的名字保留,但实现已映射到 InfiniOps ReshapeAndCacheFlash,K/V 缓存采用 [blocks, block_size, heads, dim],与 FlashAttention 的读取一致。不是仅修改 cache 创建端;本轮 NVIDIA flash eager 和 TP2 graph 都能生成,16 个 greedy token 一致。

本轮统一修正与验证记录:#592

{1}, infinicore::DataType::kInt32,
infinicore::Device{infinicore::Device::Type::kCpu});
*reinterpret_cast<int32_t *>(last_token_shift_cpu->data()) = -1;
last_token_shift_ = last_token_shift_cpu->to(device);
}

@pengcheng888 pengcheng888 Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

is_last_pp_stage为true时,才初始化last_token_shift_ 。

那is_last_pp_stage为false时,下面也会使用,能保证各个平台last_token_shift_ 的默认值时0么?
应该为is_last_pp_stage为false时也执行,把last_token_shift_ 置0.

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

非末级 PP 不会执行到这里:forward() 在 model_->forward() 之后立即检查 !is_last_pp_stage() 并返回 {空 logits, hidden_states},last_token_shift_ 的读取在这个 return 之后。因此它在非末级保持空 Tensor 即可,不依赖任何平台把未初始化内容置零。

本轮统一修正与验证记录:#592

end_offsets, last_token_shift_);
auto packed_hidden = hidden_states->view(
{hidden_states->size(1), hidden_states->size(2)});
lm_head_input = infinicore::op::embedding(

@pengcheng888 pengcheng888 Sep 17, 2026 •

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

infinicore::op::embedding的作用是把一个token id转为一个tensor向量。

修改后的pr,通过last_token_positions调用infinicore::op::embedding筛选hidden_states。这样做超出了infinicore::op::embedding算子的能力范围。反而没有之前的 infinicore::op::select_last_token_hidden_好。

这里为什么要修改。

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

但貌似也行。
但感觉 infinicore::op::embedding出现在这里怪怪的。
如果没有比较合适的理由,还是建议用之前的 infinicore::op::select_last_token_hidden_。

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Embedding 的计算语义是按整数索引取二维表的行,并不限定表一定是词向量:这里把 hidden_states 视为 [total_tokens, hidden_dim],用 offsets[1:]-1 取各请求末行,结果正是 last-token hidden。原 select_last_token_hidden_ 属于旧算子 API,现代 InfiniOps 没有对应入口;此处复用已支持的 Add+Embedding,避免重新引入旧 InfiniOP 内核。

本轮统一修正与验证记录:#592

->view({batch_size, seq_len, num_heads_ * value_head_dim}); // [bs, seq_len, n_q_head * value_head_dim]
}

infinicore::Tensor StaticAttentionImpl::forward_graph_(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里的 StaticAttentionImpl::forward_graph_没有必要加。

之前的版本中, static attn 本身不支持graph。感觉为了强行支持,却调用了paged attention, 不太合适。

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

已移除 StaticAttentionImpl::forward_graph_ 及借用 paged cache 的实现。static attention 执行 eager;编译器不生成 static graph,直接请求捕获 static attention 时也会明确拒绝。

本轮统一修正与验证记录:#592

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

同意

@@ -40,14 +42,19 @@ infinicore::Tensor PagedAttentionImpl::forward(const AttentionLayer &layer,
const size_t value_head_dim = value->size(value->ndim() - 1);
infinicore::Tensor attn_output = infinicore::Tensor::empty({seq_len, num_heads_, value_head_dim}, query->dtype(), query->device());
if (is_prefill) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

为什么把 infinicore::op::paged_attention_prefill_修改为infinicore::op::mha_varlen_算子。

paged_attention_prefill_不再使用了么

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

paged_attention_prefill_ 仍保留为兼容入口,内部最终也是 MhaVarlen。推理引擎已有 GPU cumulative offsets 和在 CPU 计算的真实 max_query/max_sequence_length,直接传给 mha_varlen_ 可以避免每层 D2H 读取长度、重建 cumulative K。已修正该调用不再拿总 token 数/表容量当最大序列长度;缺少标量元数据时才走 paged_attention_prefill_。

本轮统一修正与验证记录:#592

Comment thread csrc/layers/mlp/mlp.cpp
// 3. Project down
// GateUpParallelLinear produces the packed [gate, up] layout expected here.
auto gate_up = gate_up_proj_->forward(hidden_states_mutable);
auto intermediate = infinicore::op::silu_and_mul(gate_up);

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里将 infinicore::op::swiglu修改为了 infinicore::op::silu_and_mul, 那 infinicore::op::silu_and_mul是各个平台都支持么?
为什么要换,是 infinicore::op::silu_and_mul算子性能更好么

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里将 infinicore::op::swiglu修改为了 infinicore::op::silu_and_mul, 那 infinicore::op::silu_and_mul是各个平台都支持么? 为什么要换,是 infinicore::op::silu_and_mul算子性能更好么

开源框架好像主流是用silu and mul

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

计算语义相同,都是 SiLU(gate)*up。标准 MLP 的 gate_up_proj 已输出连续 [gate, up],SiluAndMul 直接消费它,避免先拆分再拼接的中间拷贝;不是仅靠改名断言 kernel 更快。具体平台仍取决于已构建的 InfiniOps provider,本轮 NVIDIA eager 和 graph 数值对照测试通过,不据此宣称所有平台通过。

本轮统一修正与验证记录:#592

}

return infinicore::op::linear_w4a16_awq(input_contiguous->contiguous(), qweight, scales, qzeros, bias_opt);
throw std::runtime_error(

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

量化是都没实现么

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

量化是都没实现么

看起来是

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

当前 InfiniLM 迁移分支确实还没有完整 AWQ/GPTQ/MXFP4/W8A8 与 INT8 KV 执行链,不能把 InfiniOps 存在部分量化基础算子等同于端到端已接通。已恢复配置解析和参数布局,保留实际执行入口的明确缺失错误;本 PR 不宣称量化推理通过。

本轮统一修正与验证记录:#592

infinicore::op::distributed::allreduce_(
output, output, INFINICCL_SUM, communicator);
output, output, infinicclSum, communicator);
if (has_bias) {

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这里为什么要把has_bias单独拆出来,做一次add。 有的量化算子是可以将gemm+add一起算的。

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

这段是 RowParallel 的 all-reduce 路径:各 TP rank 产生的是局部 matmul 和,共享 bias 应在 SUM 后加一次。如果每个 rank 都先融合同一个完整 bias,归约会得到 sum(partials)+tp_size*bias。普通 forward 仍传 has_bias,可由 linear/GEMM 融合;这里分开是保证 TP 数学语义。

本轮统一修正与验证记录:#592

voltjia and others added 23 commits September 27, 2026 09:53
* feat(iluvatar): enable modern Infini stack

* feat(iluvatar): enable canonical attention adapters

* feat(iluvatar): enable flash attention backend

* fix(iluvatar): address review feedback

* fix(model-loading): mmap zip-format pytorch checkpoints

* fix(bench): reuse paged cache during warmup
Register Hygon with the canonical InfiniOps bridge, greedy sampling, RoPE cache, and FlashAttention adapters.

Extend the integration builder with Hygon architecture and RCCL wiring. Keep the platform-specific InfiniOps operator selection outside the repository and require it through --operator-config, matching the Iluvatar workflow.
* perf(runtime): make stream access constant time

* perf(mlp): consume packed gate-up output

* perf(ops): cache default infiniops implementation

* perf(paged): reuse decode metadata buffers

* perf(cache): skip unused paged cache scale upload

* perf(graph): avoid eager output snapshots

* perf(inference): optimize reviewed execution paths

* perf(speculative): reuse graphs for token verification

* perf(paged): capture ChatGLM short decode graph

* test(runtime): cover optimized inference paths

* style: format optimized inference paths

* refactor(graph): centralize reviewed profile checks
@wooway777
wooway777 force-pushed the refactor/adopt-modern-infini-stack branch from 0ea02da to d2a8504 Compare September 27, 2026 03:37
@wooway777

Copy link
Copy Markdown
Collaborator

处理如下:

  1. 已 rebase 最新 main(270feb3e),当前 head 为 d2a8504a,GitHub 显示 mergeable。
  2. mha_varlen / mha_kvcache 保持直接调用 InfiniOps 的 FlashAttnVarlenFunc / FlashAttnWithKvcache;本轮补上 dense mha 的 InfiniOps 适配,将 [B,S,H,D] 打包成 varlen 形式并保留 is_causal 语义。旧的 flash_attention_adaptor.hpp、*_flashattn.cc、mha*/hygon 旧链接实现已删除。
  3. 旧 paged_attention / paged_attention_prefill 的 public header、实现和 pybind bridge 已删除,ops.hpp 不再暴露这两个不可用入口;paged KV 路径统一走 mha* 接口。
  4. Ali 这点不能直接加白名单:当前 InfiniRT Device::Type 没有 Ali,当前 InfiniOps linked closure 也没有 Ali FlashAttention provider。旧版支持不能等价迁移;等底层两个项目补齐 device/provider 后再恢复入口,避免 InfiniLM 声称支持实际不可执行的后端。
  5. graph.hpp 中减少的 replay callback 和 host-int-array binding 是有意迁移:现在由 GraphTensor::SnapshotPolicy 管理静态元数据快照/原地回放更新,Graph 持有 runtime/allocation lease 并按 capture safety 分段;这替代了旧的 GraphReplayStage 回调和 bind_host_int_array 注册表,同时避免 host metadata 在回放期间失效。
  6. NVIDIA graph-safe FlashAttnWithKvcache provider index 修正为当前 active metadata 的 slot 16。

验证:test/static 108/108 通过;infinicore_runtime、_infinicore、_infinilm 全部构建并安装通过;9g_8b_thinking --warmup ... --enable-paged-attn --attn=flash-attn --enable-graph 端到端通过,TTFT 80.72 ms、Avg ITL 12.38 ms、Decode 323.13 tok/s。

@wooway777

wooway777 commented Sep 27, 2026 •

Copy link
Copy Markdown
Collaborator

补充:CI 使用 pip 版 clang-format 与本地 LLVM 包输出不同;已按 CI 同版本 21.1.8 追加格式化提交,当前 head 为 f94d83d,Check Format / ruff 均已通过(ci job 按仓库条件 skipping)。

@wooway777
wooway777 marked this pull request as ready for review September 28, 2026 07:49
@wooway777
wooway777 requested a review from a team September 28, 2026 07:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants