Conversation
7c36e2f to
077867b
Compare
03d7533 to
dcfbebc
Compare
2a44095 to
5de9a02
Compare
| throw std::invalid_argument("infinilm::config::ConfigFactory::createConfig: Unsupported model config type: " + model_type); | ||
| } | ||
|
|
||
| static const std::unordered_set<std::string> kModernModelTypes{ |
There was a problem hiding this comment.
这里的第二份模型白名单确实多余,已删除。配置创建以 get_model_config_map() 注册表为准;这次恢复的是注册机制,不会把未编译、未注册的模型声明为可运行。
本轮统一修正与验证记录:#592
| const std::string quant_method = quantization_config.value("quant_method", ""); | ||
|
|
||
| // Determine the quantization scheme from the JSON config | ||
| if (quant_method == "compressed-tensors") { |
There was a problem hiding this comment.
为啥在这儿就都不支持了,我感觉是不是到了算子调用再拦住比较好。不然回头把东西补回来的时候又要一串一串改。
以及之前的全量测试好像确实忘记加量化相关的东西了
There was a problem hiding this comment.
已恢复 compressed-tensors、AWQ、GPTQ、Quark/MXFP4 的配置解析和参数布局,不再在 QuantConfig 中统一拒绝。未接通的量化执行/权重处理仍在实际使用入口明确报错;新增原生配置测试覆盖各量化类型,不能据此宣称量化推理已恢复。
本轮统一修正与验证记录:#592
| throw std::invalid_argument("infinilm::config::ConfigFactory::createConfig: Unsupported model config type: " + model_type); | ||
| } | ||
|
|
||
| static const std::unordered_set<std::string> kModernModelTypes{ |
There was a problem hiding this comment.
所以现在是要求显示列举支持的模型类型了么?本来应该是不需要的
There was a problem hiding this comment.
前面有 it = config_map.find(model_type);
如果it有值,就说明支持这个model_type。 可以不用枚举
There was a problem hiding this comment.
不需要维护第二份模型列表。已删除 kModernModelTypes,保留已有 config_map.find(model_type) 的注册表检查,并用自定义注册模型的原生测试确认无需修改工厂即可扩展。
本轮统一修正与验证记录:#592
There was a problem hiding this comment.
同上,这样以后恢复起来岂不是很费劲。由缺失算子支持造成的问题还是建议直接暴露在算子层,而不是直接从基建里把痕迹都移除了
There was a problem hiding this comment.
已恢复 KV dtype 的配置记录和 INT8/FP16/BF16 scheme 映射,缺失能力在实际执行时报告。INT8 KV 的量化/反量化和 attention provider 仍是未完成的执行能力,配置可解析并不代表 INT8 KV 推理通过。
本轮统一修正与验证记录:#592
| }; | ||
|
|
||
| std::optional<CompiledResult> compiled_short_decode_b1_; | ||
| std::optional<CompiledResult> compiled_baichuan_prefill_b1_s10_; |
There was a problem hiding this comment.
已删除 compiled_baichuan_prefill_b1_s10_ 和按模型 profile 构建的 short-decode 分支。通用 PagedCompiler 只保留按 batch/输入兼容性选择的 decode graph。
本轮统一修正与验证记录:#592
There was a problem hiding this comment.
这些新增内容中,Baichuan 固定长度 prefill、ChatGLM/InternLM 专用 short-decode profile 属于超出迁移范围的特化,已全部移除。保留通用 decode 捕获、输入校验和 replay 存储更新;NVIDIA 双卡 TP 图模式已验证。
本轮统一修正与验证记录:#592
| }; | ||
|
|
||
| std::optional<CompiledResult> compiled_short_decode_b1_; | ||
| std::optional<CompiledResult> compiled_baichuan_prefill_b1_s10_; |
There was a problem hiding this comment.
反正这个通用文件不应该出现baichuan独占的内容
There was a problem hiding this comment.
同意,已移除通用头文件和实现中的全部 Baichuan 专属字段、环境变量、profile 与固定输入判定,使用通用 decode graph/eager 路径。
本轮统一修正与验证记录:#592
There was a problem hiding this comment.
已移除这份会话产生的过时规划文档;仓库边界和构建步骤以 README/CONTRIBUTING 为准。
本轮统一修正与验证记录:#592
| size_t num_blocks_per_layer = config.num_blocks(); | ||
| size_t block_size = config.block_size(); | ||
|
|
||
| infinicore::Shape kv_shape; |
There was a problem hiding this comment.
这里相当于paged attention也使用了对flash attention友好的kv cache排布,paged attention算子是否有相关修改?
There was a problem hiding this comment.
这里不该修改的。
可能是之前的InfiniLM中cache创建,attn计算,cache更新,这三个行为的组织方式不好。
ai看着容易混淆了。
There was a problem hiding this comment.
后续可以尝试调整到一起去。
static,paged, flash三个attn计算,在自己命名空间里,都各自对应了一份cache创建,attn计算,cache更新的函数。
There was a problem hiding this comment.
是一起修改的:cache 使用 [K/V, blocks, block_size, heads, dim];paged_caching_infiniops.cc 调 ReshapeAndCacheFlash;paged_attention.cc 的 decode 转到 MhaKVCache;prefill 转到 MhaVarlen,三者消费同一布局。这里保留统一布局,若只恢复 cache 的旧排布,会与新的写入和读取接口不匹配。本轮 NVIDIA paged 两请求及 flash TP2 graph 推理已验证。
本轮统一修正与验证记录:#592
| @@ -97,8 +97,8 @@ std::tuple<infinicore::Tensor, infinicore::Tensor> FlashAttentionImpl::do_kv_cac | |||
| auto k_cache_layer = kv_cache->narrow({{0, 0, 1}})->squeeze(0); | |||
| auto v_cache_layer = kv_cache->narrow({{0, 1, 1}})->squeeze(0); | |||
| infinicore::op::paged_caching_( | |||
| k_cache_layer->permute({0, 2, 1, 3}), // permute to BHSD for paged_caching_ | |||
There was a problem hiding this comment.
这里是修改了paged caching算子所用的kv cache排布么?
There was a problem hiding this comment.
那这样修改后flash attention还能说话么
There was a problem hiding this comment.
是的,paged_caching_ 的名字保留,但实现已映射到 InfiniOps ReshapeAndCacheFlash,K/V 缓存采用 [blocks, block_size, heads, dim],与 FlashAttention 的读取一致。不是仅修改 cache 创建端;本轮 NVIDIA flash eager 和 TP2 graph 都能生成,16 个 greedy token 一致。
本轮统一修正与验证记录:#592
| {1}, infinicore::DataType::kInt32, | ||
| infinicore::Device{infinicore::Device::Type::kCpu}); | ||
| *reinterpret_cast<int32_t *>(last_token_shift_cpu->data()) = -1; | ||
| last_token_shift_ = last_token_shift_cpu->to(device); | ||
| } |
There was a problem hiding this comment.
is_last_pp_stage为true时,才初始化last_token_shift_ 。
那is_last_pp_stage为false时,下面也会使用,能保证各个平台last_token_shift_ 的默认值时0么?
应该为is_last_pp_stage为false时也执行,把last_token_shift_ 置0.
There was a problem hiding this comment.
非末级 PP 不会执行到这里:forward() 在 model_->forward() 之后立即检查 !is_last_pp_stage() 并返回 {空 logits, hidden_states},last_token_shift_ 的读取在这个 return 之后。因此它在非末级保持空 Tensor 即可,不依赖任何平台把未初始化内容置零。
本轮统一修正与验证记录:#592
| end_offsets, last_token_shift_); | ||
| auto packed_hidden = hidden_states->view( | ||
| {hidden_states->size(1), hidden_states->size(2)}); | ||
| lm_head_input = infinicore::op::embedding( |
There was a problem hiding this comment.
infinicore::op::embedding的作用是把一个token id转为一个tensor向量。
修改后的pr,通过last_token_positions调用infinicore::op::embedding筛选hidden_states。这样做超出了infinicore::op::embedding算子的能力范围。反而没有之前的 infinicore::op::select_last_token_hidden_好。
这里为什么要修改。
There was a problem hiding this comment.
但貌似也行。
但感觉 infinicore::op::embedding出现在这里怪怪的。
如果没有比较合适的理由,还是建议用之前的 infinicore::op::select_last_token_hidden_。
There was a problem hiding this comment.
Embedding 的计算语义是按整数索引取二维表的行,并不限定表一定是词向量:这里把 hidden_states 视为 [total_tokens, hidden_dim],用 offsets[1:]-1 取各请求末行,结果正是 last-token hidden。原 select_last_token_hidden_ 属于旧算子 API,现代 InfiniOps 没有对应入口;此处复用已支持的 Add+Embedding,避免重新引入旧 InfiniOP 内核。
本轮统一修正与验证记录:#592
| ->view({batch_size, seq_len, num_heads_ * value_head_dim}); // [bs, seq_len, n_q_head * value_head_dim] | ||
| } | ||
|
|
||
| infinicore::Tensor StaticAttentionImpl::forward_graph_( |
There was a problem hiding this comment.
这里的 StaticAttentionImpl::forward_graph_没有必要加。
之前的版本中, static attn 本身不支持graph。感觉为了强行支持,却调用了paged attention, 不太合适。
There was a problem hiding this comment.
已移除 StaticAttentionImpl::forward_graph_ 及借用 paged cache 的实现。static attention 执行 eager;编译器不生成 static graph,直接请求捕获 static attention 时也会明确拒绝。
本轮统一修正与验证记录:#592
| @@ -40,14 +42,19 @@ infinicore::Tensor PagedAttentionImpl::forward(const AttentionLayer &layer, | |||
| const size_t value_head_dim = value->size(value->ndim() - 1); | |||
| infinicore::Tensor attn_output = infinicore::Tensor::empty({seq_len, num_heads_, value_head_dim}, query->dtype(), query->device()); | |||
| if (is_prefill) { | |||
There was a problem hiding this comment.
为什么把 infinicore::op::paged_attention_prefill_修改为infinicore::op::mha_varlen_算子。
paged_attention_prefill_不再使用了么
There was a problem hiding this comment.
paged_attention_prefill_ 仍保留为兼容入口,内部最终也是 MhaVarlen。推理引擎已有 GPU cumulative offsets 和在 CPU 计算的真实 max_query/max_sequence_length,直接传给 mha_varlen_ 可以避免每层 D2H 读取长度、重建 cumulative K。已修正该调用不再拿总 token 数/表容量当最大序列长度;缺少标量元数据时才走 paged_attention_prefill_。
本轮统一修正与验证记录:#592
| // 3. Project down | ||
| // GateUpParallelLinear produces the packed [gate, up] layout expected here. | ||
| auto gate_up = gate_up_proj_->forward(hidden_states_mutable); | ||
| auto intermediate = infinicore::op::silu_and_mul(gate_up); |
There was a problem hiding this comment.
这里将 infinicore::op::swiglu修改为了 infinicore::op::silu_and_mul, 那 infinicore::op::silu_and_mul是各个平台都支持么?
为什么要换,是 infinicore::op::silu_and_mul算子性能更好么
There was a problem hiding this comment.
这里将 infinicore::op::swiglu修改为了 infinicore::op::silu_and_mul, 那 infinicore::op::silu_and_mul是各个平台都支持么? 为什么要换,是 infinicore::op::silu_and_mul算子性能更好么
开源框架好像主流是用silu and mul
There was a problem hiding this comment.
计算语义相同,都是 SiLU(gate)*up。标准 MLP 的 gate_up_proj 已输出连续 [gate, up],SiluAndMul 直接消费它,避免先拆分再拼接的中间拷贝;不是仅靠改名断言 kernel 更快。具体平台仍取决于已构建的 InfiniOps provider,本轮 NVIDIA eager 和 graph 数值对照测试通过,不据此宣称所有平台通过。
本轮统一修正与验证记录:#592
| } | ||
|
|
||
| return infinicore::op::linear_w4a16_awq(input_contiguous->contiguous(), qweight, scales, qzeros, bias_opt); | ||
| throw std::runtime_error( |
There was a problem hiding this comment.
当前 InfiniLM 迁移分支确实还没有完整 AWQ/GPTQ/MXFP4/W8A8 与 INT8 KV 执行链,不能把 InfiniOps 存在部分量化基础算子等同于端到端已接通。已恢复配置解析和参数布局,保留实际执行入口的明确缺失错误;本 PR 不宣称量化推理通过。
本轮统一修正与验证记录:#592
| infinicore::op::distributed::allreduce_( | ||
| output, output, INFINICCL_SUM, communicator); | ||
| output, output, infinicclSum, communicator); | ||
| if (has_bias) { |
There was a problem hiding this comment.
这里为什么要把has_bias单独拆出来,做一次add。 有的量化算子是可以将gemm+add一起算的。
There was a problem hiding this comment.
这段是 RowParallel 的 all-reduce 路径:各 TP rank 产生的是局部 matmul 和,共享 bias 应在 SUM 后加一次。如果每个 rank 都先融合同一个完整 bias,归约会得到 sum(partials)+tp_size*bias。普通 forward 仍传 has_bias,可由 linear/GEMM 融合;这里分开是保证 TP 数学语义。
本轮统一修正与验证记录:#592
* feat(iluvatar): enable modern Infini stack * feat(iluvatar): enable canonical attention adapters * feat(iluvatar): enable flash attention backend * fix(iluvatar): address review feedback * fix(model-loading): mmap zip-format pytorch checkpoints * fix(bench): reuse paged cache during warmup
Register Hygon with the canonical InfiniOps bridge, greedy sampling, RoPE cache, and FlashAttention adapters. Extend the integration builder with Hygon architecture and RCCL wiring. Keep the platform-specific InfiniOps operator selection outside the repository and require it through --operator-config, matching the Iluvatar workflow.
* perf(runtime): make stream access constant time * perf(mlp): consume packed gate-up output * perf(ops): cache default infiniops implementation * perf(paged): reuse decode metadata buffers * perf(cache): skip unused paged cache scale upload * perf(graph): avoid eager output snapshots * perf(inference): optimize reviewed execution paths * perf(speculative): reuse graphs for token verification * perf(paged): capture ChatGLM short decode graph * test(runtime): cover optimized inference paths * style: format optimized inference paths * refactor(graph): centralize reviewed profile checks
0ea02da to
d2a8504
Compare
|
处理如下:
验证: |
|
补充:CI 使用 pip 版 clang-format 与本地 LLVM 包输出不同;已按 CI 同版本 21.1.8 追加格式化提交,当前 head 为 f94d83d,Check Format / ruff 均已通过(ci job 按仓库条件 skipping)。 |
Summary
mainat80bb09ecebc9aabf198b9b866a89456bca1df946.Current head:
5de9a0283ec66ed30006372c535a75983407cdfb.The branch contains 43 commits on top of current
mainand remains mergeable.Related to InfiniTensor/InfiniCore#1373.
Migration and Runtime Changes
MhaKVCachework in host segments while the remaining operators run in device-graph segments; TP paged decode no longer falls back to a wholly eager engine.beta=0plus broadcast Add, including correct row/column-parallel placement and pre-transposed weights.infinicclCommInitRank. This avoids heap corruption caused by destroying a short-lived InfiniCore Runtime in each worker thread.Current Upstream Stack
55cfe5e6761c4ebb5e8eb77b65301479a7b6032c0cdbb16967e15f2e055dea1ec9641617bf3b6cf670e50081f181d8a3c4f9a6226c8524a648ed90c4c2d76051ca21774f0b515dfd8afe59537e200ac2e8ccc0cb23ca5b1b1d63be29ba61ea031de72807f74ce8033ce601dbb0bbf31277755d033693768c#506 uses AllGather, Send, and Recv, so #57/#58/#59 remain required. #69 is an independent
master-based prerequisite; validation combined itsMARCH_TYPE=310behavior with the API stack.Newly Validated Moore Capability
InfiniOps #819/#962 provide the missing paged prefill/decode attention closure. InfiniCCL #69 passes the actual MUSA architecture into MCCL so S5000's existing BF16 collective support is visible.
The selected formal matrix passed 13/13 commands:
Every row reported segmented graph execution with
host_segments > 0; none used whole-engine eager fallback. D14 and D17 additionally prove TP4/TP2 communicator setup and BF16 AllReduce. D14 completed in 191.391s and D17 in 74.904s after the two single-purpose fixes.The formal operator smoke also passed D64/D128 paged prefill/decode coverage:
4 passed, 8 deselected.All selected Moore commands are greedy/default sampling. Non-greedy sampling remains gated because the Moore manifest does not yet include a supported
top_k_top_p_sampling_from_logitsimplementation.Preserved NVIDIA Validation
CUDA_LAUNCH_BLOCKING.Gates Intentionally Retained
sliding_window, so a short prompt is not sufficient semantic validation.Verification
git diff --check: passed.git range-diff.Type of Change
Landing Order
master.masterbefore updating the final component pin.This PR remains draft until the upstream component PRs and final pins land, but it is ready for code review against the dependency order above.