Repository navigation
Conversation
|
Fix-up for the Moore build/runtime failures in \19c6490d:\n\n- Load TVM-FFI's DLPack header before ATen's vendored DLPack header, so the TVM-FFI ABI sees \DLManagedTensorVersioned\ and the current \DLTensor\ layout.\n- Keep ATen/MATE runtime state in each .cc\ via a forward-declared \MateState; generated MUSA dispatch units therefore only consume a lightweight provider declaration.\n- Use PyTorch's official \�t::Tensor\ pybind caster, mark the MATE loader lambdas mutable, and convert external softmax-LSE views before handing them to MATE.\n\nValidated on the new Moore S5000 host using Torch 2.7.1, MATE 0.2.5, and \�pache_tvm_ffi 0.1.9.post3+musa.1. I first built current InfiniRT master with \WITH_CPU=ON -DWITH_MOORE=ON, then built this PR with \WITH_LINKED=ON\ and the two FlashAttention operators selected. Compilation and linking passed.\n\nExisting tests, no new smoke tests: \pytest tests/test_flash_attn_varlen_func.py tests/test_flash_attn_with_kvcache.py --devices moore -k 'musa-16 and (q_lens0 or dense or paged)'\ -> 13 passed, 5 skipped, 94 deselected. The skipped cases are existing capability-specific skips. |
Summary
__tvm_ffi_*entry points.Validation
This is a draft while Moore hardware validation is in progress.