Repository navigation
Conversation
Codecov Report❌ Patch coverage is
🚀 New features to boost your workflow:
|
…nsoralloc_contract` Sparse block tensors were allocated empty and filled lazily on the heap, while `tensorfree!` released every block through the allocator. Allocate exactly the blocks that will be written instead, and pre-insert the product blocks into a sparse destination so that temporaries derived from it are complete. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- Contract into a sparse `C` through a temporary holding only the product blocks unless `C` is a valid BLAS destination, instead of copying all blocks of `C`. - Allocate missing blocks of sparse destinations lazily through the operation's allocator (as non-temporaries), so that adding a differently-patterned tensor into an allocator-backed temporary stays consistent. - Deduplicate the block keys in `tensoralloc_add`, which projects several source blocks onto the same output block for traces. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`TO.tensoralloc` returns an empty sparse tensor that remembers whether it is a temporary. Blocks added later are requested from the allocator with the matching `istemp`, and `tensorfree!` (and `zerovector!`) only release blocks of temporaries, so `tensorfree!` never receives arrays that were not allocated as temporaries. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`TO.tensoralloc` for block tensors kept the block type of the inputs, so blocks from allocators with a different storage type (e.g. `PtrArray` temporaries of `ManualAllocator`) were silently converted on insertion: the copy was passed to `tensorfree!`, and the original was never released. Infer the block type from the allocator instead, for both dense and sparse block tensors. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Replace the `istemp` flag by the allocator itself, so that kernels which do not receive an allocator (`mul!` from TensorKit's `blas_contract!`/`planarcontract!`, `trace_permute!` from `tensortrace!`) can still obtain missing blocks of a temporary through it. Allocating the blocks up front in `tensoralloc_add` and `tensoralloc_contract` is now an optimization rather than required for correctness. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Sparse block tensors were allocated empty and filled with heap-allocated blocks.
tensorfree!then gave those blocks to the allocator, which had never handed them out. After this PR,tensorfree!only receives arrays that were allocated as temporaries, and every one of them is freed.nothingfor other tensors). Any kernel that needs a missing block of a temporary, includingmul!andtrace_permute!, which TensorKit calls without an allocator, gets it from there.tensorfree!only releases the blocks of temporaries.TO.tensorallocstores blocks as the type the allocator hands out, e.g.PtrArraytemporaries fromManualAllocator. Before, those were silently copied into the input's storage type: the copy got freed and the original leaked. This applies to dense and sparse block tensors.tensoralloc_add/tensoralloc_contractallocate the blocks an operation will write up front. This is an optimisation; blocks that are still missing, e.g. in a sum like(D + 2E), are allocated on demand.tensorcontract!into aCthat needs permuting goes through a temporary that holds only the product blocks, instead of a copy of all ofC. With 256 blocks inCand 32 in the product, this goes from 329 ms to 42 ms.New tests use a strict tracking allocator that hands out temporaries as
PtrArrays. It fails on blocks that get converted, on frees of arrays that weren't handed out as temporaries, and on temporaries that are never freed.Follow-up:
@tensorallocates the temporary for a sum from its first term only. A TensorOperations hook that sees every term would let these be allocated completely up front.🤖 Generated with Claude Code