Repository navigation
Conversation
The ElementarySpace fallbacks built and sorted a sector Set on every structure-cache lookup keyed by a SumSpace HomSpace. The component-wise key is finer than necessary but sound. axes(S, c) used an unimported flatten on an untyped vector. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
blocks(t) now builds each sub-tensor's BlockIterator once and assembles the mortar per sector; foreachblock and == for block tensors go through it. Missing sparse entries get a zero block of the block size instead of a full zero tensor, also in subblock. setindex! with a block array uses a single view per block, and no longer calls getindex! on dense parents. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- add copy_blocks! (blockwise dense -> BlockMatrix copy, avoids scalar indexing) and use it for all dense write-backs - truncate_domain!/truncate_codomain!: do not copy the destination that is fully overwritten - norm: one AbstractBlockTensorMap method, no temporary; also pass p to the inner norms - add == and tr for AbstractBlockTensorMap over nonzero keys - fix undefined S in real/imag error message Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nalg.jl Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`_copy_subblocks!` looped over destination fusion trees and, for each, over all source blocks calling `v[f₁, f₂]` (a structure cache lookup) even when the slice was empty, rebuilding `blockedrange` axes per tree. The reverse constructors `BlockTensorMap(t, space)` / `SparseBlockTensorMap(t, space)` indexed the block tensor per fusion tree, building a mortar over all blocks. Both directions now share one generic body (branching on `issparse`): loop over the blocks once, iterate each block's `subblocks` once, and copy to/from the matching slice of the full tensor's subblock, looked up through a single `subblocks` iterator. Per-leg block offsets are precomputed per uncoupled sector. Trivial sectors take the cheap `subblock` path to avoid a structure lookup per block. Sparse targets keep the previous semantics of storing every block with a nonempty space. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Blockwise add/scale/permute/braid/transpose/tensoradd!/trace_permute! into a SparseBlockTensorMap used to read missing entries through getindex, which allocates a zero-filled block that is then added into. Missing entries are now allocated uninitialized (TO.tensoralloc, honouring the allocator) and the kernel is called with β = Zero(); present entries receive β directly. Entries of the destination that the (sparse) source does not touch are scaled up front, or dropped for a sparse destination when β == 0, so preallocated destinations keep their buffers. VI.scale!(ty, tx, α) is now add!(ty, tx, α, Zero()), and VI.add allocates its destination with similar instead of zerovector, since every entry is overwritten. Remove the 8-argument TK.add_transform! methods: TensorKit 0.17.2's add_transform! takes a conjsrc argument, so these were unreachable (and called kernels that no longer exist). Merge the AdjointTensorMap-source permute!/ transpose!/braid! methods into the generic ones; the adjoint permute!/transpose! referenced undefined p₁/p₂ and indexed with p instead of the linear permutation. Fix SparseTensorArray numout/numin on types (referenced undefined A), and route braid!(::TensorMap, ::BlockTensorMap) through braid! instead of deprecated add_braid!. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cover sparse/dense destinations with missing and present entries,
β ∈ {Zero(), 0, One(), random}, adjoint sources, fermionic sectors, and
that β = 0 overwrites NaN-filled destinations.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s` iterators `subblocks(t)` now caches each entry's subblock structure, mirroring `blocks(t)`; `block`/`subblock` share the same per-entry accessors, and the conversion code reuses them instead of its own getter. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The component-wise versions were sound as a cache key but finer than TensorKit's "same set of sectors" meaning, reporting e.g. `⊞(A, B)` and `⊞(B, A)` as unequal. Equality now checks mutual sector containment via `hassector`, and the hash combines the min and max sector hash, both set functions, without allocating. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Performance pass targeting spurious allocations and repeated TensorKit cache lookups.
Changes
Cache lookups
SumSpacegets allocation-freesectorhash/sectorequalwith TensorKit's "same set of sectors" meaning: equality checks mutual containment viahassector, and the hash combines the min and max sector hash. TheElementarySpacefallback built and sorted aSetof sectors on everysectorstructurelookup, which hitblocksectors,hasblock,fusiontreesandsubblocks.blocks(t)andforeachblocknow build each sub-tensor'sblocks(x)once. Before, they did two global-cache lookups per sub-tensor per sector.block/subblockreturn small zero blocks for missing entries instead of materializing full zero tensors.convert(TensorMap, t)andBlockTensorMap/SparseBlockTensorMap(t::AbstractTensorMap, W)iterate each block'ssubblocksonce. Before, they did one structure lookup per (fusion tree, block) pair.Generic fallbacks
normfor allAbstractBlockTensorMaps (previously dense only), plus block-wise==andtr.Copies and allocations
copy_blocks!replaces the scalar-indexingcopyto!(::BlockMatrix, ::Matrix)in the factorization, truncation andlmul!/rmul!write-backs.add!,scale!,tensoradd!,permute!/transpose!/braid!andtrace_permute!are allocated uninitialized and passedβ = Zero(), instead of being zero-filled and then added into. A shared_addblock!/_scale_untouched!handles this.VI.addallocates withsimilarinstead ofzerovector.Fixes found along the way
norm(t, p)ignoredpwithin each block, sonorm(t, p)differed fromnorm(TensorMap(t), p)for p ≠ 2.t[f₁, f₂] = vthrew aMethodErroron denseBlockTensorMaps (getindex!on anArray).permute!/transpose!/braid!methods referenced undefinedp₁/p₂and used the wrong permutation. They are now merged into the generic methods.add_transform!methods. TensorKit 0.17 only calls the 9-argument form.braid!(::TensorMap, ::BlockTensorMap)called the deprecatedTK.add_braid!.numout/numin(::Type{<:SparseTensorArray})referenced an undefined variable.axes(::SumSpace, ::Sector)usedflatten, which isn't imported.Swas interpolated in thereal/imagerror message.Benchmarks
U1 sum space with 3 components.
tqisV⊗V←V⊗V(81 blocks) andstqa sparse version (p = 0.3);t/stareV⊗V←V. Single-threaded BLAS, minimum of 60 runs, basee4ee03cvs this branch.blocksectors(space(tq))block(stq, c)collect(blocks(tq))norm(st)st ≈ stt == ttr(tq)qr_compact(tq)qr_compact(stq)svd_trunc(tq; trunc = truncrank(5))TensorMap(tq)BlockTensorMap(TensorMap(tq), space(tq))add/permute: unchanged at these block sizes, where per-block TensorKit overhead dominates. With 8× larger sector dimensions, in-place updates into a preallocated sparse destination drop from about 475 KB to about 4 KB per call.TensorMap(t)is up to about 10% slower, near noise.Tests
The full suite passes locally, Aqua included. New tests cover:
block/blocks/foreachblock/subblockagainst the dense equivalentSumSpacesector-hash consistencynormfor p ∈ {1, 2, 3, Inf},==across dense/sparse mixes, andtrpermute!/transpose!/braid!into sparse and dense destinations, with adjoint sources and β ∈ {Zero(), 0, One(), random}🤖 Generated with Claude Code