Skip to content

Reduce overhead: cached block access, SumSpace sector hashing, sparse missing entries - #83

Open
lkdvos wants to merge 12 commits into
mainfrom
optimizatio
Open

lkdvos wants to merge 12 commits into
mainfrom
optimizatio

Conversation

@lkdvos

@lkdvos lkdvos commented Oct 8, 2026 •

Copy link
Copy Markdown
Member

Performance pass targeting spurious allocations and repeated TensorKit cache lookups.

Changes

Cache lookups

  • SumSpace gets allocation-free sectorhash/sectorequal with TensorKit's "same set of sectors" meaning: equality checks mutual containment via hassector, and the hash combines the min and max sector hash. The ElementarySpace fallback built and sorted a Set of sectors on every sectorstructure lookup, which hit blocksectors, hasblock, fusiontrees and subblocks.
  • blocks(t) and foreachblock now build each sub-tensor's blocks(x) once. Before, they did two global-cache lookups per sub-tensor per sector.
  • Sparse block/subblock return small zero blocks for missing entries instead of materializing full zero tensors.
  • convert(TensorMap, t) and BlockTensorMap/SparseBlockTensorMap(t::AbstractTensorMap, W) iterate each block's subblocks once. Before, they did one structure lookup per (fusion tree, block) pair.

Generic fallbacks

  • Block-wise norm for all AbstractBlockTensorMaps (previously dense only), plus block-wise == and tr.

Copies and allocations

  • copy_blocks! replaces the scalar-indexing copyto!(::BlockMatrix, ::Matrix) in the factorization, truncation and lmul!/rmul! write-backs.
  • The truncation write-back no longer copies a destination that is fully overwritten.
  • Missing sparse entries in add!, scale!, tensoradd!, permute!/transpose!/braid! and trace_permute! are allocated uninitialized and passed β = Zero(), instead of being zero-filled and then added into. A shared _addblock!/_scale_untouched! handles this.
  • VI.add allocates with similar instead of zerovector.

Fixes found along the way

  • Changes results for p ≠ 2: norm(t, p) ignored p within each block, so norm(t, p) differed from norm(TensorMap(t), p) for p ≠ 2.
  • t[f₁, f₂] = v threw a MethodError on dense BlockTensorMaps (getindex! on an Array).
  • The adjoint-source permute!/transpose!/braid! methods referenced undefined p₁/p₂ and used the wrong permutation. They are now merged into the generic methods.
  • I removed the unreachable 8-argument add_transform! methods. TensorKit 0.17 only calls the 9-argument form.
  • braid!(::TensorMap, ::BlockTensorMap) called the deprecated TK.add_braid!.
  • Smaller fixes:
    • numout/numin(::Type{<:SparseTensorArray}) referenced an undefined variable.
    • axes(::SumSpace, ::Sector) used flatten, which isn't imported.
    • An undefined S was interpolated in the real/imag error message.

Benchmarks

U1 sum space with 3 components. tq is V⊗V←V⊗V (81 blocks) and stq a sparse version (p = 0.3); t/st are V⊗V←V. Single-threaded BLAS, minimum of 60 runs, base e4ee03c vs this branch.

op before after
blocksectors(space(tq)) 1.7 μs / 51 allocs 0.5 μs / 3
block(stq, c) 54.5 μs / 551 17.7 μs / 232
collect(blocks(tq)) 286 μs / 2419 76 μs / 423
norm(st) 63.2 μs / 747 1.1 μs / 1
st ≈ st 220 μs / 2454 15.3 μs / 220
t == t 98.4 μs / 1006 46.4 μs / 510
tr(tq) 287 μs / 2445 8.5 μs / 37
qr_compact(tq) 1012 μs / 5057 583 μs / 1872
qr_compact(stq) 1192 μs / 6335 595 μs / 1994
svd_trunc(tq; trunc = truncrank(5)) 1999 μs / 8756 1389 μs / 3037
TensorMap(tq) 469 μs / 3204 337 μs / 532
BlockTensorMap(TensorMap(tq), space(tq)) 1475 μs / 6934 399 μs / 1330
  • Sparse add/permute: unchanged at these block sizes, where per-block TensorKit overhead dominates. With 8× larger sector dimensions, in-place updates into a preallocated sparse destination drop from about 475 KB to about 4 KB per call.
  • Trivial sectors: conversions there are roughly unchanged, and TensorMap(t) is up to about 10% slower, near noise.

Tests

The full suite passes locally, Aqua included. New tests cover:

  • sparse block/blocks/foreachblock/subblock against the dense equivalent
  • SumSpace sector-hash consistency
  • norm for p ∈ {1, 2, 3, Inf}, == across dense/sparse mixes, and tr
  • factorization write-back aliasing
  • conversion round trips for U1, SU2 and Trivial sectors, including empty blocks
  • permute!/transpose!/braid! into sparse and dense destinations, with adjoint sources and β ∈ {Zero(), 0, One(), random}

🤖 Generated with Claude Code

lkdvos and others added 11 commits October 9, 2026 13:15
The ElementarySpace fallbacks built and sorted a sector Set on every
structure-cache lookup keyed by a SumSpace HomSpace. The component-wise
key is finer than necessary but sound. axes(S, c) used an unimported
flatten on an untyped vector.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
blocks(t) now builds each sub-tensor's BlockIterator once and assembles
the mortar per sector; foreachblock and == for block tensors go through
it. Missing sparse entries get a zero block of the block size instead of
a full zero tensor, also in subblock. setindex! with a block array uses
a single view per block, and no longer calls getindex! on dense parents.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- add copy_blocks! (blockwise dense -> BlockMatrix copy, avoids scalar indexing) and use it for all dense write-backs
- truncate_domain!/truncate_codomain!: do not copy the destination that is fully overwritten
- norm: one AbstractBlockTensorMap method, no temporary; also pass p to the inner norms
- add == and tr for AbstractBlockTensorMap over nonzero keys
- fix undefined S in real/imag error message

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…nalg.jl

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
`_copy_subblocks!` looped over destination fusion trees and, for each, over
all source blocks calling `v[f₁, f₂]` (a structure cache lookup) even when the
slice was empty, rebuilding `blockedrange` axes per tree. The reverse
constructors `BlockTensorMap(t, space)` / `SparseBlockTensorMap(t, space)`
indexed the block tensor per fusion tree, building a mortar over all blocks.

Both directions now share one generic body (branching on `issparse`): loop
over the blocks once, iterate each block's `subblocks` once, and copy to/from
the matching slice of the full tensor's subblock, looked up through a single
`subblocks` iterator. Per-leg block offsets are precomputed per uncoupled
sector. Trivial sectors take the cheap `subblock` path to avoid a structure
lookup per block. Sparse targets keep the previous semantics of storing every
block with a nonempty space.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Blockwise add/scale/permute/braid/transpose/tensoradd!/trace_permute! into a
SparseBlockTensorMap used to read missing entries through getindex, which
allocates a zero-filled block that is then added into. Missing entries are now
allocated uninitialized (TO.tensoralloc, honouring the allocator) and the kernel
is called with β = Zero(); present entries receive β directly. Entries of the
destination that the (sparse) source does not touch are scaled up front, or
dropped for a sparse destination when β == 0, so preallocated destinations keep
their buffers.

VI.scale!(ty, tx, α) is now add!(ty, tx, α, Zero()), and VI.add allocates its
destination with similar instead of zerovector, since every entry is overwritten.

Remove the 8-argument TK.add_transform! methods: TensorKit 0.17.2's
add_transform! takes a conjsrc argument, so these were unreachable (and called
kernels that no longer exist). Merge the AdjointTensorMap-source permute!/
transpose!/braid! methods into the generic ones; the adjoint permute!/transpose!
referenced undefined p₁/p₂ and indexed with p instead of the linear permutation.
Fix SparseTensorArray numout/numin on types (referenced undefined A), and route
braid!(::TensorMap, ::BlockTensorMap) through braid! instead of deprecated add_braid!.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Cover sparse/dense destinations with missing and present entries,
β ∈ {Zero(), 0, One(), random}, adjoint sources, fermionic sectors, and
that β = 0 overwrites NaN-filled destinations.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…s` iterators

`subblocks(t)` now caches each entry's subblock structure, mirroring `blocks(t)`;
`block`/`subblock` share the same per-entry accessors, and the conversion
code reuses them instead of its own getter.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
The component-wise versions were sound as a cache key but finer than TensorKit's
"same set of sectors" meaning, reporting e.g. `⊞(A, B)` and `⊞(B, A)` as unequal.
Equality now checks mutual sector containment via `hassector`, and the hash
combines the min and max sector hash, both set functions, without allocating.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@codecov

codecov Bot commented Oct 9, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 92.93478% with 13 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
...c/tensors/abstractblocktensor/abstracttensormap.jl 87.67% 9 Missing ⚠️
src/auxiliary/sparsetensorarray.jl 33.33% 2 Missing ⚠️
src/linalg/linalg.jl 95.65% 1 Missing ⚠️
src/tensors/indexmanipulations.jl 85.71% 1 Missing ⚠️
Files with missing lines Coverage Δ
src/auxiliary/blockarrays.jl 84.61% <100.00%> (+4.61%) ⬆️
src/linalg/factorizations.jl 96.22% <100.00%> (ø)
src/tensors/abstractblocktensor/conversion.jl 84.74% <100.00%> (+6.17%) ⬆️
src/tensors/blocktensor.jl 66.66% <100.00%> (-3.13%) ⬇️
src/tensors/sparseblocktensor.jl 68.18% <100.00%> (-0.30%) ⬇️
src/tensors/tensoroperations.jl 85.55% <100.00%> (+0.32%) ⬆️
src/tensors/vectorinterface.jl 97.82% <100.00%> (+0.45%) ⬆️
src/vectorspaces/sumspace.jl 68.29% <100.00%> (+6.60%) ⬆️
src/linalg/linalg.jl 88.53% <95.65%> (+5.80%) ⬆️
src/tensors/indexmanipulations.jl 53.42% <85.71%> (+29.89%) ⬆️
... and 2 more

... and 2 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant