Skip to content

Overhaul on error reporting of algorithms - #512

Merged
lkdvos merged 35 commits into
mainfrom
bd/tdvp-errors
Oct 2, 2026
Merged

lkdvos merged 35 commits into
mainfrom
bd/tdvp-errors

Conversation

@borisdevos

@borisdevos borisdevos commented Aug 14, 2026 •

Copy link
Copy Markdown
Member

[Relevant update deeper in conversation]

Description

The main motivation started with TDVP2 not reporting its truncation error. I noticed it could just make use of gauge2! to remove duplication. While doing this, I realised multiple parts in the code didn't report their error, or didn't clarify clearly what the error actually means. In particular for the time evolution code which isn't variational, it made me realise that ϵ could mean anything. So this PR ended up expanding massively to also documenting per algorithm where relevant what the returned error represents.

Details of the changes are mentioned in the changelog, and motivation for the errors in the docstrings or documentation. Importantly:

  • timestep/timestep!/time_evolve/time_evolve! now return (ψ, envs, ϵ). time_evolve also logs more correctly. Tests added for this.
  • Documentation of what every reported error actually means.

Something I noticed along the way with changebonds is that the meaning of its truncations differ too strongly to unify and justify returning the error. There's an argument to returning it for (VUMPS)SvdCut as the error there is genuinely a truncation error, but I didn't do that.

Checklist

  • Tests pass locally (julia --project=test test/runtests.jl, or the relevant subset)
  • Documentation updated, if this PR changes public API (docstrings, docs/src/)
  • Runic formatter is run
  • Changelog entry added under [Unreleased] in docs/src/changelog.md, if this PR is user-facing (new feature, behavior change, bug fix, deprecation, or removal)

@borisdevos borisdevos added the documentation Improvements or additions to documentation label Aug 14, 2026
@github-actions

Copy link
Copy Markdown
Contributor

After the build completes, the updated documentation will be available here

@lkdvos lkdvos left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Left some comments throughout, but as a more general idea here, since this is breaking anyways:

I think the entire idea of psi, envs, eps is probably not great to begin with, precisely because it is hard (or important) to keep that to meaning the same in every part. If we are making breaking changes anyways, it might be a convenient time to just go to a more KrylovKit-related approach where we just return an info struct, where we are then actually free to return the different quantities, and name them appropriately. This both keeps the signature the same everywhere, without unwantedly promising meaning to the number.

Thanks for taking the time to properly document many of these things by the way, this is definitely a welcome addition. There are some subtleties about the prose not lining up with the theory or the implementation, since especially for the convergence measures being practical had higher priority than being rigorous, and it seems like the language kind of mixes between the two. I'm not sure if you wanted to describe the theory or the implementation?

Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread docs/src/man/algorithms.md Outdated
Comment thread src/algorithms/groundstate/dmrg.jl Outdated
Comment thread src/algorithms/timestep/bug.jl Outdated
@codecov

codecov Bot commented Aug 14, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 56.50000% with 87 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
src/utility/algorithminfo.jl 37.93% 36 Missing ⚠️
src/algorithms/approximate/idmrg.jl 14.28% 18 Missing ⚠️
src/algorithms/statmech/idmrg.jl 0.00% 11 Missing ⚠️
src/algorithms/approximate/zipup.jl 0.00% 10 Missing ⚠️
src/algorithms/statmech/gradient_grassmann.jl 0.00% 3 Missing ⚠️
src/algorithms/approximate/approximate.jl 50.00% 2 Missing ⚠️
src/algorithms/approximate/vomps.jl 0.00% 2 Missing ⚠️
src/algorithms/changebonds/svdcut.jl 71.42% 2 Missing ⚠️
src/algorithms/statmech/leading_boundary.jl 50.00% 2 Missing ⚠️
src/algorithms/statmech/vomps.jl 66.66% 1 Missing ⚠️
Files with missing lines Coverage Δ
src/MPSKit.jl 100.00% <ø> (ø)
src/algorithms/approximate/fvomps.jl 96.00% <100.00%> (+0.16%) ⬆️
src/algorithms/derivatives/mpo_derivatives.jl 74.21% <ø> (ø)
src/algorithms/groundstate/dmrg.jl 91.17% <100.00%> (+0.36%) ⬆️
src/algorithms/groundstate/find_groundstate.jl 76.92% <ø> (ø)
src/algorithms/groundstate/gradient_grassmann.jl 89.47% <100.00%> (+1.23%) ⬆️
src/algorithms/groundstate/idmrg.jl 96.89% <100.00%> (+0.05%) ⬆️
src/algorithms/groundstate/vumps.jl 100.00% <100.00%> (ø)
src/algorithms/propagator/corvector.jl 0.00% <ø> (ø)
src/algorithms/timestep/bug.jl 93.24% <100.00%> (+6.10%) ⬆️
... and 16 more
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@borisdevos

Copy link
Copy Markdown
Member Author

Thanks for the review! I agree I was inconsistent in separating theory vs what's effectively done in the code. I think it's beneficial to give both, so I'll see how far I get in that. I'm also in favor of the info struct, so I'll try that out!

@borisdevos borisdevos changed the title Report truncation error in time evolution + clarify the role of every returned error Overhaul on error reporting of algorithms Aug 20, 2026
@borisdevos

Copy link
Copy Markdown
Member Author

Okay, a bunch has happened, and the goal of this PR has completely shifted, but the changes are better, and it's good that they're done at once (or at least shown here bunched up, there's an argument to splitting up some parts).

I introduced AlgorithmInfo as the info struct, done in such a way that every algorithm that can return an error does it via this one struct. Along the way I actually found more algorithms than TDVP2 which calculated truncation errors, but didn't return them. So I think I found them all now. I think the interface is clean, but do let me know if anything's weird about it.

I expanded on the docs a bunch more as well, correcting some of the false statements I made along the way. I think it reads more clearly now what a user could expect from these errors versus what they actually get. I think I'm less wrong than last time, but there might still be mistakes 🙃

Comment thread src/utility/algorithminfo.jl Outdated
Comment thread src/utility/algorithminfo.jl Outdated
@borisdevos

borisdevos commented Aug 28, 2026 •

Copy link
Copy Markdown
Member Author

Alright, the new AlgorithmInfo with a Dict field is now available! The most important related change is that the old info.normres is now replaced by certain "convergence keys", which I've managed to fix to 4 that occur in the current algorithms. Of course, this can be expanded, and it should be such that the new convergence_measure function can call it (as well as for pretty printing). I think the names I chose there make sense, the only one I have some doubt for is :localchange for the finite VOMPS code.

Concerning the unicode, I decided to keep them as aliases, but the ASCII versions are promoted everywhere, and it's also what you see when you show/display the info. The aliases are mentioned briefly in the docs and docstring of AlgorithmInfo.

Just to give an idea of how this displays, an example for DMRG2:

julia> eps
AlgorithmInfo:
  converged              = true after 6 iterations
  galerkin               = 6.417388341837349e-7
  max_truncation_error   = 1.2567011712421484e-6        (largest single factorisation)
  total_truncation_error = 1.6092084215067107e-6        (quadrature over 32 truncations)

I forgot to mention last time, but I haven't regenerated the docs. This should be done before merging.

Comment thread docs/src/man/algorithms.md Outdated
@lkdvos lkdvos mentioned this pull request Sep 24, 2026
7 tasks done
@lkdvos

lkdvos commented Oct 1, 2026 •

Copy link
Copy Markdown
Member

groundstate errors here will be fixed by #532.

Otherwise I think I've added the changes I was going to make, @borisdevos would you like to review my changes, see if this still makes sense in the way that you see it? Otherwise I'll merge this tomorrow, in the hope of still being ready for an MPSKit release by Monday ish

@borisdevos

Copy link
Copy Markdown
Member Author

If you don't mind, I'd like to check it out tomorrow morning. I took a quick glance though, and it looks cleaner and more user-friendly (read: less overwhelming). Thanks for taking the time on this :)

@borisdevos borisdevos left a comment

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

If I understand the changes correctly, they largely boil down to not supporting what I used to call ϵ_sq and numtrunc, and ϵ_max is replaced by the entire list of ϵ's per bond.

I'm completely fine with ϵ_sq being removed. I had originally added this precisely because this PR started with the time evolution changes, and I liked seeing this discarded-weight-into-norm-loss thing, but it's overkill for practically everything else.
I'm neutral about numtrunc: it's cool to see, but most people care about the error more, so dropping this is okay.
The only thing I'm a bit on the fence about is ϵ_max no longer being explicitly returned, though you can access it with the new truncation_errors. I won't block this PR being merged with this choice, though I'd like to understand why one would want to see the entire range of errors, instead of the one which really may influence convergence.

Something I just thought of now while writing this, but I think it would be nice in the case of algorithms which involve truncation to mention literally whether convergence was reached through the truncation error or the imposed convergence_measure. Maybe this is overkill, because you can read it off indirectly with the info provided, similar to how one can get to ϵ_max still, but just a thought.

Comment thread docs/src/man/algorithms.md Outdated
The energy variance ``\langle H^2 \rangle - \langle H \rangle^2`` is an independent and more demanding measure.
Note what it actually quantifies, namely how far the state is from being an *exact eigenstate*, which is not the same thing as the error on some other observable.
**Convergence is not accuracy.**
A converged algorithm has found a fixed point within the set of MPS of the current bond dimension, which can still be far from the true ground state: a single-site algorithm at a fixed bond dimension can converge to machine precision regardless.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Although I understand what this sentence is saying, I find the phrasing a bit off, mostly the first clause. Github doesn't let me suggest here, but maybe something like:
"For single-site algorithms which keep the bond dimension of the MPS fixed, one can get the ground state to converge to machine precision, yet still be far from the true ground state."
I'm being nitpicky here though, so feel free to ignore.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree here, so I reworded it slightly, taking inspiration from your suggestion. I still altered it a bit because the single-site algorithms now also can do some expansion (CBE-style and post-expansion style), and just mentioned fixed bond dimension MPS manifold

Comment thread docs/src/man/algorithms.md Outdated
They do not shrink together, so there is a sweet spot in `dt` rather than "smaller is better".
This is because a smaller `dt` lowers the splitting error but takes more steps to reach the same time, and every step truncates again.
These controls are not independent.
With a threshold-based `trunc` such as [`truncerror`](@extref MatrixAlgebraKit.truncerror), every step can discard weight up to that threshold, however small `dt` is, and reaching a fixed final time `T` takes `T / dt` steps.

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Maybe instead of picking out one truncation strategy, we can refer to the page referring to them all? I had previously listed them all with a brief explanation with how each of them can be useful, which I understand is overkill, though I don't like "promoting" just one of them

Comment thread src/algorithms/timestep/time_evolve.jl Outdated
@lkdvos

lkdvos commented Oct 2, 2026

Copy link
Copy Markdown
Member

The reason I wanted to rework the truncation is mostly because I don't want to make too many choices about what the correct measure is there. The square norm would be the total truncation if everything would have been independent, which it isn't, the maximal truncation error doesn't really give all that much information, so I felt like just giving all the information and leaving it up to you to inspect. In principle this shouldn't really cost anything anyways.

For the convergence measure I fully agree, and that is some of the things that would have been part of the AlgorithmsInterface stuff, but unfortunately I didn't really find the time to make that happen...

borisdevos and others added 21 commits October 2, 2026 10:34
- `AlgorithmInfo` reports `truncation_errors`, the last cut at every bond, instead of the
  aggregated max/total/count; drop `TruncationAccumulator`, the key aliases and info combining.
  `time_evolve` returns the per-step history.
- compare against tolerances with `<=` everywhere
- `GradientGrassmann` takes `converged` from `hasconverged`; `leading_boundary` now passes the
  algorithm's `hasconverged`/`shouldstop` to the optimizer
- fix multiline `IDMRG2` `approximate` writing the right-to-left edge update into the wrong row
- docs: rewrite the errors and accuracy sections, shorten `tol` docstrings, fix dangling refs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@lkdvos
lkdvos enabled auto-merge (squash) October 2, 2026 15:27
@lkdvos
lkdvos merged commit 78ef9c1 into main Oct 2, 2026
58 of 70 checks passed
@lkdvos
lkdvos deleted the bd/tdvp-errors branch October 2, 2026 17:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

documentation Improvements or additions to documentation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants