Repository navigation
Overhaul on error reporting of algorithms - #512
Conversation
|
After the build completes, the updated documentation will be available here |
lkdvos
left a comment
There was a problem hiding this comment.
Left some comments throughout, but as a more general idea here, since this is breaking anyways:
I think the entire idea of psi, envs, eps is probably not great to begin with, precisely because it is hard (or important) to keep that to meaning the same in every part. If we are making breaking changes anyways, it might be a convenient time to just go to a more KrylovKit-related approach where we just return an info struct, where we are then actually free to return the different quantities, and name them appropriately. This both keeps the signature the same everywhere, without unwantedly promising meaning to the number.
Thanks for taking the time to properly document many of these things by the way, this is definitely a welcome addition. There are some subtleties about the prose not lining up with the theory or the implementation, since especially for the convergence measures being practical had higher priority than being rigorous, and it seems like the language kind of mixes between the two. I'm not sure if you wanted to describe the theory or the implementation?
|
Thanks for the review! I agree I was inconsistent in separating theory vs what's effectively done in the code. I think it's beneficial to give both, so I'll see how far I get in that. I'm also in favor of the info struct, so I'll try that out! |
|
Okay, a bunch has happened, and the goal of this PR has completely shifted, but the changes are better, and it's good that they're done at once (or at least shown here bunched up, there's an argument to splitting up some parts). I introduced I expanded on the docs a bunch more as well, correcting some of the false statements I made along the way. I think it reads more clearly now what a user could expect from these errors versus what they actually get. I think I'm less wrong than last time, but there might still be mistakes 🙃 |
|
Alright, the new Concerning the unicode, I decided to keep them as aliases, but the ASCII versions are promoted everywhere, and it's also what you see when you show/display the info. The aliases are mentioned briefly in the docs and docstring of Just to give an idea of how this displays, an example for julia> eps
AlgorithmInfo:
converged = true after 6 iterations
galerkin = 6.417388341837349e-7
max_truncation_error = 1.2567011712421484e-6 (largest single factorisation)
total_truncation_error = 1.6092084215067107e-6 (quadrature over 32 truncations)I forgot to mention last time, but I haven't regenerated the docs. This should be done before merging. |
|
groundstate errors here will be fixed by #532. Otherwise I think I've added the changes I was going to make, @borisdevos would you like to review my changes, see if this still makes sense in the way that you see it? Otherwise I'll merge this tomorrow, in the hope of still being ready for an MPSKit release by Monday ish |
|
If you don't mind, I'd like to check it out tomorrow morning. I took a quick glance though, and it looks cleaner and more user-friendly (read: less overwhelming). Thanks for taking the time on this :) |
borisdevos
left a comment
There was a problem hiding this comment.
If I understand the changes correctly, they largely boil down to not supporting what I used to call ϵ_sq and numtrunc, and ϵ_max is replaced by the entire list of ϵ's per bond.
I'm completely fine with ϵ_sq being removed. I had originally added this precisely because this PR started with the time evolution changes, and I liked seeing this discarded-weight-into-norm-loss thing, but it's overkill for practically everything else.
I'm neutral about numtrunc: it's cool to see, but most people care about the error more, so dropping this is okay.
The only thing I'm a bit on the fence about is ϵ_max no longer being explicitly returned, though you can access it with the new truncation_errors. I won't block this PR being merged with this choice, though I'd like to understand why one would want to see the entire range of errors, instead of the one which really may influence convergence.
Something I just thought of now while writing this, but I think it would be nice in the case of algorithms which involve truncation to mention literally whether convergence was reached through the truncation error or the imposed convergence_measure. Maybe this is overkill, because you can read it off indirectly with the info provided, similar to how one can get to ϵ_max still, but just a thought.
| The energy variance ``\langle H^2 \rangle - \langle H \rangle^2`` is an independent and more demanding measure. | ||
| Note what it actually quantifies, namely how far the state is from being an *exact eigenstate*, which is not the same thing as the error on some other observable. | ||
| **Convergence is not accuracy.** | ||
| A converged algorithm has found a fixed point within the set of MPS of the current bond dimension, which can still be far from the true ground state: a single-site algorithm at a fixed bond dimension can converge to machine precision regardless. |
There was a problem hiding this comment.
Although I understand what this sentence is saying, I find the phrasing a bit off, mostly the first clause. Github doesn't let me suggest here, but maybe something like:
"For single-site algorithms which keep the bond dimension of the MPS fixed, one can get the ground state to converge to machine precision, yet still be far from the true ground state."
I'm being nitpicky here though, so feel free to ignore.
There was a problem hiding this comment.
I agree here, so I reworded it slightly, taking inspiration from your suggestion. I still altered it a bit because the single-site algorithms now also can do some expansion (CBE-style and post-expansion style), and just mentioned fixed bond dimension MPS manifold
| They do not shrink together, so there is a sweet spot in `dt` rather than "smaller is better". | ||
| This is because a smaller `dt` lowers the splitting error but takes more steps to reach the same time, and every step truncates again. | ||
| These controls are not independent. | ||
| With a threshold-based `trunc` such as [`truncerror`](@extref MatrixAlgebraKit.truncerror), every step can discard weight up to that threshold, however small `dt` is, and reaching a fixed final time `T` takes `T / dt` steps. |
There was a problem hiding this comment.
Maybe instead of picking out one truncation strategy, we can refer to the page referring to them all? I had previously listed them all with a brief explanation with how each of them can be useful, which I understand is overkill, though I don't like "promoting" just one of them
|
The reason I wanted to rework the truncation is mostly because I don't want to make too many choices about what the correct measure is there. The square norm would be the total truncation if everything would have been independent, which it isn't, the maximal truncation error doesn't really give all that much information, so I felt like just giving all the information and leaving it up to you to inspect. In principle this shouldn't really cost anything anyways. For the convergence measure I fully agree, and that is some of the things that would have been part of the AlgorithmsInterface stuff, but unfortunately I didn't really find the time to make that happen... |
…ns in time evolution
…rmat consistency [skip ci]
- `AlgorithmInfo` reports `truncation_errors`, the last cut at every bond, instead of the aggregated max/total/count; drop `TruncationAccumulator`, the key aliases and info combining. `time_evolve` returns the per-step history. - compare against tolerances with `<=` everywhere - `GradientGrassmann` takes `converged` from `hasconverged`; `leading_boundary` now passes the algorithm's `hasconverged`/`shouldstop` to the optimizer - fix multiline `IDMRG2` `approximate` writing the right-to-left edge update into the wrong row - docs: rewrite the errors and accuracy sections, shorten `tol` docstrings, fix dangling refs Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
[Relevant update deeper in conversation]
Description
The main motivation started with
TDVP2not reporting its truncation error. I noticed it could just make use ofgauge2!to remove duplication. While doing this, I realised multiple parts in the code didn't report their error, or didn't clarify clearly what the error actually means. In particular for the time evolution code which isn't variational, it made me realise thatϵcould mean anything. So this PR ended up expanding massively to also documenting per algorithm where relevant what the returned error represents.Details of the changes are mentioned in the changelog, and motivation for the errors in the docstrings or documentation. Importantly:
timestep/timestep!/time_evolve/time_evolve!now return(ψ, envs, ϵ).time_evolvealso logs more correctly. Tests added for this.Something I noticed along the way with
changebondsis that the meaning of its truncations differ too strongly to unify and justify returning the error. There's an argument to returning it for(VUMPS)SvdCutas the error there is genuinely a truncation error, but I didn't do that.Checklist
julia --project=test test/runtests.jl, or the relevant subset)docs/src/)[Unreleased]indocs/src/changelog.md, if this PR is user-facing (new feature, behavior change, bug fix, deprecation, or removal)