Status: Superseded by ADR-014
Original decision (historical): Use PyQt5 5.15.11. Chosen because the upstream project used PyQt5, PyQt5's ecosystem was more mature at the time, and migration carried risk.
Superseding decision: The project migrated to PyQt6 6.7+ in the same PR that introduced in-process AI inference. See ADR-014 for the rationale (mainly: PyQt6 eliminated the WinError 1114 DLL load-order conflict that motivated ADR-011, unblocking the subprocess removal in ADR-013).
Status: Accepted
Context: Need to integrate Segment Anything Model 2 for semi-automated annotation
Decision: Use Ultralytics library instead of direct SAM2 installation
Rationale:
- Simplifies SAM model loading (single line)
- Includes PyTorch dependencies
- Automatic model caching
- No manual model download required
- Supports both SAM 2.0 and SAM 2.1 variants
Consequences:
- ✅ Simplified installation (no separate SAM2 setup)
- ✅ Automatic model management
- ✅ Consistent API
⚠️ Dependency on Ultralytics release cycle
Status: Accepted
Context: Project files need to reference image locations
Decision: Store absolute paths to images in project JSON
Rationale:
- Images can be anywhere on filesystem
- No requirement to keep images with project file
- Simplifies project structure
Consequences:
- ✅ Flexible image locations
- ❌ Projects not portable between machines
- ❌ Moving images breaks projects
Mitigation: Export functions copy images to output directory
Status: Accepted (Technical Debt)
Context: Application is GUI-heavy with complex interactions
Decision: Rely on manual testing only
Rationale:
- PyQt testing requires significant setup (pytest-qt, fixtures)
- Visual nature of tool makes automated testing difficult
- Small development team
- Rapid iteration on features
Consequences:
- ❌ Risk of regressions
- ❌ Manual testing required for all changes
- ❌ Slower development velocity for large refactors
- ✅ Lower initial development overhead
Future Consideration: Add unit tests for utility functions (calculate_area, conversions)
Status: Accepted
Context: Projects were getting corrupted when application terminated during loading (v0.8.9 bug)
Decision: Set is_loading_project flag to disable autosave during load
Rationale:
- Autosave triggered with partially loaded state corrupts file
- Loading large projects is slow, increases risk
- Simple flag prevents the issue
Consequences:
- ✅ Prevents project corruption
- ✅ Minimal code change
⚠️ Users lose autosave protection during load window
Status: Accepted
Context: Need to merge, validate, and manipulate polygon geometries
Decision: Use Shapely library for all polygon operations
Rationale:
- Industry-standard computational geometry library
- Handles invalid polygons gracefully
- Efficient union/intersection operations
- Well-tested algorithms
Consequences:
- ✅ Robust polygon handling
- ✅ Easy merge operations
- ✅ Automatic polygon validation
⚠️ Additional dependency
Status: Accepted
Context: Need to store polygon annotations
Decision: Store as flattened list [x1, y1, x2, y2, ...] instead of nested [[x1, y1], [x2, y2], ...]
Rationale:
- Compatible with COCO JSON format
- Smaller file size
- Standard in annotation tools
Consequences:
- ✅ COCO compatibility
- ✅ Compact representation
⚠️ Must convert to/from paired format for some operations
Status: Accepted
Context: Users need annotations in different formats for various ML frameworks
Decision: Implement exporters for COCO, YOLO, Pascal VOC, labeled images, semantic labels
Rationale:
- Different frameworks have different input requirements
- YOLO and COCO are most common
- Labeled images useful for visual verification
- Semantic labels needed for segmentation models
Consequences:
- ✅ Wide compatibility
- ✅ Flexible workflow
⚠️ More code to maintain⚠️ Must keep up with format changes (e.g., YOLOv11)
Status: Accepted
Context: TIFF stacks and CZI files have multiple slices that need individual annotations
Decision: Store annotations per slice with naming convention {filename}_T{t}_Z{z}_C{c}
Rationale:
- Each slice is effectively a separate 2D image
- Simple extension of existing single-image annotation
- User can navigate and annotate independently
Consequences:
- ✅ Simple mental model (each slice = image)
- ✅ Reuses existing annotation code
⚠️ Large stacks create many entries in annotations dict⚠️ No 3D annotation support
Status: Accepted
Context: SAM and display require 8-bit images, but microscopy often uses 16-bit
Decision: Normalize 16-bit to 8-bit using percentile clipping
Rationale:
- SAM models trained on 8-bit RGB images
- Displays only show 8-bit effectively
- Percentile clipping (2nd-98th) provides better contrast than linear
Consequences:
- ✅ Better visual contrast
- ✅ SAM compatibility
⚠️ Information loss (quantization)⚠️ Different normalization per image/slice
Status: Superseded by ADR-013
Context: Both SAM 2 (via Ultralytics) and Grounding DINO (via transformers) load PyTorch into the process. On Windows + Python 3.14, importing PyQt5 first and then loading PyTorch causes WinError 1114 (DLL load order conflict between Qt and Torch native dependencies). The application is fundamentally PyQt5-based, so we cannot reorder these imports.
Decision: Run each ML model in its own subprocess script that has no PyQt5 imports — sam_worker.py for SAM and dino_worker.py for DINO. The parent GUI process speaks to each worker over stdin/stdout with JSON requests and responses.
Rationale:
- The DLL conflict only manifests when both libraries are loaded in the same process. Splitting them across processes avoids the issue entirely.
- Keeps the GUI responsive: heavy model loading doesn't block PyQt's event loop in the same address space.
- Lets us swap or upgrade torch/transformers/ultralytics versions without worrying about Qt interactions.
- The JSON-over-stdio protocol is simple, language-agnostic, and easy to debug — just inspect what the worker prints.
Consequences:
- ✅ Works reliably on Windows + Python 3.14 (the original motivating bug)
- ✅ Worker scripts are PyQt-free; they can be tested independently
⚠️ Per-inference subprocess spawn cost (~1-2 s startup + first model load)⚠️ Need UTF-8 forced on both ends of the pipe (PYTHONIOENCODING=utf-8in env,encoding="utf-8", errors="replace"on parent) — Windows cp1252 default crashes on non-ASCII bytes in torch warnings⚠️ Two near-identical worker scripts to maintain (sam_worker.pymirrors the pattern fromdino_worker.py)
Superseded by: Migrating to PyQt6 (ADR-013) eliminated the underlying DLL conflict. The subprocess hop, JSON marshalling, and check_worker_isolation.py tooling were removed in the same PR.
Related:
- Implementation (historical):
sam_utils.py/sam_worker.py,dino_utils.py/dino_worker.py - Original SAM-only version landed in #65 (Python 3.14 support)
- DINO subprocess pattern landed alongside the DINO feature
Status: Accepted
Context: Both SAM and DINO model weights are large (SAM 2 tiny ~80 MB up to large ~400 MB; Grounding DINO base ~1.9 GB) and may not exist on first run. An earlier DINO flow required an explicit "Load" button click that did the resolve-or-download dance synchronously before the user could detect anything.
Decision: Selecting a model from the dropdown only updates state. Actual downloads happen on first use (first Detect call). UI feedback in the status label distinguishes "Ready: " (weights present) from " — will download on first detection".
Rationale:
- Matches the existing SAM behaviour (
change_sam_modeljust stores the name; download happens in the worker). - Removes a redundant click — one fewer thing for users to discover.
- Selecting a model the user picked by mistake is now free; only confirmed Detect triggers the (potentially heavy) download.
Consequences:
- ✅ Consistent UX between the SAM and DINO panels
- ✅ Faster perceived startup; no spurious downloads from idle browsing
⚠️ First Detect after selection blocks the UI while download runs (~1 min for DINO base); the status label shows progress but the dialog is otherwise unresponsive⚠️ No async download progress dialog —huggingface_hubprints to stdout
Status: Accepted
Context: ADR-011 introduced a subprocess hop for every SAM and DINO inference call to work around a PyQt5 + Torch DLL load-order conflict on Windows + Python 3.14. The workaround cost a fresh python sam_worker.py / dino_worker.py spawn per inference (~1-2 s warm latency, model reloaded from disk on every call) plus a temp-PNG marshal of the image.
Migrating the GUI from PyQt5 to PyQt6 (same PR) was expected to eliminate the DLL conflict — initially verified by tools/check_pyqt6_torch_coexistence.py importing PyQt6 packages → torch cleanly. However, further testing (see ADR-017) discovered that the conflict resurfaces when Qt's platform plugin is loaded before torch, which happens inside QApplication(). The practical workaround is to import torch eagerly before creating the QApplication.
Decision: Run SAM and DINO inference directly inside the main Python process. Keep the model objects on the SAMUtils / DINOUtils singletons so they persist across calls. Wrap each inference in a short-lived QThread to keep the UI thread responsive; the public API blocks the caller via a nested QEventLoop so call sites in annotator_window.py stay synchronous-looking.
Rationale:
- The latency win is the whole point. Subprocess spawn + Python startup + model reload was ~1-2 s every call; in-process with a cached model is ~50-500 ms.
- Threading via a nested
QEventLoop(the_run_synchelper insam_utils.py) lets the calling thread keep pumping events — timers, repaints, progress dialog cancels still work — while inference runs on the QThread. Existing call sites need no refactor. - Torch and transformers are imported lazily on first inference, so app startup stays fast for users who never touch SAM/DINO.
_qimage_to_numpyalready exists; converting the QImage on the calling thread (not on the worker) keeps Qt objects single-threaded as required.
Consequences:
- ✅ Each inference is ~1-2 s faster on Windows; less dramatic on macOS/Linux but still smoother.
- ✅ Cached model survives between calls — opening a DINO model once costs once. The DINO model stays on its compute device (CPU or CUDA) for its full lifetime; the old worker shuffled CPU↔GPU per call, defeating the caching gain on PCIe. Call
DINOUtils.unload()/SAMUtils.unload()to free GPU memory explicitly. - ✅ UI stays responsive during batch DINO+SAM runs (the calling thread's
QEventLoopstill processes events). - ✅ One source of truth per model — no more keeping
sam_utils.pyandsam_worker.pyaligned. - ✅ Exceptions from the inference worker (model load failures, CUDA errors) propagate out of
_run_syncrather than being printed and silently turned intoNone. Thechange_sam_modelerror path inannotator_window.pyactually catches now. ⚠️ A crash in torch (CUDA OOM, segfault) now takes the app down where the subprocess used to absorb it. Mitigation: inference is wrapped intry/exceptat the_run_syncboundary; the user sees an error dialog instead of a frozen UI.⚠️ Model RAM stays resident until the user closes the app (or invokes theunload()method).⚠️ Re-entrancy is a real hazard, addressed with belt-and-braces:_run_syncsets a module-level_inference_in_flightflag and raisesInferenceBusyErrorif re-entered. Same-thread re-entry can happen because the calling thread pumps its event loop while waiting (a timer fire, a click on an un-disabled widget, etc.). AQMutexwould not help — same-thread re-acquisition deadlocks on a non-recursive mutex and is meaningless on a recursive one.- The known re-entry vector — the SAM debounce timer firing during an in-flight inference — is guarded at the call site:
apply_sam_predictioninannotator_window.pycarries its own_sam_inference_in_flightflag and skips. Batch DINO already disables its trigger buttons. - The two-layer design is intentional: the call-site flag handles the common case quietly; the
_run_syncflag is the safety net that surfaces unknown re-entry vectors as a real exception rather than corrupting the model with concurrent.forward()calls (torch / ultralytics / transformers model objects are not thread-safe).
Related:
- Implementation:
sam_utils.py,dino_utils.py(both refactored in the same PR that retires ADR-011). - Smoke test:
tools/check_pyqt6_torch_coexistence.py(gate that gated this whole change). - Supersedes: ADR-011.
Status: Accepted
Context: The project shipped on PyQt5 5.15+ (ADR-001) from inception. Two pressures combined to motivate a migration:
- The PyQt5 + Torch DLL load-order conflict on Windows + Python 3.14 (ADR-011) forced an entire subprocess isolation layer. It was hypothesised that Qt6's packaging would eliminate the conflict entirely, but real-world testing (see ADR-017) showed the conflict persists when Qt's platform plugin is loaded before torch, regardless of whether PyQt5 or PyQt6 is the binding. The migration still removes PyQt5-specific issues (XCB plugin paths, enum namespacing drift).
- PyQt5 is in maintenance mode. PyQt6 is the actively developed line, gets new Qt6.x features, and has better Linux native integration (XCB plugin paths in particular).
Decision: Migrate the GUI binding from PyQt5 (>=5.15.0) to PyQt6 (>=6.7.0). Land in a single PR alongside the subprocess-removal work (ADR-013), gated behind tools/check_pyqt6_torch_coexistence.py to confirm the DLL conflict is actually gone on Windows + Python 3.14.
Rationale:
- Two coupled changes share most of their cost (touching every file that imports PyQt5) so doing them in one PR avoids paying the migration tax twice.
- Most PyQt5→PyQt6 differences are enum namespacing (
Qt.AlignCenter→Qt.AlignmentFlag.AlignCenter) and module relocations (QActionmoves fromQtWidgetstoQtGui) — mechanical, codemod-able. The behavioural risk is in event APIs (event.pos()→event.position(), returningQPointFnotQPoint) and a handful of removed widgets (QDesktopWidget→QGuiApplication.primaryScreen()). - The existing test suite (65 pytest-qt tests, mostly exercising coordinate transforms) serves as the regression safety net.
Consequences:
- ✅ Subprocess workers retired; inference is in-process with cached models (see ADR-013).
- ✅ Cleaner Linux story —
libxcb-cursor0is required by Qt 6 (was optional under Qt 5), but the platform plugin path mess is gone. - ✅ Long support runway: PyQt6 is the maintained binding.
⚠️ One-time migration cost: ~30 files touched, enum namespacing acrossannotator_window.py(300+ references),event.pos()→event.position()rewrite inimage_label.py.⚠️ PyQt6 is GPLv3 / commercial like PyQt5. Switching to PySide6 (LGPL) was considered and rejected to stay close to the existingpyqtSignal/pyqtSlotAPI.- ✅ All
.exec_()call sites insrc/migrated to.exec()in the v0.9.0 fix-pack — the PyQt5 alias is gone from this codebase.
Verification:
tools/check_pyqt6_torch_coexistence.pytests both import orders. The production order (torch first, thenQApplication) must pass. The Qt-first order is the known-failing case and is checked only to document the environment. Run before merging on the Windows + Python 3.14 target.- 65 tests pass on the new binding under
QT_QPA_PLATFORM=offscreen. - Full app constructs and renders headlessly; snake-game easter egg validates the
QDesktopWidget→QGuiApplication.primaryScreen()replacement.
Related:
Status: Accepted (v0.9.0)
Context: During DINO batch / single-image review, the user has
to accept (Enter) or reject (Escape) pending masks. The keyboard
handling was originally in ImageLabel.keyPressEvent, which only
fires when the canvas has focus. In practice the user clicks slice
entries, image entries, or buttons during review — focus moves to
those widgets and Enter is consumed locally (e.g. QListWidget
emits itemActivated), never reaching the canvas. The result: Enter
and Escape silently failed during the most common review workflow.
Three options were considered:
- Force focus back to the canvas on every UI interaction — intrusive, breaks normal navigation (Tab/Arrow keys on lists), and fragile because Qt's focus chain is not always predictable.
- Global
QShortcutwith ApplicationShortcut context — fires regardless of focus but unconditionally hijacks Enter / Escape, breaking modal dialogs (Enter activates default button) and inline editing inQLineEdit/QInputDialog. - Application-wide
QObjectevent filter that intercepts only when DINO temp_annotations are pending, and only when the focused widget is not a text input and no modal dialog is active.
Decision: Option 3. Implement DINOReviewEventFilter, install it
on QApplication.instance() once at startup, and gate the
interception on three conditions: pending DINO temp_annotations,
no active modal widget, focus not on QLineEdit/QTextEdit.
Consequences:
- ✅ Enter/Escape works regardless of which widget holds focus during DINO review.
- ✅ Modal dialogs and text-input fields are unaffected.
- ✅ Pattern is reusable for any future "review pending state" feature.
⚠️ Adds a per-key-press function call cost to the entire app. The filter short-circuits in three cheap checks before any work, so the overhead is negligible (≤ a few μs per keystroke).⚠️ Single global filter means future review-state features must share it or layer additional filters; if more review modes appear, collapse them into a strategy registry rather than installing multiple top-level filters.
Related:
- Implementation:
DINOReviewEventFilterclass incontrollers/dino_controller.py(moved there in Phase 4b);installEventFiltercall inui/shortcuts.py:install_event_filters, invoked fromImageAnnotator.__init__(moved there in Phase 8). - Cross-cuts: documented in Cross-cutting Concepts → DINO Temp Annotations.
Status: Accepted (Phase 6 of the modular refactor)
Context: Before Phase 6, ImageLabel.set_main_window(main_window)
injected the orchestrator into the canvas widget, and the widget poked
~50 sites on main_window directly — both reading state
(paint_brush_size, class_mapping, current_class, scroll_area,
current_slice, image_file_name) and mutating it
(all_annotations[name] = …, add_class(…),
update_annotation_list(), save_current_annotations(),
update_slice_list_colors(), schedule_sam_prediction(),
zoom_in(), enable_tools(), etc.). The coupling made:
- ImageLabel impossible to test in isolation without a
whole-
ImageAnnotatorfixture. - Every controller extraction (Phases 3–5) leak through
main_windowdelegation pass-throughs, because deleting them would break the widget. - The Phase 7 per-tool split (paint / eraser / polygon / rectangle
handler classes) impractical, because each handler would need the
same
main_windowreference and would multiply the coupling.
Three options were considered:
- Protocol / duck-typed callback object — pass a small protocol with the methods ImageLabel needs. Strict, type-safe, but writes are still synchronous direct calls; the widget still knows the exact method names on the orchestrator.
- Defer the fix — leave
main_windowfor one more phase, accept the debt. Cheapest, but each subsequent refactor pays the cost. - Qt signals for every write + a narrow read accessor object —
ImageLabel emits typed signals; the orchestrator connects each to
a controller slot during
__init__. Reads go through aCanvasContextobject with method-style accessors.
Decision: Option 3. ImageLabel declares ~20 pyqtSignals covering
annotation lifecycle, SAM, class, tool/UI state, navigation, and
batch finalisation. Reads go via a CanvasContext instance passed in
through set_context(ctx). The previous set_main_window /
self.main_window field is removed entirely.
The connection block lives in ImageAnnotator._connect_image_label_signals,
called once at the end of __init__ after every controller exists.
CanvasContext wraps the main window rather than copying state, so
the source of truth stays on ImageAnnotator and controllers see
their writes reflected on the next read.
Consequences:
- ✅ ImageLabel has zero
main_windowreferences; signals form the documented public write surface at the top of the class. - ✅ ImageLabel is now testable in isolation by connecting signals to stub slots; no controller fixture needed.
- ✅ Phase 7 (per-tool handlers) can carve
mousePressEvent/mouseMoveEventetc. without each handler needing the orchestrator. - ✅ Signal connections are explicit and grep-able — searching for
il.annotationCommitted.connectfinds the single wiring site. ⚠️ Two parallel mechanisms (signals for writes,CanvasContextfor reads) need to be kept in step. The widget's signal block and_connect_image_label_signalsmust stay in sync; a missing connection is a silent no-op write.⚠️ Signal connections rely on Qt's defaultAutoConnectionsemantics, which is synchronous within a single thread. Consumers that depend on a write taking effect before the next read (e.g.classRequestedemit followed by_ctx.class_id(name)read) must stay on the GUI thread.⚠️ The synchronous batch-save signal (annotationsBatchSaved) preserves the original O(1)-save-per-batch behaviour. Replacing it with per-annotation save would silently turn paint commits into O(N) saves. Future refactors must keep the batch boundary.
Pattern for adding a new ImageLabel → orchestrator interaction:
- Add a
pyqtSignal(<args>)toImageLabel. - Add a slot method on a controller (or main window) with matching signature.
- Wire it in
_connect_image_label_signals. - Replace the previous direct call site in ImageLabel with
self.<signal>.emit(<args>).
Pattern for adding a new read accessor:
- Add a method on
CanvasContextreturning the value. - Use
self._ctx.<accessor>()at the read site in ImageLabel.
Related:
- Implementation:
widgets/canvas_context.py,widgets/image_label.py(signal block lines 42–70),annotator_window.py:_connect_image_label_signals. - Cross-cuts: documented in Cross-cutting Concepts → Canvas Decoupling.
- Predecessor pattern: ADR-015 (DINO event filter) showed that ImageLabel can't reliably observe global keyboard state without help; ADR-018 generalises "explicit interaction surface, narrow read surface" to all canvas ↔ orchestrator traffic.
Status: Accepted (Phase 7 of the modular refactor)
Context: After Phase 6, ImageLabel no longer held a back-reference
to ImageAnnotator, but it still embedded four distinct annotation
tools (polygon, rectangle, paint_brush, eraser) as if/elif branches
spread across six event methods (mousePressEvent, mouseMoveEvent,
mouseReleaseEvent, mouseDoubleClickEvent, keyPressEvent,
paintEvent). Each tool also owned helper methods on the widget
(start_painting, commit_paint_annotation, commit_eraser_changes,
finish_polygon, cancel_current_annotation, …). Adding a new tool
meant touching all six event methods plus the widget's helper layer,
and the file had reached ~1,240 LOC.
Three options were considered:
- Keep tools as if/elif branches — cheapest, but the widget keeps accruing every new tool's behaviour.
- Per-tool widget subclass (one
QWidgetper tool, swap on tool change) — too heavy: tool switches would require teardown of the pixmap, scroll context, zoom factor, and the SAM/DINO/edit-mode sub-states that cut across tool selection. - Per-tool handler classes with a thin dispatcher on the widget.
Plain Python objects (not QObjects); the widget keeps a
_tools: dict[str, ToolHandler]and routes events toactive_tool_handler. Tools emit through the widget's existing Phase 6 signals.
Decision: Option 3. Each tool becomes a subclass of ToolHandler
in widgets/tools/. The contract:
- Event hooks return
Truewhen consumed:on_mouse_press,on_mouse_move,on_mouse_release,on_double_click,on_enter,on_escape. paint_overlay(painter)renders in-progress state (paint mask, eraser mask, polygon-in-progress, rectangle preview).has_unsaved_state()/commit()/discard()participate in the widget'scheck_unsaved_changes()dialog.deactivate()runs when the user switches away from this tool; default is no-op (matches the pre-Phase-7 "silently drop temp state mid-stroke" behaviour).
Deliberate non-decision: state ownership. Tool handlers contain
only behaviour; the temp-state fields (current_rectangle,
current_annotation, temp_paint_mask, temp_eraser_mask,
drawing_polygon, drawing_rectangle, is_painting, is_erasing)
remain on ImageLabel. Reason: AnnotationController.finish_rectangle
and finish_polygon (Phase 5a) read mw.image_label.current_rectangle
and mw.image_label.current_annotation directly. Moving the state
onto the handlers would have required a parallel controller refactor.
Handlers mutate self.label.X for those fields; pure-tool state
(e.g. future tool-internal counters) can live on the handler. See
the architectural-smell note below.
What stays on ImageLabel (intentional non-extraction):
- Navigation (zoom, pan, offset, scaled pixmap) — cross-cutting.
- SAM bbox / points state — activates from any tool via the SAM-box / SAM-points toggles, cuts across the main tools.
- Polygon edit mode (
editing_polygon,handle_editing_click,handle_editing_move,draw_editing_polygon) — modal state orthogonal to tool selection; setscurrent_tool = Nonewhile active. Promoting this to a handler would tangle the modal flow. - DINO
temp_annotations+accept_temp_annotations— cross-cutting; already touched by ADR-015's event filter. draw_tool_size_indicator— small enough that splitting it across paint/eraser handlers buys nothing.
paintEvent overlay pass. Iterates all handlers'
paint_overlay(), not just the active one. Reason: pre-Phase-7 the
temp paint mask, temp eraser mask, and polygon-in-progress rendered
whenever their state was populated, regardless of current_tool.
Each handler's paint_overlay short-circuits when its state is empty,
so the iteration is cheap and the user can switch tools mid-stroke
without losing visual feedback.
Consequences:
- ✅
image_label.pyshrinks from 1,239 to ~960 LOC. Adding a new tool now means: create one file inwidgets/tools/, register it in_tools, wire a button inannotator_window.py. No event-method edits. - ✅ Each tool can be unit-tested by instantiating the handler with
a stub
labelcarrying signals and_ctx— no controller fixture needed. - ✅ Phase 6's signal contract (ADR-018) is unchanged: handlers emit
via
self.label.<signal>.emit(...). ⚠️ State leak across the widget boundary. Handlers reach intoself.label.Xfor state. The contract drifts toward "handler is a namespaced function bag." Mitigation: revisit if/when controllers are updated to ask the handler (e.g.polygon_tool.points()) instead of reading the widget's field.⚠️ deactivate()is no-op by default. If you make itdiscard()later, audit the three call sites that still writecurrent_tool = Nonedirectly (ImageLabel.clear(),ImageLabel.start_polygon_edit, three locations inSAMController) — they bypassset_active_tooland therefore the hook.⚠️ check_unsaved_changesnow iterates all handlers, not just paint/eraser. Polygon participates viahas_unsaved_state() = len > 2(sub-3-point polygons are silently discarded on switch — they can't be saved anyway).
Pattern for adding a new mouse-driven tool:
- Create
widgets/tools/foo_tool.pywithclass FooTool(ToolHandler):. - Override the event hooks you need; emit via
self.label.<signal>.emit(...)and read viaself.label._ctx.X(). - Register in
ImageLabel.__init__'s_tools = {…, "foo": FooTool(self)}. - Add a button in
ui/sidebar.py:build_sidebarnext to the existing tool buttons, register it inwindow.tool_group, and connectclickedtowindow.toggle_tool. Then add a branch inImageAnnotator.toggle_toolthat callsself.image_label.set_active_tool("foo")for that button (since Phase 8 the UI building lives inui/sidebar.py, not on the orchestrator).
Related:
- Implementation:
widgets/tools/base.py,widgets/tools/{rectangle,polygon,paint,eraser}_tool.py,widgets/image_label.py:set_active_tool,widgets/image_label.py:paintEventoverlay-iteration block. - Predecessor: ADR-018 (Phase 6 signal decoupling) made this safe by
removing the
main_windowreference; handlers don't need an orchestrator handle. - Cross-cuts: documented in Cross-cutting Concepts → Canvas Decoupling (extended to describe the tool dispatcher).
Status: Accepted
Context: During Phase 1 of the modular refactoring (2025-06-10), 25 modules were moved into core/, dialogs/, inference/, io/, ui/, widgets/ subpackages. The smoke tests (test_smoke.py) verified that every module could be imported at top-level. All 30 smoke tests passed. However, four stale inline imports inside method bodies were missed:
# annotator_window.py — inside function bodies, NOT top-level
from .dino_utils import GDINO_MODEL_PATHS # moved to .inference.dino_utils
from .annotation_statistics import ... # moved to .dialogs.annotation_statistics
from .project_details import ... # moved to .dialogs.project_details
from .project_search import ... # moved to .dialogs.project_searchThese imports were deferred until the specific UI action triggered the function (e.g. picking a DINO model from the dropdown). The smoke tests, which only import modules, never execute function bodies and therefore never resolved the inline from .dino_utils reference. The bug surfaced only in manual QA when selecting a DINO model.
Decision: Add a static AST analysis test (test_annotator_window_inline_imports_are_resolvable) that parses annotator_window.py, extracts every bare relative import (from .module), and asserts the module still exists in the package root. The test fails with the exact line number for any stale import, preventing silent runtime-only regressions from reaching CI.
Rationale:
- Top-level import rewrites are mechanical and easy to verify via module import.
- Inline imports inside method bodies are invisible to module-level import tests.
- Manual QA is the fallback for behaviour, not for mechanical import correctness.
- AST inspection is cheap (~1 ms), zero false positives for this codebase, and runs in every CI build along with smoke tests.
Consequences:
- 🛑 Regression now impossible: the 30th smoke test would have failed the PR before merge.
- 🔧 No runtime cost — purely static analysis.
⚠️ Only coversannotator_window.py. If other files use the same inline-import pattern, the test should be generalized (or each file that contains inline imports gets its own AST check). In this codebase,annotator_window.pyis the only file with significant inline imports.⚠️ Doesn't catch dynamic imports (__import__,importlib.import_module), but we don't use those.
Related:
- Implementation:
tests/integration/test_smoke.py(test_annotator_window_inline_imports_are_resolvable). - Cross-cuts:
CLAUDE.md"Testing Checklist" updated to reference this test as a mandatory CI gate.
Status: Accepted
Context: ADR-011 and ADR-014 both discussed a DLL load-order conflict on Windows when PyQt and PyTorch share a process. The conflict was first observed with PyQt5 (ADR-011) and later claimed to be resolved by migrating to PyQt6 (ADR-014):
"Qt6's packaging reshuffle eliminates it." — ADR-014
"...verified by
tools/check_pyqt6_torch_coexistence.pyimporting PyQt6 → torch → transformers → ultralytics cleanly in one process..." — ADR-013
This claim was based on testing at the time, but it tested the wrong order: importing PyQt6 packages before torch works even in Qt5. The actual failure mode is triggered only when Qt's native platform plugin is loaded, which happens inside QApplication.__init__(), not at import PyQt6. The earlier verification script did not call QApplication(), so it never exercised the real failure path.
Real-world testing with torch 2.11.0+cu126 + PyQt6 6.10.2 + Python 3.14.2 on Windows 11 shows the conflict still surfaces when Qt's platform DLLs (e.g. qwindows.dll) are loaded BEFORE torch's c10.dll. The error is OSError: [WinError 1114] A dynamic link library (DLL) initialization routine failed.
Root cause analysis: Qt and torch both ship native DLLs that load into the same process. On Windows the DLL load order and address-space layout matter. When Qt's platform plugin claims certain memory slots or loads conflicting CRT libraries before torch does, torch's c10.dll init fails. The conflict is NOT between PyQt5 and torch per se — it is between Qt platform plugins and torch, regardless of whether the binding is PyQt5 or PyQt6.
Decision: Two complementary changes:
- In
main.py, eagerlyimport torch(with anImportErrorfallback) before importingQApplicationand creating the app. This ensures torch's DLLs claim their slot first. - In
__init__.py, replace eager toplevel imports ofannotator_window,image_label, andsam_utilswith a__getattr__-based lazy loader. The package init runs beforemain.pywhen launched via thesreeniconsole script (digitalsreeni_image_annotator.main:main). If__init__.pyeagerly imports modules that transitively import PyQt6 (e.g.annotator_window), Qt loads first and theimport torchinmain.pycrashes with the same WinError 1114. Lazy loading defers the Qt import until someone actually accessespkg.ImageAnnotator, which only happens after the torch-first guard has run.
Verification:
tools/check_pyqt6_torch_coexistence.pynow tests both orders:torch→QApplication(production order) — PASS.QApplication→torch(the claimed-safe order) — FAIL on Windows with torch 2.11.0.
- Exit code 0 means production order works; exit code 1 means even torch-first fails and subprocess isolation (ADR-011) must be restored.
- Smoke test
test_public_api_exportspasses:__getattr__correctly resolves all five public names.
Consequences:
- ✅ SAM and DINO model loading works on Windows + Python 3.14 + PyQt6 without subprocess overhead.
- ✅ App startup cost is negligible — torch import adds ~0.5-1 s before the splash window appears, which is acceptable for a desktop annotation tool.
⚠️ tests/integration/test_smoke.pycannot importmain.pybecause the pytest-qt test process already has Qt loaded; importing torch afterward triggers the same WinError 1114.main.pyis therefore excluded from the module-import list and is validated by CLI smoke tests instead.⚠️ Future Qt upgrades may change DLL packaging and make this unnecessary, butcheck_pyqt6_torch_coexistence.pywill detect that automatically.⚠️ Any new public name added to__init__.pymust also be wired through__getattr__or it will transitively pull in PyQt6 and break the torch-first guard.
Related:
- Supersedes (in spirit): ADR-014's claim that PyQt6 eliminates the conflict.
- Unblocks: ADR-013 in-process inference on the affected Windows environment.
- Implementation:
src/digitalsreeni_image_annotator/main.py. - Gate:
tools/check_pyqt6_torch_coexistence.py.
Status: Accepted
Context: The low-vision accessibility feature (continuous UI font
zoom, 8–24pt) needed (a) the chosen size to survive app restarts and
(b) canvas overlay elements — annotation labels, SAM point markers,
pen widths — to grow with the setting. UI preferences were previously
reset on every launch, and the .iap project file was the only
persistence mechanism in the app.
Decision:
- Introduce the app's first QSettings usage
(
QSettings("DigitalSreeni", "ImageAnnotator"), moduleapp_settings.py) forui/font_ptandui/dark_mode. These are per-user preferences, so they do not go into the.iapfile — a project opened by a different user must not impose a font size. - A single integer
ImageAnnotator.ui_font_ptis the source of truth; the named presets and the step shortcuts both funnel throughtheme.set_font_pt(clamp → apply → persist → menu sync). - Canvas overlay sizes derive from
ui_scale = ui_font_pt / 10.0(10 = the legacy default, so the default renders pixel-identical to the pre-feature code).ImageLabelreceives the value via a plain setter fromapply_theme_and_font, not via CanvasContext — consistent with the existing directimage_label.setFontcall, and avoids a paint-before-context-set window.
Alternatives considered:
- Storing prefs in the
.iapfile — rejected: project files are shared artifacts; accessibility settings are personal. - Templating the static stylesheets per font size — rejected: appended QSS override rules (later rules win at equal specificity) achieve the same with zero churn in the two stylesheet strings.
Consequences:
- ✅ Font size and dark mode persist across restarts.
- ✅ Tests stay hermetic: every
app_settingsfunction accepts an injectableQSettings(INI temp file) instance. ⚠️ Any new scalable UI metric should useImageLabel._pen_w/_overlay_fontor the appended-override block intheme.apply_theme_and_font— hardcoded px values won't follow the setting (see "UI Font Zoom" in08_crosscutting_concepts.md).⚠️ Deliberately-compact widgets (DINO threshold table / phrase panel) don't hardcode their small font inline; the appended block owns it via type/objectName selectors (ClassThresholdTable,PhraseEditorPanel …,#dino_phrase_hint) so compact ≠ unscaled. Follow that pattern for new compact widgets.⚠️ Known debt:dino_merge_dialog.pystill carries hardcodedfont-size:Npxtokens and acolor:#444dark-mode contrast issue, so it doesn't scale. Tracked, not an oversight; fix when that dialog is next touched.
Status: Accepted
Context: Users annotating domain-specific imagery (microscopy, medical, materials) get generic SAM masks that need heavy correction. We want to let them fine-tune SAM 2 / 2.1 on their own annotations and reuse the result in the existing SAM-box / SAM-points workflow (upstream issue #73).
The obvious approach — mirror the YOLO trainer's model.train(...) —
does not work: Ultralytics registers only a predictor for SAM's
segment task (SAM.task_map), so SAM(...).train() raises
NotImplementedError (verified on ultralytics 8.4.51).
Decision: Fine-tune with a custom PyTorch loop that reuses
Ultralytics' own forward path. SAM(...).model is a plain
SAM2Model nn.Module; its SAM2Predictor exposes the forward in
reusable pieces — get_im_features (image encoder) and
prompt_inference / _inference_features (prompt encoder + mask
decoder). These are not wrapped in inference_mode unless reached
via the public __call__, so calling them directly under
torch.enable_grad() yields differentiable mask logits. The engine
(training/sam_trainer.py) adds focal+dice loss (≈20:1) + AdamW +
backward. Default freeze policy: train only sam_mask_decoder
(image + prompt encoders frozen); an optional flag also unfreezes the
image encoder.
Checkpoints are saved as {"model": state_dict} — the exact shape
Ultralytics' _load_checkpoint reads (it rebuilds the architecture
from the filename suffix and load_state_dicts the nested model
key). Consequently a fine-tuned file must keep its base token in the
name (e.g. myrun_sam2_t.pt), enforced by make_custom_filename;
build_sam selects the architecture by ckpt.endswith(token). Every
save is round-trip-verified by reloading through SAM(out_path) and
running one forward — failing loudly rather than producing a file that
won't reload (cf. facebookresearch/sam2#337 key-mismatch failures).
Alternatives considered:
- facebookresearch/sam2 training code — rejected: heavy extra
dependency overlapping Ultralytics' bundled SAM2, and its checkpoints
need state-dict conversion to reload into our
SAM()inference path. - Export dataset + train externally — rejected as the default (less
"integrated"), though
Prepare SAM Dataset+ folder training give a similar offline path for users who want it.
Consequences:
- ✅ No new runtime dependency; fine-tuned models drop straight into the existing SAM selector and inference path.
- ✅ Exposure to Ultralytics internals is confined to a few
already-exercised predictor methods, guarded by
test_sam_finetuning.py::TestUltralyticsAPI(fails on an upgrade that renames them). ⚠️ The trainer loads its ownSAMinstance on itsQThread(it does not touchSAMUtils._model), and must not usesam_utils._run_sync(its re-entry guard is GUI-thread-local). The real hazard is two SAM models (resident inference + training) on one CUDA context, soSAMTrainControllerlocks the SAM inference UI (tools + model selector + the fine-tune menu) for the duration — re-enabled intraining_finishedon both the success and error paths.⚠️ Decoder fine-tuning is realistically GPU-only; a CPU-only box is hard-warned before a run (resolve_torch_device), and the device is pinned so an incompatible GPU is honoured as CPU instead of crashing.⚠️ Encoder features are recomputed per epoch (bounded memory) rather than cached across epochs; revisit if large datasets need the speedup.⚠️ Loss must use the inference coordinate frame. SAM2 letterboxes the image (LetterBox(1024, center=False), pad bottom/right) and inference maps masks back withops.scale_masks(..., padding=False), which crops that padding before upsampling. The training loss therefore runs the decoder logits through the sameops.scale_masksbefore comparing to the GT mask — a naiveF.interpolateover the full low-res mask bakes the padding into the target and the decoder learns masks shifted by the pad (a downward shift on non-square images, caught only during GUI testing because the e2e tests used square images). The landscape regression test (test_landscape_no_mask_shift) and theops.scale_masksAPI-drift guard protect this.
Status: Accepted (issue #75)
Context: Selecting an existing annotation was only possible through the
bottom-left annotation list (already ExtendedSelection) or by double-clicking
a mask on the canvas — which immediately enters vertex-edit mode. There was no
single-click select, no box/multi-select on the image, and canvas Delete worked
only while in vertex-edit mode. Issue #75 asked for single-click select (without
entering edit), rubber-band box select, modifier multi-select, and multi-delete —
all directly on the canvas.
Decision: Add an idle-mode selection layer to ImageLabel and route it
through the existing annotation-list selection so delete/merge/change-class are
reused unchanged:
- Idle activation. Selection is live only in
_is_select_mode()— no drawing tool, not editing, not SAM, no temp review. Picking any tool restores drawing. No new tool button (matches the user's "a single click should select" ask). - Gestures. Plain click selects the smallest mask under the cursor (covers segmentation and bbox); click on empty space clears; drag draws a rubber band and selects every annotation whose bounds intersect it; Shift makes a click toggle and a drag additive. Double-click is unchanged (still vertex edit).
- Ctrl stays pan. Ctrl+drag pan (with its carefully tuned reference frame) is left untouched; multi-select uses Shift instead of Ctrl.
- One selection, two surfaces. The canvas emits
canvasSelectionChanged(annotations, mode);AnnotationController.apply_canvas_selectioncomputes the new set (replace/add/toggle), setsimage_label.highlighted_annotations, and mirrors it onto the list with signals blocked.Deleteon the canvas reusesdelete_selected_annotations(which reads the list selection).
Consequences:
- Delete / Merge / Change-Class need no new logic — they already operate on the list selection, which the canvas now drives.
⚠️ Matching between the canvas and list relies on dict value-equality, like the rest of the selection code (image_label.annotationsis a deepcopy ofall_annotations, and PyQt round-tripsUserRoledicts as copies, so identity is never stable). Value-equal duplicate masks would select together — a pre-existing, accepted limitation. See the crosscutting "Canvas selection ↔ list selection" section.⚠️ The list mirror must blockitemSelectionChangedwhile selecting, or it recurses back throughupdate_highlighted_annotationsand overwrites the set.
Selection is rendered class-colour-independent (amendment). The first cut
drew the selected mask in solid red — invisible on a red-class mask, and the
default palette assigned red as the first class colour. Selection is now an
overlay drawn in a final pass on top of every mask, independent of class colour
and modelled on the sibling open-garden-planner app's CAD selection: a dashed
selection-blue bounding-box marquee (_SELECTION_COLOR = QColor(0, 120, 215, 220)) plus bright opaque-blue handle squares at the 4 corners + 4 edge
midpoints, white-cased and fixed on-screen size (_draw_selection_overlay in
widgets/image_label.py). The handles are what make selection unmistakable
regardless of mask colour (a single thin dashed outline was too faint; an earlier
marching-ants + marquee was too busy). The handles are now grab targets for
resize/move of any selected shape (see ADR-023). The mask keeps its
class colour; the rubber-band rect uses the same blue dashed style. Separately,
the default class palette
(core/constants.py::DEFAULT_CLASS_COLORS / default_class_color) was reordered
so red is last (no fresh project starts on red) and muted, and the default
fill opacity dropped to 0.2 (DEFAULT_FILL_OPACITY) so masks don't bury the
image. Existing projects keep their persisted class colours.
Status: Accepted (issue #40)
Context: "bbox"-keyed annotations (from COCO/YOLO import and detectors)
were not editable at all — start_polygon_edit only matches "segmentation",
so double-click vertex edit skipped them. ADR-022 draws 8 handle squares around
any selected annotation, but they were visual-only. A first cut wired them up
for "bbox"-typed annotations only — but almost everything in this app is stored
as "segmentation" (drawn rectangles, polygons, SAM/DINO masks all are), so the
handles looked grabbable on every shape yet did nothing on the shapes users
actually have. The handles must act on any selected shape.
Decision: Wire the handles up as direct-manipulation resize/move of the
single selected shape, modelled on the sibling open-garden-planner app's
ResizeHandle. No new mode, no double-click — it works off the existing idle-mode
selection:
- Single-shape, any kind. Handles are draggable when exactly one annotation
with a bounding box is selected (
_single_selected_shape()); a multi-select leaves them visual. The press handler resolves to the live object (_live_annotation) and recordskind—"seg"(polygon/mask) or"bbox"(box-only import) — which picks the geometry the handles drive (_begin_shape_edit). - Anchor-from-handle. A corner/edge drag computes the new bounding box
(
_resize_bbox: replaces the dragged coordinate, opposite side fixed, normalised, ≥ 1px). A"bbox"shape sets[x, y, w, h]directly; a polygon scales every vertex from the old box to the new one (_scale_segmentation), so the outline resizes proportionally. Per-handle resize cursors match OGP (_BBOX_HANDLE_CURSORS;SizeAllover the interior). - Move is drag-gated. A press inside the shape starts a pending move that
promotes only once the drag clears the
3px/zoomthreshold — so a plain click still falls through to selection (preserving nested-mask click-through). Move translates the box ([x,y,w,h]) or all vertices (_translate_segmentation). The geometry mutates in place so the canvas + overlay redraw live. - Bbox key stays in sync. Imported annotations carry both
segmentationandbbox; editing the polygon recomputes thebboxkey (_sync_bbox_key) so export/training stay consistent. Drawn shapes have no bbox key and gain none. - Commit / cancel. Release clamps into the image (ADR-024 — move slides the
intact shape back inside, resize trims/clamps) and emits
bboxEditCommitted→AnnotationController.commit_bbox_edit(save + list rebuild + re-mirror the selection). Escape restores the original geometry.
Consequences:
- The handles you see are exactly the grab targets —
_draw_selection_overlayand_bbox_handle_atshare_bbox_handle_points, so visual and hit geometry can't drift — and now they work on every selected shape, not just imported boxes. - Resizing a polygon scales it (handles drive the bounding box); reshaping a
polygon vertex-by-vertex is still double-click vertex edit. A
"bbox"shape stays rectangular by construction. ⚠️ The shape-drag branches sit before the rubber-band branch in the idle-mode mouse dispatch; both are gated on_is_select_mode()so a tool/edit/SAM state still wins. (Internal names keep thebbox_edit/bboxEditCommittedprefix — they denote editing via the bounding-box handles, whatever the underlying geometry.)
Status: Accepted (issues #32, #36)
Context: Annotation coordinates could be persisted outside the image
rectangle and silently poison training data. Drawn shapes were already safe
(finish_polygon/finish_rectangle shapely-intersect with the image boundary),
but two paths weren't: manual edits (polygon vertex drag; the new bbox drag)
clamped nothing, and the Image Augmenter wrote rotated/zoomed/flipped polygons
verbatim.
Decision: Add three pure helpers in utils.py and apply the right one per
path:
- Clamp manual edits with
clamp_segmentation/clamp_bbox— per-coordinate snap into[0, w] × [0, h]. Per-coordinate (not a shapely cut) is deliberate: it preserves the vertex count and ordering, so a polygon being dragged never loses or splits points mid-edit. Applied in place at edit commit (polygon Enter; bbox release), persisting through the existing save-by-reference path. - Clip augmented data with
clip_polygon_to_bounds— a shapely intersection (largest resulting polygon;buffer(0)first to repair self-intersections an affine augmentation can introduce). Geometric trimming is correct here because an augmented shape genuinely extends past the frame and should be cut at the edge, not have stray vertices snapped onto it. A polygon left fully outside returnsNoneand is dropped by the augmenter loop.
Consequences:
- One vocabulary, two semantics: clamp (cheap, count-preserving, for live edits) vs clip (exact, may drop/split, for batch augmentation). The choice is about whether vertex correspondence must survive, not about which is "more correct".
⚠️ clip_polygon_to_boundscan return fewer/more vertices than the input and may returnNone; callers must handle the drop (the augmentercontinues).- The existing
finish_polygon/finish_rectangleinline clips were left as-is to keep the diff contained; they could later delegate toclip_polygon_to_bounds.
Status: Accepted (issue #24)
Context: SAM/DINO masks are stored as raw dense polygons — _mask_to_polygon
returns the flattened cv2.findContours boundary with no simplification, so a
single mask can carry hundreds of vertices, bloating label files. Issue #24 asked
for a "mask complexity — less ↔ more points" control. The point add/remove half of
#24 was already covered by the SAM-points tool; this is the remaining piece.
Decision: A per-annotation, reversible Detail % control, surfaced as a column in the Annotations panel:
- Detail % (1–100, 100 = raw).
utils.simplify_polygon(raw, pct)thins via Douglas-Peucker (cv2.approxPolyDP), binary-searching the epsilon for the richest polygon whose vertex count is still ≤round(raw_count × pct/100). - Reversible via a preserved raw. The dense original is lazy-captured into
segmentation_rawthe first time a mask is thinned (nothing simplifies it before that, so the livesegmentationis the raw at capture). 100 % copiessegmentation_rawback intosegmentationexactly. No edits to the SAM/DINO/ manual accept paths were needed. - Two new annotation keys (
segmentation_raw,detail_pct) ride along: they round-trip through.iapfor free (project save doesann.copy()→convert_to_serializable→ JSON), and exports read only the effectivesegmentation, so the simplified polygon is what's exported. Imported/old annotations have neither key → handled by lazy-init. - Live + in place. The change handler resolves the selected row to the live
drawn object by value-equality (
image_label._live_annotation, reused from #40), mutatessegmentationin place, refreshes the Area cell + the row's UserRole, redraws, and saves. Thebboxkey (if present) is recomputed.
The Annotations panel became a QTableWidget (ID | Class | Area | Detail %),
mirroring dialogs/dino_phrase_editor.ClassThresholdTable (per-row spinbox via
setCellWidget, SelectRows, NoEditTriggers, stylesheet-only header). This
re-homes the #75 canvas↔list selection bridge onto a table:
- The annotation dict lives in column 0's UserRole (the value-equality marker).
count()/item(i)/selectedItems()→rowCount()/item(r, 0)/row-deduped selectedIndexes(); the mirror usessetRangeSelected(additive) becauseselectRow()replaces the selection in ExtendedSelection mode and would drop all but the last row.blockSignals+ value-equality are preserved verbatim.
Consequences:
- Closing #24 with a small, contained change: the feature is the table UI + one controller handler + one pure util; the accept paths are untouched.
- ✅ Fully reversible per annotation: Detail %=100 restores
segmentation_rawexactly. Exception: reshaping a polygon with the #40 handles invalidates the baseline —_clamp_edited_shapedropssegmentation_rawand resetsdetail_pct=100, so the edited geometry becomes the new raw (the old dense outline no longer describes the reshaped polygon, and a later 100 % must not silently revert the edit). The detail handler also re-pointshighlighted_annotationsat the mutated object so the overlay + a subsequent handle drag stay coherent. ⚠️ The spinboxvalueChangedis connected after the initialsetValue, so building/rebuilding the table never fires the simplification handler.⚠️ The deadcore/annotation_utils.pystill references the old QListWidget API but is unimported (confirmed) — left as-is to keep the diff contained.
Status: Accepted
Context: Annotation edits (create, delete, merge, move/scale, change class, detail %, paint, eraser, SAM/DINO accept) were all irreversible. The only safety net was a confirmation dialog on delete and a keep/delete prompt on merge — both of which broke flow (delete also popped a success dialog). The justification for those dialogs was "you can't undo," so removing them required a real undo/redo.
The mutation surface is wide and subtle: every operation writes both
image_label.annotations (the live working copy) and all_annotations[key], and
the two share inner list objects via the shallow-copy save. Annotations are
matched by value-equality, not identity (ADR-022/025), numbers are reassigned
on most edits (renumber_annotations), and Detail % carries a lazily-captured
segmentation_raw (ADR-025). A fine-grained command-per-operation design would
have to reproduce every one of these invariants in its undo path.
Decision: Snapshot the whole per-image annotation dict before each edit;
undo restores a snapshot wholesale. Restoring the entire dict sidesteps all the
value-equality / renumbering / selection-rehoming / segmentation_raw
subtleties — there is nothing to reconcile, only a deep copy to install.
controllers/annotation_history.AnnotationHistoryholds per-image-key undo/redo stacks (key =current_slice or image_file_name), so Ctrl+Z acts on the image on screen and never reaches an image you can't see. Stacks are retained across navigation and cleared on clear-all / new-project / project open. Depth is capped (50) and the symmetric model needs no separate baseline:record(before)pushes the pre-edit state,undo(current)swaps current onto redo and returns the popped before-state,redo(current)is the mirror.- One choke-point,
AnnotationController.record_history(), called before each synchronous mutation (finish polygon/rectangle, delete, merge, change class, eraser replace, SAM accept, DINO accept). It is not hooked ontosave_current_annotations()— that also fires on navigation and runs after mutation, so it can neither be filtered to real edits nor capture a clean "before." - Deferred gestures (bbox move/scale, paint stroke, polygon vertex edit)
notify the controller only after mutating in place. They capture the baseline
at gesture start via a new
ImageLabel.editBaselineRequestedsignal →capture_edit_baseline, and push it at commit (commit_edit_baseline, called fromcommit_bbox_edit,commit_polygon_edit, and theannotationsBatchSavedhandler). A deep-equality dedup inrecord()drops aborted gestures (Esc'd drag, empty stroke) so they leave no entry.- Vertex edit also got a save-discipline fix. Its Enter-commit historically
only refreshed the list and relied on a later save to persist (and Esc did
not revert the in-place drags).
commit_polygon_editnow callssave_current_annotations, and Esc restores the segmentation from a snapshot taken at edit-mode entry — so the commit is both persisted and undoable, and Esc truly cancels.
- Vertex edit also got a save-discipline fix. Its Enter-commit historically
only refreshed the list and relied on a later save to persist (and Esc did
not revert the in-place drags).
- Detail-% coalescing. The spinbox fires
valueChangedper step; a whole drag on one annotation records once (token = key + number + class), so one Ctrl+Z reverts the entire drag includingdetail_pctandsegmentation_raw. - Shortcuts are
QShortcuts withApplicationShortcutcontext (Ctrl+Z; Ctrl+Y and Ctrl+Shift+Z for redo) — the annotation-listQTableWidgetwould otherwise consume Ctrl+Z. Undo/redo are no-ops during project load, while a modal is open, while a text field has focus, or while a draw/edit gesture is in flight (_undo_blocked). Undo persists viaauto_save— the net must survive reopen.
Delete and merge dialogs removed. With undo as the net, delete_selected_annotations
drops both the confirmation and the success dialog; merge_annotations drops the
keep/delete prompt (originals are always replaced by the union) and the success
dialog. Validation warnings stay.
Consequences:
- ✅ Every annotation edit is reversible; destructive ops are instant and flow-friendly.
- ✅ Robust against the value-equality/renumber/raw subtleties because it restores whole dicts rather than replaying operations.
⚠️ Memory is a bounded deep copy per edit per image (annotations are small; depth-capped at 50).⚠️ Undo clears the current selection rather than trying to re-resolve it by value across a list rebuild — the safe, predictable choice.
Status: Under Consideration
Proposal: Add unit tests for non-GUI utilities (calculate_area, coordinate conversions, export functions)
Pros:
- Catch regressions in utility functions
- Build confidence for refactoring
- Document expected behavior
Cons:
- Setup overhead
- Maintenance burden
- May not catch most bugs (which are in GUI)
Status: Under Consideration
Proposal: Copy images to project folder, store relative paths
Pros:
- Portable projects
- Self-contained
Cons:
- Disk space duplication
- Slow for large image sets
- Export already copies images