Tools for recovering data from failed Storage Spaces Direct (S2D) clusters by parsing the undocumented SPACEDB on-disk metadata format.
S2D on-disk metadata is not publicly documented. This tool bypasses the need for a running cluster by directly parsing the SPACEDB B-tree structures embedded in pool member disks to recover the slab allocation map — which physical disk blocks belong to which virtual disk.
Based on Junho Kim's StorageSpaceReconstructor format documentation (Kim 2022, Wiley doi:10.1111/1556-4029.14992).
SPACEDB metadata parser and virtual disk reconstruction engine.
Features:
- Parses SPACEDB/SDBC/SDBB B-tree structures from raw disks or VHDX files
- Extracts physical disk (PD) and virtual disk (VD) definitions
- Recovers the complete slab allocation map (VD → PD block mapping)
- Handles 3-column striping with 2-way mirroring across 2-node clusters
- Full VHDX format support (MS-VHDX specification)
- GPT partition table parsing with Storage Spaces partition type detection
- Virtual disk reconstruction from slab allocations
- NTFS/MDF signature scanning in reconstructed images
- JSON export of allocation maps
# Analyze SPACEDB metadata from a VHDX file
spacedb_explorer --file path/to/disk.vhdx
# Filter to specific virtual disk
spacedb_explorer --file path/to/disk.vhdx --vd 154
# Export allocation map as JSON
spacedb_explorer --file path/to/disk.vhdx --export slab_map.json
# Reconstruct a virtual disk image
spacedb_explorer --file path/to/metadata.vhdx --reconstruct --vd 154 \
--source /path/to/vhdx/directory --output /path/to/output.img
# Scan reconstructed image for SQL Server databases
spacedb_explorer --scan /path/to/reconstructed.img --output /path/to/extracted/
# Analyze raw physical disk (requires admin/root)
spacedb_explorer --disk 2Parallel disk scanner for recovering SQL Server data from S2D pool member disks.
Features:
- Scans multiple physical disks concurrently (1 thread per disk, rayon)
- Aho-Corasick multi-pattern matching (ASCII + UTF-16LE)
- SQL Server row structure parsing (8KB pages, TextMix, variable-length columns)
- PostgreSQL backend for crash-safe incremental scanning
- SIMD-accelerated zero-slab detection (AVX2)
- Double-buffered I/O with delta fingerprinting for skip detection
- TUI progress dashboard with per-disk throughput and ETA
- Near-miss logging for pattern discovery
# Incremental scan (skips previously scanned slabs)
s2d_scanner
# Fresh scan (drops tables, rescans everything)
s2d_scanner --fresh
# Scan specific disk range
s2d_scanner --disks 2-8
# Scan a reconstructed virtual disk image (no admin required)
s2d_scanner --image /path/to/reconstructed.img
# Show scan summary
s2d_scanner --probeVHDX → GPT → SBL Cache Partition (512MB) → SPACEDB region (at partition end)
SPACEDB Header (4KB): pool UUID, disk UUID, format timestamp
→ SDBC Header (512B): entry_size=64, entry_count, modified_time
→ SDBB Entries (64 bytes each): B-tree leaf records
Type 1: Pool validation
Type 2: Physical disk definitions (PD ID, UUID, block count)
Type 3: Virtual disk definitions (VD ID, UUID, block count, name)
Type 4: Slab allocation records (VD→PD block mapping)
The tool handles 3-column striping with 2-way mirroring across 2-node clusters:
- 3 columns (indices 0, 1, 2): Data is striped across 3 physical disks per write
- 2 copies (mirror 0, mirror 1): Each stripe is mirrored across nodes
- Block derivation:
virtual_block = stripe_row × num_columns + column_index - Column index extracted from SDBB Type 4
field_vd_id(Layout B encoding)
cargo build --releaseBinaries will be in target/release/.
Both tools read PG_URL from .env or environment for PostgreSQL connection:
PG_URL=postgresql://user:password@127.0.0.1:5432/recovery- Rust 1.75+ (2021 edition)
- PostgreSQL 14+ for scan data storage
- Administrator/root privileges for raw physical disk access
- Windows for PhysicalDrive access (VHDX/image file analysis works on any platform)
- Slab size: 256 MB (S2D storage unit)
- SQL page size: 8 KB
- Physical disk access uses
NO_BUFFERINGfor direct I/O on Windows - VHDX reader implements enough of the MS-VHDX spec to translate virtual → file offsets
- Pattern matching uses dual encoding (ASCII + UTF-16LE) for SQL Server nvarchar data
MIT