Skip to content
GodsQuantumPublic

About

Self-hosted subtitle production for video — transcribe, edit, style, automate and burn accurate animated subtitles with Rust, SvelteKit and FFmpeg.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Repository files navigation

AutoSubs logo

AutoSubs

Transcribe it. Fix it. Style it. Burn it — or automate the whole folder.
A self-hosted subtitle production workbench built around Rust, SvelteKit, FFmpeg/libass and your own transcription provider.

MIT license Rust backend SvelteKit UI FFmpeg English and French UI GHCR image

🇫🇷 README en français


AutoSubs is for the part that happens after you have a video: get word timings from Whisper/Speaches or import an existing subtitle file, clean up the text, repair bad timing, style it once, preview the right canvas, and let FFmpeg burn it with libass. For recurring work, a Workflow watches a folder and runs the same pipeline automatically.

It is deliberately not a browser-only subtitle toy. The Rust backend owns timing normalization, segmentation, persistence, rendering and workflow state. The web UI is a client of that API, so manual jobs and automated jobs go through the same rules.

📸 Screenshots

AutoSubs production queue

Queue — local/resumable uploads and server-side files feed the same persistent job system.

AutoSubs subtitle editor

Editor — seekable video preview, subtitle list, canonical regrouping, timing tools and subtitle-only exports.

AutoSubs responsive mobile interface

The interface is designed for desktop, tablet and phone — not just squeezed into a smaller viewport.

✨ What it does

  • Video ingest without a tiny extension allow-list — AutoSubs asks ffprobe whether a stable file actually contains video.
  • Resumable browser uploads — tus 1.0-style HEAD/PATCH uploads resume after network loss; re-selecting the same file after a reload resumes from the server offset.
  • Server-side picker + favorites — use files already mounted into the container instead of uploading them again; favorite frequently used folders once and reopen them from a compact shortcut list.
  • Sidecar-first workflow — import .ass, .ssa, .srt or AutoSubs JSON; attach, replace or detach a sidecar before rendering.
  • Transcription providers — OpenAI-compatible transcription endpoints plus an optional local provider/fallback such as Speaches.
  • Optional forced alignment — keep transcription and timing refinement separate: AutoSubs can validate word boundaries returned by a WhisperX-compatible HTTP aligner and safely fall back to native timestamps when alignment is unavailable or invalid.
  • Optional LLM correction — spelling/punctuation correction after import/transcription while preserving line count and timings.
  • Canonical subtitle timing engine — one Rust implementation repairs invalid ranges, overlaps, gaps and word timings. The browser does not maintain a competing copy.
  • Unicode-aware line grouping — grapheme counting, Unicode line-break opportunities and French no-break rules instead of UTF-8 byte counting.
  • ASS/libass styling — Pop, Highlight, Bounce, Karaoke, Word by word, Fade, Slide-up and plain subtitle modes; custom fonts, outline, shadow, placement and highlight colors.
  • Durable word timing — the canonical word-by-word timeline survives regrouping and visual edits; timed animations use the original word timings. Word-by-word display keeps apostrophe prefixes and split hyphen compounds attached to their word.
  • French-safe visual layout — French syntax-aware segmentation and hard maxLines enforcement use explicit line breaks, so rendered captions never gain hidden extra lines.
  • Complete job lifecycle — edit, split, merge, delete, retranscribe and re-render jobs from the queue/editor. Deleting a job keeps its source and final output files.
  • Source geometry invariant — Source + Preserve keeps the primary video dimensions and aspect ratio without scale, pad, crop, or black bars.
  • Adaptive format profiles — Source, 9:16, 16:9, 1:1, 4:5 and custom canvases. Matching source ratios keep their original pixel resolution; ratio changes use the largest exact, even canvas that fits inside the source dimensions, avoiding gratuitous upscale/downscale. contain, cover and stretch remain explicit choices.
  • Brands — group logo/outro assets and choose a default preset per output format.
  • Workflows — independent watch/output/archive folders, Brand/preset resolution, native filesystem events plus periodic reconciliation for NFS-mounted folders, with an explicit Video only or Video + SRT output policy.
  • Bundle archive after success — the source and same-stem companion files move together only after the selected outputs are published successfully; prefix collisions such as clip2.mp4 are not captured by a clip.mp4 job.
  • Transactional output publication — existing outputs are kept recoverable until the new video + SRT + ASS + JSON set is committed.
  • Persistent jobs — SQLite keeps queue state, settings, workflows, assets and events across restarts. Active jobs become interrupted after an unexpected restart instead of pretending they completed.
  • Real cancellation — waiting jobs and running FFmpeg/network work use cancellation tokens; a cancelled job does not later consume a freed encode slot.
  • Machine-readable render progress — progress comes from FFmpeg's -progress protocol, not regexes against human stderr output.
  • Runtime-benchmarked hardware encoder discovery — AutoSubs does not trust ffmpeg -encoders or a trivial one-frame probe. NVENC, QSV, VA-API, Vulkan and AMF are stress-tested at 2160×3840 for 360 frames, timed, and ranked per machine; auto tries validated hardware in measured order, then falls back to libx264.
  • Outro normalization — main video and outro are normalized into one concat graph so different dimensions/FPS/audio layouts do not require a fragile stream-copy concat.
  • EN / FR UI — instant browser-local language switch. This is separate from the transcription language setting.

🚀 Install

The pre-built image is published as ghcr.io/godsquantum/autosubs:latest.

mkdir autosubs && cd autosubs
curl -O https://raw.githubusercontent.com/GodsQuantum/AutoSubs/main/compose.example.yaml
curl -O https://raw.githubusercontent.com/GodsQuantum/AutoSubs/main/.env.example
cp .env.example .env
mkdir -p config data fonts media
# adjust MEDIA_PATH / bind mounts as needed, then:
docker compose -f compose.example.yaml up -d

Open http://<server-ip>:3051.

Native Linux AppImage

Docker remains the recommended server/NAS deployment, but AutoSubs is also released as a native Linux AppImage for x86_64 and aarch64. Download the matching AutoSubs-<version>-<arch>.AppImage asset from the GitHub Release, verify its adjacent .sha256 file, then:

chmod +x AutoSubs-*.AppImage
./AutoSubs-*.AppImage

The AppImage starts the same Rust backend and Svelte UI on http://127.0.0.1:3051, then opens your default browser. It does not install anything as root. Mutable state uses XDG user locations:

${XDG_CONFIG_HOME:-$HOME/.config}/autosubs        SQLite/config
${XDG_DATA_HOME:-$HOME/.local/share}/autosubs     data + custom fonts
${XDG_STATE_HOME:-$HOME/.local/state}/autosubs    graphical-launch logs

Use ./AutoSubs-*.AppImage --background (or --no-browser) when you want the backend and enabled folder workflows to keep running without opening a browser. Closing the browser never stops workflows; stopping the AppImage process does.

The AppImage carries a known-good FFmpeg/libass + fontconfig command-line baseline. By default those bundled media tools are preferred for reproducibility. Set AUTOSUBS_USE_SYSTEM_MEDIA_TOOLS=1 before launch to prefer your distro tools instead, useful when a rolling distro exposes newer GPU encoders or drivers. Existing AUTOSUBS_* variables still override AppImage defaults; XDG_CONFIG_HOME, XDG_DATA_HOME and XDG_STATE_HOME control its normal per-user storage locations.

Storage rule that matters

/config contains autosubs.db and must be local storage. AutoSubs uses SQLite WAL and refuses known network filesystems such as NFS/CIFS/SSHFS for the database path. Your videos, watch folders and outputs can absolutely live on NFS; mount them separately under an allowed media root.

A typical layout is:

/config             local SSD / host filesystem — SQLite only
/data               local or fast app working data — uploads/jobs/renders
/fonts              app-managed/custom fonts — writable when UI font import is enabled
/srv/media/...        large source/output/archive trees

NAS / external media deployment

If your media already lives on a NAS, mount the NAS once and expose only the roots AutoSubs is allowed to browse:

services:
  autosubs:
    image: ghcr.io/godsquantum/autosubs:latest
    container_name: autosubs
    init: true
    restart: unless-stopped
    user: "1000:1000"
    ports:
      - "3051:3000"
    environment:
      TZ: UTC
      AUTOSUBS_CONFIG_DIR: /config
      AUTOSUBS_DATA_DIR: /data
      AUTOSUBS_ALLOWED_ROOTS: /data:/srv/media
      AUTOSUBS_MAX_RENDER_JOBS: "2"
      AUTOSUBS_MAX_TRANSCRIPTION_JOBS: "2"
      AUTOSUBS_LOCAL_TRANSCRIPTION_ENABLED: "true"
      AUTOSUBS_LOCAL_TRANSCRIPTION_URL: http://transcriber:8000/v1
    volumes:
      - ./config:/config
      - ./data:/data
      - ./fonts:/fonts
      - /srv/media:/srv/media

Provider environment variables bootstrap an empty database only. After first start, Settings in the UI are authoritative. That avoids a container restart unexpectedly overwriting a key/URL you changed from the UI.

For OpenAI-compatible providers, use the provider base URL ending in /v1 (for example http://speaches:8000/v1 or https://api.groq.com/openai/v1). AutoSubs derives /models for discovery and /audio/transcriptions for transcription. A full .../audio/transcriptions URL is also accepted for backward compatibility.

For migration, the old SPEACHES_URL variable is still accepted as a first-boot alias for AUTOSUBS_LOCAL_TRANSCRIPTION_URL.

🎬 Manual production flow

  1. Add a local video, choose a server-side video, or pair a video with .srt / .ass / .ssa / .json.
  2. AutoSubs probes the media and imports or generates subtitle word timings.
  3. Optional LLM correction runs on text only.
  4. The canonical Rust engine normalizes timings and grouping.
  5. The job reaches Ready. Nothing has been re-encoded yet unless you explicitly chose immediate render.
  6. Review/edit/split/merge/delete/search/replace/regroup/shift timings in Editor. You can insert explicit visual line breaks without retiming words and nudge individual word boundaries by 10 ms within adjacent-word constraints. Use Remove final periods for short-form caption cleanup; it preserves commas, !, ?, and ellipses and can be undone before saving. The canonical word timeline remains available for later regrouping.
  7. Export SRT/ASS/JSON without touching the video, or choose Auto / Fast / Quality / Compact, inspect the resolved encoder + ETA range, then click Render video.
  8. FFmpeg/libass renders to a .partial staging file. .partial is internal media staging and is never treated as a subtitle file.
  9. Video and sidecars publish together; an optional source archive happens last. Existing jobs can be retranscribed or re-rendered from the queue.

That Ready step is intentional: correcting three words should not cost another full encode just to inspect the result.

Brands, presets and formats

A Preset owns visual subtitle behavior: typography, animation, colors, placement, segmentation limits, fit mode, target format and optional outro override.

A Brand can own a logo, default outro and one default preset for each format. A Workflow resolves its style in this order:

explicit workflow preset
        ↓
brand default for the workflow format
        ↓
global/default preset resolution

Applying a preset to a Job adopts the preset format, fit mode and segmentation limits by default. Explicit edits made after applying it remain possible.

Each Job inherits its outro from the selected Preset (then its Brand). The Editor can override that choice for one Job with another video asset or no outro.

🔄 Watch folders

Workflows combine low-latency filesystem events with periodic reconciliation. This matters on NFS: a remote write does not necessarily produce the local inotify event you expected.

Before claiming a candidate AutoSubs checks that its size/mtime stay stable, probes it with ffprobe, deduplicates it persistently, then looks for matching sidecars in priority order:

.ass / .ssa  →  .srt  →  .json  →  transcription

If any preparation or render step fails, the original source stays where it was.

FFmpeg and hardware acceleration

The runtime image pins a dated Debian snapshot with FFmpeg 9.0.2, libass and Mesa VA-API/Vulkan 26.2.3 on all supported architectures; the amd64 image also includes Intel media VA-API 26.2.4. This keeps the multimedia stack reproducible while still providing current 2026 hardware-video fixes without installing x86-only drivers on arm64.

At startup AutoSubs does not trust ffmpeg -encoders. Every H.264 hardware backend is tested with a sustained 2160×3840 / 360-frame encode and a 20-second guard. Successful backends are timed and ranked; Settings shows the measured milliseconds and the backend that Auto will select. This catches drivers that initialize successfully on a tiny frame but fail under a real 4K workload.

For Intel/AMD Linux acceleration, expose /dev/dri to the container and add the host video/render groups as required by your distro. The image includes Mesa VA-API and Vulkan drivers on every supported architecture, plus Intel's media VA-API driver on amd64. For NVIDIA, use the NVIDIA Container Toolkit and expose the GPU in your Compose stack. Hardware access is intentionally not enabled by default in the example Compose.

auto ranks validated H.264 hardware backends by measured runtime. Differences within 5% are treated as benchmark noise and resolved with a stability-first preference (NVENC, QSV, VA-API, Vulkan, AMF); outside that margin the genuinely faster backend wins. h264_vulkan is therefore selected when it survives the sustained 2160×3840 stress benchmark and is meaningfully faster on that machine. If a backend later fails on a real file, AutoSubs tries the next validated hardware backend before falling back to libx264. Explicit encoder selections remain explicit. For 4K software rendering, avoid hard 1 GiB container limits: HEVC decode + libass + libx264 can transiently exceed that.

Optional precision word alignment

Native Whisper/faster-whisper word timestamps remain fully supported. For stricter word-by-word timing, enable the alignment stage in Settings and provide a compatible HTTP endpoint. AutoSubs sends the accepted transcript plus the original mono audio, validates returned lexical identity/count and monotonic boundaries, and uses native timings automatically if the provider fails validation.

A tested CPU-only WhisperX reference provider lives in deploy/whisperx-aligner/. It is optional and deliberately separate from transcription: you can keep Speaches or another OpenAI-compatible ASR provider and use WhisperX only for forced alignment.

⚙️ Configuration

Core runtime variables:

Fonts: AutoSubs lists both system/fontconfig faces and app-managed fonts. The app font directory defaults to /fonts and is configurable with AUTOSUBS_FONTS_DIR. Mount it writable if you want UI font imports; a read-only mount is still valid when you only consume pre-provisioned fonts.

Variable Default Purpose
AUTOSUBS_PORT 3000 HTTP port inside the container.
AUTOSUBS_CONFIG_DIR /config Local SQLite/config directory.
AUTOSUBS_DATA_DIR /data Upload/job/render working data.
AUTOSUBS_FONTS_DIR /fonts App-managed font directory; writable for UI font imports. System fonts remain discoverable through fontconfig.
AUTOSUBS_ALLOWED_ROOTS /data:/media in the image Colon-separated roots exposed by the server picker/workflows. If omitted outside Docker, AutoSubs falls back to its data directory.
AUTOSUBS_MAX_RENDER_JOBS 2 Concurrent expensive render slots.
AUTOSUBS_MAX_TRANSCRIPTION_JOBS 2 Concurrent transcription slots.
AUTOSUBS_MAX_QUEUED_JOBS 256 Maximum number of active queued/processing jobs admitted at once.
AUTOSUBS_WORKFLOW_SCAN_SECONDS 5 Periodic workflow reconciliation interval (also covers NFS writes missed by native events).
AUTOSUBS_FILE_STABILITY_MS 2000 Required stable-size/mtime window before a watched file is accepted.
AUTOSUBS_MAX_UPLOAD_BYTES 53687091200 Maximum resumable upload size (50 GiB).

First-database bootstrap variables:

Variable Purpose
AUTOSUBS_TRANSCRIPTION_LANGUAGE Initial transcription language.
AUTOSUBS_TRANSCRIPTION_URL Initial external transcription endpoint.
AUTOSUBS_TRANSCRIPTION_MODEL Initial external model.
AUTOSUBS_LOCAL_TRANSCRIPTION_ENABLED Enable the local provider initially.
AUTOSUBS_LOCAL_TRANSCRIPTION_URL Initial Speaches/local endpoint.
AUTOSUBS_LOCAL_TRANSCRIPTION_MODEL Initial local model.
AUTOSUBS_LOCAL_FALLBACK_ENABLED Allow local fallback after external failure.
AUTOSUBS_TRANSCRIPTION_API_KEY Optional initial external-provider key.
AUTOSUBS_LOCAL_TRANSCRIPTION_API_KEY Optional initial local-provider key.

Provider keys can be bootstrapped through environment variables or configured later in Settings. Stored secrets are never echoed back to the browser.

API

The current API is versioned under /api/v1:

GET          /api/v1/health
GET          /api/v1/capabilities
POST         /api/v1/preview/frame                  Authoritative FFmpeg/libass preview frame
GET          /api/v1/events                         SSE

GET          /api/v1/jobs
POST         /api/v1/jobs/from-path
GET/PUT/DELETE /api/v1/jobs/{id}                    Delete keeps source/final media
GET/HEAD     /api/v1/jobs/{id}/video                 Compatibility: output, then source
GET/HEAD     /api/v1/jobs/{id}/video/source          Original source only
GET/HEAD     /api/v1/jobs/{id}/video/output          Rendered output only
POST         /api/v1/jobs/{id}/cancel
POST         /api/v1/jobs/{id}/prepare
GET          /api/v1/jobs/{id}/render-options       Profiles, resolved encoder and ETA
POST         /api/v1/jobs/{id}/render
POST         /api/v1/jobs/{id}/retranscribe
PUT/DELETE   /api/v1/jobs/{id}/sidecar
POST         /api/v1/jobs/{id}/sidecar/upload
GET/PUT      /api/v1/jobs/{id}/subtitles
POST         /api/v1/jobs/{id}/regroup
GET          /api/v1/jobs/{id}/subtitles/{srt|ass|json}

GET/POST     /api/v1/fonts                          System + app font catalog / app font import
GET          /api/v1/fonts/css                      Browser @font-face stylesheet
GET          /api/v1/fonts/{id}/content             Safe font content endpoint

POST         /api/v1/uploads                        tus create
HEAD/PATCH   /api/v1/uploads/{id}                   tus resume/upload

GET/POST     /api/v1/presets
DELETE       /api/v1/presets/{id}
GET/POST     /api/v1/brands
DELETE       /api/v1/brands/{id}
GET/POST     /api/v1/workflows
DELETE       /api/v1/workflows/{id}
GET/PUT      /api/v1/settings
GET          /api/v1/browse
GET/PUT      /api/v1/browse/favorites
GET/POST     /api/v1/assets
DELETE       /api/v1/assets/{id}

The server-side picker canonicalizes requested paths after symlink resolution and rejects paths outside AUTOSUBS_ALLOWED_ROOTS.

Development

The release build uses Node only as a frontend build stage. The runtime image contains no Node.js toolchain.

# Rust
cargo fmt --all -- --check
cargo clippy --locked --all-targets --all-features -- -D warnings
cargo test --locked --all-features

# Frontend
cd frontend
npm ci
npm run check
npm test
npm run build

# Full image
docker build -t autosubs:dev .

main is gated by CI for Rust, SvelteKit and Docker. Dependency updates cover Cargo, npm, Docker and GitHub Actions. Tagged releases publish multi-architecture GHCR images with SBOM/provenance metadata.

Contributions are welcome — see .github/CONTRIBUTING.md. For usage help, see .github/SUPPORT.md.

Project layout

src/api/          HTTP/SSE/tus endpoints
src/jobs.rs       persistent queue + job runner
src/media/        ffprobe, transcription, FFmpeg plans and rendering
src/subtitle/     SRT/ASS, timing normalization, segmentation, LLM correction
src/workflows.rs  watcher supervisor + periodic reconciliation
frontend/         SvelteKit static client + frontend tests
docs/             docs and versioned UI previews

Rust regression tests live next to the modules they exercise under `#[cfg(test)]`.

🔒 Security

AutoSubs can read and write mounted media paths and can invoke FFmpeg on them. Do not expose it directly to the public Internet. Put it behind your normal authenticated reverse proxy/VPN and mount only the directories it actually needs.

See security policy for vulnerability reporting and deployment notes.

License

MIT — see LICENSE.

About

Self-hosted subtitle production for video — transcribe, edit, style, automate and burn accurate animated subtitles with Rust, SvelteKit and FFmpeg.

Topics

Resources

Code of conduct

Contributing

Security policy

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages