narratty: setup & implementation plan (rev 2.3)¶
Planning document
This is the original plan. Some of it is not built yet: chapters, the
pronunciation lexicon, manifest.json, the requirements image and golden-frame
checks. The user guide describes what exists.
Renamed from shellcast on 2026-09-29: that name is taken on PyPI and npm and used by several terminal projects and a commercial iOS app.
narratty is the standalone terminal recorder for the Video-as-Code pipeline and
the symmetric peer to demo-machine. It turns a .narratty.yaml spec into a
narrated terminal video, doing its own TTS, timing, VHS rendering and muxing. It is
usable on its own, and the top-level videogen CLI also generates its spec from the
combined tutorial file.
What changed in rev 2 (vs. the pasted draft) 1. Two run modes: native (host-installed Python CLI + host tools) and sandboxed (the same CLI re-executes itself inside
narratty-basevia Docker/Podman). See Run modes. 2. A concrete modern Python stack (uv, Typer, Pydantic v2, pydantic-settings, Rich, ruff, mypy, nox, pytest + syrupy). See Tech stack. 3. Shell completion for bash/zsh/fish/PowerShell with dynamic completers for voices, providers, scene ids and spec files. See Shell completion. 4. Sync model fix: the draft appendedSleep <audio>after the actions but concatenated audio back-to-back, so narration drifts ahead of the video by the typing time of every scene. Rev 2 computes an explicit timeline and places each clip at its scene's start. See Sync model. 5. New actionswaitandhiddenscenes, new commands (init,schema,voices,tape,cache), and proposed answers to all open questions. 6. Kokoro viakokoro-onnx(thekokoropackage is PyTorch-based, not ONNX) and a licensing flag on Piper. See TTS. 7. Writable workspaces (rev 2.1): the read-only repo mount is replaced by aworkspaceblock withsnapshot(default),rwandromodes, persistent build caches, artifact export and per-spec network, so demos can run builds that write into the repo. See Workspace & build outputs. 8. UX & flexibility (rev 2.1): draft previews, subtitles, pronunciation lexicon, deterministic prompt/env, hooks, spec reuse, plugins, and more. See UX & flexibility. 9. Sandbox permissions in the spec (rev 2.2): network (none/full/allowlistof host:port, for licence servers and remote caches), env passthrough and extra mounts, with a consent prompt. See Sandbox permissions. 10. Layered pronunciation lexicon (rev 2.2): built-in, user, project, spec. See Pronunciation lexicon. 11. Tooling aligned with repo-env (rev 2.2), plus the decisions confirmed so far. See Decisions.
Goals¶
- One command turns a
.narratty.yamlinto a narrated.mp4. - Voiceover-synchronized: every narration clip starts exactly when its scene starts on screen, and a scene never ends before its narration does.
- Pluggable TTS with two supported local providers: Piper and Kokoro.
- Reproducible: pinned toolchain, hash-based audio cache, deterministic tapes.
- Runs sandboxed (container, spec declares tools and paths) or natively
(after
uv tool install narratty), with the same CLI and the same spec. - First-class CLI ergonomics: rich errors with YAML line numbers, shell completion, JSON Schema for editor autocompletion of specs.
Non-goals¶
- No browser/GUI automation (that is demo-machine / Track A).
- No authoring of the combined multi-engine file (that is the
videogenCLI). - No cloud TTS.
- No native Windows support in v1 (Windows users run sandboxed mode or WSL).
Run modes: native and sandboxed¶
One Python package, one CLI. The mode only decides where the pipeline runs.
| native | sandboxed (docker / podman) |
|
|---|---|---|
| Install | uv tool install narratty (or pipx install narratty); narratty[kokoro] for Kokoro |
Same CLI on the host, plus Docker or Podman |
| Where commands run | Your shell, your user, your filesystem | Inside narratty-base, non-root, no network |
Tools (vhs, ffmpeg, ttyd, requires.tools) |
Must be on PATH; doctor checks and prints install hints |
Baked into the image or the generated requirements layer |
workspace |
Same modes; snapshot is a scratch clone on the host | Same modes; the chosen dir is bind-mounted at /work |
| Reproducibility | Best effort; doctor warns on version drift from pinned versions |
Pinned by image digest |
| Use for | Fast authoring loop, trusted specs, machines without Docker | CI, reference renders, specs you did not write |
Selecting the mode¶
Precedence: CLI flag > NARRATTY_RUNTIME env > user config
(~/.config/narratty/config.toml) > default auto.
auto resolves to: native when already inside the container
(NARRATTY_IN_CONTAINER=1), otherwise docker, then podman if either is
available, otherwise native with a one-line warning that the run is not
sandboxed. The resolved mode is always printed on the first line of output.
How sandboxed mode works¶
The host CLI acts as a thin launcher:
- Parse and validate the spec on the host (fast failure, good errors, no container start for typos).
- Resolve the image:
narratty-base:<cli-version>@<digest>, or the hash-tagged requirements layer whenrequires.toolsexceeds the base (built with network, then run without). - Run one container, re-invoking the same subcommand inside it with
--runtime native:
docker run --rm \
--user "$(id -u):$(id -g)" \
--read-only --tmpfs /tmp --tmpfs /home/narratty \
--cap-drop ALL --security-opt no-new-privileges \
--network none \
-e NARRATTY_IN_CONTAINER=1 \
-v "$WORKSPACE_DIR:/work:rw" \
-v "$PWD/out:/out:rw" \
-v "$CACHE_DIR:/cache:rw" \
-v narratty-cache-bazel:/home/narratty/.cache/bazel \
-w /work \
narratty-base:0.1.0@sha256:... \
build /work/demo.narratty.yaml -o /out/demo.mp4 --runtime native
$WORKSPACE_DIRis the snapshot dir (default), the repo itself (rw), or the repo with:ro.--networkfollowssandbox.network(see Sandbox permissions). One named volume per entry inworkspace.caches.- The audio cache is a host directory (
platformdirsuser cache dir) mounted into the container, so native and sandboxed runs share cache hits. - VHS drives headless Chromium; under
--cap-drop ALLit needs its own sandbox disabled (VHS_NO_SANDBOX=true). The container is the sandbox. - Podman runs rootless with the same flags (
--userns keep-idinstead of--user). - The CLI version and the image tag are released together, so a given
narrattyalways runs the image it was tested with.
Safety note for native mode¶
A spec's type_command actions are real commands executed in a real shell as your
user. Native mode is for specs you trust. narratty build in native mode prints the
list of commands on first run of a spec and --yes / NARRATTY_ASSUME_YES skips the
prompt (CI). Sandboxed mode needs no prompt.
Workspace & build outputs¶
Many demos run a build (bazel build //..., cmake --build, npm run build) that
writes into the repo: bazel-* symlinks, build/, node_modules/, lockfile
updates. A read-only repo mount breaks those demos, and a plain read-write mount has
two problems: it dirties your real checkout, and the second render starts from a
different state than the first (the build is already done, output looks different).
So the workspace gets an explicit mode:
workspace.mode |
What /work is |
Use for |
|---|---|---|
snapshot (default) |
A disposable copy of the repo, writable, deleted after the render | Anything that builds or edits files; every render starts from the same clean state |
rw |
The real repo, bind-mounted read-write | When the demo must leave results behind, or the repo is too big to copy |
ro |
The real repo, read-only | Pure tours (ls, cat, git log); fastest, safest |
How snapshot works (identical in native and sandboxed mode):
1. git clone --local --no-checkout into a scratch dir under the narratty cache
(objects are hard-linked, so it is fast even for big repos), then check out HEAD.
2. With include_uncommitted: true (default), copy modified and untracked,
non-ignored files on top (git ls-files -m -o --exclude-standard), so you can
demo work in progress without committing.
3. Non-git directories fall back to a copy honouring .gitignore/.narrattyignore.
4. Mount or cd into it, render, copy artifacts to out/, delete the snapshot
(--keep-workspace keeps it for debugging and prints its path).
Build caches. Demo builds should be quick on screen. workspace.caches maps a
name to a path inside the session (e.g. ~/.cache/bazel, ~/.npm, ~/.m2). In
sandboxed mode each becomes a named Docker volume narratty-cache-<name>; in native
mode the real cache dirs are used. Combine with a hidden warm-up scene or a
before hook (see UX below) so the recorded build shows a fast incremental run.
Ownership. Containers run with the caller's UID/GID (--user on Docker,
--userns keep-id on Podman), so anything written to an rw workspace or out/ is
owned by you, not root.
Network, secrets, extra mounts. Declared in the spec's sandbox block, next to
workspace; see Sandbox permissions.
Guard rails. rw mode requires a clean git tree unless --allow-dirty is passed,
and prints which paths the render changed afterwards (git status --short), so a demo
never silently mixes with your own edits.
Sandbox permissions¶
Some demos need to reach the outside: a licence server for a commercial tool, a
Bazel/ccache remote cache, a package registry. Like workspace.mode, this is declared
in the spec, because the spec author knows what the demo needs.
sandbox:
network: none | allowlist | full # default: none
allow_hosts: [host:port, ...] # only for allowlist
env_passthrough: [NAME, ...] # copied from your shell; values never live in the spec
env: { NAME: value } # fixed, non-secret values
extra_mounts: [{ host, container, mode }] # e.g. ~/.netrc, a licence file
ssh_agent: false # forward SSH_AUTH_SOCK (e.g. git over ssh)
docker: false # mount the engine socket (Docker outside of Docker)
How allowlist works. The container sits on an internal Docker network with no
route out. A small forwarder sidecar (part of narratty, same image) listens on each
allowed host:port and passes TCP through to the real host; --add-host maps each
name to the sidecar. Because it forwards plain TCP, it works for raw protocols
(FlexLM-style licence servers) and for TLS (remote caches, SNI and certificates
intact), and no proxy settings are needed inside the demo. Anything not listed fails
to resolve or connect.
Who decides. The spec declares what it needs; you stay in control:
- Consent: the first run of a spec that asks for more than network: none, any
env_passthrough, extra_mounts, ssh_agent or docker shows a short summary and asks
(questionary prompt). The approval is remembered per spec path and a hash of its
sandbox block, so a changed spec asks again. --yes for CI.
- Overrides: --network none|allowlist|full and --allow-host on the command
line; a user policy in ~/.config/narratty/config.toml can cap what any spec may
get (e.g. max_network = "allowlist", allow_env = ["LM_LICENSE_FILE"]). A spec
that needs more than the cap fails with a clear message rather than running
half-broken.
- Native mode cannot enforce any of this. narratty says so once and runs with your
normal environment and network.
- manifest.json records the effective permissions of every render, and doctor
warns that network-dependent renders are less reproducible.
Milestones: none, full, env passthrough and extra mounts in v1; allowlist right
after (it needs the sidecar).
Tech stack¶
Target Python ≥ 3.12. Heavy dependencies are lazy-imported so --help and
completion stay fast.
| Concern | Choice | Why |
|---|---|---|
| Project/env/lock | uv (uv.lock, uv tool install) |
Same as repo-env |
| Build backend / version | hatchling + hatch-vcs | Version from git tags, same as repo-env |
| Task runner | nox (uv backend) | Same session layout as repo-env (lint, check_types, tests, integration) |
| Prompts | questionary | Consent prompts; same as repo-env |
| CLI | Typer (Click underneath) | Type-hint CLI, Rich help, built-in shell completion |
| Terminal output | Rich | Progress per stage, tables for voices/doctor, pretty errors |
| Spec models | Pydantic v2 | Validation, discriminated unions for actions, JSON Schema export |
| Config/env | pydantic-settings | NARRATTY_* env + TOML user config, same models |
| YAML | ruamel.yaml (safe loader) | Keeps line/column so validation errors point at the YAML line |
| Paths | platformdirs | Correct cache/config dirs on Linux and macOS |
| Downloads (voices/models) | httpx + sha256 verification | Pinned model files, resumable, no network in container |
| Audio | ffprobe/ffmpeg subprocess; soundfile for WAV I/O in Kokoro |
No heavyweight audio stack |
| Piper | piper-tts (Python package, provides piper) |
Same package in native and container; no separate binary download |
| Kokoro | kokoro-onnx + onnxruntime (optional extra narratty[kokoro]) |
ONNX, no PyTorch; ~300 MB model instead of multi-GB torch |
| Lint/format | ruff | Single fast tool |
| Types | mypy (strict on src/) |
Same as repo-env |
| Docs | mkdocs-material + mkdocstrings, versioned with mike | Same as repo-env |
| Commits/releases | commitizen, CHANGELOG, release workflow | Same as repo-env |
| Tests | pytest, syrupy (snapshot tapes), hypothesis (timeline math) | Golden tapes as snapshots |
| Hooks/CI | pre-commit (ruff, mypy, codespell), GitHub Actions with a uv cache | |
| Code quality | vulture, fawltydeps | Same as repo-env |
| Container | Multi-stage Dockerfile, uv sync --frozen in the image |
Same lockfile as native |
pyproject.toml sketch:
[project]
name = "narratty"
dynamic = ["version"]
license = { text = "MIT" }
requires-python = ">=3.12"
dependencies = [
"typer>=0.12", "rich>=13", "questionary>=2.0", "pydantic>=2.7",
"pydantic-settings>=2.3", "ruamel.yaml>=0.18", "platformdirs>=4", "httpx>=0.27",
"piper-tts>=1.3",
]
[project.scripts]
narratty = "narratty.cli:app"
[build-system]
requires = ["hatchling", "hatch-vcs"]
build-backend = "hatchling.build"
[project.optional-dependencies] # grouped like repo-env
kokoro = ["kokoro-onnx>=0.4", "soundfile>=0.12"]
test = ["pytest", "pytest-cov", "pytest-xdist", "syrupy", "hypothesis"]
lint = ["ruff"]
type-check = ["mypy"]
docs = ["mkdocs-material", "mkdocstrings[python]", "mike"]
dev = ["nox", "commitizen", "narratty[test,lint,type-check,docs,kokoro]"]
Shell completion¶
Typer provides completion for bash, zsh, fish and PowerShell:
narratty --install-completion # detects the current shell and installs it
narratty --show-completion zsh # print the script (for dotfiles / packaging)
Dynamic completers (autocompletion= callbacks) on top of the built-in ones:
| Argument / option | Completes |
|---|---|
<spec> |
*.narratty.yaml / *.narratty.yml in the current tree |
--provider |
piper, kokoro (Kokoro only if the extra is installed) |
--voice |
voice ids for the provider in the spec or --provider (from the curated catalog, no model loading) |
--scene / --from-scene |
scene ids read from the spec on the command line |
--runtime |
auto, native, docker, podman |
--theme |
VHS theme names (static list shipped with the package) |
Rules that keep completion usable: - Completers never import onnxruntime, piper or kokoro and never start a container; they read static catalogs and the spec file only. Target < 150 ms. - In sandboxed mode completion still runs on the host CLI, so it works identically. - A test runs each completer through Click's completion test harness.
Project layout¶
narratty/
├── pyproject.toml # package + entry point (narratty = narratty.cli:app)
├── uv.lock
├── README.md # this plan, later the user guide
├── Dockerfile # narratty-base (VHS + ttyd + ffmpeg + Piper + CLIs)
├── Dockerfile.kokoro # FROM narratty-base, adds kokoro-onnx + model
├── src/narratty/
│ ├── __init__.py
│ ├── cli.py # Typer app: commands + completers
│ ├── settings.py # pydantic-settings: env, config file, defaults
│ ├── model.py # Pydantic models for .narratty.yaml
│ ├── parser.py # ruamel load → model; errors with line/col
│ ├── schema.py # JSON Schema export
│ ├── timeline.py # scene start/end times from actions + audio
│ ├── tape.py # timeline → byte-stable VHS .tape
│ ├── render.py # run VHS → silent video
│ ├── mux.py # place clips on the timeline, mux onto video
│ ├── workspace.py # snapshot/rw/ro prep, caches, artifact export, cleanup
│ ├── runtime/
│ │ ├── base.py # Runtime protocol: run(cmd, spec, settings)
│ │ ├── native.py # subprocess on the host
│ │ └── container.py # docker/podman launcher, mounts, image resolve
│ ├── tts/
│ │ ├── base.py # TtsProvider protocol + registry
│ │ ├── normalize.py # shared text normalization
│ │ ├── catalog.py # curated voices: id, url, sha256, license
│ │ ├── piper.py
│ │ └── kokoro.py
│ ├── cache.py # content-addressed audio cache
│ ├── media.py # ffprobe helpers (durations, stream checks)
│ ├── doctor.py # preflight per runtime
│ └── data/
│ ├── voices.toml # curated catalog
│ └── themes.txt # VHS theme names for completion
└── tests/
├── unit/ # parser, timeline, tape snapshots, cache keys
├── contract/ # each TtsProvider against the protocol
├── cli/ # CliRunner + completion tests
└── integration/ # end-to-end build (marked, runs in container CI)
The .narratty.yaml spec¶
# yaml-language-server: $schema=https://raw.githubusercontent.com/ditschi/narratty/main/schema/v1.json
version: 1
meta:
title: "Repository layout tour"
tts:
provider: piper # piper | kokoro
voice: en_US-lessac-medium # provider-specific voice id
# piper: { length_scale: 1.0 }
# kokoro: { speed: 1.0 }
timing:
narration_buffer_ms: 500 # silence after each narration (matches demo-machine)
lead_in_ms: 300 # before the first scene
tail_ms: 1000 # after the last scene
terminal:
width: 1200
height: 700
theme: "Dracula"
font_size: 22
typing_speed_ms: 40 # overridable per scene
shell: bash
requires: # drives the container; checked by doctor in native mode
tools: [bat, eza, yazi, bazelisk]
workspace: # see "Workspace & build outputs"
source: . # repo to demo, relative to the spec file
mode: snapshot # snapshot (default) | rw | ro
include_uncommitted: true # snapshot: carry over modified + untracked files
caches: # persistent across renders, never part of the repo
bazel: ~/.cache/bazel
artifacts: # copied to out/ after the render
- bazel-bin/app/app
sandbox: # what the demo may reach; see "Sandbox permissions"
network: allowlist # none (default) | allowlist | full
allow_hosts:
- license.example.com:27000 # e.g. a licence server (raw TCP)
- cache.example.com:443 # e.g. a Bazel remote cache
env_passthrough: [LM_LICENSE_FILE, BAZEL_REMOTE_CACHE_TOKEN] # names only
env:
TERM: xterm-256color
extra_mounts:
- { host: ~/.netrc, container: ~/.netrc, mode: ro }
scenes:
- id: setup
hidden: true # runs but is not recorded (VHS Hide/Show)
actions:
- run: "cd /work && clear"
- id: intro
narration: >
This is a quick tour of the repository layout.
actions:
- run: "eza --tree --level=1 ."
- wait: { screen: '\$ $', timeout_ms: 5000 } # wait for the prompt to return
- id: wrap
narration: >
That is the high level map.
Changes vs. draft: top-level version, a timing block, hidden scenes, wait
action, shell, and a schema comment so editors autocomplete and validate specs
(narratty schema > narratty.schema.json).
Actions (v1)¶
Modelled as a Pydantic discriminated union; unknown keys are errors.
| Action | Meaning | VHS emission |
|---|---|---|
run: "…" |
type a command, press Enter, pause timing.run_hold_ms |
Type "…", Enter, Sleep |
type_command: "…" |
type text at the scene's typing speed | Type "…" |
enter |
press Enter (legacy; write key: Enter) |
Enter |
ctrl_sequence: "C-c" |
control chord | Ctrl+C |
hold: auto \| <ms> |
pause; auto = fill to the end of the narration |
Sleep <n>ms |
wait: <regex> or wait: { screen: <regex>, timeout_ms } |
block until the regex matches the last line | Wait+Screen@<t> /<regex>/ |
key: "<VHS key>" |
VHS key, case-insensitive, optional repeat count | <key> |
narration is scene-level (one voiceover per scene).
Pipeline¶
flowchart LR
S[.narratty.yaml] --> P[parser]
P --> T[tts provider<br/>piper or kokoro]
T <--> C[(audio cache)]
T --> D[ffprobe durations]
D --> L[timeline]
P --> L
L --> G[tape generator]
G --> R[VHS render<br/>silent mp4]
L --> M[mux: place clips<br/>at scene starts]
T --> M
R --> M
M --> V[verify durations]
V --> OUT[final .mp4]
- parse: load and validate into typed models (host side, always).
- tts: one clip per narrated scene; cache hit skips synthesis. Scenes are
synthesized in parallel (thread pool,
--jobs). - durations:
ffprobeeach clip for exact seconds. - timeline: compute each scene's start and length (below).
- tape: emit the byte-stable
.tape. - render:
vhs scene.tape→ silentvideo.mp4. - mux: build the narration track by placing each clip at its scene start
(
adelay+amix, or a generated silence-padded concat), then mux onto the video. - verify: ffprobe the result; fail if total video length differs from the timeline by more than a threshold (default 250 ms), warn above 100 ms.
Sync model¶
The draft emitted actions then Sleep <audio> per scene, but concatenated audio
back to back. Each scene's video was therefore typing + audio long while its audio
was only audio long, so narration drifted earlier with every scene.
Rev 2 makes the timeline explicit:
- Narration starts when the scene starts, so the voice talks while the command is
typed (natural screencast feel). Option
narration_start: after_actionsper scene for the draft's behaviour. action_ms= deterministic time of the scene's actions: characters × typing speed, plus fixed key costs, plus literalholdvalues.waitcounts its expected time as 0 and relies on the fill below.scene_ms = max(action_ms, audio_ms + narration_buffer_ms); the tape appendsSleep (scene_ms - action_ms)after the actions.hold: autois where that fill goes if it appears mid-scene.- Scenes without narration use their literal timing.
hiddenscenes contribute 0 ms. - The mux places clip n at
lead_in_ms + Σ scene_ms[0..n-1], so an error in one scene never accumulates into later scenes' audio. - Known limit: a command whose output takes longer than its scene's fill (or a
waitthat blocks long) stretches the video.verifycatches this; the fix in the spec is a largerholdor moving the slow command into ahiddenscene.
TTS: pluggable Piper + Kokoro¶
# tts/base.py
class TtsProvider(Protocol):
name: ClassVar[str]
model_version: str
def synthesize(self, text: str, voice: str, out: Path, opts: ProviderOpts) -> Path: ...
def list_voices(self) -> list[VoiceInfo]: ...
def check(self) -> list[DoctorFinding]: ... # used by doctor
| Provider | How | Voice ids | Notes |
|---|---|---|---|
| Piper | piper-tts package (piper CLI / Python API), text on stdin |
en_US-lessac-medium |
CPU, tiny models, fast; default and CI provider |
| Kokoro | kokoro-onnx + onnxruntime → 24 kHz WAV |
af_heart, am_adam |
82M model, better prosody; optional extra and separate image |
- Selection:
tts.provider+tts.voice; validation checks the voice against the catalog (see answers below). - Cache key:
sha256(provider, voice, normalized_text, model_version, sorted opts). Content-addressed under the user cache dir;narratty cache prune --older-than 30d. - Normalization: shared by both providers (whitespace, numbers, punctuation) so keys and prosody stay consistent.
- Parity with demo-machine: same provider + voice + normalization gives identical
audio on both tracks. Share
normalize.pyandvoices.tomlwith demo-machine if possible. - Failure handling: missing package, model or voice fails in
doctor/validate, never mid-render. - Licensing flag: Piper development moved to OHF-Voice/piper1-gpl and current
piper-ttsreleases are GPL-3.0; both Piper and Kokoro rely on espeak-ng (GPL) for phonemization. Individual voices have their own licences (recorded per voice invoices.toml). Calling Piper as a subprocess instead of importing it keeps narratty's own licence independent (e.g. MIT or Apache-2.0); importing it as a library would pull narratty towards GPL-3.0.
VHS tape generation¶
- Header from
terminal:Output,Set Width/Height/FontSize/Theme/Shell/TypingSpeed. - Per scene: a
# scene: <id>comment, thenType/Enter/Ctrl+…/Wait/Sleep. hiddenscenes are wrapped inHide…Show.- Durations are emitted in integer milliseconds; ordering is deterministic, so the tape is byte-stable for a given spec and audio set (snapshot-tested).
narratty tape <spec>prints the tape without rendering, for debugging.
Asciicast output¶
build --format cast records without VHS. render/script.py turns spec and timeline
into steps (Type, Press, Sleep, WaitScreen, Hide, …); render/tape.py
prints them as a tape, render/cast.py plays them in a pseudo-terminal and writes
asciicast v2 events. Its clock stops while hidden, so hidden output lands at the
moment recording resumes. Visible scenes become markers; their recorded starts
place the clips. render/player.py writes the page (asciinema-player with
audioUrl, cast and MP3 embedded).
Packaging & container¶
Images¶
narratty-base: pinneddebian:stable-slimdigest, Python 3.12, thenarrattypackage installed withuv sync --frozen,vhs+ttyd+ffmpeg(pinned), Chromium for VHS, fontconfig + one monospace Nerd Font, Piper with one default voice, and the curated CLIs (bat,eza,yazi,fd,ripgrep,git,jq). Non-root usernarratty.ENTRYPOINT ["narratty"].narratty-base:<ver>-kokoro:FROM narratty-base, addskokoro-onnx, onnxruntime and the pinned Kokoro model. Pulled only whentts.provider: kokoro.- Both published to GHCR with the CLI release.
Long-tail tools¶
If requires.tools names something outside the base, the launcher generates a thin
FROM narratty-base Dockerfile that installs only those tools with a pinned
installer (mise by default, version-pinned apt as fallback) and tags it
narratty-req:<sha256(requires + base digest)>. Cached locally, rebuilt only when the
requirements or the base change. Network is enabled only for that build.
In native mode the same list is only checked (doctor reports what is missing and how
to install it).
Paths, mounts, sandboxing¶
- The workspace (see Workspace & build outputs) is bind
mounted at
/work, the output dir at/out; repo data is never baked into an image. - Non-root, read-only rootfs with tmpfs for
/tmpand$HOME,--cap-drop ALL,no-new-privileges,--network none. videogencallsnarratty buildand gets the same runtime selection for free.
CLI surface¶
Global options: --runtime, --image, --cache-dir, --jobs, -v/-q, --json
(machine-readable output for videogen), --yes, --install-completion,
--show-completion, --version.
| Command | Does |
|---|---|
narratty init [path] |
Scaffold a commented .narratty.yaml with the schema line |
narratty validate <spec> |
Parse + schema-validate; check voice, actions, requires |
narratty schema |
Print the JSON Schema for the spec |
narratty voices [--provider] |
List catalog voices, marking which are installed locally |
narratty voices pull <id> |
Download and verify a voice/model (native mode) |
narratty plan <spec> |
Print the timeline and total length without rendering (estimates if audio is not cached) |
narratty tts <spec> |
Synthesize or cache-hit all clips; print durations and the timeline |
narratty tape <spec> |
Print the generated .tape |
narratty render <spec> |
TTS + tape + VHS → silent video |
narratty build <spec> -o out.mp4 |
Full pipeline + verify; --scene, --draft, --workspace-mode, --keep-workspace, --network, --var k=v |
narratty doctor [<spec>] |
Preflight for the chosen runtime: tools and versions, voices, mounts, container runtime, image availability |
narratty cache {info,prune} |
Inspect or prune the audio cache, snapshots and build-cache volumes |
Exit codes: 0 ok, 1 unexpected error, 2 usage error, 3 validation error, 4 missing
dependency (doctor), 5 render/mux failure, 6 sync verification failure, 7 a recorded
command exited contrary to its expect_exit.
Determinism & caching¶
- Content-addressed audio cache (key above), shared between native and sandboxed runs.
- Byte-stable
.tapefor a given spec + audio set. - Pinned: base image digest, Python lockfile, Piper and Kokoro model files (sha256 in
voices.toml), VHS, ttyd, ffmpeg, font. build --check: rebuild and compare sampled frames (one per scene, at scene midpoint) against a golden set, like demo-machine's golden frames. Only meaningful in sandboxed mode.
Testing¶
- Unit: parser errors with line numbers; action union; timeline math (hypothesis:
clip n always starts at the sum of previous scene lengths, no scene shorter than
its narration); tape snapshots (syrupy); cache key stability; settings precedence;
runtime
autoresolution; container command construction (snapshot of thedocker runargv). - CLI:
CliRunnerfor every command; completion callbacks return expected values and do not import heavy modules. - Provider contract: Piper and Kokoro satisfy
TtsProviderand produce a readable, non-empty WAV for a fixture line (Kokoro tests skipped when the extra is missing). - Integration: 3-scene fixture (one hidden) →
buildin the container → assert video + audio streams and duration ≈ timeline total; run once natively on a Linux CI runner with tools installed to keep native mode honest. - CI runs without network; models are baked into the images, and native CI caches them.
UX & flexibility considerations¶
Grouped by where they help. Items marked v1 are worth building in the first release; the rest are cheap to add later if the design leaves room now.
Fast authoring loop¶
- v1
--draft: half resolution, low fps, audio from a length estimate (words ÷ speaking rate) instead of TTS. Lets you iterate on actions in seconds. - v1
--scene intro --scene wrapand--from-scene: render a subset. The shell state still needs earlier scenes, so skipped scenes runhidden, not dropped. - v1
narratty plan <spec>: prints the timeline table (scene, narration length, action length, total) and the total video length before anything renders. narratty preview: play the result (or open the draft) when done.
Output quality and accessibility¶
- v1 subtitles: the timeline already knows each narration's start and length, so
emit
.srt/.vttfor free;--burn-subtitlesto hard-code them. - v1 chapters: scene ids/titles as MP4 chapter markers.
- v1 pronunciation lexicon: layered built-in, user, project and spec lists; see Pronunciation lexicon. Technical words are where TTS sounds worst.
- More formats:
-o demo.webm,-o demo.gif(GIF without audio, for READMEs). - Per-scene voice/speed overrides, for a two-speaker dialogue or emphasis.
- Background music track with automatic ducking under narration.
Pronunciation lexicon¶
Implemented without ipa and --no-user-config; the user guide's Pronunciation section
describes the shipped behaviour.
A mix of all levels, merged in this order (later wins):
| Level | Where | Holds |
|---|---|---|
| 1. Built-in | shipped with narratty, versioned | Common tech terms: kubectl, nginx, sudo, YAML, CLI, GUI… |
| 2. User | ~/.config/narratty/lexicon.toml |
Your personal preferences |
| 3. Project | narratty.lexicon.toml next to the spec or up to the repo root, or pulled in via extends |
Terms of the subject being demoed, shared by a series of videos |
| 4. Spec | tts.lexicon: in .narratty.yaml |
One-offs for this video |
# narratty.lexicon.toml
Bazel = "Bay-zel"
kubectl = "cube control"
[nginx]
say = "engine x" # respelling, works with every provider
ipa = "ˈɛndʒɪn ˈɛks" # used where the provider accepts phonemes
- Matching is whole-word; case-sensitive for entries with capitals, so
Bazeland a variable calledbazelcan differ. - Reproducibility: only the entries that actually change a narration go into that
clip's cache key and into
manifest.json, so a user entry can't silently change someone else's render;--no-user-configignores level 2 (default in CI). - Sandboxed mode: the host CLI resolves all levels and hands the container a fully merged spec, so user and project files never need mounting.
narratty lexicon show <spec>prints the merged list with each entry's source;narratty lexicon check <spec>lists likely problem words in the narration (all-caps, camelCase, words with digits or dots) that no level covers.
Predictable terminal content¶
- v1 deterministic session: clean
HOME, fixedPS1(configurable, e.g.prompt: "$ "or a starship config), fixedTZ/LANG, fixed hostname/user in the container,sandbox.envfor everything else. Stops real usernames, paths and timestamps from leaking into videos and golden frames. - v1 fail on command errors: the prompt embeds an invisible exit-code marker;
a
waitor scene end that sees a non-zero exit fails the render with the scene id (opt out per action withallow_fail: true). Otherwise a broken demo is only noticed when someone watches the video. - Redaction:
redact: ["ghp_[A-Za-z0-9]+", "$HOME"]masks patterns in typed text.
Flexible specs¶
- v1 hooks:
before/aftershell commands that run in the same workspace but are not recorded and do not open a VHS session (warm caches, seed files, fetch deps).hiddenscenes stay for things that must happen in the recorded shell (cd,export). - v1 variables:
${VAR}and${VAR:-default}interpolation fromvars:and the environment, so one spec serves several branches/targets (--var target=//app:all). extends: ../common.narratty.yamlto share terminal, TTS and lexicon settings across a series of videos; project-levelnarratty.tomlfor defaults.- Scene snippets/includes for recurring sequences (e.g. "build and run").
narratty import demo.tape: turn an existing VHS tape into a spec skeleton.
Extensibility¶
- TTS providers and actions registered through Python entry points
(
narratty.tts,narratty.actions), so a team can add a provider without forking. Piper and Kokoro themselves are registered that way. - Stable
--jsonoutput and amanifest.jsonnext to every render (spec hash, tool and model versions, image digest, per-scene timing, artifact list), whichvideogenand CI can consume.
Output layout¶
out/<spec-name>/
├── <spec-name>.mp4
├── <spec-name>.srt / .vtt
├── manifest.json
├── artifacts/ # workspace.artifacts
└── debug/ # --keep-intermediates: tape, silent video, audio clips
Behaviour people notice¶
- Ctrl+C stops the container, removes the snapshot and leaves no half-written
.mp4(write to a temp name, rename on success). - Every error names the scene id and the YAML line; unknown keys get "did you mean".
- A progress line per stage with timings;
-qfor CI,-vshows the exactdocker run/vhs/ffmpegcommands so users can reproduce by hand.
Milestones¶
- Scaffold: uv project with repo-env's layout (hatch-vcs, nox, ruff, mypy, pytest,
pre-commit, mkdocs, commitizen, CI and release workflows, MIT licence), Typer app with
--versionand working shell completion. - Spec + validate: Pydantic models, ruamel parser with line numbers,
validate,schema,init,doctor(native tool checks). - TTS layer:
TtsProvider, Piper backend, normalization, cache, durations,voices,tts. - Timeline + tape + render: timeline math,
tape, VHS render to silent video (native). - Mux + verify: timeline-placed audio, mux, verify → first narrated
.mp4natively. - Workspace + sandboxed mode:
snapshot/rw/roworkspaces, caches, artifacts;narratty-baseimage, container runtime, hardening flags, shared cache;--runtime auto. 5a. Sandbox permissions:sandboxblock, consent prompt, user policy caps;none/full, env passthrough, extra mounts. Thenallowlistwith the sidecar. 5b. UX v1:plan,--draft, subtitles, chapters, layered lexicon, deterministic session, exit-code checks, hooks, variables,manifest.json. - Kokoro: backend,
[kokoro]extra,-kokoroimage, voice validation for both. - Requirements layer: hash-tagged thin image from
requires.tools. - Reference render: the Bazel-basics storyboard end to end in both modes, plus
--checkgolden frames.
Native mode lands first (milestones 1–4) because it gives the fastest feedback loop; sandboxed mode then wraps a pipeline that already works.
Open questions: proposed answers¶
- Kokoro size in the base image → Make it optional. Separate
narratty-base:<ver>-kokoroimage (pulled only whentts.provider: kokoro) and anarratty[kokoro]extra for native mode. Keeps the default image and CI lean and still gives one-command Kokoro renders. Revised in 0.2: Kokoro is the default provider and a regular dependency, and the one image contains it. The fp16 model (177 MB) keeps the size down; Piper stays for languages Kokoro lacks, such as German. - Voice catalog: allowlist vs. discovery → Both, layered. Ship a curated
voices.toml(id, provider, language, download URL, sha256, licence) that drives validation,voicesand completion, and accept any locally discovered model too (with a warning that it is uncatalogued and not pinned). Sandboxed mode accepts only voices baked into the image. - Buffer / lead-in defaults → Match demo-machine:
narration_buffer_ms: 500. Addlead_in_ms: 300andtail_ms: 1000, all overridable intimingand per scene. Ideally both tools read the defaults from one shared place.
Decisions (confirmed)¶
Confirmed by the user on 2026-09-29:
- Default runtime: auto, meaning Docker (or Podman) automatically when installed,
otherwise native with a warning.
- Repo home: a public GitHub repo under ditschi; images on GHCR, package on PyPI.
- macOS: optional. Native mode on macOS is best-effort, with a non-blocking macOS
smoke job in CI; sandboxed mode works there through Docker Desktop or Podman anyway.
- Workspace and sandbox permissions are configured in the spec, with CLI and user
policy overrides.
- Licence: MIT (same as repo-env). Piper stays a subprocess so its GPL-3.0 does
not apply to narratty.
- Default workspace mode: snapshot.
Still open¶
Nothing blocking; the plan is ready to implement.