# Explainer Studio — Production Render Pipeline

The Explainer Studio is **frontend-only**. The "Studio render" button in the
studio does not fake a video file: completed MP4s are produced by the site
operator with the same pipeline as `/media/demo/` (built with
`~/workspace/demo-work/assemble.py`), and each completed video gets a page
under `media/explainer/videos/<slug>.html`.

## What unlocks a studio render

1. **A TTS provider API key** (e.g. ElevenLabs), held by the site operator as
   a secret. The operator generates one MP3 per narration scene — production
   voice, not the browser's built-in voice.
2. **An ffmpeg render runner** (operator's machine or CI). Pipeline:
   approved script → per-scene TTS audio → 1920×1080 slide PNGs from the
   script's `[visual directions]` → ffmpeg segment assembly (still + audio,
   subtle zoom, fades) → concat → MP4.
3. **Publishing:** copy `videos/_template.html` to `videos/<slug>.html`,
   drop the MP4 at `videos/assets/<slug>.mp4`, add an entry to
   `videos/index.json`, fill in the script text + element tags + the
   "AI-generated explainer — script reviewed by [name], [date]" label, and
   redeploy the site. The Media Hub's explainer grid reads `index.json`
   automatically.

Nothing in the browser studio writes to the server — this is deliberate.
Until the operator runs the pipeline, queue items stay `queued`.

## Fully-worked example

Operator's working directory for one video: `~/workspace/explainer-renders/<slug>/`
with subfolders `audio/`, `slides/`, `segments/`.

### 1. Split the approved script into scenes

Each scene = one narration paragraph + its `[visual directions]` bracket.
For a 2-minute video at ~140 wpm you want ~280 words of narration total.

### 2. Generate TTS audio per scene (ElevenLabs example)

```bash
ELEVEN_API_KEY="$(op read 'op://secrets/elevenlabs/api_key')"   # operator's secret store
VOICE_ID="YOUR_VOICE_ID"                                          # operator's chosen voice

# scene 1 narration saved to narration-01.txt, etc.
for i in 01 02 03 04 05; do
  curl -sS -X POST "https://api.elevenlabs.io/v1/text-to-speech/$VOICE_ID" \
    -H "xi-api-key: $ELEVEN_API_KEY" \
    -H "Content-Type: application/json" \
    -d "$(jq -Rs '{text: ., model_id: "eleven_multilingual_v2",
                   voice_settings: {stability: 0.5, similarity_boost: 0.75}}' \
                   < narration-$i.txt)" \
    -o audio/scene-$i.mp3
done
```

Verify each MP3 plays and check durations:

```bash
for f in audio/scene-*.mp3; do
  ffprobe -v error -show_entries format=duration -of csv=p=0 "$f"
done
```

### 3. Build the slide PNGs (1920×1080)

One slide per scene, reflecting its `[visual directions]` — title card,
icon card, or diagram. Any image tool works; the demo used hand-built PNGs.
A minimal Pillow script for branded title/text cards:

```python
from PIL import Image, ImageDraw, ImageFont
NAVY, GOLD, WHITE = (14, 76, 87), (232, 199, 106), (255, 255, 255)
scenes = [
    ("Standard 3", "Academic and Learning Environments"),
    ("What the standard requires", "Professional, respectful, intellectually stimulating environments"),
    ("Your role", "You shape the learning environment every day"),
]
font_big = ImageFont.truetype("/usr/share/fonts/truetype/dejavu/DejaVuSerif-Bold.ttf", 110)
font_sm  = ImageFont.truetype("/usr/share/fonts/truetype/dejavu/DejaVuSans.ttf", 54)
for i, (kicker, title) in enumerate(scenes, 1):
    img = Image.new("RGB", (1920, 1080), NAVY)
    d = ImageDraw.Draw(img)
    d.rectangle([0, 0, 1920, 14], fill=GOLD)                      # gold rule
    d.text((120, 380), kicker.upper(), font=font_sm, fill=GOLD)
    d.text((120, 480), title, font=font_big, fill=WHITE)
    d.text((120, 990), "SI MedEd · unofficial explainer", font=font_sm, fill=(200, 215, 218))
    img.save(f"slides/scene-{i:02d}.png")
```

### 4. Assemble one MP4 segment per scene (mirrors `demo-work/assemble.py`)

```bash
FPS=30
i=01
AUD="audio/scene-$i.mp3"
IMG="slides/scene-$i.png"
T=$(python3 -c "import subprocess;print(float(subprocess.run(['ffprobe','-v','error','-show_entries','format=duration','-of','csv=p=0','$AUD'],capture_output=True,text=True).stdout.strip()) + 0.5)")
FRAMES=$(python3 -c "print(round($T * $FPS))")

ffmpeg -y -v error \
  -loop 1 -framerate $FPS -i "$IMG" \
  -i "$AUD" -filter_complex \
  "[0:v]scale=1920:1080,zoompan=z='min(1+0.0005*on,1.05)':x='iw/2-(iw/zoom/2)':y='ih/2-(ih/zoom/2)':d=${FRAMES}:s=1920x1080:fps=${FPS},fade=t=in:st=0:d=0.5,fade=t=out:st=$(python3 -c "print(round($T-0.5,2))"):d=0.5,format=yuv420p[v];[1:a]apad=whole_dur=$(python3 -c "print(round($T,2))"),aresample=44100,aformat=channel_layouts=stereo[a]" \
  -map "[v]" -map "[a]" -c:v libx264 -preset medium -crf 20 -r $FPS \
  -c:a aac -b:a 128k -frames:v $FRAMES -t $(python3 -c "print(round($T,2))") \
  "segments/scene-$i.mp4"
```

Repeat for every scene (the demo's `assemble.py` loops a PLAN table doing
exactly this), then concat:

```bash
printf "file 'segments/scene-%s.mp4'\n" 01 02 03 04 05 > segments/list.txt
ffmpeg -y -v error -f concat -safe 0 -i segments/list.txt -c copy final.mp4
ffprobe -v error -show_entries format=duration -of csv=p=0 final.mp4
```

### 5. Publish

```bash
SLUG="standard-3-learning-environments"
cp final.mp4 ~/workspace/simeded/simeded/media/explainer/videos/assets/$SLUG.mp4
cp ~/workspace/simeded/simeded/media/explainer/videos/_template.html \
   ~/workspace/simeded/simeded/media/explainer/videos/$SLUG.html
# then edit the page: title, script text, element tags, reviewed-by line,
# and add { "slug": "...", "title": "...", "duration": "2 min", "elements": [...] }
# to videos/index.json
```

### 6. Deploy

Redeploy the `simeded` Cloudflare Pages project (preview only — no custom
domain for the studio). The hub grid at `/media/` picks up the new entry
from `videos/index.json` automatically.

## Checklist before publishing a video page

- [ ] Script text on the page matches the approved, human-reviewed version
- [ ] Every `Element N.N` / `Standard N` tag verified against the script
- [ ] "AI-generated explainer — script reviewed by [name], [date]" line filled in
- [ ] Unofficial / no-affiliation labels present
- [ ] MP4 plays (spot-check first/last 5 seconds) and poster renders
