| 1 | # Video-delivery Design Reference Manual |
| 2 | |
| 3 | Conditional design guidance for presentations whose intended use is a recorded, |
| 4 | self-running, or video delivery. |
| 5 | |
| 6 | **Trigger**: load this reference when the effective delivery purpose is video, |
| 7 | recorded narration, or unattended playback. Also load its script rules for an |
| 8 | explicit final/literal narration input. Speaker notes, animation, or audio |
| 9 | requested for an otherwise ordinary deck do not activate it alone. Explicit |
| 10 | video/MP4 delivery does; Quick additionally activates §3's direct-delivery |
| 11 | contract. |
| 12 | |
| 13 | **Ownership**: this is a conditional Generate reference, not a profile or a new |
| 14 | artifact route. Default keeps its Strategist and confirmation flow; Quick keeps |
| 15 | its one-pass active-context flow. Existing notes, animation, audio, and native |
| 16 | PowerPoint export stages retain their schemas and commands. |
| 17 | When Beautify is active, its wording/page/order invariants still bind; apply |
| 18 | this reference only inside the design and motion freedom that profile permits. |
| 19 | |
| 20 | --- |
| 21 | |
| 22 | ## 1. Intake and Script State |
| 23 | |
| 24 | Classify supplied spoken material before planning: |
| 25 | |
| 26 | | Material | Treatment | |
| 27 | |---|---| |
| 28 | | Ordinary source or rough transcript | Use as source material; edit, condense, and reorganize under the selected route's normal content-divergence contract | |
| 29 | | Explicit final/literal narration script | Preserve every spoken word and its order; segment only at semantic scene boundaries | |
| 30 | | SRT used to generate new TTS | Preserve cue text when it is explicitly final; use source timecodes only as pacing evidence because the new synthesis timing becomes authoritative | |
| 31 | | SRT bound to an existing recording | Preserve its text/audio timing authority; do not regenerate TTS or pretend that one long recording was split automatically | |
| 32 | | Already page-separated final script | Preserve the supplied page boundaries unless the user explicitly permits restructuring | |
| 33 | | Target platform, canvas, or duration | Use the existing canvas registry; resolve scene granularity, page count, and notes length together | |
| 34 | |
| 35 | **Hard rule — final means explicit**: freeze wording only when the user identifies |
| 36 | the script as final, literal, or verbatim. Never promote ASR output, subtitles, |
| 37 | or a draft transcript into a literal contract by inference. |
| 38 | |
| 39 | **Default — semantic segmentation (may override for a user-authored page |
| 40 | plan)**: one scene represents one coherent visual state or mental-map step, not |
| 41 | one sentence, subtitle cue, or effect. Several cues may share a scene; one scene |
| 42 | may contain several ordered reveals. |
| 43 | |
| 44 | **Final-script production input**: after the page roster is final but before SVG |
| 45 | authoring, write the resolved per-slide script once to `notes/total.md`. Use |
| 46 | `# Slide <number>` headings and `---` separators so the file can exist before SVG |
| 47 | filenames do; preserve each body segment verbatim. It is a production input, not |
| 48 | a storyboard or substitute Design Spec. Run `total_md_split.py` only after the |
| 49 | SVG roster exists. |
| 50 | |
| 51 | --- |
| 52 | |
| 53 | ## 2. Scene and Page Planning |
| 54 | |
| 55 | **Default — quality follows purpose (may override)**: explanation prioritizes |
| 56 | understanding; promotion or brand work may prioritize emotion, recall, or |
| 57 | impact. Give every change a communication job. |
| 58 | |
| 59 | | Narrative relationship | Page treatment | |
| 60 | |---|---| |
| 61 | | Several lines explain one idea | Keep one page/scene and reveal only the semantic units needed for that explanation | |
| 62 | | One system persists | Before roster/notes freeze, derive states from the prior composition; keep orienting cues and change the semantic delta | |
| 63 | | New evidence expands a known map | Retain orienting cues; adapt the active region and context as needed | |
| 64 | | The same object changes position, scale, containment, or state | Consider compatible Morph endpoints when movement improves orientation | |
| 65 | | The audience must adopt a genuinely new mental map | Start a new composition and make the transition explicit | |
| 66 | |
| 67 | **Default — stable visual anchors (may override when the mental map resets)**: |
| 68 | within one explanation, preserve recognizable roles, relationships, or spatial |
| 69 | cues. Position, scale, and style may change while identity and orientation |
| 70 | remain legible; reset for a new map. |
| 71 | |
| 72 | **Default — one semantic focus change per beat (may override for one inseparable |
| 73 | idea)**: change several elements together only for one communication unit. Do |
| 74 | not alter unrelated regions merely for busyness. |
| 75 | |
| 76 | **Default — scene chrome earns its place (may override for navigation, identity, |
| 77 | attribution, or fidelity)**: for newly authored recorded, self-running, or video |
| 78 | scenes, do not carry a report-style fixed header, footer, or page number merely |
| 79 | by deck convention. Let the semantic title participate in the scene composition, |
| 80 | and omit nonessential running chrome especially on cover, ending, and breathing |
| 81 | scenes. Retain source or template chrome when the active profile's fidelity |
| 82 | boundary requires it, and retain new chrome when it genuinely orients the |
| 83 | audience or carries required identity or attribution. |
| 84 | |
| 85 | **Default — screen for orientation, notes for speech (may override for literal |
| 86 | on-screen copy)**: place keywords, structure, evidence, and relationships on the |
| 87 | slide; keep full explanation in notes. Do not duplicate the narration script as |
| 88 | body copy. |
| 89 | |
| 90 | **Page-count rule**: derive page/notes boundaries from scenes, mental-map arcs, |
| 91 | endpoints, and duration—not cues or sentences. Profile-fixed count/order/content, |
| 92 | including 1:1/fidelity, permits only existing-neighbor evaluation; never alter |
| 93 | those invariants for motion. |
| 94 | |
| 95 | --- |
| 96 | |
| 97 | ## 3. Default and Quick Planning Handoff |
| 98 | |
| 99 | **Default**: Stage 1 confirms the existing open-text `delivery_context`; it does |
| 100 | not ask a separate video question. When the confirmed value identifies |
| 101 | recorded/self-running/video delivery, load this reference before authoring the |
| 102 | three Stage-2 whole solutions. Apply its scene grammar to every direction; it |
| 103 | does not add a style catalog or confirmation field. Record delivery context and |
| 104 | afterlife in §I, visible states and optional motion jobs in §IX, and script/notes |
| 105 | policy plus target duration in §X. When the final-script branch is active, |
| 106 | create the frozen `notes/total.md` after the approved roster/lock is final and |
| 107 | before Step 5 or split-mode handoff. |
| 108 | |
| 109 | **Default — reading mode (may override for durable close reading)**: recorded |
| 110 | explanation leans `presentation`; choose `balanced` when close-reading afterlife |
| 111 | materially outweighs video delivery. |
| 112 | |
| 113 | **Quick**: there is no Stage 1 or separate video-purpose confirmation. Explicit |
| 114 | video/recorded/self-running intent activates this reference after source |
| 115 | sufficiency is known and before the one-pass roster, resource, and motion |
| 116 | decisions; absent that intent, keep ordinary Quick behavior. Load the script |
| 117 | rules alone when an explicit final/literal narration will become notes/audio. |
| 118 | Keep the applicable scene grammar and final-script handling in active context. |
| 119 | A pre-SVG `notes/total.md` is an enabled production artifact, not a forbidden |
| 120 | planning checkpoint; Quick still creates no root Design Spec, lock, |
| 121 | confirmation payload, or storyboard. |
| 122 | |
| 123 | **Hard rule — Quick video Custom Animations**: when Quick generates a PPTX for |
| 124 | recorded, self-running, or video delivery, enable Custom Animations before SVG |
| 125 | authoring and complete the custom-animation stage before base export. Use |
| 126 | semantic groups and page-specific choreography; deck-wide `-a auto` and page |
| 127 | transitions do not satisfy this requirement. Individual pages or groups may |
| 128 | remain static, so this is not an animation-coverage quota. A validated |
| 129 | `animations.json` is required unless the user explicitly requests static or |
| 130 | page-transition-only playback. |
| 131 | |
| 132 | **Mandatory — Quick direct video input**: when Quick must deliver a narrated |
| 133 | video or MP4 rather than only a deck for later recording, enable Speaker Notes, |
| 134 | Narration Audio, and video export; write the complete per-scene narration to |
| 135 | `notes/total.md` before P01 and use it as page-design input. After the SVG |
| 136 | roster, only agent-authored wording may be finalized; final/literal input remains |
| 137 | verbatim. Before audio, complete the required Custom Animations configuration |
| 138 | and decide whether narration governs any group timing. |
| 139 | |
| 140 | **Production outcomes**: |
| 141 | |
| 142 | | Need | Decision | |
| 143 | |---|---| |
| 144 | | Spoken delivery or a supplied final script | Enable Speaker Notes | |
| 145 | | User asks the workflow to synthesize narration | Enable Narration Audio; Speaker Notes is its dependency | |
| 146 | | Progressive reveal, continuing geometry, or timed emphasis materially aids explanation | Enable/load the appropriate animation capability | |
| 147 | | Quick generates a PPTX for recorded, self-running, or video delivery | Enable Custom Animations before SVG authoring and validate `animations.json` before base export | |
| 148 | | Quick directly delivers a narrated video or MP4 | Also enable Speaker Notes, Narration Audio, and video export; resolve narration-governed timing before audio, requiring timestamped page-local SRT for cue sync or subtitle delivery | |
| 149 | | The user explicitly requests static playback or disables object motion | Keep object animation off; retain the remaining notes/audio/video outcomes as requested | |
| 150 | |
| 151 | **Capability boundary**: Default generation does not force object animation or |
| 152 | generated audio merely because a deck may later be recorded. Quick with an |
| 153 | effective recorded/self-running/video delivery purpose does require Custom |
| 154 | Animations, while explicit user instructions for static or |
| 155 | page-transition-only playback remain authoritative. This requirement selects |
| 156 | the capability, not motion coverage or one effect for every page. |
| 157 | |
| 158 | --- |
| 159 | |
| 160 | ## 4. SVG, Notes, and Motion Realization |
| 161 | |
| 162 | When §3 created `notes/total.md` before SVG, read it once before the first SVG |
| 163 | and design each page around its corresponding spoken segment. Give every |
| 164 | independently narrated or timed semantic unit a descriptive direct-root `<g |
| 165 | id>`; keep inseparable units grouped. Preserve a final/literal script exactly; |
| 166 | agent-authored direct-video narration may change only during its final-SVG |
| 167 | validation before audio. |
| 168 | |
| 169 | **Hard rule — script/design consistency**: a final script is literal content. |
| 170 | If the finished visual page introduces an independent claim or relationship the |
| 171 | script does not explain, repair the page or return to planning; never rewrite or |
| 172 | pad the final script during the late notes pass. Conversely, every spoken idea |
| 173 | that requires visual orientation must have a visible state or deliberate |
| 174 | speech-only treatment. |
| 175 | |
| 176 | **Motion readiness**: load `animations.md` before SVG authoring whenever the |
| 177 | plan needs compatible Morph endpoints or page/object-specific motion. Author |
| 178 | every required start/end state and real semantic group before the final checker; |
| 179 | post-processing cannot invent missing visual endpoints or target IDs. |
| 180 | |
| 181 | **Motion restraint**: use transitions, reveals, emphasis, and Morph only for a |
| 182 | named communication job. There is no motion-coverage quota, and `effect: none` |
| 183 | remains valid. Auto-running narration uses `after-previous` / `with-previous`, |
| 184 | never `on-click`. |
| 185 | |
| 186 | **Mandatory when narration governs object motion**: before SVG authoring, load |
| 187 | `animations.md` and preserve real semantic groups; before audio, create and |
| 188 | validate canonical `animations.json`. After the base PPTX/report and timestamped |
| 189 | page audio/SRT, map timed groups in `narration_timing.json`, derive |
| 190 | `narration_animations.json`, and export the narrated PPTX/MP4. Only derived |
| 191 | triggers/delays wait for SRT; object identity, effect, and order do not. `-a |
| 192 | auto` or inherited fixed stagger is not semantic synchronization. For an |
| 193 | explicit user-selected static/page-transition-only Quick exception, or for |
| 194 | ordinary Default narration-independent deck-wide motion, omit these sidecars |
| 195 | and the object-sync claim. |
| 196 | |
| 197 | **Sound effects**: exclude them from this pass and planning artifacts. After |
| 198 | final SVG/motion, animation post-processing owns on-demand selection and native |
| 199 | PPTX configuration; otherwise remain silent. For direct narrated MP4 delivery, |
| 200 | `generate-audio` owns the selected sound-delivery branch: native PowerPoint |
| 201 | encoding plus triggered post-export mix, or an explicitly requested real-time |
| 202 | PowerPoint slideshow capture. Video gain and limiting never enter |
| 203 | `animations.json`; capture uses the balance actually heard during Slide Show. |
| 204 | |
| 205 | **Production sequence**: after the final SVG check, validate any pre-SVG |
| 206 | narration against the visible pages; ordinary draft-source runs instead use the |
| 207 | final-SVG-grounded notes generation. Split notes, execute the resolved motion |
| 208 | path, and export the editable PPTX. Direct Quick video continues |
| 209 | through audio and, when required, timestamped SRT. Custom Animations use the |
| 210 | narrated-sidecar flow when narration governs group timing; |
| 211 | narration-independent custom motion exports its canonical timing without an |
| 212 | object-sync claim before the narrated PPTX and MP4. |
| 213 | |
| 214 | --- |
| 215 | |
| 216 | ## 5. Delivery Boundary |
| 217 | |
| 218 | **Canonical artifact**: the editable PPTX remains canonical. `generate-audio` |
| 219 | owns provider/voice/rate selection, page audio/SRT generation, semantic |
| 220 | narration timing, narrated PPTX export, optional native PowerPoint video export, |
| 221 | the explicit slideshow-capture handoff, and the triggered sound-effects mix for |
| 222 | direct MP4 delivery. |
| 223 | |
| 224 | **Conditional MP4**: run `powerpoint_video.py --check` only for the native-export |
| 225 | branch. If native Windows PowerPoint export is unavailable, keep the narrated |
| 226 | PPTX as the successful upstream artifact. An explicit slideshow-capture choice |
| 227 | may hand that artifact to a user-operated Windows PowerPoint recorder; it is not |
| 228 | complete until the capture is returned and accepted. Do not substitute |
| 229 | screenshots, HTML, or a third-party renderer and call it equivalent. |
| 230 | |
| 231 | **Hard rule — choose one PowerPoint video sound boundary**: |
| 232 | |
| 233 | | Delivery branch | Sound contract | |
| 234 | |---|---| |
| 235 | | Native encoder | PowerPoint supplies visual animation and narration but may omit transition/object sounds. With resolved cues, treat its MP4 as raw and require the verified `video_sound_mix.py` output. | |
| 236 | | Real-time slideshow capture | PowerPoint remains the renderer and audio player; a recorder captures the full-screen Slide Show and exactly one application/system-audio source. The accepted capture must contain narration and every configured cue once, and must not enter `video_sound_mix.py`. | |
| 237 | |
| 238 | The branches are mutually exclusive because mixing a capture would duplicate |
| 239 | its cues. Keep the native cue configuration in the canonical PPTX. Slideshow |
| 240 | capture is explicit and human-audited; it does not inherit the native mix |
| 241 | receipt or become an automatic fallback. |
| 242 | |
| 243 | **Current boundary**: importing and automatically splitting one long finished |
| 244 | recording is unsupported. Require page-level audio or an explicit page/time map; |
| 245 | otherwise deliver the designed deck and frozen notes without claiming audio |
| 246 | integration. |
| 247 |