返回 ppt-master
video-design.md
根目录 / skills / ppt-master / references / video-design.md
1 # Video-delivery Design Reference Manual
2
3 Conditional design guidance for presentations whose intended use is a recorded,
4 self-running, or video delivery.
5
6 **Trigger**: load this reference when the effective delivery purpose is video,
7 recorded narration, or unattended playback. Also load its script rules for an
8 explicit final/literal narration input. Speaker notes, animation, or audio
9 requested for an otherwise ordinary deck do not activate it alone. Explicit
10 video/MP4 delivery does; Quick additionally activates §3's direct-delivery
11 contract.
12
13 **Ownership**: this is a conditional Generate reference, not a profile or a new
14 artifact route. Default keeps its Strategist and confirmation flow; Quick keeps
15 its one-pass active-context flow. Existing notes, animation, audio, and native
16 PowerPoint export stages retain their schemas and commands.
17 When Beautify is active, its wording/page/order invariants still bind; apply
18 this reference only inside the design and motion freedom that profile permits.
19
20 ---
21
22 ## 1. Intake and Script State
23
24 Classify supplied spoken material before planning:
25
26 | Material | Treatment |
27 |---|---|
28 | Ordinary source or rough transcript | Use as source material; edit, condense, and reorganize under the selected route's normal content-divergence contract |
29 | Explicit final/literal narration script | Preserve every spoken word and its order; segment only at semantic scene boundaries |
30 | SRT used to generate new TTS | Preserve cue text when it is explicitly final; use source timecodes only as pacing evidence because the new synthesis timing becomes authoritative |
31 | SRT bound to an existing recording | Preserve its text/audio timing authority; do not regenerate TTS or pretend that one long recording was split automatically |
32 | Already page-separated final script | Preserve the supplied page boundaries unless the user explicitly permits restructuring |
33 | Target platform, canvas, or duration | Use the existing canvas registry; resolve scene granularity, page count, and notes length together |
34
35 **Hard rule — final means explicit**: freeze wording only when the user identifies
36 the script as final, literal, or verbatim. Never promote ASR output, subtitles,
37 or a draft transcript into a literal contract by inference.
38
39 **Default — semantic segmentation (may override for a user-authored page
40 plan)**: one scene represents one coherent visual state or mental-map step, not
41 one sentence, subtitle cue, or effect. Several cues may share a scene; one scene
42 may contain several ordered reveals.
43
44 **Final-script production input**: after the page roster is final but before SVG
45 authoring, write the resolved per-slide script once to `notes/total.md`. Use
46 `# Slide <number>` headings and `---` separators so the file can exist before SVG
47 filenames do; preserve each body segment verbatim. It is a production input, not
48 a storyboard or substitute Design Spec. Run `total_md_split.py` only after the
49 SVG roster exists.
50
51 ---
52
53 ## 2. Scene and Page Planning
54
55 **Default — quality follows purpose (may override)**: explanation prioritizes
56 understanding; promotion or brand work may prioritize emotion, recall, or
57 impact. Give every change a communication job.
58
59 | Narrative relationship | Page treatment |
60 |---|---|
61 | Several lines explain one idea | Keep one page/scene and reveal only the semantic units needed for that explanation |
62 | One system persists | Before roster/notes freeze, derive states from the prior composition; keep orienting cues and change the semantic delta |
63 | New evidence expands a known map | Retain orienting cues; adapt the active region and context as needed |
64 | The same object changes position, scale, containment, or state | Consider compatible Morph endpoints when movement improves orientation |
65 | The audience must adopt a genuinely new mental map | Start a new composition and make the transition explicit |
66
67 **Default — stable visual anchors (may override when the mental map resets)**:
68 within one explanation, preserve recognizable roles, relationships, or spatial
69 cues. Position, scale, and style may change while identity and orientation
70 remain legible; reset for a new map.
71
72 **Default — one semantic focus change per beat (may override for one inseparable
73 idea)**: change several elements together only for one communication unit. Do
74 not alter unrelated regions merely for busyness.
75
76 **Default — scene chrome earns its place (may override for navigation, identity,
77 attribution, or fidelity)**: for newly authored recorded, self-running, or video
78 scenes, do not carry a report-style fixed header, footer, or page number merely
79 by deck convention. Let the semantic title participate in the scene composition,
80 and omit nonessential running chrome especially on cover, ending, and breathing
81 scenes. Retain source or template chrome when the active profile's fidelity
82 boundary requires it, and retain new chrome when it genuinely orients the
83 audience or carries required identity or attribution.
84
85 **Default — screen for orientation, notes for speech (may override for literal
86 on-screen copy)**: place keywords, structure, evidence, and relationships on the
87 slide; keep full explanation in notes. Do not duplicate the narration script as
88 body copy.
89
90 **Page-count rule**: derive page/notes boundaries from scenes, mental-map arcs,
91 endpoints, and duration—not cues or sentences. Profile-fixed count/order/content,
92 including 1:1/fidelity, permits only existing-neighbor evaluation; never alter
93 those invariants for motion.
94
95 ---
96
97 ## 3. Default and Quick Planning Handoff
98
99 **Default**: Stage 1 confirms the existing open-text `delivery_context`; it does
100 not ask a separate video question. When the confirmed value identifies
101 recorded/self-running/video delivery, load this reference before authoring the
102 three Stage-2 whole solutions. Apply its scene grammar to every direction; it
103 does not add a style catalog or confirmation field. Record delivery context and
104 afterlife in §I, visible states and optional motion jobs in §IX, and script/notes
105 policy plus target duration in §X. When the final-script branch is active,
106 create the frozen `notes/total.md` after the approved roster/lock is final and
107 before Step 5 or split-mode handoff.
108
109 **Default — reading mode (may override for durable close reading)**: recorded
110 explanation leans `presentation`; choose `balanced` when close-reading afterlife
111 materially outweighs video delivery.
112
113 **Quick**: there is no Stage 1 or separate video-purpose confirmation. Explicit
114 video/recorded/self-running intent activates this reference after source
115 sufficiency is known and before the one-pass roster, resource, and motion
116 decisions; absent that intent, keep ordinary Quick behavior. Load the script
117 rules alone when an explicit final/literal narration will become notes/audio.
118 Keep the applicable scene grammar and final-script handling in active context.
119 A pre-SVG `notes/total.md` is an enabled production artifact, not a forbidden
120 planning checkpoint; Quick still creates no root Design Spec, lock,
121 confirmation payload, or storyboard.
122
123 **Hard rule — Quick video Custom Animations**: when Quick generates a PPTX for
124 recorded, self-running, or video delivery, enable Custom Animations before SVG
125 authoring and complete the custom-animation stage before base export. Use
126 semantic groups and page-specific choreography; deck-wide `-a auto` and page
127 transitions do not satisfy this requirement. Individual pages or groups may
128 remain static, so this is not an animation-coverage quota. A validated
129 `animations.json` is required unless the user explicitly requests static or
130 page-transition-only playback.
131
132 **Mandatory — Quick direct video input**: when Quick must deliver a narrated
133 video or MP4 rather than only a deck for later recording, enable Speaker Notes,
134 Narration Audio, and video export; write the complete per-scene narration to
135 `notes/total.md` before P01 and use it as page-design input. After the SVG
136 roster, only agent-authored wording may be finalized; final/literal input remains
137 verbatim. Before audio, complete the required Custom Animations configuration
138 and decide whether narration governs any group timing.
139
140 **Production outcomes**:
141
142 | Need | Decision |
143 |---|---|
144 | Spoken delivery or a supplied final script | Enable Speaker Notes |
145 | User asks the workflow to synthesize narration | Enable Narration Audio; Speaker Notes is its dependency |
146 | Progressive reveal, continuing geometry, or timed emphasis materially aids explanation | Enable/load the appropriate animation capability |
147 | Quick generates a PPTX for recorded, self-running, or video delivery | Enable Custom Animations before SVG authoring and validate `animations.json` before base export |
148 | Quick directly delivers a narrated video or MP4 | Also enable Speaker Notes, Narration Audio, and video export; resolve narration-governed timing before audio, requiring timestamped page-local SRT for cue sync or subtitle delivery |
149 | The user explicitly requests static playback or disables object motion | Keep object animation off; retain the remaining notes/audio/video outcomes as requested |
150
151 **Capability boundary**: Default generation does not force object animation or
152 generated audio merely because a deck may later be recorded. Quick with an
153 effective recorded/self-running/video delivery purpose does require Custom
154 Animations, while explicit user instructions for static or
155 page-transition-only playback remain authoritative. This requirement selects
156 the capability, not motion coverage or one effect for every page.
157
158 ---
159
160 ## 4. SVG, Notes, and Motion Realization
161
162 When §3 created `notes/total.md` before SVG, read it once before the first SVG
163 and design each page around its corresponding spoken segment. Give every
164 independently narrated or timed semantic unit a descriptive direct-root `<g
165 id>`; keep inseparable units grouped. Preserve a final/literal script exactly;
166 agent-authored direct-video narration may change only during its final-SVG
167 validation before audio.
168
169 **Hard rule — script/design consistency**: a final script is literal content.
170 If the finished visual page introduces an independent claim or relationship the
171 script does not explain, repair the page or return to planning; never rewrite or
172 pad the final script during the late notes pass. Conversely, every spoken idea
173 that requires visual orientation must have a visible state or deliberate
174 speech-only treatment.
175
176 **Motion readiness**: load `animations.md` before SVG authoring whenever the
177 plan needs compatible Morph endpoints or page/object-specific motion. Author
178 every required start/end state and real semantic group before the final checker;
179 post-processing cannot invent missing visual endpoints or target IDs.
180
181 **Motion restraint**: use transitions, reveals, emphasis, and Morph only for a
182 named communication job. There is no motion-coverage quota, and `effect: none`
183 remains valid. Auto-running narration uses `after-previous` / `with-previous`,
184 never `on-click`.
185
186 **Mandatory when narration governs object motion**: before SVG authoring, load
187 `animations.md` and preserve real semantic groups; before audio, create and
188 validate canonical `animations.json`. After the base PPTX/report and timestamped
189 page audio/SRT, map timed groups in `narration_timing.json`, derive
190 `narration_animations.json`, and export the narrated PPTX/MP4. Only derived
191 triggers/delays wait for SRT; object identity, effect, and order do not. `-a
192 auto` or inherited fixed stagger is not semantic synchronization. For an
193 explicit user-selected static/page-transition-only Quick exception, or for
194 ordinary Default narration-independent deck-wide motion, omit these sidecars
195 and the object-sync claim.
196
197 **Sound effects**: exclude them from this pass and planning artifacts. After
198 final SVG/motion, animation post-processing owns on-demand selection and native
199 PPTX configuration; otherwise remain silent. For direct narrated MP4 delivery,
200 `generate-audio` owns the selected sound-delivery branch: native PowerPoint
201 encoding plus triggered post-export mix, or an explicitly requested real-time
202 PowerPoint slideshow capture. Video gain and limiting never enter
203 `animations.json`; capture uses the balance actually heard during Slide Show.
204
205 **Production sequence**: after the final SVG check, validate any pre-SVG
206 narration against the visible pages; ordinary draft-source runs instead use the
207 final-SVG-grounded notes generation. Split notes, execute the resolved motion
208 path, and export the editable PPTX. Direct Quick video continues
209 through audio and, when required, timestamped SRT. Custom Animations use the
210 narrated-sidecar flow when narration governs group timing;
211 narration-independent custom motion exports its canonical timing without an
212 object-sync claim before the narrated PPTX and MP4.
213
214 ---
215
216 ## 5. Delivery Boundary
217
218 **Canonical artifact**: the editable PPTX remains canonical. `generate-audio`
219 owns provider/voice/rate selection, page audio/SRT generation, semantic
220 narration timing, narrated PPTX export, optional native PowerPoint video export,
221 the explicit slideshow-capture handoff, and the triggered sound-effects mix for
222 direct MP4 delivery.
223
224 **Conditional MP4**: run `powerpoint_video.py --check` only for the native-export
225 branch. If native Windows PowerPoint export is unavailable, keep the narrated
226 PPTX as the successful upstream artifact. An explicit slideshow-capture choice
227 may hand that artifact to a user-operated Windows PowerPoint recorder; it is not
228 complete until the capture is returned and accepted. Do not substitute
229 screenshots, HTML, or a third-party renderer and call it equivalent.
230
231 **Hard rule — choose one PowerPoint video sound boundary**:
232
233 | Delivery branch | Sound contract |
234 |---|---|
235 | Native encoder | PowerPoint supplies visual animation and narration but may omit transition/object sounds. With resolved cues, treat its MP4 as raw and require the verified `video_sound_mix.py` output. |
236 | Real-time slideshow capture | PowerPoint remains the renderer and audio player; a recorder captures the full-screen Slide Show and exactly one application/system-audio source. The accepted capture must contain narration and every configured cue once, and must not enter `video_sound_mix.py`. |
237
238 The branches are mutually exclusive because mixing a capture would duplicate
239 its cues. Keep the native cue configuration in the canonical PPTX. Slideshow
240 capture is explicit and human-audited; it does not inherit the native mix
241 receipt or become an automatic fallback.
242
243 **Current boundary**: importing and automatically splitting one long finished
244 recording is unsupported. Require page-level audio or an explicit page/time map;
245 otherwise deliver the designed deck and frozen notes without claiming audio
246 integration.
247
247 lines MARKDOWN