返回 ppt-master
image-to-pptx.md
根目录 / skills / ppt-master / workflows / profiles / image-to-pptx.md
1 ---
2 description: Quick-only Generate profile for reconstructing one or more source images into layered, editable PPTX slides.
3 ---
4
5 # Image to PPTX Profile
6
7 > Quick-only Generate profile, not a top-level route. Normalize one or more
8 > supplied images into the represented page roster, then rebuild each page as
9 > native text, identity-faithful source graphics, and independently placeable
10 > image layers.
11
12 Reconstructs approved mockups, rendered slide images, contact sheets, or
13 flattened page visuals into a new PPTX. It does not redesign the deck and does
14 not run Strategist or confirmation. The visible result is the reference truth,
15 but the output is not a screenshot skin: image content is rebuilt into the
16 smallest useful background / foreground / subject stack, visible text is
17 restored as native text, and source graphics remain visually exact.
18
19 **Trigger**: the user supplies one or more raster visuals and explicitly asks
20 to restore the represented pages as a PPTX. Ordinary photos, illustrations, and
21 moodboards used only as resources for a new story do not activate this profile.
22
23 **Support boundary — Codex required**: this profile is currently documented
24 and validated for Codex because it depends on Codex's native reference-image
25 generation/editing capability and direct inspection of every derived layer.
26 Other hosts are not adapted or supported by this workflow; they may happen to
27 work, but the repository makes no compatibility claim and defines no alternate
28 host or generic image-backend fallback for this profile.
29
30 **Hard rule — Quick only**: this profile always loads
31 [`quick-generate.md`](./quick-generate.md) and never
32 [`generate-pptx.md`](../generate-pptx.md). The user does not need to say
33 "Quick" separately. Skip Strategist, Confirm UI, template selection,
34 `design_spec.md`, `spec_lock.md`, and the Default first-page gate. The current
35 main agent decides the reconstruction directly, prepares all resources, hand-
36 authors SVG pages, runs the lockless final checker, and exports.
37
38 **Hard rule — source surface, not template application**: stay in Quick free
39 design and do not install or apply a Brand/Style/Layout/Deck workspace for this
40 profile. A template would compete with the canonical page geometry and visual
41 identity. A reusable-system request routes to Create Template instead.
42
43 ---
44
45 ## 1. Routing and Output Boundary
46
47 | Request shape | Behavior |
48 |---|---|
49 | One or more image files represent pages that must become a final editable PPTX | Activate this profile and Quick Generate |
50 | One file contains several clearly separated slide frames | Split it into the ordered page roster first, then reconstruct each frame |
51 | Images are ordinary content assets or inspiration for a new deck | Use ordinary Generate |
52 | Images should define a reusable Brand/Style/Layout/Deck workspace | Use [`create-template.md`](../create-template.md) |
53 | A semantic source PPTX exists and only its layout should improve | Use [`beautify-pptx.md`](./beautify-pptx.md) |
54
55 **Hard rule — mutually exclusive fidelity profiles**: Image to PPTX and
56 Beautify never compose. Beautify preserves semantic PPTX content while
57 redesigning layout; Image to PPTX preserves a rendered page surface while
58 rebuilding its object and image-layer boundaries.
59
60 **Hard rule — flat final deck**: output stays `pptx_structure.mode: flat`. Page
61 pixels do not prove Master/Layout identity, placeholders, theme ancestry,
62 hidden objects, notes, animations, chart data sources, or authoring history.
63 Do not infer them. A reusable-system request routes to Create Template instead.
64
65 ---
66
67 ## 2. Normalize Source Images into Pages
68
69 🚧 **GATE**: an ordered canonical page-image roster exists before image-layer
70 decisions or SVG authoring.
71
72 | Source form | Normalization |
73 |---|---|
74 | One file containing one complete page | Keep it as one canonical page image |
75 | Several files containing one page each | Preserve the explicit or filename-natural order |
76 | One or more regular contact sheets | Split row-major with `slice_images.py --grid`, without alpha removal |
77 | A file containing several non-grid page frames | Record each visibly bounded page bbox and crop it into a separate lossless canonical page image |
78 | Boundaries or order are genuinely ambiguous | Mark the roster blocked instead of silently merging, dropping, or reordering pages |
79
80 Archive original files under `sources/`; keep normalized page images under
81 `images/source-pages/` or another project-local source-page folder. Never
82 overwrite an original file.
83
84 **Hard rule — one normalized frame, one slide**: frame count, not input-file
85 count, owns slide count. Every normalized page frame maps to one output slide
86 in the same order. Preserve the frame's aspect ratio. Mixed aspect ratios are
87 blocked until the current agent resolves one explicit whole-deck treatment.
88
89 **Mandatory — inspect every canonical page**: ordinary image-resource
90 inspection limits do not apply to this page roster. Inspect each normalized
91 page once to identify text, source graphics, scene-image regions, overlap,
92 region-level source sufficiency, boundary completeness, occlusion, and the
93 minimum useful layer stack. Reopen only the current page or a specifically
94 unresolved region afterward.
95
96 ---
97
98 ## 3. Reconstruct by Content Family
99
100 Classify visible regions by what they are, not by how easy they are to crop.
101
102 | Content family | Default realization | Non-negotiable boundary |
103 |---|---|---|
104 | Editable text | `native_text` | Restore exact visible wording, line grouping, alignment, emphasis, and approximate font metrics; do not bake ordinary slide text into generated images |
105 | Source graphic | `source_graphic` | Logos, icons, badges, vector-like ornaments, and decorative marks preserve visible identity. Use an exact vector, deterministic redraw, sufficient source pixels, or Codex reference reconstruction according to the quality ladder below; never substitute a merely similar graphic |
106 | Data graphic | `native_chart`, `native_table`, or exact `source_graphic` | Preserve every visible value, label, relationship, and geometry. Rebuild natively only when the source is legible enough to verify; otherwise use an exact crop/vector or mark `manual_required`. Never ask a generative model to recreate chart/table/data content |
107 | Simple exact geometry | `native_shape` | Use a native shape only when fill, stroke, geometry, and layering can be matched faithfully; otherwise prepare an identity-faithful source-graphic asset |
108 | Scene image | `image_layer` | Photos, people, characters, products, environmental backgrounds, textures, and complex illustrations may be reference-edited or regenerated as registered layers |
109 | Unreadable or unsafe region | `manual_required` | Block rather than invent wording, identity, values, or a visually different replacement |
110
111 **Hard rule — separate layer need from realization**: source clarity never
112 decides whether a required editable, movable, or overlapping object becomes a
113 separate layer; it decides only how that layer is prepared.
114
115 **Mandatory — assess source sufficiency per region**: judge each region at final
116 display size without a page-wide score or threshold. Inspect detail,
117 contamination/occlusion, and whether identity, geometry, lettering, or data
118 remain verifiable.
119
120 | Source evidence for a required independent image layer | Realization |
121 |---|---|
122 | Complete, cleanly separable, and sufficient at final display size | Prepare a source-derived crop or RGBA layer at the recorded geometry |
123 | Contaminated, occluded, incomplete, or too low-resolution; identity and geometry remain verifiable | Reference-edit or reconstruct the layer and exposed background from that evidence |
124 | Required identity, wording, values, or geometry cannot be verified | Mark `manual_required`; do not invent authoritative content |
125
126 **Graphic identity is authoritative; source pixel bytes are not**: use an exact
127 known vector when available. Deterministically redraw a simple, fully legible
128 graphic as SVG/native geometry. Reuse source pixels only when they are complete
129 and sufficient at final display size. When a complex logo, icon, badge,
130 ornament, or wordmark is visibly identifiable but too low-resolution, use its
131 source crop as the Codex reference and reconstruct a higher-resolution asset
132 that preserves the same silhouette, proportions, colors, lettering, bbox, and
133 z-order. Never merely interpolate low-resolution pixels, redesign the brand,
134 replace it with a similar library icon, or invent unreadable identity. If those
135 properties cannot be verified, mark the graphic `manual_required`.
136
137 **Visible-surface authority**: preserve every legible string, number, label,
138 relative position, crop, z-order, color relationship, and emphasis visible in
139 the source. Do not improve the layout, rewrite copy, correct claims through
140 research, reveal invented semantics, or silently replace branded graphics.
141
142 ---
143
144 ## 4. Build the Minimum Useful Layer Stack
145
146 For each page, decide the smallest stack that makes the intended objects
147 independent. Do not split a page merely to maximize layer count.
148
149 Typical bottom-to-top order:
150
151 1. `base` — clean full-canvas background with all planned removable subjects,
152 foreground objects, and editable text removed; hidden background pixels are
153 reconstructed where necessary.
154 2. `midground-*` — optional scene layers that must sit between the base and
155 primary subject.
156 3. `subject-*` — people, characters, products, props, or other independently
157 movable cutouts.
158 4. `foreground-*` — effects, foliage, particles, framing objects, or other
159 scene elements that cross the subject or native slide objects.
160 5. `source-graphic-*` — exact or identity-faithfully reconstructed logos,
161 icons, badges, and ornamental marks, plus exact/native data graphics at
162 their visible z-order.
163 6. `native-text-*` and exact native shapes.
164
165 **Registered-group rule**: every base/midground/subject/foreground layer in a
166 group stays registered to the same canonical page or scene bbox. A
167 source-derived member retains recorded geometry; every Codex-derived member
168 starts from that canonical source. Preserve canvas, position, scale,
169 pose, lighting, and style. Do not trim registered full-canvas layers;
170 transparent pixels retain alignment.
171
172 When one or more scene layers require reference editing or reconstruction, use
173 [`image-generator.md`](../../references/image-generator.md) §4.4's registered
174 reconstruction group as the primitive:
175
176 - create one clean base by removing **all** scene subjects/foreground objects,
177 source/data graphics, and editable text planned for separate realization,
178 then reconstruct only the newly exposed background;
179 - create at least one independent subject/foreground output from the same
180 canonical source whenever the page contains scene content that must be
181 independently editable; the base plus that output are the minimum two
182 independently prepared image layers, while §3 decides whether each layer
183 retains sufficient source pixels or requires reference reconstruction;
184 - derive every additional layer independently from the canonical source, never
185 from the base or another generated layer;
186 - preserve the original pose, scale, and coordinates on RGBA transparency;
187 - repeat only for additional layers that genuinely need independent movement,
188 overlap, or animation.
189
190 **Batch non-overlapping objects**: one object does not imply one generation.
191 When several subjects, props, effects, or source-graphic reconstructions have
192 pairwise-disjoint padded bboxes—including visible shadows and effects—and can
193 share one isolation treatment, ask Codex for one `layer-plate` containing all
194 of them with clear separation. Use either:
195
196 - a full-canvas registered plate that keeps the source positions, then create
197 one nested-SVG picture crop per recorded bbox; or
198 - a regular isolated-cell sheet when source coordinates are unnecessary, then
199 use `slice_images.py --grid ... --names ... --trim --alpha` and place the
200 resulting assets at their recorded source bboxes.
201
202 Both paths yield independent PPT picture objects from one generated output.
203 If transparency is unavailable, use one exact flat key color for the whole
204 plate and remove it once; never regenerate a separate green-background image
205 for each object. Objects that overlap one another or require different z-order
206 must use separate plates/layers.
207
208 The reference-image CLI does not inherit source dimensions automatically. Pass
209 an explicit matching aspect ratio/size, then verify that every member of one
210 registration group has the same final pixel canvas. In SVG, place the base and
211 all full-canvas layers at identical `x`, `y`, `width`, and `height` with
212 `no-crop` behavior.
213
214 **Reconstruct for final resolution**: apply the §3 source-sufficiency decision
215 per region. Retaining complete source pixels is valid only when they remain
216 sharp enough at final display size; a clear source may still require reference
217 reconstruction when separation needs hidden or uncontaminated pixels. When
218 detail is insufficient, use Codex reference reconstruction; interpolation alone
219 does not recover detail.
220
221 **Reference-edit, not reinterpretation**: reconstruction prompts name the
222 canonical source page/region and ask to preserve the visible composition and
223 style. They may inpaint hidden scene pixels or complete an occluded subject,
224 but must not redesign the scene, change a character/person, introduce text,
225 substitute or alter a logo, or invent extra decorative graphics.
226
227 When Codex cannot return transparency, generate the isolated layer or shared
228 plate on one exact flat key color and use `slice_images.py` as a `1x1` sheet
229 with `--alpha` and **without** `--trim`, preserving full-canvas registration.
230 Several plate members share this single keyed output.
231
232 ---
233
234 ## 5. Source Evidence without a Quick Plan
235
236 Before deciding layers, write source evidence to:
237
238 ```text
239 <project_path>/analysis/reconstruction_inventory.json
240 ```
241
242 The inventory records what is visibly present, not a resumable implementation
243 plan. Keep it limited to:
244
245 - original file and normalized page path;
246 - page order, source-frame bbox, SHA-256, and pixel dimensions;
247 - visible regions with stable ids, source bboxes, observed family
248 (`text`, `graphic`, `image`, or `unknown`), verbatim text when applicable,
249 and confidence;
250 - observed source sufficiency, boundary completeness, occlusion/contamination,
251 and identity/data verifiability at final placement;
252 - overlap/z-order observations and unresolved evidence.
253
254 Do **not** put final layer choices, generation prompts, output filenames, or SVG
255 bindings into this inventory. The current main agent keeps those decisions in
256 active context and writes only required operational image manifests/evidence.
257 Context loss restarts the Quick run; the inventory is not a resume artifact.
258
259 Low-confidence visible text, an uncertain page boundary, or an unidentified
260 branded/data graphic is unresolved evidence and blocks successful delivery.
261
262 ---
263
264 ## 6. Image Preparation
265
266 When any `image_layer` or low-resolution `source_graphic` requires reference
267 editing or generation, load
268 [`image-base.md`](../../references/image-base.md) and
269 [`image-generator.md`](../../references/image-generator.md). The current Codex
270 main agent resolves the layer stack directly, uses Codex's native
271 reference-image capability, and finishes every required layer before SVG
272 authoring. Do not adapt `image_gen.py`, its generic manifest, or provider
273 backends for this profile.
274
275 - Use `text_policy: none` for scene reconstruction layers. Use `embedded` only
276 when an exact visible wordmark/letterform is integral to a reconstructed
277 source graphic; ordinary slide text always remains native.
278 - Exhaust the available Codex image path automatically; block before export if a
279 required layer remains `Needs-Manual`.
280 - Preserve each prepared image layer's source page/region, source hash,
281 realization method, operation, output path/hash, registration group, and
282 z-order in the applicable operational evidence; include prompt and
283 backend/model when the layer was reference-edited or reconstructed.
284 - Re-run `analyze_images.py` after assets change.
285 - A generated candidate is not usable until its expected file exists, it has
286 been inspected once, and its registration group or plate has been checked
287 against the canonical page.
288 - Inspect the recomposed page once after all generated layers, plate crops,
289 source graphics, native shapes, and native text are in place. This narrow
290 readback is mandatory fidelity validation, not resource reselection.
291
292 ---
293
294 ## 7. SVG Authoring and Release Gate
295
296 Follow Quick Generate after source normalization and resource preparation.
297 Hand-author pages serially from the prepared base, registered scene layers,
298 identity-faithful source graphics, native shapes, and native text. Give independently
299 movable layers stable direct-root group ids so later animation can target them.
300
301 **Forbidden — screenshot skin**: do not use the complete source page as the
302 sole full-slide picture and add token editable text above it. The source page
303 is a comparison reference, not a hidden backing layer in the delivered slide.
304
305 Verify each page against its canonical image:
306
307 | Final check | Required evidence |
308 |---|---|
309 | Page roster | Every normalized frame becomes one slide in the same order and canvas treatment |
310 | Native text | Every legible string/number is present verbatim and remains editable |
311 | Source graphics | Logos, icons, and decorative graphics preserve the original identity and geometry at adequate final resolution; no similar substitute or unverified redesign appears |
312 | Data graphics | Every chart/table/data value and relationship is native-and-verified or retained from an exact source asset; none is generatively recreated |
313 | Layer registration | Base, subject, foreground, and other generated layers share the expected canvas/placement and show no jumps, seams, halos, or independent-crop drift |
314 | Visible image fidelity | The recomposed scene preserves the source's visible subject identity, pose, crop, lighting, color relationships, and z-order |
315 | Honest reconstruction | AI-recovered hidden pixels are identified as reconstruction, not claimed as original source detail |
316 | Independent objects | Every layer requested for editing or animation is a distinct SVG/PPT picture object; non-overlapping members may originate from one shared generated plate |
317 | Reference exclusion | Canonical full-page source images remain comparison evidence and are not referenced or packaged as delivered slide media |
318 | Package quality | Quick's lockless final SVG checker and PPTX postflight pass |
319
320 If a generated layer drifts, retry from the canonical reference with a narrower
321 edit instruction. Do not compensate by changing native text/graphics or by
322 flattening the full page. If the Codex image path is exhausted,
323 mark the affected layer `Needs-Manual` and block successful export.
324
325 ```markdown
326 ## ✅ Image to PPTX Complete
327
328 - [x] Source files were normalized into the complete ordered page roster
329 - [x] Visible text is native and verbatim
330 - [x] Source graphics preserve identity and are sharp enough at final size
331 - [x] Required background / foreground / subject layers are independent and registered
332 - [x] Shared plates were split/cropped into the required independent objects
333 - [x] Recombined pages match the supplied visual references
334 - [x] Canonical full-page source images are absent from delivered slide media
335 - [x] Quick's SVG quality gate and PPTX postflight pass
336 - [ ] **Next**: Report the PPTX and identify native, exact-source, and AI-reconstructed objects
337 ```
338
338 lines MARKDOWN