This system takes one property listing and turns it into a short vertical video: a presenter walks through the home and talks to camera, with cutaway clips of the rooms in between. The two kinds of clip are made by two separate prompt writers that run side by side: one for the presenter (A-roll), one for the rooms (B-roll). They share a start, a toolkit and a finish. Most of the work is planning and writing instructions for an AI video model called Seedance. A person approves every step that costs money. Each part has a short codename; click any box to see what it does.
Show or hide by status:
Honest note: the parts are built, but no reel with the current presenter has yet gone all the way from listing to finished video and been approved by a person. The current presenter recipe has never rendered a frame, and no room clip has been collected under today's spending checks.
A shared start, then two lanes that run at the same time, then a shared finish. Codenames follow a film-set theme: each part is named after the crew job it does. On a phone everything stacks top to bottom.
One presenter, one room photo, one block of the voice, one continuous take with its own speech. A model writes the acting; code builds the rest.
No person, no sound. Code prints the camera instructions from rules; a reviewer model reads them against the photos before any money is spent.
How the two prompt writers compare. "Shared" means one piece of the system serves both lanes, not two copies.
| Presenter lane (A-roll) | Room-clip lane (B-roll) | Same? | |
|---|---|---|---|
| Inputs used | One cleaned room photo, the photo's item list, the outfit photo, the presenter's silent video, optionally six face photos, the voice recording and its word-by-word transcript, and the director's draft shot card. | Cleaned room photos (a pack can cut between several rooms), the landmarks and keep-or-drop list from the photo reading, and the order of cutaways. Photos only: no person, no sound. | Partly: same cleaned photos and photo reading |
| Who writes the prompt | Mostly code: a fixed template filled from typed choices. One creative writing call (Opus 5.5) writes the acting, after a vision check has confirmed every object it names. | Code only. The words are printed from rules; no model writes them. A reviewer model (Sonnet) only reads and passes or fails them. (A separate draft-preview route that Jerel asks for by name does use a model call, for drafts only.) | Different |
| Checks before spend | Object-by-object photo check (a hard stop), spoken claims checked against the listing facts, reference fingerprints, the "does the presenter move" check, then a read-only rerun of every submit check. | Prompt rules check (length, banned camera words), headroom check on every turn, "is each shot new" check, the reviewer model (fails closed, at most two rewrite rounds), a "was this exact shot already paid for" check, then a read-only rerun of the sealed bundle. | Different checks, same final gate |
| Audio | The clip carries its own generated speech, driven by the voice recording. The presenter video is muted first so the voice is the only sound. | Silent clip. It plays under the voice as a cutaway, or against music when the reel has no presenter. | Different |
| Shot structure | One continuous take, 4 to 10 seconds, one block. No cuts, no multi-shot. | Multi-shot: 2-second shots joined by hard cuts. 6 s = 3 shots, 8 s = 4, 16 s = 8. At least two shots per room. One exception: a single unbroken 4-second take. | Different |
| Camera moves allowed | Hold, follow, lead, snap zoom. No turns, orbits or cuts. | Push in, turn left or right, slide left or right (the slides have never been rendered, so they are unproven). Banned: orbits, drone rises, whips, zooms, shake, pulling back. | Different menus, one shared rulebook |
| Lens, height, camera spot | Wide (84°) or normal (47°). Chest height by default, eye height allowed. Can later step onto a checked spot on the visible floor. | The photo's own perspective, no chosen lens. Standing eye height. Always the photo's own camera spot. | Different |
| What is shared | The cleaned photos (PLATES) and the one photo reading (ATLAS); the viewpoint logic (SURVEYOR), which already prints the room clips' crop, lens and move wording and lists every valid angle for both lanes; the request clerk (CLERK); the spending gate (TOLLGATE); the dashboard (VILLAGE); the same Seedance model and account; the cut on the voice clock (CUTTING ROOM). | Shared | |
| What the in-progress work changes | A lot. The director now hands over a draft shot card (STORYBOARD). The one typed shot card (CALL SHEET), the automatic checks (CONTINUITY) and the revision loop (RETAKE) are all being built for this lane first. | Little so far. The move onto the shared viewpoint logic is done and changed no output. The room clips do not get the shot card, the new checks or the revision loop yet; that is SECOND UNIT ROADMAP. | Presenter lane first |
Everything that can change inside one shot taken from one listing photo. The choices differ by lane, so pick one. The choices multiply; tick or untick a row to see the raw count. Raw counts are illustrative: they show how choices multiply before the checks throw out the impossible ones (a turn toward a side of the photo with nothing in it, say).
| What is chosen | Room clip (B-roll) | Presenter shot (A-roll) |
|---|---|---|
| Aim: which slice of the photo the tall phone frame shows (six positions, left to right) | 6 | 6 |
| Lens | 1 (photo's own) | × 2 (wide, normal) |
| Camera height (a low camera is refused: it would see surfaces the photo never showed) | 1 (eye) | 1 (chest) |
| Camera spot (today) | 1 (photo's own) | 1 (photo's own) |
| Camera move | × 5 | × 4 |
| Raw camera combinations, before checks (illustrative) | 30 | 48 |
| Camera angles SURVEYOR actually lists after its checks (real) | 14 | 24 |
| What the presenter does: talk in place, point something out, walk toward camera, walk away | n/a | × 4 |
| Camera feel: steady rig or handheld, none, slight or noticeable shake BEING BUILT | n/a (always smooth) | × 6 |
| Where the presenter stands, 3 across × 3 deep ROADMAP | n/a | × 9 |
| New camera spots on the visible floor, say 3 plus the original ROADMAP | n/a | × 4 |
The real list is not a simple slice of the raw count: it counts each landmark a move can head for as its own angle, then drops everything a check refuses (for example every turn toward a side with nothing photographed past it). Across all 18 photos of the test listing, SURVEYOR lists 312 room-clip angles and 518 presenter-shot angles. Those real counts cover the camera only; presenter actions, camera feel and standing spots multiply on top, and those multiplied totals are illustrative, not a number the system reports.
| Codename | Where | What it does | Status |
|---|
Real saved runs, not mock-ups: what the planner decided, then how each prompt writer turned that into the text sent to the video model. Every piece is labelled CODE (fixed rules or typed choices), MODEL (written by an AI model, named where the run recorded it) or JEREL.
District 15 semi-D run · line 1 of the script
Jerel supplied the finished words. The director does not change them; it decides what the viewer sees under each line.
This is one of the cheapest freehold semi-D in Singapore.
One model call read the photos and the script and split each 5-second line into shots: presenter or room, which photo, what it shows and why. Code then checked the shots cover the line with no gaps. This is line 1 of 17.
What it shows: Open on the supplied living-room plate: place Ernest on the visible glossy floor left of the sectional, medium-wide toward camera with the chandeliers, seating, and glazing behind; verify placement before motion.
Why: Open the reel and establish the accepted price-ranking statement through narration.
Planned A-roll over the supplied living-room plate, no generated face: stage Ernest on the visible open glossy floor left of the sectional, medium-wide toward camera with the living room and glazing behind; deliver the market line naturally with one small open-palmed gesture.
What it shows: Start wide on the white facade and gate; make one slow upward drift across the balconies and glazing, ending on the stacked architecture without emphasizing the address plaque. photo 00 remains excluded for the supplied residual watermark lettering.
Why: Ground the opening claim in the actual property exterior.
Silent 2.0-second still cutaway from the supplied exterior plate: start wide on the full white facade and boundary gate, make one slow upward drift toward the stacked balconies, and end on the upper glazing; keep the address plaque secondary, with no captions or invented objects.
What it shows: Return to the same visible living-room floor left of the sectional, medium-wide toward camera with the glazing behind; finish with a small closing gesture.
Why: Complete the accepted opening claim without asking the exterior still to prove it.
Planned A-roll over the same supplied living-room plate, no generated face: return Ernest to the open floor left of the sectional, medium-wide with the glazing behind; finish the sentence with a calm direct look and small closing gesture.
The two presenter shots above share a photo, so they became one take. A movement planner picked from a fixed menu: how to move, stand, gesture and where the camera goes. Code only accepts menu words.
| Kind of take | walk, direction forward |
|---|---|
| Stance / gesture / manner | feet staggered · open-palm talking gestures · face three-quarters to camera |
| Camera | camera backs away as the presenter walks toward it |
| Why this room | The bright living room has visible glossy tiled flooring and open glazed areas, supporting a forward introduction through the seating scene. |
| Start | Begin with the sectional seating, coffee table, and glazed openings in view. |
| End | Advance across the visible tiled floor with the seating and outdoor area beyond the glass anchoring the scene. |
| Standing spot | the far open section of the glossy white tiled living-room floor between the glazed openings and the sectional sofa Marked on the photo by GPT-5.6 Luna, then checked by a second reviewer (GPT-6 through Codex). |
Same run, same line. Mostly code: a fixed template filled with typed choices, plus one creative model call for the performance.
From step 3 above, for line 1:
walk forward, camera backs away, glance once, feet staggered, in the living-room photo, while Ernest says This is one of the cheapest freehold semi-D in Singapore.
Also carried from the director's first note: a medium-wide view toward camera with the chandeliers, seating and glazing behind.
CODE Code builds a request from typed choices and settings:
MODEL-WRITTEN GPT-5.6 Luna Then one creative call writes the six performance lines. It was given the photo and this instruction, which code assembled from the planner's words:
Keep walking. Previs composition: Start: Begin with the sectional seating, coffee table, and glazed openings in view; end: Advance across the visible tiled floor with the seating and outdoor area beyond the glass anchoring the scene; direction: forward; camera intent: camera backs away; camera status: placement not yet fixed.
| First frame and where he stands | Ernest is already walking forward toward the camera on the open glossy tiled floor between the glazed openings and the sectional seating, feet naturally staggered, with the sectional seating, coffee table and glazed openings in view. Kept word for word in the final prompt. |
|---|---|
| How he moves for the whole take | Ernest stays on the open glossy tiled floor between the glazed openings and sectional seating for the whole take, continuing one unhurried forward walk toward the camera with alternating steps and naturally staggered feet; at the last frame one foot carries the weight while the other advances. Kept word for word in the final prompt. |
| What the viewer sees him do | The forward walk carries Ernest across the open glossy tiled floor between the glazed openings and the sectional seating, shifting weight into the next step and ending this beat still advancing with his feet staggered. As the walk continues, his eyes make one brief three-quarter glance toward the glazed openings, then return to the lens while his head follows the glance and turns back toward the camera. Kept word for word in the final prompt. |
| How much of him is in frame | Keep Ernest head to toe with both feet on the photographed glossy tiled floor, at a medium-wide scale that leaves the sectional seating, coffee table and glazed openings visible around him. Kept word for word in the final prompt. |
| What happens after the last word | the forward walk carries on through the final frame, with the feet continuing into the next alternating step. Kept word for word in the final prompt. |
| What the camera does | The handheld phone keeps a level horizon with natural small operator shake and backs away at the presenter's forward walking pace, keeping his face and mouth readable as the photographed room stays within the frame; the retreat continues through the final frame. Kept word for word in the final prompt. |
It also listed the tensions it noticed and how it settled them:
MODEL CHECK Walking check. A Codex reviewer (model name not recorded) read the finished prompt: moves = yes, direction = forward, holding a phone = no, wording that risks a still take = none. Result: passed.
MODEL CHECK Item checks: none needed. Item checks are a vision look at the photo for each object the presenter touches. This take names no object to touch, so none ran.
1,462 words. Highlights show where each piece came from; plain text is the fixed template.
【GENERATION GOAL】 Generate Ernest's speech and synchronized lip movements from @Audio1 while preserving the photographed source plate as one continuous take. Generate Ernest's speech and synchronized lip movements, matching the voice, delivery, pronunciation and cadence in @Audio1. This is a generation, not an edit of any reference. 【REFERENCE ASSET ROLES】 @Image1 — the LOCATION: Bright furnished interior seating area with glossy light-colored tiled flooring, large glazed openings, and an outdoor area visible beyond the glass. The plate alone supplies the visible location, geometry, materials and atmosphere; it is an illustrative room or exterior view. It supplies geography, architecture, materials and light. LOCK the lighting to this reference: light direction, time of day and atmosphere come from @Image1 alone and never change. @Image2 — the ATTIRE STILL: Plain attire photograph documenting a black tee, black trousers and black shoes. It supplies wardrobe and build only, not identity. Do not use its background, lighting or scene. @Video1 — the IDENTITY AND BEHAVIOR REFERENCE: real footage of Ernest speaking to camera. It supplies real Ernest presenter with natural facial articulation and relaxed on-camera mannerisms, skin texture and appearance, plus the presenter's facial articulation, mouth amplitude, expression transitions, blink cadence and relaxed head-and-shoulder mannerisms. Its setting, framing, wardrobe, words, lighting, exposure, colour grade, contrast and softness are not part of this shot — @Video1 contributes identity and motion only, never picture style. @Audio1 — the SPEECH: Local WAV reference from the exact accepted voice-timing window for this block; it carries Ernest's speech timing and a quiet padded tail only. Only its sound enters this generation: no music, no sound effects. 【SUBJECTS】 Ernest maps to @Video1 — real Ernest presenter with natural facial articulation and relaxed on-camera mannerisms — wearing the wardrobe of @Image2: black tee, black trousers and black shoes. One Ernest, existing once. The presenter's articulation and mannerisms follow @Video1. The scene is @Image1 and nothing else: Bright furnished interior seating area with glossy light-colored tiled flooring, large glazed openings, and an outdoor area visible beyond the glass. Visible items: L-shaped dark leather sectional sofa on the right; Additional dark sofa seating near the center-right; Low dark coffee table in front of the seating; Small side tables and stacked items beside the seating; Light wood sideboard along the left wall; Two crystal chandeliers hanging from the ceiling; Ceiling-mounted cassette air-conditioning unit; Recessed ceiling lights; Full-height curtains around the glazed openings; Framed artwork on the right wall; Dark interior door on the rear wall; Outdoor table, chairs, and umbrella visible through the glass; Trees, landscaping, and neighboring structures visible outside. Every object in frame appears in @Image1 or on Ernest's references. One generic living room. ENVIRONMENTAL MOTION CONTRACT — Nothing in the room moves or changes state unless a beat operates it, and the effect of operating it stays within that object. Ernest and the presenter's own cast shadow move naturally. 【THE SPEECH — @Audio1 IS THE SOUND LAW】 CARRIER: @Audio1 is a voice audio file — it has no picture. Every frame of the output shows this generic living room and Ernest at full brightness, first frame to last, opening and closing on the picture with no fade. MASTER AUDIO: Ernest says exactly: "This is one of the cheapest freehold semi-D in Singapore." WORDS OF THIS BLOCK (as heard in @Audio1 — the FILE is the truth): "This is one of the cheapest freehold semi-D in Singapore." PRONUNCIATION — Say natural Singapore English delivery, preserving the accepted words exactly. ROOM VOICE — Natural quiet room tone under Ernest's generated speech; no music or sound effects. MOUTH OWNERSHIP: every audible word belongs to Ernest, and that is the only mouth in frame. The presenter speaks these words and nothing else — no ad-libs, no "uhm", no second voice, no echoed repeat of the words. TAIL: when the last word ends, the presenter's mouth closes and the face stays alive — the blink cadence continues and the eyes go where VISIBLE ACTION puts them — and the forward walk carries on through the final frame, with the feet continuing into the next alternating step. The picture fills the requested duration without truncating speech. FORMAT — ONE CONTINUOUS TAKE. The speech runs the length of @Audio1, and any remaining picture after the stem ends stays quiet to the final frame. Real time. LOCATION MAP — The visible source layout places the visible glossy floor left of the dark sectional in relation to the chandeliers, seating and glazing. Preserve the source edges, depth and screen-left/screen-right relationships. Frame-left stays frame-left within this one continuous take. FRAMING — 9:16 vertical inside the photographed room. a medium-wide view toward camera with the chandeliers, seating and glazing behind. Keep the composition within the photographed source plate. The visible scene is limited to this source inventory: Bright furnished interior seating area with glossy light-colored tiled flooring, large glazed openings, and an outdoor area visible beyond the glass. Visible items: L-shaped dark leather sectional sofa on the right; Additional dark sofa seating near the center-right; Low dark coffee table in front of the seating; Small side tables and stacked items beside the seating; Light wood sideboard along the left wall; Two crystal chandeliers hanging from the ceiling; Ceiling-mounted cassette air-conditioning unit; Recessed ceiling lights; Full-height curtains around the glazed openings; Framed artwork on the right wall; Dark interior door on the rear wall; Outdoor table, chairs, and umbrella visible through the glass; Trees, landscaping, and neighboring structures visible outside. Keep only edges, structure and floor that @Image1 actually shows; never invent an object, surface, wall or depth to satisfy the composition. @Image1 remains the authority for geometry, materials, light and atmosphere. OFF-SCREEN LOCK — Ernest is the only person in frame. Ernest is the only person on camera and the only audible speaker. HEIGHT RULER — Keep Ernest head to toe with both feet on the photographed glossy tiled floor, at a medium-wide scale that leaves the sectional seating, coffee table and glazed openings visible around him. FIRST FRAME AND BLOCKING — Ernest is already walking forward toward the camera on the open glossy tiled floor between the glazed openings and the sectional seating, feet naturally staggered, with the sectional seating, coffee table and glazed openings in view. The audio alone determines mouth state at frame zero; the mouth follows the actual start of speech, including speech at time zero. MOVEMENT CONTRACT — Ernest stays on the open glossy tiled floor between the glazed openings and sectional seating for the whole take, continuing one unhurried forward walk toward the camera with alternating steps and naturally staggered feet; at the last frame one foot carries the weight while the other advances. VISIBLE ACTION — The forward walk carries Ernest across the open glossy tiled floor between the glazed openings and the sectional seating, shifting weight into the next step and ending this beat still advancing with his feet staggered. As the walk continues, his eyes make one brief three-quarter glance toward the glazed openings, then return to the lens while his head follows the glance and turns back toward the camera. Blinks, breath and micro-saccades remain naturally alive through the final words and quiet tail, taking their character from @Video1. The selected actions govern hands, gaze and travel; @Video1 supplies their natural character, not a different route. CAMERA — The handheld phone keeps a level horizon with natural small operator shake and backs away at the presenter's forward walking pace, keeping his face and mouth readable as the photographed room stays within the frame; the retreat continues through the final frame. Keep all revealed edges, surfaces and depth within @Image1. Natural phone rendering, straight lines straight at the frame edges. PHYSICS — Natural human body motion; the photographed architecture, surfaces and visible objects retain their source geometry throughout the take. Skin keeps its own texture and pores. LIGHTING & VISUAL CONSISTENCY — Light, softness, contrast, exposure and colour balance derive strictly from @Image1 and nowhere else: match the photo's interior light exactly — @Image1 alone defines the light direction, time of day, softness, exposure, colour balance and atmosphere., one coherent direction. No relight, no added sun, no beauty key, no rim light. Lighting follows @Image1, the location — do not inherit lighting from @Image2 or @Video1. Exposure holds on the presenter's face through the take. Do not import the other references’ lighting, colour grade, contrast, softness or sharpness onto the presenter or the scene — @Video1 is identity and motion only; the generated picture's entire photographic character comes from @Image1. AUDIO — The soundtrack is Ernest's speech from @Audio1 — the same voice saying the same words with the same timing, heard in the natural acoustic of this generic living room — and nothing else. No music, no sound effects, no added ambience bed, no subtitles.
Farrer Park condo sample, a different listing. All code: rules plus camera viewpoint wording. A separate critic model then checks the prompt against the photos before any money is spent.
This run had no director step. It was a quick sample built from a supplied line and photo pack, so nothing planned the story. What stood in for it:
CODE No model writes this prompt. Code picks each shot's move, target, lens and crop from typed options, then drops them into fixed rules:
| When | Room | Move | Ends on | Lens |
|---|---|---|---|---|
| 0s to 2s | balcony | turn left toward | curved condo towers | 84° wide rectilinear |
| 2s to 4s | bedroom | slide right past | timber hexagon wall panels | 47° standard-normal |
| 4s to 6s | living room | walk slowly toward | bed with cream padded headboard | 84° wide rectilinear |
MODEL CHECK Claude Sonnet Critic, round 1: rejected. Before any money is spent, a separate model reads the prompt next to the photos. It found the living-room shot described its edges wrongly:
REJECTED end landmark: corner window wall — invented geometry / frame edge not in photo (shot 2); choose a different end landmark. photo evidence: In the living-room photo the balcony seating beyond the glass sits at the extreme screen-LEFT margin, immediately behind and below the corner window wall (dark rattan chair and dark table occupying roughly the leftmost quarter of the frame, below mid-height). It is collocated with the corner window wall, not opposite it, so it cannot be 'upper screen-right' while the window wall is 'screen-left', and a left crop whose left edge is the source left boundary cannot have its right edge cut it. The prompt's own ROOM MAP lists 'dark rattan seating and a table on the balcony visible through the left glazing' under SCREEN-LEFT, contradicting this shot line.
Code added the ban never end on: corner window wall and rewrote the whole prompt. The living-room shot became a slow walk toward the bed, which is visible in that open-plan photo. The first draft was not kept on disk, so only the critic's description of it survives.
MODEL CHECK Claude Sonnet Critic, round 2: pass, 0 problems found. 2 critic calls in total, both logged.
4s-6s: WIDEST full-height crop of @Image3; top-left vertical edge is the @Image3 screen-left boundary; farthest frame-right is the @Image3 screen-right boundary; screen-top/bottom are @Image3 source frame boundaries. Start: corner window wall upper screen-left, balcony seating beyond the glass upper screen-left, looped chrome pendant screen-center. Slow gimbal walk toward the bed with cream padded headboard; keeping the bed with cream padded headboard in frame; do not fly to a new subject; stop before its near face; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: bed with cream padded headboard remains at screen-center with surrounding room still visible.
1,588 words, one 6-second clip of three rooms. Every word is placed by code. The purple words were written earlier by the photo-reading model, and code copies them in.
REFERENCES AND DURATION Inputs: @Image1 (balcony listing photo; the only reference for this viewpoint); @Image2 (bedroom listing photo; the only reference for this viewpoint); @Image3 (living room listing photo; the only reference for this viewpoint) — 3 photographed viewpoints. Duration: 6.0 seconds total, 3 shots, 2 hard cuts. 9:16 vertical. SCENE CONTEXT Daylight multi-room listing walkthrough covering balcony (@Image1) and bedroom (@Image2) and living room (@Image3); hard cuts between photographed rooms, no invented corridors. ACTIVE REFERENCES @Image1: source of truth for balcony geometry, furniture, materials, colors, views, and lighting. @Image1 is the only reference for this viewpoint; infer no unseen geometry and stay inside the source photo's visible space. @Image2: source of truth for bedroom geometry, furniture, materials, colors, views, and lighting. @Image2 is the only reference for this viewpoint; infer no unseen geometry and stay inside the source photo's visible space. @Image3: source of truth for living room geometry, furniture, materials, colors, views, and lighting. @Image3 is the only reference for this viewpoint; infer no unseen geometry and stay inside the source photo's visible space. ROOM MAP @Image1 BALCONY: - SCREEN-LEFT (Background): two tall curved condominium towers of stacked balconies filling the left half of the view. - CENTER (Foreground-to-Mid): a frameless glass balustrade with a slim metal handrail running across the balcony edge, a black rattan table with a glass top standing in the middle of the balcony, two black rattan armchairs with cream seat pads and dark floral-print scatter cushions, one either side of the table. - OVERHEAD / CEILING (Foreground): a black conical pendant light hanging from the balcony soffit at the top of the frame. The curved condo towers ends at @Image1 screen-left boundary; nothing exists beyond it. The photographed extent ends at the @Image1 screen-right boundary; nothing exists beyond it. The photographed extent ends at the @Image1 screen-top boundary; nothing exists beyond it. The photographed extent ends at the @Image1 screen-bottom boundary; nothing exists beyond it. @Image2 BEDROOM: - SCREEN-LEFT (Mid): a scattered arrangement of timber hexagon panels climbing the wall above the headboard, a wide cream padded leather headboard running along the left wall. - FAR CENTER (Background): a pale timber-slatted panel wall beside the wardrobe. - CENTER (Foreground-to-Mid): a made double bed under a creased white bedspread with a dark piped edge and a long bolster at the head, a small black bracket mounted on the wall at the head of the bed with a cable running down from it. The timber hexagon wall panels ends at @Image2 screen-left boundary; nothing exists beyond it. The photographed extent ends at the @Image2 screen-right boundary; nothing exists beyond it. The photographed extent ends at the @Image2 screen-top boundary; nothing exists beyond it. The photographed extent ends at the @Image2 screen-bottom boundary; nothing exists beyond it. @Image3 LIVING ROOM: - SCREEN-LEFT (Foreground-to-Far): a black-framed corner window wall on the left with grey curtains, looking onto neighbouring condo towers, dark rattan seating and a table on the balcony visible through the left glazing, a dark square table in the foreground holding a small white teapot and two cups. - CENTER (Foreground): a low dark storage cabinet on castors tucked under the square table. - OVERHEAD / CEILING (Mid): a looped-ribbon pendant on a slim cord hanging high in the upper left of the frame. The corner window wall ends at @Image3 screen-left boundary; nothing exists beyond it. The photographed extent ends at the @Image3 screen-right boundary; nothing exists beyond it. The photographed extent ends at the @Image3 screen-top boundary; nothing exists beyond it. The photographed extent ends at the @Image3 screen-bottom boundary; nothing exists beyond it. Reflective surfaces show reflections only; screens stay off and dark. SOURCE FRAMING Each interval is a full-height 9:16 crop of its source photo; discard width, never squeeze. Top/bottom are source frame boundaries; interval edges are named below. Keep circles, lines, and scale true. MULTI-ANGLE FROM ONE PHOTOGRAPH Each shot runs at most 2 seconds and ends on a hard cut at a named whole-second mark. The only exception is one continuous take with no cuts, which runs 4 seconds. Two shots of one photograph are two framings of the same photographed viewpoint, never a second camera position and never a second room. Each shot names its own framing and its own camera direction, and no two shots of one photograph share both. Each shot names what its left vertical edge cuts, what its right vertical edge cuts, and what sits at screen-top and screen-bottom. Every named edge is an object visible in that photograph. Each shot names what enters frame and from which edge, and names the furthest the frame may travel; nothing exists beyond it. One pace word per shot. Do not repeat a speed or distance restraint inside a shot. Reveal only photographed room content. Do not invent unseen walls, floors, rooms, furniture, décor or reverse sides, and no new room detail or hidden viewpoint appears. POV AND VISIBILITY LOCK Start at each source photo's entry viewpoint at standing eye height; move only through visible space in that photo. Keep behind-camera space and space beyond source edges out of view. No new doors, windows, walls, or rooms; no ceiling above or floor below the captured extent. Never interpolate through unphotographed space between rooms. FORMAT MODE CONTROLLED MULTI-SHOT SEQUENCE. New named-edge crop at each cut; each cut changes crop position or camera direction; landmark per interval. HARD CUT only. No fades or dissolves. CAMERA PATH 0s-2s: Fresh left full-height crop of @Image1; top-left vertical edge is the @Image1 screen-left boundary; farthest frame-right cuts the glass balustrade; screen-top/bottom are @Image1 source frame boundaries. Start: curved condo towers upper screen-left, glass balustrade screen-right. Already moving in a slow in-place pan left toward the curved condo towers; keeping the curved condo towers in frame; the curved condo towers is the furthest the camera turns and nothing beyond it enters frame; do not fly to a new subject; finishes on the curved condo towers; do not close in; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: curved condo towers remains readable with surrounding photographed room still visible. At the 2-second mark: HARD CUT to @Image2; do not interpolate viewpoint. 2s-4s: Fresh left full-height crop of @Image2; top-left vertical edge is the @Image2 screen-left boundary; farthest frame-right cuts the cream padded headboard; screen-top/bottom are @Image2 source frame boundaries. Start: timber hexagon wall panels screen-left, cream padded headboard screen-right. Slow slide right past the timber hexagon wall panels; the camera travels straight sideways while the lens holds its forward bearing and does not pan to track the timber hexagon wall panels; do not fly to a new subject; do not close in; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: timber hexagon wall panels remains at screen-left with surrounding room still visible. At the 4-second mark: HARD CUT to @Image3; do not interpolate viewpoint. 4s-6s: WIDEST full-height crop of @Image3; top-left vertical edge is the @Image3 screen-left boundary; farthest frame-right is the @Image3 screen-right boundary; screen-top/bottom are @Image3 source frame boundaries. Start: corner window wall upper screen-left, balcony seating beyond the glass upper screen-left, looped chrome pendant screen-center. Slow gimbal walk toward the bed with cream padded headboard; keeping the bed with cream padded headboard in frame; do not fly to a new subject; stop before its near face; move at one even constant speed from first frame to last, with no acceleration, deceleration, ease-in, ease-out, or speed ramp. End: bed with cream padded headboard remains at screen-center with surrounding room still visible. OPTICS 0s-2s: 84° wide rectilinear, deep focus; 2s-4s: 47° standard-normal, deep focus; 4s-6s: 84° wide rectilinear, deep focus. Level camera, straight verticals; no fisheye, dutch angle, or lens drift. MOTION Gimbal-stabilized and level, with no shake, sway, or footstep bounce. Each shot's camera move holds one even constant speed from first frame to last, bounded by the photographed space. No acceleration or deceleration. No ease-in. No ease-out. No speed ramp. The camera holds its stated move; photographed architecture stays true and stable. No zoom, optical or digital. Indoor: framing height and horizon stay constant for the whole shot. LIGHTING Match each source photo's daylight, shadows, color temperature, fixture states, curtains, and window state. Exposure stays constant; no relighting, grade, flare, or pumping. AUDIO No audio, no BGM, no subtitles. FIDELITY LOCKS Furniture, materials, colors, and floor pattern stay fixed to the active source photo. Screens stay off with a dark blank screen. Window views and reflections stay consistent. No added furniture, décor, plants, rugs, or staging. No added people. Do not invent pets or animals. No on-screen text, logos, or watermarks. Room proportions and scale stay fixed. WORLD MOTION If something already visible in the active source photo would move in real life, let it move at a natural pace. Do not freeze it. Do not add movers that are not already in the photo. Cable cars or gondolas on a visible cable travel along that cable. Vehicles on a visible road keep driving. Boats or ships on visible water keep moving with the water. Water has a light natural ripple where water is shown. A fan or any other spinning object already in the photo keeps spinning. A real pet already in the photo (dog, cat, bird) may move like a living animal. A stuffed toy, plush, statue, or printed animal stays still. Do not invent cable cars, traffic, boats, people, or animals.
This run is from 22 September. Since 23 September the camera wording (moves, lenses, edges) lives in one shared viewpoint toolkit (SURVEYOR) that both writers call. The wording is the same kind, but it is no longer copied inside the room writer.