From a property listing to a finished reel

This system takes one property listing and turns it into a short vertical video: a presenter walks through the home and talks to camera, with cutaway clips of the rooms in between. The two kinds of clip are made by two separate prompt writers that run side by side: one for the presenter (A-roll), one for the rooms (B-roll). They share a start, a toolkit and a finish. Most of the work is planning and writing instructions for an AI video model called Seedance. A person approves every step that costs money. Each part has a short codename; click any box to see what it does.

Words used on this page

Prompt
The written instructions sent to the video model for one clip.
Reference
A photo, video or sound file handed to the video model alongside the prompt, so it copies the real thing instead of inventing it.
A-roll
Presenter shots: the presenter on screen, talking, in one photographed room, in one unbroken take.
B-roll
Room clips: camera moves through a room with no person in it, shown while the voice keeps playing.
Pack
One B-roll video job: several 2-second room shots joined by hard cuts, 6, 8 or 16 seconds long.
Block
One sentence group of the voice, 4 to 10 seconds. One presenter take per block; room clips cut in over it.
Viewpoint
Where the camera stands, which way it points, how wide it sees, and how it moves.
Jerel
The human director who approves words, prompts and spending.

Status

BUILTWorking today on the test version, or finished and saved on the work branch.
BEING BUILTIn progress today.
ROADMAPPlanned for later, or waiting on a decision from Jerel.

Show or hide by status:

Honest note: the parts are built, but no reel with the current presenter has yet gone all the way from listing to finished video and been approved by a person. The current presenter recipe has never rendered a frame, and no room clip has been collected under today's spending checks.

The flow

A shared start, then two lanes that run at the same time, then a shared finish. Codenames follow a film-set theme: each part is named after the crew job it does. On a phone everything stacks top to bottom.

Shared start
↓ The plan splits here: presenter beats go left, room cutaways go right. Both lanes run in parallel. ↓

Presenter lane (A-roll)

One presenter, one room photo, one block of the voice, one continuous take with its own speech. A model writes the acting; code builds the rest.

Room-clip lane (B-roll)

No person, no sound. Code prints the camera instructions from rules; a reviewer model reads them against the photos before any money is spent.

↓ Both lanes meet at one review screen and one spending gate ↓
Shared finish

Same or different?

How the two prompt writers compare. "Shared" means one piece of the system serves both lanes, not two copies.

Presenter lane (A-roll)Room-clip lane (B-roll)Same?
Inputs used One cleaned room photo, the photo's item list, the outfit photo, the presenter's silent video, optionally six face photos, the voice recording and its word-by-word transcript, and the director's draft shot card. Cleaned room photos (a pack can cut between several rooms), the landmarks and keep-or-drop list from the photo reading, and the order of cutaways. Photos only: no person, no sound. Partly: same cleaned photos and photo reading
Who writes the prompt Mostly code: a fixed template filled from typed choices. One creative writing call (Opus 5.5) writes the acting, after a vision check has confirmed every object it names. Code only. The words are printed from rules; no model writes them. A reviewer model (Sonnet) only reads and passes or fails them. (A separate draft-preview route that Jerel asks for by name does use a model call, for drafts only.) Different
Checks before spend Object-by-object photo check (a hard stop), spoken claims checked against the listing facts, reference fingerprints, the "does the presenter move" check, then a read-only rerun of every submit check. Prompt rules check (length, banned camera words), headroom check on every turn, "is each shot new" check, the reviewer model (fails closed, at most two rewrite rounds), a "was this exact shot already paid for" check, then a read-only rerun of the sealed bundle. Different checks, same final gate
Audio The clip carries its own generated speech, driven by the voice recording. The presenter video is muted first so the voice is the only sound. Silent clip. It plays under the voice as a cutaway, or against music when the reel has no presenter. Different
Shot structure One continuous take, 4 to 10 seconds, one block. No cuts, no multi-shot. Multi-shot: 2-second shots joined by hard cuts. 6 s = 3 shots, 8 s = 4, 16 s = 8. At least two shots per room. One exception: a single unbroken 4-second take. Different
Camera moves allowed Hold, follow, lead, snap zoom. No turns, orbits or cuts. Push in, turn left or right, slide left or right (the slides have never been rendered, so they are unproven). Banned: orbits, drone rises, whips, zooms, shake, pulling back. Different menus, one shared rulebook
Lens, height, camera spot Wide (84°) or normal (47°). Chest height by default, eye height allowed. Can later step onto a checked spot on the visible floor. The photo's own perspective, no chosen lens. Standing eye height. Always the photo's own camera spot. Different
What is shared The cleaned photos (PLATES) and the one photo reading (ATLAS); the viewpoint logic (SURVEYOR), which already prints the room clips' crop, lens and move wording and lists every valid angle for both lanes; the request clerk (CLERK); the spending gate (TOLLGATE); the dashboard (VILLAGE); the same Seedance model and account; the cut on the voice clock (CUTTING ROOM). Shared
What the in-progress work changes A lot. The director now hands over a draft shot card (STORYBOARD). The one typed shot card (CALL SHEET), the automatic checks (CONTINUITY) and the revision loop (RETAKE) are all being built for this lane first. Little so far. The move onto the shared viewpoint logic is done and changed no output. The room clips do not get the shot card, the new checks or the revision loop yet; that is SECOND UNIT ROADMAP. Presenter lane first

The moving pieces (variables)

Everything that can change inside one shot taken from one listing photo. The choices differ by lane, so pick one. The choices multiply; tick or untick a row to see the raw count. Raw counts are illustrative: they show how choices multiply before the checks throw out the impossible ones (a turn toward a side of the photo with nothing in it, say).

Raw versions of one shot (illustrative): 0

Worked example: the test listing's living-room photo

What is chosenRoom clip (B-roll)Presenter shot (A-roll)
Aim: which slice of the photo the tall phone frame shows (six positions, left to right)66
Lens1 (photo's own)× 2 (wide, normal)
Camera height (a low camera is refused: it would see surfaces the photo never showed)1 (eye)1 (chest)
Camera spot (today)1 (photo's own)1 (photo's own)
Camera move× 5× 4
Raw camera combinations, before checks (illustrative)3048
Camera angles SURVEYOR actually lists after its checks (real)1424
What the presenter does: talk in place, point something out, walk toward camera, walk awayn/a× 4
Camera feel: steady rig or handheld, none, slight or noticeable shake BEING BUILTn/a (always smooth)× 6
Where the presenter stands, 3 across × 3 deep ROADMAPn/a× 9
New camera spots on the visible floor, say 3 plus the original ROADMAPn/a× 4

The real list is not a simple slice of the raw count: it counts each landmark a move can head for as its own angle, then drops everything a check refuses (for example every turn toward a side with nothing photographed past it). Across all 18 photos of the test listing, SURVEYOR lists 312 room-clip angles and 518 presenter-shot angles. Those real counts cover the camera only; presenter actions, camera feel and standing spots multiply on top, and those multiplied totals are illustrative, not a number the system reports.

Codename legend (every part on one list)
CodenameWhereWhat it doesStatus