Script breakdown

CODEX

The script becomes data before week one ends.

Hand Codex the screenplay and get back every scene, every character, and every element as data your whole production can use. The breakdown that used to take days of highlighters is structured, searchable, and repeatable from day one.

What changes
01

Breakdowns in hours, not weeks

Upload the screenplay PDF. Codex returns the full breakdown: scenes, cast, props, stunts, effects. Prep starts with data on day one instead of a stack of marked-up pages.

02

No invented elements

The parse never uses an LLM. It works from curated vocabularies and context rules, so it cannot invent a prop it never read. Same script in, same breakdown out, every time.

03

Every department reads the same script

Cast rollups, INT/EXT and day-night counts, dialogue tallies: one shared picture of the script for the AD team, producers, and every department head.

04

The story feeds the whole studio

Pair Codex with Fabric and the breakdown links into your production graph. Search and reporting then understand your story, not just your files, and reuse gets easier every season.

How it does it

Full scene structure, preserved

Per scene: scene number, heading, INT/EXT, time of day, set and sub-location, characters present, and dialogue blocks with V.O., O.S. and CONT'D parentheticals kept intact. Script-level output adds title, writer, page count, and summary statistics. Delivered as structured JSON.

12-category schema, 9 auto-tagged

Props, vehicles, animals, wardrobe, sounds, SFX, VFX, stunts, and extras tag automatically today. Makeup and hair, set dressing, and special equipment sit in the schema, ready for your tagging workflow.

Cast rollups per character

Every character gets a scene list, dialogue counts, first appearance, and a speaking versus non-speaking flag, ordered by first appearance. Casting and scheduling read straight from the data.

A taxonomy your team manages

436 default terms across 7 managed categories, plus 4 genre packs adding 109 terms and custom pack creation. 17 context-disambiguation rules keep idioms honest: "riding shotgun" never becomes an armorer requisition. Paste in a page and the dry-run tester shows exactly what would extract.

AI kept in its place

AI never touches the parse. It only suggests new taxonomy terms, each with a confidence score and context snippet, into a human accept-or-reject queue with threshold-based bulk accept. Your team decides what enters the vocabulary.

For your engineers

The technical depths.

Counts and limits below are verified against shipping code. The full capability datasheet is available on request.

Codex specifications
Input
Screenplay PDF up to 10 MB, validated by magic bytes, MIME sniff and content scan before parsing
Output
Structured JSON: per-scene records plus script-level title, writer, page count, and scene, character, dialogue and location counts with INT/EXT ratio
Scene fields
Scene number · heading · INT/EXT · time of day · set and sub-location · characters present · dialogue blocks with V.O./O.S./CONT'D parentheticals preserved
Tagging schema
12 categories, 9 auto-tagged today: props, vehicles, animals, wardrobe, sounds, sfx, vfx, stunts, extras · makeup_hair, set_dressing, special_equipment in schema
Cast rollups
Per character: scene list, dialogue counts, first appearance, speaking vs non-speaking, ordered by first appearance
Taxonomy
436 default terms across 7 managed categories · 4 genre packs (109 terms) · custom pack creation · 17 context-disambiguation rules
Vocabulary tools
Full create/edit control for per-customer terms and rules · whole-vocabulary import/export · cross-category term search · paste-in dry-run tester
Parse engine
No LLM in the parse: curated vocabularies and context rules enriched by a statistical language model, time-capped per scene with graceful degrade · AI suggests taxonomy terms only, with confidence scores into a human accept/reject queue
Integration
Embeds as a panel in your dashboards · SSO for people · managed API-key lifecycle (create, list, revoke) for pipelines
Questions
Does Codex use an LLM to break down the script?

No. The parse itself never calls an LLM. Codex extracts elements with curated vocabularies and context rules, enriched by a time-capped statistical language model, so results are repeatable and it never invents an element it did not read. AI is used only to suggest new taxonomy terms, and every suggestion waits in a human accept-or-reject queue.

What does Codex produce from a screenplay?

Codex turns a screenplay PDF into structured JSON: full scene structure including scene number, heading, INT/EXT, time of day, set, characters present and dialogue blocks, an element breakdown on a 12-category schema with 9 categories auto-tagged today, cast rollups for every character, and script-level statistics such as scene, character, dialogue and location counts.

How does Codex fit the tools we already use?

Codex embeds as a panel in your dashboards, uses your SSO for people, and manages API keys for pipelines. Paired with Fabric, breakdown entities link into the production graph for search and reporting. In the four-week Asset Foundry pilot, your script can become the production's context brain in week one.