Processing & generation

FORGE

Any format in. Any media out. Governed.

Forge converts what you already own and creates what you need next: video, images, documents, audio and 3D, at fleet scale. Every output can be watermarked and audited, on GPU fleets that ignite on demand and go dark at idle.

What changes
01

Dead formats come back to work

Your archive is full of files nothing opens. Forge converts video, images, documents and audio into the formats your teams use today, so finished work gets reused instead of re-made.

02

New media, made on demand

Concept frames, temp audio, narration, turntables: generated through one governed door instead of a dozen tools and logins. Your teams ask, Forge makes, and every job is tracked.

03

Governed and traceable by default

Outputs carry visible watermarks on video, images and documents. Image outputs can be signed with C2PA Content Credentials, including an AI-source assertion, so downstream teams can see what was made and how.

04

No render farm to babysit

GPU fleets scale from zero to the work and back to zero at idle. Capacity follows the queue, not a hardware budget, and no one babysits the farm.

How it does it

13 job types behind one API

Four processing types cover video, image, document and 3D model work. Nine generative types cover image, video, audio, audio FX, audio ML, text, 3D, vision analysis and transcription. Every one is callable via REST or MCP, the open agent-connector standard, so people and agents use the same door.

60+ models, one contract

A central registry puts 60+ generative model routes behind one API, including current frontier image and video model families. Each model carries capability flags, size and duration limits, and approximate per-unit cost for attribution, with graceful fallback chains. Generation can be pinned to local models or to providers you approve.

Deep in every medium

ProRes and DNxHR deliverables with GPU-accelerated H.264/H.265 encoding. Images to 8192 px with ACES tone mapping and face-aware smart crop. OCR with per-word confidence and AI-ready extraction sidecars. Stem separation up to 6 stems, text-to-speech, and speech-to-text with speaker diarization. 3D turntables up to 360 frames.

Watermarking and provenance built in

Visible text or logo watermarks on video, image and document outputs, with 5 positions plus opacity, tiling, scale, font and color control. Watermark detection on images. C2PA Content Credentials signing on image outputs, with an AI-source assertion and a sidecar-only mode.

Pipelines, not one-offs

A visual workflow builder chains jobs into pipelines: outputs feed inputs by token, playlist stages merge 2 to 10 clips, and generative parameters degrade gracefully to the nearest supported value. HMAC-signed webhooks fire on 6 job lifecycle events, and batch submission handles up to 100 files per call.

For your engineers

The technical depths.

Counts and limits below are verified against shipping code. The full capability datasheet is available on request.

Forge specifications
Job types
13 total: process_video, process_image, process_document, process_3d_model, plus genai_image, genai_video, genai_audio, genai_audio_fx, genai_audio_ml, genai_text, genai_3d, genai_vision, genai_transcript. Callable via REST or MCP; watermarking available as an MCP tool and as parameters on video, image and document jobs.
Video
Containers mp4, mov, avi, mkv, webm plus ProRes and DNxHR deliverable profiles (LB/SQ/HQ 8-bit 4:2:2, HQX 10-bit) · codecs h264, h265, prores, dnxhd, vp9, av1, mjpeg, gif with a GPU-accelerated H.264/H.265 encode path · trims, proxies, conforms, crops · thumbnail extraction by frame or timecode · storyboard/contact sheets up to 1,000 frames · playlist concatenation of 2-10 clips with per-clip trim · media probe to metadata JSON.
Image
14 selectable output formats incl. exr, hdr, dpx, tiff, webp, avif, heic, jxl · 8/16/32-bit depth, ICC profiles, DPI control · tone mapping (ACES, filmic and more) and color-space conversion · smart crop incl. face-aware mode · background removal with a choice of 4 models and alpha matting · multi-rendition variants in one job · analysis operations emitting JSON sidecars · generation to 8192 px.
Documents
26 input extensions across office, markup, scanned images and PDF · outputs: pdf, archival PDF/A, txt, json, docx, md, and per-page png/jpg renders to 600 DPI with a 200-page ceiling · OCR with multi-language support, deskew, scan cleanup and per-word confidence · AI-ready extraction sidecars: plain text, tables JSON with page and bounding-box provenance, per-page layout JSON · PDF toolbox: merge, split, extract, rotate, compress, encrypt, decrypt · author a PDF from inline text, markdown or HTML.
Audio & transcription
Text-to-speech and music generation (10-300 s) via leading voice and music providers, with narration and dialogue modes; wav, mp3, flac, ogg out · stem separation up to 6 stems (vocals, drums, bass, guitar, piano, other) · karaoke split, beat timing, forced lyric alignment, sheet-music engraving · speech-to-text with translation, speaker diarization and word-level timestamps · loudness normalize, audio merge, audio-to-video mux, waveform render, auto-subtitle.
3D
Turntable renders of 1-360 frames at up to 60 fps with GPU path-traced sampling · format conversion to glb, gltf, obj, fbx, stl, ply, dae · camera framing auto, intelligent, or fixed.
Generative catalog
60+ model routes across image, video, audio, text, 3D, vision analysis and transcription behind one API · central registry with per-model capability flags (reference images, inpaint/outpaint with masks, first/last-frame video, seed control), per-model size and duration limits, approximate per-unit cost metadata, and graceful fallback chains · limits: images 1-4 per request to 8192 px; video at 720p, 1080p and 4K.
Watermarking & provenance
Visible text or logo watermarks on video, image and document outputs: 5 positions with opacity, tiling, scale, font and color control · watermark detection on images · C2PA Content Credentials signing on image outputs with an AI-source assertion and sidecar-only mode.
Agents, workflows & scale
~20 MCP tools: job submission (single and batches up to 100 files), status, long-poll wait and results, watermarking, S3 browse and signed-URL minting, capability discovery, and self-describing API access · visual workflow builder with token-passing between steps · account-level webhooks on 6 job lifecycle events, HMAC-signed, JSON or XML, with delivery history and endpoint testing · separate GPU fleets per job type; capacity follows queue depth from zero to fleet and back to zero at idle.
Questions
What is Forge in Fortify's Asset Foundry?

Forge is the processing and generation service of Fortify Media's Asset Foundry. It runs 13 job types, 4 for processing and 9 generative, and puts 60+ generative models behind one API, covering video, images, documents, audio and 3D. It runs in your AWS account or on-prem, and in production deployments content stays in your environment.

Does Forge watermark AI-generated content?

Yes. Forge applies visible text or logo watermarks to video, image and document outputs, and signs image outputs with C2PA Content Credentials, including an AI-source assertion. The platform is SOC 2 Type II aligned and designed for MPA content-security best practices.

Can we control which AI models Forge uses?

Yes. Forge generation can be pinned to local models or to providers you approve. Forge's central registry tracks each model's capabilities, size and duration limits, and approximate per-unit cost for attribution, and falls back gracefully when a model is unavailable.