01
Text-to-video generation from natural language prompts
02
Image-to-video conditioning: upload a JPEG, PNG, or WebP as the first frame
03
Video extension: continue a completed clip to build longer scenes
04
Sora 2: 720p portrait (720x1280) and landscape (1280x720)
05
Sora 2 Pro: 720p, 1024p portrait (1024x1792) and landscape (1792x1024)
06
Up to 20 seconds per generation (4, 8, 12, and 16-second presets supported)
07
Synchronized speech, sound effects, and ambient soundscapes
08
Non-human character upload (mascots, animals, objects) for multi-clip consistency
09
World-state persistence: objects maintain spatial relationships across cuts
10
Asynchronous Videos API with job ID, status polling (queued, in_progress, completed, failed), and webhook support
11
MP4 output plus downloadable thumbnail and spritesheet
12
Signed download URLs valid for 1 hour after generation
13
Batch API submission available for large offline render queues at reduced cost
14
C2PA metadata and SynthID-style watermarking on all generated outputs