Skip to main content
POST
Generate video with ControlNet
Generate a video from a prompt while keeping the structure of a control video: edges, depth, pose and similar signals. Send the ControlNet video as init_video, and set controlnet_type to the kind of signal it carries.
model_id is required and must be h3-minimax-controlnet. controlnet_type must be one of the ControlNet models listed on ModelsLab, such as canny, depth, openpose, lineart or scribble.
init_video should be a ControlNet video, meaning a video that already shows the control signal, such as a canny edge map, a depth map or an OpenPose skeleton. It is not the source footage. Its signal must match controlnet_type; for example, send a canny edge video with controlnet_type: canny.
The maximum resolution is 1440 pixels: width and height must each be between 320 and 1440.

Request

Make a POST request to below endpoint and pass the required parameters in the request body.
curl

Body

json
Keep convert_to_controlnet_input set to false when init_video is a ControlNet video, which is the expected input. Setting it to true makes the model extract the controlnet_type signal from an ordinary video first.

Response

The request is queued and answers right away with processing. The future_links field holds the URLs where the video will appear. Poll Fetch Video with the returned id, or pass a webhook to be notified when it’s ready.
json

Body

application/json
key
string
required

Your API Key used for request authorization

model_id
enum<string>
required

The model to use. This endpoint only serves h3-minimax-controlnet.

Available options:
h3-minimax-controlnet
prompt
string
required

Text prompt describing the video to generate.

controlnet_type
string
required

The ControlNet model the control video represents, e.g. canny, depth, openpose, lineart, scribble. Must be one of the ControlNet models listed on ModelsLab.

init_video
string<uri>
required

URL of the ControlNet video: a video that already shows the control signal (e.g. a canny edge map, depth map or OpenPose skeleton) matching controlnet_type, not the source footage. The output follows its structure.

convert_to_controlnet_input
boolean
default:false

Set to true to send an ordinary video and let the model extract the controlnet_type signal from it. Keep false when init_video is a ControlNet video, which is the expected input.

duration
integer
default:5

Output length in seconds. Values outside 5–15 are clamped.

Required range: 5 <= x <= 15
width
integer
default:1280

Output width in pixels, from 320 to 1440.

Required range: 320 <= x <= 1440
height
integer
default:704

Output height in pixels, from 320 to 1440.

Required range: 320 <= x <= 1440
num_inference_steps
integer
default:40

Number of denoising steps. Higher values may improve quality but increase processing time.

seed
integer | null

Seed for reproducible results. Random if omitted.

instant_response
boolean
default:true

Return future links immediately instead of waiting for the result.

temp
boolean
default:false

Store the output in temporary storage.

webhook
string<uri>

URL to receive a POST call once generation completes.

track_id
string

ID returned in the webhook payload to identify this request.

Response

ControlNet video generation response

status
enum<string>

Status of the video generation

Available options:
success,
processing,
error
generationTime
number

Time taken to generate the video in seconds

id
integer

Unique identifier for the video generation

output
string<uri>[]

Array of generated video URLs

Array of proxy video URLs

Array of future video URLs for queued requests

meta
object

Metadata about the video generation including all parameters used

eta
integer

Estimated time for completion in seconds (processing status)

message
string

Status message or additional information

tip
string

Additional information or tips for the user

fetch_result
string<uri>

URL to fetch the result when processing