Workflow authoring (create, update, validate, lifecycle — headless)
Every chapter so far treats workflows as something someone else built in the editor. This chapter is about building them through the API: creating a workflow from a JSON definition, reading a definition back, replacing it, validating it before you commit, estimating its cost, managing its metadata, and driving the full lifecycle (publish, archive, revert, clone) — all without ever opening the visual editor.
The center of gravity is the public definition schema 1.0: a portable JSON document that
describes a workflow’s graph using the human-readable catalog keys from
09-catalog.md (node type "ai/lifestyle", model "flux-2-dev-edit", preset
"minimalist") instead of internal numeric IDs. The same document shape is what you send to
create, what you get back when you read, and what you submit to update — so a definition is also
the natural export/import format between workspaces and environments.
Permissions. Reads (
GET /workflows,GET /{id},GET /{id}/definition,GET /{id}/versions) require thews:workflows:readpermission; writes (create, update, validate, estimate, delete, and the lifecycle operations) requirews:workflows:manage. Choose the level when you create the API key.
1. The definition document (schema 1.0)
Section titled “1. The definition document (schema 1.0)”A complete example — a product photo and an optional style reference go into an AI lifestyle
node, which feeds an output node. Port names (product_image, reference_image, result) and
parameter names (mood, aspect_ratio, custom_prompt) are exactly what the catalog declares
for these node types — always read GET /api/v1/node-types/{category}/{name} before wiring:
{ "schema_version": "1.0", "nodes": [ { "id": "product_photo", "type": "input/image", "label": "Product Photo" }, { "id": "style_ref", "type": "input/image", "label": "Style Reference", "parameters": { "required": false } }, // optional workflow input { "id": "hero", "type": "ai/lifestyle", "label": "Hero Image", "parameters": { "mood": "minimalist", // preset, by catalog code "aspect_ratio": "16_9", // preset, by catalog code "custom_prompt": "soft morning light, editorial studio look" }, "model": { "model": "flux-2-dev-edit", // model, by public slug "params": { "num_inference_steps": 28 } } }, { "id": "hero_out", "type": "output/image", "label": "Hero" } ], "connections": [ { "from_node": "product_photo", "from_output": "image", "to_node": "hero", "to_input": "product_image" }, { "from_node": "style_ref", "from_output": "image", "to_node": "hero", "to_input": "reference_image" }, { "from_node": "hero", "from_output": "result", "to_node": "hero_out", "to_input": "image" } ], "layout": { // OPTIONAL — see §1.5 "nodes": { "product_photo": { "x": 0, "y": 0 }, "hero": { "x": 640, "y": 120, "size_preset": "lg" } } }}1.1 Nodes
Section titled “1.1 Nodes”| Field | Meaning |
|---|---|
id |
You choose it. Unique within the workflow, pattern ^[a-zA-Z0-9][a-zA-Z0-9_-]{0,63}$. Pick readable slugs (hero, brand_brief); editor-generated GUIDs are also accepted, so round-trips never fail. |
type |
A node type code from the catalog, always category/name (GET /api/v1/node-types). |
label |
Optional display label. For input/output nodes it also drives the derived interface port names (see 03-workflows.md §3). |
parameters |
Values for the parameters the node type declares (see §1.2). |
disabled |
Optional; a disabled node is skipped at execution. |
model |
Optional AI model override (see §1.3). Only meaningful on AI nodes. |
1.2 Parameters: presets, effects, and everything else
Section titled “1.2 Parameters: presets, effects, and everything else”Each node type declares its parameters (name, type, constraints) — read them with
GET /api/v1/node-types/{category}/{name}. Three families matter:
- Preset parameters take a preset code from the parameter’s category:
"mood": "minimalist". Valid codes:GET /api/v1/presets?category=mood. - Effect parameters take an effect code, either plain (
"effect": "foil_gold") or with an intensity:"effect": { "effect": "foil_gold", "intensity": 60 }(integer; meaningful only for effects withsupports_intensity: true). Valid codes:GET /api/v1/effects. - Everything else (
text,number,boolean,enum, …) is passed as a plain JSON value.
Nodes that fill a document template — design/template_render, document/pdf and
aggregate/pdf — select it with template_id (a tpl_… from
GET /api/v1/design-templates) and an optional template_revision,
never with documentId or the other internal document fields the editor stores (rejected with
internal_template_parameter). Their input ports are the template’s field codes, typed like the
fields (text, number, boolean, JSON list for a Repeat, image); read them from the template’s contract,
not from the node type. design/template_render always fixes a published revision (the current one
when template_revision is omitted). document/pdf and aggregate/pdf are not pinned: without
template_revision they render the template’s current draft at run time — set it for a
reproducible workflow. In document/pdf and aggregate/pdf an unconnected field renders the template’s
default value, and a field the template marks required with no default must be connected, or the
definition is rejected with template_field_not_connected (every run would fail). design/template_render
can also receive a field through its data objects, so it checks required fields when it renders
(missing_policy). Reading a definition back returns the same
portable form, also for nodes configured in the editor.
Unknown parameters are not errors. A parameter name the node type does not declare is
discarded with a warning (code unknown_parameter in meta.warnings). This keeps your
definitions tolerant to catalog evolution — but check the warnings: a typo in a parameter name
means the value silently does nothing.
1.3 The model override
Section titled “1.3 The model override”"model": { "model": "flux-2-dev-edit", // public slug; omit to use the workspace/system default "params": { "guidance_scale": 3.5 }, // model parameters (GET /models/{model}/params-schema) "capability_models": { // for multi-capability nodes: per-capability override "outpaint": "flux-2-dev-edit" }}Omit model entirely (or just model.model) and the node uses the effective default for your
workspace — the same one reported by GET /api/v1/node-types/{category}/{name}/models. The model
must support one of the node type’s capabilities, or the write is rejected with
model_not_allowed_for_capability.
params require an explicit model. Model parameters are validated against the declared
schema of the model you name (GET /models/{model}/params-schema): unknown keys are dropped with
a warning (unknown_model_parameter). If you send params without naming a model — or the model
declares no parameter schema — the whole params bag is discarded with a
model_params_schema_unavailable warning: nothing undeclared is ever persisted or forwarded to a
provider. The same filter applies on export: GET /definition only emits declared parameters.
1.4 Connections
Section titled “1.4 Connections”Flat list; every entry names the source node + output port and the target node + input port. Port
names come from the node type’s inputs[]/outputs[] in the catalog. All four fields are
required; a reference to a missing node or port is a blocking error (unknown_node,
unknown_port).
Compact array inputs
Section titled “Compact array inputs”When an array input declares "compact": true, connect individual producers through indexed input
names. This example combines three portable checks with utility/quality_gate:
{ "nodes": [ { "id": "image_check", "type": "media/technical_info" }, { "id": "audio_check", "type": "media/technical_info" }, { "id": "video_check", "type": "media/technical_info" }, { "id": "gate", "type": "utility/quality_gate", "parameters": { "profile": "all_must_pass", "mode": "route" } } ], "connections": [ { "from_node": "image_check", "from_output": "quality_check", "to_node": "gate", "to_input": "checks_0" }, { "from_node": "audio_check", "from_output": "quality_check", "to_node": "gate", "to_input": "checks_1" }, { "from_node": "video_check", "from_output": "quality_check", "to_node": "gate", "to_input": "checks_2" } ]}Indices must be unique and below max_count. The editor keeps them contiguous when connections are
removed. Headless authors should do the same for predictable round-trips. If an upstream node already
returns a json[], connect it once to the base name (checks) instead; the indexed and base-array forms
cannot be mixed.
utility/quality_gate combines evidence; it does not inspect media or call AI. route completes and
exposes the real passed verdict, warn completes while preserving a false verdict and report, and
fail fails the gate node so no dependent delivery node runs. Because Madoo preserves work already
completed, a workflow stopped by a downstream fail-fast gate may finish as partial_success; use the
absence of expected delivery outputs and the node diagnostics to distinguish this from an accepted
delivery.
Audio signal quality
Section titled “Audio signal quality”audio/quality_report measures an authorized storage-backed audio asset without modifying it or calling a
provider. Profiles provide transparent baseline thresholds and any explicit numeric parameter overrides
the profile value. The node exposes both the detailed measurements and a portable check:
{ "nodes": [ { "id": "voice", "type": "input/audio" }, { "id": "audio_quality", "type": "audio/quality_report", "parameters": { "profile": "delivery", "check_id": "voice-delivery", "target_lufs": -16, "maximum_true_peak_dbtp": -1.5, "fail_on_violation": false } }, { "id": "gate", "type": "utility/quality_gate", "parameters": { "profile": "all_must_pass", "mode": "route" } } ], "connections": [ { "from_node": "voice", "from_output": "audio", "to_node": "audio_quality", "to_input": "audio" }, { "from_node": "audio_quality", "from_output": "quality_check", "to_node": "gate", "to_input": "checks_0" } ]}LUFS measures perceived programme loudness; dBTP describes reconstructed true peaks; dBFS is used for the
decoded sample threshold. clippedSampleCount is the actual number of decoded samples at or above
clipping_threshold_dbfs, not a boolean inferred from the container. Internal silent regions are reported
as dropout candidates: the node does not claim to know whether an intentional pause is a defect.
fail_on_violation=false is the composable default: the node completes with passed=false and retains all
evidence. Set it to true only when this check itself must stop dependent nodes. The operation is local and
has zero provider-credit cost, but it still participates in the normal execution, retry and settlement
flow.
Measured loudness normalization
Section titled “Measured loudness normalization”audio/loudness_normalize creates a new audio asset at a consistent delivery loudness. It is not an alias
for audio/volume: in the recommended two_pass mode it first measures the complete programme using the
node’s actual LUFS, true-peak and loudness-range targets, then feeds those measurements into the second
EBU R128 pass. Madoo probes and measures the produced file again before publishing it.
{ "id": "normalize_voice", "type": "audio/loudness_normalize", "parameters": { "profile": "podcast", "check_id": "podcast-delivery", "mode": "two_pass", "output_format": "wav", "fail_on_target_miss": true }}Profiles are visible defaults, not hidden processing modes: podcast starts at -16 LUFS/-1.5 dBTP,
broadcast at -23 LUFS/-1 dBTP and social at -14 LUFS/-1 dBTP. Explicit target_lufs,
true_peak_dbtp, loudness_range_lu and loudness_tolerance_lu values override them. single_pass is
available only when selected explicitly; it is never a silent fallback when two-pass measurement fails.
The editor labels inherited numeric values as profile defaults and lets the author clear an explicit
override to return to the selected profile without persisting a duplicate value.
Across the Audio Finishing pack, advanced=true is the canonical authoring hierarchy rather than an
Editor-only convention. REST, MCP, Agent and Assistant clients should place stable check IDs and technical
profile overrides behind an optional fine-tuning surface. Omitting an override preserves inheritance; it
must not be serialized as the parameter’s numeric minimum.
With fail_on_target_miss=true, an output outside the measured tolerance is not committed. Set it to
false when the workflow must retain and route the failed quality_check. Local normalization uses zero
provider credits while retaining normal storage admission, retries and settlement behaviour.
Measured audio duration fitting
Section titled “Measured audio duration fitting”audio/fit_duration fits one storage-backed clip to one explicit time slot. A connected
target_duration number input, expressed in seconds, overrides the saved parameter. The required tempo
rate is original duration / target duration; Madoo compares it with minimum_rate and maximum_rate
before processing.
{ "id": "fit_localized_cue", "type": "audio/fit_duration", "parameters": { "target_duration": 4.3, "check_id": "localized-cue-duration", "strategy": "tempo_then_pad", "minimum_rate": 0.85, "maximum_rate": 1.2, "preserve_pitch": true, "preserve_formants": true, "pad_position": "end", "tolerance_ms": 80, "on_out_of_range": "fail", "output_format": "wav" }}tempo performs only time stretching. tempo_then_pad may add declared silence when a clamped result is
short; tempo_then_trim may remove the tail when it is long. With an out-of-range required rate, fail
stops before processing, clamp applies the nearest bound and lets the selected correction strategy act,
and warn applies the requested rate while emitting an explicit warning. The produced file is probed
again: actual_duration, fit_report, and quality_check always describe measured output rather than an
assumed filter result. Rubber Band is preferred for pitch/formant preservation; atempo is a declared
fallback. This local operation uses zero provider credits.
Provider-neutral TTS cues
Section titled “Provider-neutral TTS cues”ai/text_to_speech produces one speech clip and a generation_metadata receipt conforming to
madoo.tts-generation/v1. Its required text can be authored or connected from the current translated cue.
Optional input ports language, voice, speaking_rate, style, seed, and reference_audio are
per-cue overrides: a connected value wins over the matching model_config.model_params value. The selected
model must advertise the control; otherwise execution stops before the paid provider call with
TTS_PARAMETER_UNSUPPORTED. F5 and Index TTS require reference_audio; ElevenLabs exposes the main
Italian-dubbing controls without a reference clip.
Use GET /api/v1/models/{model}/params-schema for static model controls and
GET /api/v1/node-types/ai/text_to_speech for the dynamic ports. The receipt records model/provider,
effective non-secret parameters, provider-returned duration/sample rate when available, and generation
time/RTF when measurable.
Named stem separation
Section titled “Named stem separation”ai/stem_separation separates a mixed recording into four mandatory audio outputs: vocals, drums,
bass, and other. These are semantic ports, not positional variants. A provider response missing any stem
fails explicitly with STEM_OUTPUT_INCOMPLETE; it is never published as a partial success. Preserve the
stems_manifest receipt (madoo.stem-separation/v1) for the model, algorithm and effective controls.
For dubbing, use vocals as the dialogue reference and reconstruct the complete background through explicit
audio/mix nodes over drums + bass + other. The other output contains remaining instruments, not the
whole residual background. The Demucs baseline is billed from Madoo’s measured input duration and accepts at
most 90 seconds in one execution; split long-form recordings into bounded segments and restore them on their
original timeline. Use ai/audio_isolation when only a cleaned voice is required and background recovery is
not needed.
Absolute dubbing timeline aggregation
Section titled “Absolute dubbing timeline aggregation”aggregate/audio_timeline closes the per-cue TTS branch and restores speech on the absolute master
timeline. Connect its required clips, cue_id, start, and end inputs to aligned iteration outputs from
the same enumeration root; speaker is optional and must remain aligned when used. The clip may come
directly from TTS or from audio/fit_duration. Pass the measured master duration in seconds to the optional
timeline_duration input when available; it overrides the saved fallback parameter.
{ "id": "localized_speech_timeline", "type": "aggregate/audio_timeline", "parameters": { "timeline_duration": 3600, "overlap_policy": "fail", "missing_cue_policy": "fail", "gap_fill": "silence", "output_format": "wav" }}The node does not merely concatenate clips: every uncovered interval remains silence, preserving pauses in
the source call. fail is the safe default for overlaps and missing iterations. mix deliberately mixes
overlapping speech; trim_previous cuts the earlier cue at the next start. missing_cue_policy=warn renders
the absent interval as silence and records it in both outputs. Rendering is bounded to 16 local inputs per
operation and intermediates are concatenated hierarchically, so long-form timelines do not require every
clip in memory or on local disk at once. Persist timeline_manifest (madoo.audio-timeline/v1) for audit and
route quality_check to utility/quality_gate. Use aggregate/audio_merge only when pauses should collapse.
For caption-driven fitting, connect each cue’s start and end directly to
audio/fit_duration.target_start and target_end. This keeps caption timing as the only source of truth;
the saved target_duration remains a non-caption fallback.
Auditable dubbing delivery manifest
Section titled “Auditable dubbing delivery manifest”After timeline rendering, mix or reconstruct the background, normalize loudness, measure the final audio,
replace the master video’s audio and probe the produced video. Then add utility/dubbing_manifest. Build its
cue_evidence with one aggregate/json row per cue using these exact properties:
{ "cueId": "cue-001", "generation": { "schema": "madoo.tts-generation/v1" }, "fitReport": { "schema": "madoo.audio-processing-report/v1" }}Connect the complete source and localized madoo.caption-track/v1 documents, the cue evidence array,
madoo.audio-timeline/v1, master/final media facts, final audio measurements and every applicable Quality
Check. Set background_mode to provided, separated, or none; separated also requires the
madoo.stem-separation/v1 receipt. Route the emitted quality_check into a blocking Quality Gate and persist
manifest as JSON.
For reusable dubbing, expose a true-default input/boolean such as preserve_background, gate the segmented
Demucs branch with utility/filter, set downstream audio/duck
background_absent_behavior=pass_foreground, and connect the same boolean to
utility/dubbing_manifest.preserve_background. False skips stem separation, passes the dialogue timeline
through without another FFmpeg mix, and records background_mode=none. Do not infer that choice from silence,
loudness, or transcript gaps: they do not reliably distinguish clean speech from quiet ambience or music under
continuous dialogue.
The resulting madoo.dubbing-manifest/v1 records source/target text, absolute cue timing, effective voice and
model, generated/fitted durations, fit state, speaker mapping, stable asset references and processing
provenance. It is built from measured evidence and cannot be supplied as free-form AI JSON.
Measured sidechain ducking
Section titled “Measured sidechain ducking”audio/duck uses foreground speech as an RMS sidechain detector, attenuates the background, and then
mixes both tracks through a peak-safety limiter. Both inputs must be storage-backed audio:
{ "id": "voice_over_mix", "type": "audio/duck", "parameters": { "profile": "voice_over", "check_id": "voice-over-ducking", "output_duration": "foreground", "output_format": "wav", "measure_result": true }}Profiles are transparent baselines: voice_over, subtle, aggressive, or custom. Explicit
threshold_db, ratio, attack_ms, release_ms, background_gain_db, and foreground_gain_db
values override the selected profile. output_duration chooses foreground, longest, or shortest.
With measurement enabled, the processing report compares the base-gain-adjusted background loudness
with the isolated ducked background and reports the observed reduction; the final mix is independently
probed and measured before publication. This local operation uses zero provider credits.
Local speech enhancement
Section titled “Local speech enhancement”audio/speech_enhance cleans one storage-backed recording locally before delivery processing. Start with a
profile and add only deliberate overrides:
{ "id": "clean_voice", "type": "audio/speech_enhance", "parameters": { "profile": "podcast", "check_id": "speech-cleanup", "denoise_technique": "auto", "preserve_ambience": true, "output_format": "wav" }}The light, podcast, and dialogue profiles progressively combine bandwidth filtering, broadband
denoise, bounded click/clip and sibilance repair, gating, and mild dynamics control. custom starts from the
light baseline. Explicit denoise, declick, declip, deesser, high_pass_hz, low_pass_hz, gate,
and compression values override the chosen profile. Strong gate/declip settings can sound unnatural;
preserve_ambience=true keeps denoise and gating gentler.
The output is re-probed and measured. processing_report declares every actual FFmpeg filter and effective
numeric value, while quality_check validates duration preservation and peak safety. The report deliberately
does not invent an absolute SNR estimate. This node does not remove arbitrary room reverberation and does not
separate voice from competing music; use ai/stem_separation when both voice and background are needed, or
ai/audio_isolation when only the cleaned voice is needed, and place
audio/loudness_normalize after cleanup when a delivery loudness target is required. The local operation uses
zero provider credits.
Long-form transcription and caption translation
Section titled “Long-form transcription and caption translation”Long media uses domain-specific enumerator/aggregator pairs rather than a generic JSON loop:
video/extract_audio → enumerate/audio_segments → ai/speech_to_text → aggregate/transcriptaggregate/transcript.caption_track → enumerate/caption_blocks → ai/caption_translation → aggregate/caption_trackaggregate/transcript needs five aligned connections: Speech to Text segments, plus Audio
Segments index, segment_id, start and end. aggregate/caption_track needs both the iterated
translation output and the original caption_track connected directly to source_track. A cue is
one timed subtitle event: translation changes its text but never its ID or timing. See the complete
authoring guide and the runnable
large-media-english-italian-captions.json.
For synchronized natural dubbing request Speech to Text granularity=word with a word-capable model such
as ElevenLabs Scribe V2 or Whisper; Wizper is segment-only. Enable
aggregate/transcript.speech_ready. It groups observed word boundaries into short readable phrases without
crossing speaker changes. After natural-speech finalization keep delivery_captions for technical-slot and
manifest evidence, but connect speech_aligned_captions to video/subtitles; end padding is excluded from
the subtitle display interval.
For a reusable localization workflow, expose the target BCP 47 tag with input/text and connect it to
every ai/caption_translation.target_language; a connected language overrides the saved parameter. Expose
runtime choices with input/boolean and route its strict boolean output through utility/filter to gate
optional branches. For natural speech, connect the measured media/technical_info.duration in seconds to
both utility/speech_timing_plan.duration and aggregate/audio_timeline.timeline_duration, rather than
hard-coding the source length.
For nested composition, a workflow/sub node may call a workflow that contains further workflow/sub
nodes. The complete execution chain must remain acyclic and is limited to five levels. Use
input/json_value when a complete caption track, timing state, quality check, plan, manifest or other JSON
document crosses the child boundary. It preserves exactly one scalar workflow value, including an array root,
and never creates fan-out. Do not use input/json or enumerate/json for this purpose: both are iteration
sources. Child outputs are bridged only after the child completes, so keep independent expensive phases as
parallel sibling sub-workflows rather than accidentally serializing them.
When a consumer in a later phase must authenticate evidence persisted by an earlier child execution, place
utility/execution_reference in the producing workflow. Expose its server-minted execution_id through an
output/text boundary and receive it through input/text in the consuming workflow, alongside the typed
evidence document. The run_ ID is a correlation reference, not a credential: the consumer still validates
tenant ownership, terminal state and the evidence itself. Do not hard-code a previous run ID into a reusable
workflow definition.
For deliver-first translation, set the first ai/caption_translation.failure_policy to best_effort.
aggregate/caption_track then emits the usable track plus pending_captions, containing only cues that still
use source-text fallback, pending_count, and complete. Add one bounded repair corridor:
pending_captions → enumerate/caption_blocks (small blocks, normally at most four cues) →
ai/caption_translation → a second aggregate/caption_track. Connect the original source track to the second
aggregator’s source_track, the first translated track to baseline_track, and the repair iterations to
translations. If the first pass is complete, pending_captions is Absent and the paid repair corridor skips
automatically. Only the second track feeds speech timing and TTS. After that one round, any remaining fallback
is delivered with complete=false and explicit warnings instead of repeating STT, the full translation, or
already valid TTS.
To expose the finished tracks in one player without transcoding the master, add
video/caption_tracks. Its required tracksConfig is a step-shape parameter: each row declares a
stable dynamic input/output port, BCP 47 language, viewer label, captions|subtitles meaning,
auto|caption_track|srt|vtt input format and whether it is the single default track. Fetch the
node’s full catalog detail before authoring; connect the original video plus one timed source per
configured row, then connect bundle to output/json. Madoo stores only tenant-relative paths in
the manifest and signs every referenced asset when an authorized viewer opens it.
Long-video social highlights
Section titled “Long-video social highlights”Treat highlight extraction as a typed, bounded edit pipeline. Expose the requested final length with
input/number (an authored value of 60 is the recommended default), inspect the master once, and
fan it out with enumerate/video_segments. Wire each segment plus its segment_id, index, start
and end to ai/video_highlight_analysis; failure_policy=best_effort lets one failed analysis yield
an empty typed candidate set instead of losing a long-running production.
Connect the iterated candidates directly to aggregate/video_highlight_plan, together with the
measured master duration and numeric target. The planner—not the provider—validates ranges, removes
overlaps, respects the duration ceiling and restores chronological order. If a complete
madoo.caption-track/v1 is available, connect it too so cuts can snap near speech boundaries. For spoken
content, request Speech to Text granularity=word and aggregate the same result twice: use
speech_ready=false as the planner’s technical word-boundary track, and speech_ready=true as the readable
caption track used for subtitle projection. The planner then prefers sentence endings, reliable speaker changes
and strong pauses, allowing only the configured small snap margin around a per-clip duration cap. A short breath
after a comma or other continuation punctuation is not a clean ending; if an AI proposal stops inside an
unfinished thought, the planner can backtrack to a recent complete sentence. After a complete
sentence it may retain up to 0.4 seconds of observed room tone, stopping before the next spoken word; this
keeps a hard cut from landing on the last phoneme. Render
the selected clips by sending the original master and typed plan to video/edit_timeline; its measured
madoo.video-edit-manifest/v1 is the authoritative source-to-output time map for
utility/caption_timeline_project.
For optional burned-in subtitles, expose input/boolean and use utility/filter twice: first to gate
the source media before extraction/transcription, then to gate the projected caption track before
video/subtitles. The false path therefore spends no Speech-to-Text credits and still returns the base
highlight video; the trade-off is that planning cannot snap cuts to caption boundaries. Persist the
highlight plan, edit manifest and quality reports as reviewable outputs. For single-subject social edits use
speaker_labels=hide. In general auto shows multiple speakers only when their identities are global and
reliable; labels scoped to independently transcribed segments are hidden, because speaker_0 in one segment
cannot be assumed to identify the same person in another. Forced labels are humanized (speaker_0 becomes
Speaker 1) in both burned-in captions and sidecars. V1 deliberately uses hard cuts,
because transitions need explicit timing semantics before audio, video and projected captions can remain
frame-aligned. See Video Highlights and the runnable
highlights demo.
1.5 Layout is optional — and faithfully round-tripped
Section titled “1.5 Layout is optional — and faithfully round-tripped”The layout section carries the visual placement (positions, sizes, collapsed state) keyed by
node id, plus per-interface editor layouts under layout.interfaces. It is never required:
| Scenario | Behaviour |
|---|---|
Create without layout |
Works. Madoo assigns default grid positions (one column per topological level), so the workflow opens cleanly in the editor. |
GET /definition of an editor-built workflow |
Full layout returned — the export is faithful. |
PUT /definition without layout |
The stored layout is preserved for nodes that survive the update; new nodes get default positions. |
PUT /definition with layout |
Taken as-is (entries for unknown node ids are dropped with a warning). |
Semantics never depend on layout: two definitions that differ only in layout run identically.
1.6 Custom interfaces
Section titled “1.6 Custom interfaces”The optional interfaces[] section defines custom interfaces
in the same shape the read API exposes, plus the mapping that connects each field to a node input
or parameter (mapping.target_type, mapping.node_id, mapping.parameter_key). Invalid mappings
are blocking errors (invalid_interface). Note that custom interfaces may be gated by your plan.
1.7 Limits
Section titled “1.7 Limits”| Limit | Value | On violation |
|---|---|---|
| Nodes per definition | 200 | 422 (too_many_nodes) |
| Connections per definition | 1000 | 422 (too_many_connections) |
| Request body size | 1 MB | 413 |
2. Creating a workflow
Section titled “2. Creating a workflow”POST {BASE_URL}/api/v1/workflows{ "name": "Newsletter Hero", // required "description": "Hero image pipeline", // optional "tags": ["newsletter"], // optional "definition": { /* schema 1.0 — required */ }}One call, one workflow — there is no “empty shell then fill it” step. Semantics:
- Atomic. Any blocking error →
422(validation_error) with the fullerrors[]list, and nothing is created. All errors are collected in one pass, not just the first. - Warnings don’t block. The workflow is created and the warnings are echoed in
meta.warnings[]. - Always a draft. Publishing is an explicit, separate step (
POST /{id}/publish— see §7). - Plan limits apply: exceeding your plan’s workflow cap returns
402(plan_limit_exceeded).
Safe retries: the Idempotency-Key header
Section titled “Safe retries: the Idempotency-Key header”Network timeouts make “did my create go through?” a real question. Send an optional
Idempotency-Key header (any string up to 255 chars, e.g. a UUID or your job ID) and retries
become safe: the same key, within the same workspace, always resolves to the same workflow. The
first call creates it; a retry with the same payload returns the already-created workflow
with 201 and the response header Idempotency-Replayed: true, instead of a duplicate. The
replay is decided before validation, so a true retry replays even if the catalog changed in
between; concurrent duplicates (two in-flight requests with the same key) are also safe — one
creates, the other replays it.
curl -s -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -H "Idempotency-Key: import-job-7f3a" \ -d @workflow.json "$BASE_URL/api/v1/workflows"Reusing a key with a different payload is rejected with 409 (idempotency_conflict) — it is
almost always a bug in the caller. One key, one logical create; use a new key for a new workflow.
A successful create returns 201 with the same shape as GET /workflows/{id} — including the
derived interface, so you can immediately see the I/O contract your input/output nodes
produced — plus the authoring envelope:
{ "id": "wf_9c1d…", "name": "Newsletter Hero", "status": "draft", "version": 1, "interface": { "inputs": [ /* … */ ], "outputs": [ /* … */ ] }, "meta": { "warnings": [ { "code": "unknown_parameter", "message": "Parameter 'colour' is not declared by node type 'ai/lifestyle' and was discarded.", "node_id": "hero", "parameter": "colour", "path": "$.nodes[2].parameters.colour", "suggestion": "See the declared parameters with GET /api/v1/node-types/ai/lifestyle." } ], "migrations": [] // reserved; always empty in V1 }}The error/warning shape
Section titled “The error/warning shape”Every issue — blocking or not — carries structured locators and a JSONPath into the document you submitted, plus an actionable suggestion when there is one:
{ "code": "unknown_model", "message": "Model 'flux-99' does not exist or is not enabled.", "node_id": "hero", "parameter": null, "path": "$.nodes[2].model.model", "suggestion": "List the selectable models with GET /api/v1/node-types/ai/lifestyle/models." }Blocking errors (422, in errors[]) |
Non-blocking warnings (in meta.warnings[] / warnings[]) |
|---|---|
unsupported_schema_version |
unknown_parameter (dropped) |
too_many_nodes, too_many_connections |
unknown_model_parameter (dropped) |
invalid_node_id, duplicate_node_id |
model_params_schema_unavailable (whole params bag dropped) |
unknown_node_type |
model_ignored (model on a non-AI node, dropped) |
unknown_preset, unknown_effect, invalid_parameter_value |
unknown_layout_node, unknown_layout_interface (dropped) |
unknown_model, model_not_allowed_for_capability, unknown_capability |
invalid_interface_layout (dropped) |
unknown_node, unknown_port, invalid_connection |
disconnected_node, no_output_nodes, unused_input_node, dead_end_branch |
missing_template_id, invalid_template_id, invalid_template_revision, template_not_found, template_revision_not_found, template_pin_failed, internal_template_parameter, template_field_not_connected (template nodes — see 08 §3) |
|
connection_type_mismatch (port data types incompatible — same rules the visual editor enforces: equal types, any, single output into an array slot, pdf↔document) |
|
cycle_detected |
unresolved_reference (export only — see §4) |
invalid_interface, duplicate_interface_id, multiple_default_interfaces |
text_template_placeholder_unconnected, unsupported_inline_node_reference |
3. Validating without saving
Section titled “3. Validating without saving”POST {BASE_URL}/api/v1/workflows/{id}/validateTwo modes:
- With a body
{ "definition": { … } }— validates that document (full catalog resolution + graph checks) without touching the stored one. The pre-flight you run before aPUT. - Without a body — validates the workflow’s stored definition.
Validation problems are the payload here, not an error status: the response is always 200
with valid, authoring_ready, the same errors[]/warnings[] shape as above, and a graph summary.
valid means the graph is structurally executable. authoring_ready is the stronger completion
gate for automated authors: it is false for high-confidence semantic problems such as unused
founder inputs, dead branches, unbound template placeholders or invented inline node references.
Human clients may still save an incremental Draft while this value is false; agents should repair
and re-validate until both values are true.
{ "valid": false, "authoring_ready": false, "errors": [ { "code": "cycle_detected", "message": "…" } ], "warnings": [ { "code": "disconnected_node", "node_id": "stray", "message": "…" } ], "summary": { "total_nodes": 5, "total_connections": 4, "input_nodes": 2, "output_nodes": 1, "has_cycles": true, "node_types_used": ["input/text", "ai/lifestyle", "output/image"], "disconnected_nodes": ["stray"] }}4. Reading a definition
Section titled “4. Reading a definition”GET {BASE_URL}/api/v1/workflows/{id}/definitionGET {BASE_URL}/api/v1/workflows/{id}/definition?version=3Returns the definition in the public schema, wrapped with the version it belongs to:
{ "workflow_version": 4, "definition": { "schema_version": "1.0", /* … */ } }Historical versions are immutable — pinning ?version= always returns the same document.
The export is faithful: an editor-built workflow comes back with its full layout, GUID node
ids and all, and can be re-submitted as-is. That makes GET /definition → POST /workflows the
supported way to copy a workflow across workspaces or environments (the definition is the portable
format; there is no separate export/import endpoint pair in the public API).
The response also carries an ETag header — the fingerprint of the definition content. Hold
on to it: it is what you send back in If-Match to make your next PUT safe against concurrent
edits (see §5). (The header is a concurrency token for PUT, not a caching tag: conditional GETs
with If-None-Match are not processed on this endpoint.)
Unresolvable references degrade safely. If a stored definition references catalog entries
that no longer resolve (a deleted preset, a retired model), the export omits those values —
internal identifiers are never emitted in their place — and reports each omission in a
warnings[] field (code unresolved_reference). Check it before re-importing an old definition.
Listing the versions:
GET {BASE_URL}/api/v1/workflows/{id}/versions{ "workflow_version": 4, "versions": [ { "version": 4, "is_current": true }, { "version": 3, "is_current": false }, { "version": 2, "is_current": false }, { "version": 1, "is_current": false } ]}There is no dedicated “restore” endpoint, because none is needed: restore = read + write.
GET /definition?version=2, then PUT that document back — the old content becomes the new
current version (a new version number; history is never rewritten).
5. Updating a definition
Section titled “5. Updating a definition”PUT {BASE_URL}/api/v1/workflows/{id}/definitionBody: { "definition": { … } }. Full replacement — there is no partial patch of a definition.
Same atomic 422 envelope as create; on success you get 200 with the updated resource and
meta.warnings, exactly like create.
Two behaviours to know:
- Versioning is copy-on-write. Updating a published workflow (or a draft whose current
version is referenced by past executions) increments
versionand writes a new immutable blob; executions pinned to older versions are unaffected. - Concurrency: use
If-Match. Without it, the later save silently wins (last-write-wins).
Optimistic concurrency with ETag / If-Match
Section titled “Optimistic concurrency with ETag / If-Match”The scenario the mechanism protects: your integration reads the definition, a colleague saves a
change from the editor in the meantime, your PUT would silently wipe their work. To prevent it:
GET /definitionreturns anETagheader — a fingerprint of the definition content.- Send it back on the
PUTin anIf-Matchheader. - If the stored definition still matches, the write goes through (
200, with the new ETag in the response — chain it into your next edit). If someone changed it in between, you get412(precondition_failed): re-read, re-apply your change, retry.
ETAG=$(curl -sI -H "Authorization: Bearer $TOKEN" \ "$BASE_URL/api/v1/workflows/$WF/definition" | grep -i '^etag:' | cut -d' ' -f2 | tr -d '\r')
curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -H "If-Match: $ETAG" \ -d @candidate.json "$BASE_URL/api/v1/workflows/$WF/definition"The tag tracks the definition content only: renaming the workflow (PATCH), publishing, or a
copy-on-write version bump that re-writes identical content do not invalidate it — only actual
node/connection/layout changes do. The precondition is checked against an authoritative,
cache-bypassing read while a per-workflow SQL lock is held through the versioned blob and
pointer write, so it holds across server instances and concurrent saves. If-Match: * is accepted (means “the
workflow exists”); a single tag per header is supported (no RFC multi-tag lists). If-Match is
optional, but recommended whenever the editor and your integration may touch the same workflow.
(Layout preservation on PUT without layout: see §1.5.)
6. Metadata, deletion, and finding your drafts
Section titled “6. Metadata, deletion, and finding your drafts”Metadata — name, description, tags — changes without touching the definition:
PATCH {BASE_URL}/api/v1/workflows/{id}{ "name": "New name", "tags": ["v2"] } // absent fields are left unchangedDeletion is a soft delete:
DELETE {BASE_URL}/api/v1/workflows/{id} → 204If the workflow has executions still running you get 409 (executions_in_progress) — cancel or
wait for them first. Deleting a published workflow is allowed; if consumers may still depend on
it, prefer archiving (POST /{id}/archive, §7) — it is reversible.
Listing by status. The workflow list defaults to published-only (unchanged for existing
integrators), but now accepts a filter, and every workflow carries a status field:
GET {BASE_URL}/api/v1/workflows?status=draft // draft | published | archived | allGET /workflows/{id} likewise resolves drafts and archived workflows (a key with read permission
sees the whole workspace surface).
7. Lifecycle: publish, archive, revert, clone
Section titled “7. Lifecycle: publish, archive, revert, clone”A workflow moves between three statuses — draft, published, archived — and the public API
drives every transition. All four endpoints take no meaningful body except clone, and return the
updated resource.
POST {BASE_URL}/api/v1/workflows/{id}/publishPOST {BASE_URL}/api/v1/workflows/{id}/archivePOST {BASE_URL}/api/v1/workflows/{id}/revertPOST {BASE_URL}/api/v1/workflows/{id}/clonePublish makes the workflow executable. Two things to know:
- It is gated on a valid definition. If the stored definition has blocking errors
(a cycle, an invalid interface mapping, …), publish is refused with the same
422validation_errorenvelope as create, and nothing changes. The gate is enforced atomically on the definition snapshot being published (a concurrent invalid save cannot slip through), and it holds platform-wide — the editor’s publish is gated by the same rule. RunPOST /{id}/validatefirst if you want to check without attempting. - Re-publishing increments the version. Publishing a draft makes it executable; publishing an already-published workflow (after definition edits) creates a new immutable version, so integrations pinned to the previous version are unaffected.
Archive takes the workflow out of service: it stops being executable and disappears from the
default list. It is idempotent (archiving twice is fine) and reversible only via clone — an
archived workflow can be neither published nor reverted (422 invalid_state).
Revert sends a published workflow back to draft — it stops being executable until
published again. Reverting an archived workflow is 422 (invalid_state).
Clone creates a brand-new draft (201, new wf_ ID) with the source’s current definition.
It works on any status — it is also the way to “resurrect” an archived workflow’s logic:
// POST /workflows/{id}/clone{ "name": "Newsletter Hero v2", "description": "…", "tags": ["…"] } // name required; // description/tags inherited if absent| Transition | Allowed from | Refused with |
|---|---|---|
publish |
draft, published (re-publish) | 422 invalid_state (archived), 422 validation_error (invalid definition) |
archive |
any (idempotent) | – |
revert |
published (draft = no-op) | 422 invalid_state (archived) |
clone |
any | – (plan workflow cap applies: 402) |
8. Estimating cost (credits)
Section titled “8. Estimating cost (credits)”POST {BASE_URL}/api/v1/workflows/{id}/estimateAnswers “what will a run of this cost?” before you execute — in credits, per node and in
total. Like validate, it has two definition modes: without definition it estimates the stored
definition; with { "definition": { … } } it estimates that document pre-save (blocking mapping
errors → the usual 422 envelope). The optional inputs and interface fields describe the intended
runtime context and use the same input shape as execution.
For example, pass an uploaded dataset when its row count controls paid fan-out:
{ "inputs": { "product_dataset": { "asset_path": "org/workspace/assets/products.xlsx" } }}For the canonical input/data → enumerate/data_rows topology, Madoo counts with the same format,
table, header, locale, delimiter and row-window configuration used at runtime. The selected row count
therefore becomes the exact multiplier for downstream paid nodes before the hold is created. This is a
count-only preflight: the projector does not materialize iterations or invoke providers.
{ "total_estimated_credits": 8, "total_expected_milli_credits": 8000, "total_reserved_milli_credits": 8300, "projection_status": "exact", "is_approximate": true, "nodes": [ { "node_id": "hero", "type": "ai/lifestyle", "model_name": "FLUX.2 Dev Edit", "estimated_credits": 5, "cardinality": 1, "unit_expected_milli_credits": 5000, "unit_reserved_milli_credits": 5000, "expected_milli_credits": 5000, "reserved_milli_credits": 5000, "projection_status": "exact", "is_approximate": false }, { "node_id": "copy", "type": "ai/text_generation", "model_name": "Claude Sonnet", "estimated_credits": 3, "cardinality": 1, "unit_expected_milli_credits": 3000, "unit_reserved_milli_credits": 3300, "expected_milli_credits": 3000, "reserved_milli_credits": 3300, "projection_status": "exact", "is_approximate": true, "warning": "Token-based model: actual cost depends on prompt and output length." } ], "warnings": []}Only billable nodes appear (AI and projected sub-workflow liability; inputs, outputs and ordinary
transformations are free). cardinality is the number of executions propagated through the graph:
distinct enumerator roots multiply, while a shared root is counted once. Milli-credit totals are the
authoritative values; total_estimated_credits and each estimated_credits are rounded compatibility
views. The unit_* fields are the shared runtime price for one invocation; the non-unit liability
fields are that price multiplied by cardinality.
projection_status: "exact" means all paid cardinalities are known. "deferred" means at least one
node cannot be fully priced yet; deferral_reason on that node explains which of the two kinds it is,
and in both cases the top-level milli totals are only the known subtotal, not a full quote.
Deferred kind 1 — the count is unknown (deferral_reason reports a fan-out cause, e.g.
dynamic_enumerator). The node runs an unknown number of times, so cardinality and the total
liability fields are null, while unit_expected_milli_credits and unit_reserved_milli_credits
remain populated whenever the provider model is resolvable: a client can display the server-owned
per-invocation price and only the multiplier is missing.
For enumerate/data_rows, runtime_input_required usually means the estimate did not receive the
dataset asset (or that input/data is fed by another runtime node). Submit the same asset_path you
intend to execute. dataset_inspection_unavailable is retryable storage/inspection unavailability;
dataset_tenant_context_required is reserved for an internal projection surface without workspace
context. Invalid datasets or mappings fail preflight instead of being converted into a deferred zero.
Deferred kind 2 — the unit price is unknown (deferral_reason: "runtime_unit_price"). This is the
mirror image: the node runs a known number of times, but its price depends on a media fact the server
measures immediately before dispatch (today: the duration of the input video for a runtime-priced
model such as Gemini Omni Flash /edit). So cardinality is populated while estimated_credits,
both unit_* fields and both total liability fields are null.
{ "node_id": "restyle", "type": "ai/video_transform", "model_name": "Gemini Omni Flash Edit", "estimated_credits": null, "cardinality": 1, "unit_expected_milli_credits": null, "unit_reserved_milli_credits": null, "expected_milli_credits": null, "reserved_milli_credits": null, "projection_status": "deferred", "deferral_reason": "runtime_unit_price" }Client contract. These money fields are always present and null, never omitted and never
0. A null means “not priced yet”, which is different from “free” — a client that coalesces it to zero will show a run as costless and then be charged. Branch onprojection_status/deferral_reasonbefore reading any amount, and treatestimated_creditsas nullable: it became so whenruntime_unit_pricewas introduced, and a client that types it as a non-nullable number will fail on a workflow containing such a node. The authoritative price for these nodes is settled at dispatch and is visible on the execution, not on the estimate. When executing such a workflow, provide the recommended operational bounds documented in 04-executions §1.2. They stop work earlier than the unconditional balance gate; public REST/MCP callers may omit them, while agent-initiated runs require them.
For deterministic per-unit models expected and reserved match. Token-priced models remain
approximate (is_approximate: true + warning), so reservation can exceed expected while actual charge
settles on real usage. A workflow with no billable nodes returns zero exact totals.
Estimates use the same model resolution as execution (your workspace’s defaults and overrides), so the figure reflects your configuration — and it is always expressed in credits; the public API never exposes the underlying USD economics.
9. The full authoring loop, end to end
Section titled “9. The full authoring loop, end to end”# 1. Discover the building blocks (chapter 09)curl -s -H "Authorization: Bearer $TOKEN" "$BASE_URL/api/v1/node-types" | jq '.data[].type'
# 2. Create (draft) — idempotent against retriesWF=$(curl -s -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -H "Idempotency-Key: my-import-001" \ -d @workflow.json "$BASE_URL/api/v1/workflows" | jq -r '.id')
# 3. Iterate: pre-flight a change, then apply it under If-Matchcurl -s -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -d @candidate.json "$BASE_URL/api/v1/workflows/$WF/validate" | jq '.valid, .authoring_ready, .errors, .warnings'ETAG=$(curl -sI -H "Authorization: Bearer $TOKEN" \ "$BASE_URL/api/v1/workflows/$WF/definition" | grep -i '^etag:' | cut -d' ' -f2 | tr -d '\r')curl -s -X PUT -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -H "If-Match: $ETAG" \ -d @candidate.json "$BASE_URL/api/v1/workflows/$WF/definition" | jq '.version, .meta.warnings'
# 4. Check the price tag, then go live (include runtime inputs when they determine fan-out)curl -s -X POST -H "Authorization: Bearer $TOKEN" -H "Content-Type: application/json" \ -d "{\"inputs\":{\"product_dataset\":{\"asset_path\":\"$DATASET_PATH\"}}}" \ "$BASE_URL/api/v1/workflows/$WF/estimate" | jq '.projection_status, .total_estimated_credits'curl -s -X POST -H "Authorization: Bearer $TOKEN" \ "$BASE_URL/api/v1/workflows/$WF/publish" | jq '.status, .version'
# 5. Execute (chapter 04)Next: 09-catalog.md is the companion reference for every key you write in a definition; 04-executions.md for running what you built.