Skip to content

Catalog discovery (node types, models, presets, effects)

The endpoints in the previous chapters are about running workflows someone has already built. This chapter is about the catalog: the building blocks workflows are made of. The catalog endpoints are read-only and answer four questions:

  1. What kinds of nodes exist? → node types (GET /api/v1/node-types)
  2. Which AI models can a node use, and which one is the default? → models per node type (GET /api/v1/node-types/{category}/{name}/models)
  3. What models and capabilities exist overall? → GET /api/v1/models, GET /api/v1/capabilities
  4. What curated style options can I reference? → presets and effects (GET /api/v1/presets, GET /api/v1/effects)

The catalog is useful for transparency (e.g. showing your users which model powers a node, and what it costs in credits), and it is the companion reference for the workflow authoring API (10-authoring.md): every node type, model slug, and preset/effect code you write in a definition must come from this catalog.

Conventions. All endpoints require the standard bearer token and workspace context (01-authentication.md), are rate-limited like the rest of the V1 API, and return the standard paginated envelope for lists. The catalog is small, so by default list responses are returned whole (has_more: false) — fetch once and work from the full set. Every catalog list also accepts optional limit (max 100) + starting_after if you prefer paging: the cursor is the item’s stable key (type for node types, model for models, {category}/{code} for presets and effects, capability for capabilities) and pages carry next_cursor like every other V1 list. The catalog is English-only in V1: responses ignore Accept-Language and always carry Content-Language: en.

Catalog keys are readable slugs, not opaque IDs. Resources you own (workflows, runs) use prefixed opaque IDs (wf_…, run_…). Catalog entries are identified by human-readable, stable keys instead — node type "ai/lifestyle", model "flux-2-dev-edit", preset "minimalist" — because they are meant to be written by hand (and by agents) inside workflow definitions.


A node type is a kind of step you can place in a workflow: an input (input/image), an AI generation step (ai/lifestyle), a transformation, an output (output/image), and so on. The type identifier is always exactly two segments: category/name.

GET {BASE_URL}/api/v1/node-types
GET {BASE_URL}/api/v1/node-types?category=AI
GET {BASE_URL}/api/v1/node-types?search=video%20merge
Query param Type Description
category string Optional filter by category (AI, Input, Output, Image Processing, …), case-insensitive.
search string Optional free-text query. The query is split into words (on spaces, /, -, _, .); a node matches when every word is found in at least one of its fields (type, name, description, when-to-use, category, capabilities), so multi-word queries like video merge or split-video work where a single contiguous substring would not. Results are ranked by relevance (then type); without search the catalog keeps its category/type order. This is the same matching/ranking the MCP search_node_types tool uses, so both return the same ordering. Use few, broad English outcome words.

Tip for agents: when a multi-word search returns nothing, retry with fewer/broader words rather than a translated phrase — every word must match some field.

Each item describes the node type’s contract — its input/output ports and its parameters:

{
"data": [
{
"type": "ai/lifestyle",
"version": "1.0.0",
"name": "Lifestyle Scene",
"description": "Places a product photo in a generated lifestyle environment.",
"when_to_use": "Use when a product image must be placed in a generated lifestyle scene.",
"category": "AI",
"sub_category": "Generation",
"status": "stable", // "stable" | "preview" | "deprecated" | "latest"
"supports_effects": true,
"enumerable": false, // true when the node produces an iterable output
"inputs": [
{ "name": "product_image", "type": "image", "required": true, "label": "Product Image",
"multi": false, "enumerable": false },
{ "name": "reference_image", "type": "image", "required": false, "label": "Reference",
"multi": false, "enumerable": false }
],
"outputs": [
{ "name": "result", "type": "image", "required": false, "multi": false, "enumerable": false }
],
"parameters": [
{ "name": "mood", "type": "preset", "category": "mood", "required": false,
"advanced": false, "exposable": true },
{ "name": "custom_prompt", "type": "text", "required": false,
"advanced": false, "exposable": true }
],
"icon": "sparkles",
"color": "#FF8800"
}
],
"has_more": false,
"total_count": 42
}

Reading the contract:

  • inputs / outputs are the node’s ports — what flows in and out when the workflow runs. Array ports use a [] type suffix (e.g. "image[]") with optional min_count / max_count. An input may list accepts: further source types it takes besides its own type. output/json’s json port has "accepts": ["number", "boolean"], so one computed number (utility/calculate) or one decision (utility/predicate) can be a workflow result; it reads back as that number or boolean in .value.

  • parameters are the node’s configuration knobs. The type tells you what kind of value the parameter takes; two types deserve attention:

    • "preset" — the value is a preset code from the category given in the parameter’s category field (see §4);
    • "effect" — the value is an effect code (see §4).

    A parameter with visible_when ({ "parameter": "symbology", "value": "qr" }) applies only when the other parameter has that value — or one of the values, when value is a list ({ "parameter": "symbology", "value": ["ean13", "code128"] }); the other parameter’s default counts when it is not set. Editors hide it otherwise.

  • status is the lifecycle stage. Avoid building new integrations on deprecated types.

  • when_to_use explains the practical intent and nearby alternatives. Agents should use it together with the typed contract instead of choosing from the node name alone.

GET {BASE_URL}/api/v1/node-types/{category}/{name}

e.g. GET /api/v1/node-types/ai/lifestyle. Same shape as the list item, plus the capabilities array: the abstract abilities this node needs from an AI model (for AI nodes), e.g. ["image_to_image"]. Capabilities are the link between node types and models — which brings us to the key endpoint of this chapter.

A JSON port or structured parameter can carry schema_ref, for example madoo.quality-check/v1, madoo.quality-policy/v1 or madoo.audio-quality-report/v1. This is a versioned Madoo contract, not merely a descriptive label: use it to understand and validate the exact document shape before composing a workflow. Different ports with the generic type json are compatible only when their declared contracts and the intended transformation agree.

design/template_render publishes three structured JSON outputs. pages_manifest uses madoo.design-pages-manifest/v1 to list selected page images and their durable references; layout_report uses madoo.template-layout-report/v1 to explain page selection, repeated content, image loading and render timings; quality_check reuses madoo.quality-check/v1 for a portable health decision. Resolve each schema_ref through GET /api/v1/json-schemas/{schema_ref}. The editor uses the same catalogue for output-port hover previews and its Output structures inspector.

Node parameters marked hidden: true are server-managed technical metadata. Authoring clients should present them read-only, not as normal editing controls. In design/template_render, document/pdf and aggregate/pdf this includes the numeric document/version IDs, GUIDs and hashes the editor stores; the detail lists instead the public selector parameters template_id (required) and template_revision, with their meaning for that node (MCP get_node_type and the agent’s catalog.get_node_type show the same two and omit the internal ones).

To choose a published DesignDocument template and discover the placeholder codes for its data input, use the design template discovery API. REST v1 and MCP return the same immutable version contract and portable tpl_ identifier.

An array input with compact: true uses one logical array handle in the editor but accepts several individual connections. In a workflow definition those connections are named by appending a zero-based index to the catalog port name: checks_0, checks_1, … up to max_count - 1. min_count and max_count apply to the number of connected values. A producer that already emits one JSON array may instead connect directly to the base port (checks); never mix the base-array and indexed forms in the same node instance.

audio/quality_report is a zero-provider-credit example with two structured outputs. metrics uses madoo.audio-quality-report/v1 for loudness, true/sample peak, near-clipped sample count and silence; quality_check uses the common madoo.quality-check/v1 contract and can be connected directly to utility/quality_gate. Catalog consumers should preserve both references: the detailed measurement report and the portable gate check serve different purposes.

audio/loudness_normalize adds a media output and three structured views of the transformation. Both before and after use madoo.audio-quality-report/v1; processing_report uses madoo.audio-processing-report/v1; quality_check uses madoo.quality-check/v1. The distinction is intentional: the audio is the deliverable, the before/after reports are measured facts, the processing report explains what was applied, and the check is the compact decision consumable by Quality Gate.

audio/fit_duration accepts a required audio, optional dynamic target_duration, and the paired target_start / target_end inputs. For caption-driven dubbing, connect start and end from the same cue; their difference is authoritative. Otherwise the dynamic duration takes precedence over the saved parameter. Its actual_duration output is the measured result in seconds; fit_report uses madoo.audio-processing-report/v1, and quality_check uses madoo.quality-check/v1. Catalog clients must retain the fail, clamp, and warn values of on_out_of_range: they are materially different authoring choices, and no tempo-rate excursion is implicit.

aggregate/audio_timeline is the domain-specific barrier for localized speech. Its clips, cue_id, start, and end inputs must be aligned iteration outputs from one root enumeration; speaker is optional but follows the same rule. A dynamic timeline_duration in seconds overrides the saved fallback. Unlike aggregate/audio_merge, it restores absolute gaps as silence and makes missing cues and overlaps explicit. The timeline_manifest output uses madoo.audio-timeline/v1; quality_check uses madoo.quality-check/v1. Catalog clients must preserve the fail|mix|trim_previous overlap policy and the fail|warn missing-cue policy rather than reducing them to a generic merge option.

utility/dubbing_manifest is the deterministic final barrier for a localized delivery. It consumes complete source/target caption tracks, an exact per-cue evidence array, the typed timeline manifest, measured master and final media, stable media assets and applicable Quality Checks. Its manifest output uses madoo.dubbing-manifest/v1; quality_check uses madoo.quality-check/v1. Catalog clients must preserve the background_mode, speaker_voice_source, require_final_video and advanced tolerance/check-id controls. The node is not a generic JSON generator: it rejects missing or duplicate cue evidence and emits only stable storage:// references, never signed delivery URLs.

audio/duck exposes two required audio inputs: foreground is both the audible foreground and the sidechain detector, while background is attenuated and mixed. ducking_report uses madoo.audio-processing-report/v1 and contains the resolved profile, FFmpeg’s linear threshold, observed integrated background reduction, durations and final signal measurements. quality_check uses madoo.quality-check/v1. Numeric ducking controls intentionally have no catalog default: clients should show the selected profile’s published values until the author explicitly creates an override.

audio/speech_enhance accepts one storage-backed audio input and returns cleaned audio, a processing_report using madoo.audio-processing-report/v1, and a portable quality_check using madoo.quality-check/v1. Its light, podcast, dialogue, and custom profiles are transparent baselines; numeric cleanup controls intentionally have no catalog default so an authoring client can show the inherited profile value without persisting it as an override. denoise_technique=auto currently resolves to the qualified FFT chain and the report publishes the resolved value and exact filters. The node does not claim dereverberation or source separation; those are separate capabilities.

For every Audio Finishing node, catalog clients must preserve the parameter advanced flag. The primary authoring surface is intentionally limited to workflow decisions such as profile, processing mode, output format and failure behaviour; stable check identifiers and technical profile overrides are fine-tuning controls. An omitted optional numeric override is not the parameter minimum: clients should display the inherited profile value when one exists, or an unset state when the selected profile intentionally defines no value. Humanized labels may be shown, but serialized enum values such as two_pass and voice_over must remain unchanged.

Some nodes don’t have a fixed port list — their ports depend on a configuration parameter. The detail response advertises this with roles (e.g. ["dynamic_outputs"], ["aggregator","dynamic_inputs"]), a has_dynamic_ports flag, a dynamic_ports descriptor array, and a value_schema on the driving parameter. The derivations:

  • derivation: "named" — the entries of a composite-config array become ports. The descriptor names the parameter you write, the items_path array inside it, and the name_field/type_field that become each port’s name/type (e.g. enumerate/json → jsonConfig.properties[].outputPort; aggregate/json → jsonAggregateConfig.columns[].inputPort). Write the object described by the parameter’s value_schema; the outputPort/inputPort values you choose become the wireable ports. When every port has the same fixed type the descriptor carries a constant port_type instead of a type_field — e.g. text/template, whose placeholders config ({ "placeholders": [ { "inputPort": "..." } ] }) turns each inputPort into a text input port. The port name IS the literal [inputPort] marker the node replaces in its template text; the reserved name template (the node’s static template input) is rejected. A text/template node needs either a literal template parameter or an inbound connection to its template input, or authoring fails (TEXT_TEMPLATE_MISSING_TEMPLATE).

    Computed columns. A CSV or JSON enumerator mapping may compute its value instead of reading one field: { "outputPort": "amount", "outputType": "number", "isComputed": true, "formula": "round(qty * price, 2)" }. The column’s outputType chooses the language. A number column takes a utility/calculate expression (exact decimals; a wrapping {{ }} is accepted). Any other column takes a Scriban text template such as {{ name | string.upcase }} ({{ sku }}), where every value is text, + joins text and arithmetic operators are rejected. Variables are the row’s fields: CSV headers, JSON keys with . written as _ (pricing.price → pricing_price). In a number formula a CSV field mapped as a number column is read with that column’s separators. A formula that does not compile is rejected at authoring (invalid_dynamic_config, path …[i].formula). A row that cannot be computed (unknown field, text in arithmetic, empty cell, division by zero) fails the node with FORMULA_ERROR, naming the column and the row; it is never an empty value.

  • derivation: "positional" — a parameter produces indexed ports. The descriptor carries name_pattern, index_base, port_type, min_count/max_count and a mode. Two families:

    • Image variants (ai/text_to_image, ai/ad_generator, ai/product_photography, ai/lifestyle, ai/showroom_visualizer, ai/virtual_tryon, ai/background_swap, ai/magic_swap, ai/style_transfer, ai/magic_box) — driven by the integer variants parameter (name_pattern: "variant_{i}", index_base: 1, port_type: "image", 1..4, mode: "single_result_else_variants"):

      • variants: 1 (or omitted) → a single result image output.
      • variants: N (2–4) → variant_1 … variant_N image outputs; the static result is replaced, so wiring from result at N>1 is rejected. Wire from variant_1/variant_2/… instead.
      • variants outside 1..4, or a non-integer value, is rejected at authoring (invalid_dynamic_config).
    • Split (video/split, audio/split) — driven by the segments array parameter (name_pattern: "segment_{i}", index_base: 0, port_type: "video"/"audio", 1..20). Write segments as an array of { "start": "...", "end": "..." } time strings (seconds or HH:MM:SS); you get segment_0 … segment_{N-1} outputs, one per range. At least one segment is required (an empty/absent segments, a non-array, a non-object item, or a non-string start/end is rejected at authoring), max 20.

  • derivation: "expression_variables" — every variable of an expression becomes a required input port of port_type (at most max_count). Used by utility/calculate: write expression, e.g. round(subtotal * (1 + vat_rate), 2) or sum(quote.items[*].amount), and connect a number or a JSON value to each variable (subtotal, vat_rate, quote). Functions (round, floor, ceil, abs, min, max, sum, count, avg, if), true/false and the fields after a . are not variables. An expression that does not parse is rejected at authoring (invalid_dynamic_config, path expression, with the character position). Every variable must be connected (CALC_UNCONNECTED_VARIABLE otherwise), and only from a number, json, boolean or any output: text or a file is rejected with CONNECTION_TYPE_MISMATCH (a CSV or JSON enumerator column feeding a variable must be number-typed). The arithmetic is exact decimal; a text value is never read as a number; division by zero, a missing JSON field or an empty list fails the node with a CALC_* code instead of producing an empty value.

  • Template field ports (design/template_render, document/pdf, aggregate/pdf) — the input ports are the field codes of the template chosen with template_id, typed like the fields. They depend on the template, not on the node type, so the node detail cannot list them: read them from the template’s contract (GET /api/v1/design-templates/{tpl_id}/versions/{revision}/contract, MCP get_design_template). A connection to a code the template does not have fails with unknown_port, whose message lists the valid codes. aggregate/pdf keeps its static introPdf / outroPdf inputs next to them.


2. Models for a node type (the key endpoint)

Section titled “2. Models for a node type (the key endpoint)”
GET {BASE_URL}/api/v1/node-types/{category}/{name}/models

For each capability the node type requires, this returns the models you can select and the default that will be used if you select none:

{
"node_type": "ai/lifestyle",
"capabilities": [
{
"capability": "image_to_image",
"display_name": "Image to Image",
"default_model": "flux-2-dev-edit", // the EFFECTIVE default for YOUR workspace
"default_source": "workspace_override", // "system" | "workspace_override"
"models": [
{
"model": "flux-2-dev-edit",
"display_name": "FLUX.2 Dev Edit",
"provider": "Fal.ai",
"credit_cost": 8,
"credit_cost_milli_credits": 8000,
"cost_unit": "per_image",
"cost_basis": "per_unit",
"quality_tier": 4,
"speed_tier": 3,
"supports_vision": false,
"is_default": true
}
// … more models …
]
}
]
}

Three things to understand:

  • The default is effective, per workspace. A workspace admin can override the system default model for any node type in the Madoo dashboard. This endpoint resolves the default for the workspace your token belongs to and tells you where it came from (default_source). Two API keys from different workspaces can legitimately see different defaults.
  • Economics are in credits, always. credit_cost_milli_credits is the exact base cost per cost_unit (1 credit = 1,000 milli-credits); prefer it for calculations and display. credit_cost remains the backward-compatible whole-credit view. The units are (per_image, per_second, per_1k_tokens, per_1k_characters, per_megapixel, per_request). For token-based (LLM) models the real cost depends on prompt/response length, so they are flagged cost_basis: "per_call_estimate" — treat their credit_cost as an estimate per call. Character-priced TTS models report per_1k_characters; when text is supplied at runtime the execution preview is deferred until that exact text is frozen before provider dispatch.
  • Multi-capability nodes return multiple entries — one per capability, each with its own model list and default.

GET {BASE_URL}/api/v1/models
GET {BASE_URL}/api/v1/models?capability=image_to_image

The same model objects as §2, with one addition: each model carries its capabilities array. Use this for a global “which models does Madoo offer” view; use §2 when you care about a specific node.

GET {BASE_URL}/api/v1/models/{model}/params-schema

e.g. GET /api/v1/models/flux-2-dev-edit/params-schema. Models can expose their own tunable parameters (temperature, guidance scale, native style options…). The schema tells you what they are, their bounds, and their defaults:

{
"model": "flux-2-dev-edit",
"schema": [
{ "name": "guidance_scale", "type": "number", "label": "Guidance Scale",
"min": 1, "max": 20, "step": 0.5, "default": 7.5, "advanced": true }
],
"defaults": { "guidance_scale": 7.5 },
"hides_node_parameters": [] // node parameters superseded when this model is selected
}

hides_node_parameters matters when a model brings a native version of something the node also offers generically: the listed node parameters are ignored while this model is selected.

GET {BASE_URL}/api/v1/capabilities

The full capability vocabulary (image_to_image, text_to_video, …) with display names and categories. Useful for building model-picker UIs or for filtering GET /api/v1/models.


Presets are curated style options (environments, surfaces, moods…) and effects are visual treatments (film grain, chrome…). Workflow nodes reference them by code in their parameters (see §1.1). The catalog gives you the codes plus display metadata — the prompt engineering behind each entry is Madoo’s, and is not exposed.

GET {BASE_URL}/api/v1/presets
GET {BASE_URL}/api/v1/presets?category=environment
GET {BASE_URL}/api/v1/effects
GET {BASE_URL}/api/v1/effects?category=metallic
Query param Type Description
category string Optional filter by category code, case-insensitive. Unknown category → empty list.
// GET /api/v1/presets?category=environment
{
"data": [
{
"code": "white_studio",
"name": "White Studio",
"description": "Clean white studio set with soft, even lighting.",
"category": "environment",
"icon": "studio",
"thumbnail": "https://…/white_studio.jpg",
"sort_order": 1
}
],
"has_more": false,
"total_count": 1
}

Effects additionally carry intensity metadata:

{
"code": "foil_gold",
"name": "Gold Foil",
"category": "metallic",
"supports_intensity": true, // can be applied with an intensity…
"default_intensity": 50, // …and this is the default (0–100)
"sort_order": 1
}

Endpoint Returns
GET /api/v1/node-types (?category=, ?search=) All active node types with ports and parameters; search filters+ranks by relevance
GET /api/v1/node-types/{category}/{name} One node type, including its capabilities
GET /api/v1/node-types/{category}/{name}/models Selectable models + effective default per capability
GET /api/v1/models (?capability=) All enabled models with capabilities
GET /api/v1/models/{model}/params-schema Tunable parameters of one model
GET /api/v1/capabilities The capability vocabulary
GET /api/v1/presets (?category=) Preset codes + display metadata
GET /api/v1/effects (?category=) Effect codes + intensity + display metadata

All list endpoints return the standard envelope with has_more: false; detail endpoints return 404 (code: not_found) for unknown or inactive entries.

For ai/text_to_speech, the node-type detail exposes provider-neutral per-cue input ports while the selected model params schema remains the authority for supported static controls. A connected TTS control overrides the saved model value and is rejected before provider submission when the selected model does not advertise it. The structured generation_metadata output uses madoo.tts-generation/v1.

For ai/stem_separation, catalog clients preserve the four semantic audio outputs (vocals, drums, bass, other) and the stems_manifest schema reference madoo.stem-separation/v1. Its controls are model-owned fine tuning; other is one instrumental stem and must not be relabeled as the full background.