xQI 3 Pro
Long prompts, text inside the image, and multi-image editing in one call.
xQI 3 Pro generates images from a text prompt and edits input images in the same request: pass `references` and the request becomes an edit, with no separate endpoint and no mode flag. It rewards a paragraph-long prompt, and it is the model to reach for when words have to render correctly inside the image or when several source images have to be combined.
- Edits and combines input images — "keep the face from image 1, the jacket from image 2"
- Several images from one request, produced together rather than one after another
- Optional prompt rewriting, and a reasoning pass that trades time for prompt adherence
- Size-based billing, where `1.5K` bills at the `1K` rate for roughly twice the pixels
Generation is asynchronous. POST answers `202` with a `requestId` and a `pollUrl`; the result arrives by polling that URL or through an HMAC-signed webhook when you pass `webhookUrl`. Signed output URLs stay valid for 23 hours.
Authentication. Add Authorization: Bearer xm_live_… to every request.
Account responsibility. Every request must include the real email of YOUR end-user via endUserEmail. You are accountable for what they generate — monitor activity and act on abuse, or your account may be suspended.
Quick example
Send a generation in 5 lines. Pick your language below.
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "A neon-lit ramen bar at night seen from the street, steam rising over the counter, two customers on stools, cinematic still, shallow depth of field"
}'Request parameters
The fields you can include in the request body. Anything not listed is ignored.
| Name | Type | Required | Default | Description |
|---|---|---|---|---|
endUserEmail | string (email) | yes | — | REQUIRED. Email of YOUR real end-user inside your product — the person who actually triggered this generation. We use it to attribute every request and give you a chain of responsibility: you can list, audit, and delete that user's generations. Do NOT pass a fake address, a shared placeholder, or someone else's email — that's a policy violation. If an end-user generates disallowed content and you do not act, your account can be suspended (see Acceptable use). Example: |
model | string | yes | — | Public model alias. Required — pass `xQI3Pro` to use this model. See the list of currently available aliases in the docs. Example: |
prompt | string | yes | — | Text prompt describing the image. This model rewards detail: subject and action first, then environment, then style, lighting and composition — a paragraph works better than a phrase. Text that must appear inside the image goes in double quotes. When you pass input images, refer to them positionally ("the woman from image 1"). Example: |
references | string | string[] | no | — | Input image(s) to edit or draw from — a single HTTPS URL or an array of 1–3. Passing any switches the request to image editing; there is no separate endpoint or flag. Order is meaningful, so name them positionally in the prompt. Each image must be JPG, JPEG, PNG, BMP, TIFF, WEBP or GIF, at most 10 MB, and works best with both sides between 384 and 2048 px. Example: |
size | string | no | "2K" | Output dimensions: preset `1K`, `1.5K` or `2K` (this model has no `4K`), or explicit pixels (e.g. `1152x2048` for a tall 9:16). Explicit values must total 262,144..4,194,304 pixels with an aspect ratio (W/H) in 1/8..8; anything outside is rejected with `400 validation_error` before any debit. A preset is always resolved to explicit pixels before generation — see `aspectRatio` for which shape you get. Billing is size-based, and an explicit `WxH` bills at the tier its pixel area falls into. Example: |
aspectRatio | "auto" | "square" | "landscape" | "portrait" | no | "auto" | Shape of a preset `size`: `square` (1:1), `landscape` (16:9), `portrait` (3:4). **Specific to this model:** `auto` does not hand the choice to the model — it resolves to the square shape of the chosen preset, because this model is always sent explicit pixels. The exact dimensions behind a preset are not a fixed contract; read the delivered size back from `images[].size`, or pass an explicit `WxH` to pin it (an explicit size wins and `aspectRatio` is then ignored). Billing follows the size tier, so the shape never changes the price. Example: |
n | integer | no | 1 | How many images to generate in this single request. 1–6. They are produced together in one pass rather than one after another, so several images take about as long as one. Each image is billed at the size-based rate. Independent of how many input images you pass — the two are separate budgets. Example: |
negativePrompt | string | no | — | What should NOT appear in the image, as plain comma-separated text — no weights and no special syntax. Effective against the usual artefacts of this model class: `text, watermark, logo, extra fingers`. Example: |
seed | integer | no | random (server-generated) | Determinism seed. Same seed + same prompt + same model + same other params = same output. If omitted, the server picks a random seed for you and returns it in `GET /v1/generations/:requestId` (`input.seed`), so any generation can be reproduced after the fact — even when you didn't pass one in. One exception, and it is this model's own: with `promptExtend: true` the prompt is rewritten before generation and the rewrite differs between runs, so the same seed no longer guarantees the same image. Leave `promptExtend` off when you need reproducibility. Example: |
watermark | boolean | no | false | Whether to add a small platform watermark in the corner. |
responseFormat | "url" | "b64_json" | no | "url" | How images are returned. Only "url" is supported — signed CDN URLs valid for 23 hours. "b64_json" is rejected with `400 validation_error` before any debit; download from the returned URL instead. |
sequential | "auto" | "disabled" | no | "disabled" | Not supported by this model — it returns independent images only. Passing `"auto"` is rejected with `400 validation_error` before any debit. Use `n` for several images in one request, or xSD 4.5 for a connected series. |
promptExtend | boolean | no | false | Let the model rewrite your prompt into a longer, more detailed one before generating. Off by default. It usually helps short or sketchy prompts and gets in the way of carefully written ones — the prompt that runs is not the prompt you sent. Note the consequence for `seed`: the rewrite differs between runs, so with this on the same seed no longer guarantees the same image. Example: |
promptExtendMode | "direct" | "agent" | no | "direct" | How the prompt is rewritten when `promptExtend` is on. `"direct"` expands it in one pass. `"agent"` is a slower, more elaborate rewrite and is **text-to-image only** — combined with `references` it is rejected with `400 validation_error` before any debit. Has no effect when `promptExtend` is off. Example: |
enableThinking | boolean | no | false | Let the model reason about the rewritten prompt before generating: better prompt adherence and composition, at a real cost in time — a request goes from a few seconds to well over a minute (measured 66-85 s), and that overhead is per request, not per image. Requires `promptExtend: true`; sent on its own it is rejected with `400 validation_error` before any debit. Example: |
webhookUrl | string (https URL) | no | — | Optional public HTTPS URL we will POST the final record to once the generation is succeeded or failed. The body is identical to what GET /v1/generations/:requestId returns. Each delivery is HMAC-signed (`X-XMode-Signature: v1=<hex>`) and timestamped (`X-XMode-Timestamp`); reject anything older than 5 minutes. Private/loopback IPs are refused (SSRF guard). See the Webhook delivery section in the Introduction for verification code. Example: |
Response
The POST answers 202 Accepted with this shape — the job is queued, not finished. Read the result from pollUrl, which returns the same record with the fields below filled in. Output URLs are signed and valid for 23 hours.
{
"requestId": "req_7Kd2mQ9xTfL4bWnR",
"model": "xQI3Pro",
"status": "queued",
"prompt": "A neon-lit ramen bar at night, steam over the counter, cinematic still",
"referencesCount": 0,
"endUserEmail": "alice@your-product.com",
"createdAt": "2026-08-21T09:12:44.108Z",
"pollUrl": "https://api.xmode.ai/v1/generations/req_7Kd2mQ9xTfL4bWnR"
}| Name | Type | Required | Default | Description |
|---|---|---|---|---|
requestId | string | yes | — | Unique id of the generation. Use it with GET /v1/generations/:requestId to poll. |
model | string | yes | — | Public alias of the model that will produce the images. |
status | "queued" | "processing" | "succeeded" | "failed" | "expired" | yes | — | Lifecycle marker. POST always returns `queued`. GET returns the current state. `expired` means the request succeeded more than 23 hours ago and signed URLs no longer work — re-run if you need the images again. |
pollUrl | string | no | — | Present on `queued` responses only. Absolute URL to GET for status. Recommended polling cadence: every 5 seconds until terminal. |
images | array | no | — | Present only when status=`succeeded`. Each item has `id`, `expiresAt`, and EITHER `url` (signed, valid 23 h, PNG) OR `error: { code: "storage_failed" }` if our storage layer dropped that single image. If every image of the request is dropped there is nothing to deliver, so the request comes back `failed` and refunded instead. The URL is re-signed on every read, so re-reading this record replaces a link that went stale on your side — but that never moves `expiresAt`, which is fixed when the generation completed. Once it passes, the record returns status `expired` with no URL and the bytes are gone. |
error | object | no | — | Present only when status=`failed`. `{ code, message }`. The same code values as the top-level error format (validation_error, content_policy, provider_error, provider_timeout, internal_error). |
cost | object | no | — | Present only when status=`succeeded`. `{ xTokens: number }`. Debited at request time (size-based); refunded automatically on `failed`. |
finishedAt | string (ISO 8601) | no | — | Present on terminal statuses (`succeeded`, `failed`, `expired`). When the generation reached its final state. |
endUserEmail | string | yes | — | Echoed back so you can confirm the grouping. Always lowercase. |
createdAt | string (ISO 8601) | yes | — | When the request was accepted (POST time). |
Use cases
Common patterns. Copy any block, replace the API key, and you have a runnable request.
Text to image
Simplest case: a prompt, the model alias, and the end-user email. Returns one image at the default size, in the square shape of that preset.
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "A neon-lit ramen bar at night seen from the street, steam rising over the counter, two customers on stools, cinematic still, shallow depth of field"
}'A wide frame at a cheaper tier
Combine `size: "1.5K"` with `aspectRatio: "landscape"` for a 16:9 frame at roughly twice the pixels of `1K` — and at the same rate as `1K`.
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "A lone hiker on a ridge at sunrise, layered mountains fading into haze",
"size": "1.5K",
"aspectRatio": "landscape"
}'Rendering text inside the image
Put the exact words in double quotes. This model places typography more reliably than most, which is what makes posters and mockups practical.
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "A minimalist coffee-shop poster, cream background, a single espresso cup, the title \"MORNING RITUAL\" in bold condensed type across the top",
"size": "2K",
"aspectRatio": "portrait"
}'Editing — combine two images
Pass 1–3 input images as `references` and describe the edit. Refer to them positionally; the order you send is the order the prompt refers to.
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "Keep the woman from image 1 exactly as she is, dress her in the jacket from image 2, and place her on a rainy street at dusk",
"references": [
"https://example.com/portrait.jpg",
"https://example.com/jacket.jpg"
],
"size": "2K"
}'Several variations in one request
Set `n` up to 6. They are generated together, so the wall-clock cost of six is close to that of one; each image is billed separately.
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "Product shot of a matte-black espresso machine on concrete, studio lighting",
"size": "1K",
"n": 6,
"negativePrompt": "text, watermark, logo, reflections of the photographer"
}'Skip polling — receive a webhook
Pass `webhookUrl` and we POST the final record there once the generation is done. The body is the same shape as GET /v1/generations/:requestId. Verify `X-XMode-Signature: v1=<hex>` against your account webhook secret (HMAC-SHA256 of `"<timestamp>.<rawBody>"`).
curl -X POST https://api.xmode.ai/v1/generations \
-H "Authorization: Bearer xm_live_<your-key>" \
-H "Content-Type: application/json" \
-d '{
"endUserEmail": "alice@your-product.com",
"model": "xQI3Pro",
"prompt": "Polaroid-style portrait of a woman with freckles, soft daylight",
"size": "2K",
"webhookUrl": "https://your-app.example.com/hooks/xmode"
}'Limits
- Max images per request (`n`): 6
- Max input images: 3
- Max input image size: 10 MB each, sides 384–2048 px
- Input image formats: JPG, JPEG, PNG, BMP, TIFF, WEBP, GIF
- Total pixels range: 0.26M – 4.19M (presets 1K ≈1 MP, 1.5K ≈2.07 MP, 2K ≈4.19 MP)
- Aspect ratio range: 1/8 – 8
- Prompt max length: 8000 characters
- Output format: PNG
- Signed URL lifetime: 23 hours