Spicy API
Hình ảnh không kiểm duyệt, tạo hình ảnh thành video và trò chuyện cho ứng dụng người lớn (18+). REST API tương thích OpenAI cùng 15 công cụ MCP: giá mô hình, ước tính chi phí, tạo nội dung với phê duyệt theo từng khoản chi, công việc, mức sử dụng, liên kết nạp tiền và giới hạn chi tiêu hàng tháng. Lời nhắc được sàng lọc trước khi tính phí và khóa sandbox không tính phí gì.
Máy chủ MCP được lưu trữ
npx add-mcp 'https://www.spicyapi.com/api/mcp'Cài vào Claude Code, Codex, Cursor và nhiều công cụ khác
Tài liệu
Spicy API documentation
Spicy API is one REST API for uncensored image, image-editing, video, video-editing, speech, live voice call, transcription, embedding and chat generation, sold per generation from a prepaid US-dollar balance: no subscription, no GPUs to run. It is OpenAI-compatible (point the OpenAI SDK at https://api.spicyapi.com/v1 with a Spicy API key) and includes ready-made sex actions for video and role-play companion models. Not to be confused with spicyapi.ai, an unrelated model aggregator.
Base URL https://api.spicyapi.com. Authenticate with Authorization: Bearer sk-spicy-... (keys from https://www.spicyapi.com/dashboard/api-keys, shown once, stored hashed). Check your balance with GET /v1/account.
Quickstart
from openai import OpenAI
client = OpenAI(api_key="sk-spicy-...", base_url="https://api.spicyapi.com/v1")
# Images: the OpenAI SDK's images.generate maps onto POST /v1/images/generations
img = client.images.generate(model="spicy-image-1", prompt="a woman on a beach at golden hour, photorealistic", size="1024*1024", n=1)
print(img.data[0].url) # durable CDN URL, reusable as image_url for edits and video
# Chat: fully compatible, streaming included
stream = client.chat.completions.create(model="spicy-chat-1", messages=[{"role": "user", "content": "hey"}], stream=True)
curl https://api.spicyapi.com/v1/videos/generations \
-H "Authorization: Bearer $SPICYAPI_KEY" -H "Content-Type: application/json" \
-d '{"model": "spicy-motion-2", "prompt": "slow build, steady rhythm", "image_url": "<a URL from /v1/images/generations>", "resolution": "720P", "duration": 5}'
# 202 {"id": "sj_...", "status": "queued", "cost_usd": 1.00} then poll every 10 to 15 seconds:
curl https://api.spicyapi.com/v1/videos/tasks/sj_... -H "Authorization: Bearer $SPICYAPI_KEY"
Models
| Model | Kind | Price | Limits |
|---|---|---|---|
spicy-image-1-pro | image | $0.09 per image | sizes 10241024, 8321216, 1216832, 12801280, 14401440, 10241536, 15361024, 10801920, 19201080, 11522048, 2048*1152; up to 6 outputs per call |
spicy-image-1 | image | $0.06 per image | sizes 10241024, 8321216, 1216832, 12801280, 14401440, 10241536, 15361024, 10801920, 19201080, 11522048, 2048*1152; up to 6 outputs per call |
spicy-image-action-1 | image | $0.15 per image | up to 4 outputs per call |
spicy-image-edit-1 | image-edit | $0.15 per image | sizes 10241024, 8321216, 1216832, 12801280, 14401440, 10241536, 15361024, 10801920, 19201080, 11522048, 2048*1152; up to 6 outputs per call; needs image_url (one of your generated images) |
spicy-motion-3 | video | $0.1 per second at 480P, $0.2 per second at 720P, $0.4 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) |
spicy-motion-3-fast | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.56 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) |
spicy-cinema-1-image | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) |
spicy-motion-2 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) |
spicy-character-video-1 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; image_url optional (text-to-video without it) |
spicy-cinema-1-character | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters |
spicy-cinema-1 | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters |
spicy-video-1 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters |
spicy-motion-draft-1 | video | $0.05 per second at 720P, $0.075 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 1,500 characters; always silent; needs image_url (one of your generated images) |
spicy-motion-1 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 10 seconds; prompt up to 1,500 characters; needs image_url (one of your generated images) |
spicy-video-edit-1 | video-edit | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 2 to 10 second clips (output up to 10 seconds) with up to 4 reference images, aspect_ratio, keep_audio |
spicy-cinema-1-edit | video-edit | $0.28 per second at 720P, $0.48 per second at 1080P | resolutions 720P, 1080P; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 3 to 30 second clips (output up to 15 seconds) with up to 5 reference images, keep_audio |
spicy-animate-1 | video-edit | $0.24 per second at standard, $0.36 per second at pro | resolutions standard, pro; motion transfer: image_url plus a motion id (GET /v1/videos/motions) or your own 2 to 30 second clip, quality standard or pro, billed per output second; needs image_url (one of your generated images) |
spicy-companion-1 | chat | $1 per 1M prompt tokens and $2.8 per 1M completion tokens (minimum $0.001 per request) | |
spicy-companion-1-flash | chat | $0.1 per 1M prompt tokens and $0.8 per 1M completion tokens (minimum $0.0005 per request) | |
spicy-chat-1 | chat | $0.8 per 1M prompt tokens and $2.4 per 1M completion tokens (minimum $0.001 per request) | |
spicy-voice-2 | speech | $0.40 per 10,000 characters of input text | input up to 5,000 characters; 2 preset voices; instructions steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside input; WAV or MP3 URL usable as audio_url on video, or stream: true for raw audio as it is synthesized |
spicy-voice-2-flash | speech | $0.30 per 10,000 characters of input text | input up to 5,000 characters; 9 preset voices; instructions steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside input; WAV or MP3 URL usable as audio_url on video, or stream: true for raw audio as it is synthesized |
spicy-voice-1-expressive | speech | $0.23 per 10,000 characters of input text | input up to 600 characters; 21 preset voices; instructions steer emotion, pace and delivery; WAV output, a URL usable as audio_url on video |
spicy-voice-1-custom | speech | $0.23 per 10,000 characters of input text; creating a voice costs $0.40 (designed) or $0.02 (cloned), once | input up to 600 characters; speaks your designed and cloned voices (vc_...); WAV output, a URL usable as audio_url on video |
spicy-voice-1 | speech | $0.20 per 10,000 characters of input text | input up to 600 characters; 45 preset voices; WAV output, a URL usable as audio_url on video |
spicy-live-1 | realtime | per 1M tokens: text in $0.46, audio in $1.86, text out $1.4, audio out $3.74, billed per turn (about half a cent per minute of conversation; opening a session needs a $0.05 balance) | 28 voices, up to 13 minutes per session, POST /v1/realtime/sessions then a WebSocket |
spicy-transcribe-1 | transcription | $0.0042 per minute of audio, billed by the second (minimum $0.0001 per request) | wav, mp3, m4a, ogg, flac, webm and more, up to 10 MB and 5 minutes |
spicy-embed-1 | embedding | $0.14 per 1M tokens | |
spicy-embed-vision-1 | embedding | $0.18 per 1M text tokens and $0.06 per 1M image tokens |
Endpoints
Core
GET /v1/models: List every model available to your account. Reference: https://www.spicyapi.com/docs/api/list-models.mdGET /v1/account: Read your balance and account status. Reference: https://www.spicyapi.com/docs/api/get-account.md
Images
POST /v1/images/generations: Generate images from a text prompt. Reference: https://www.spicyapi.com/docs/api/create-image.mdPOST /v1/images/edits: Edit an image you generated earlier, guided by a prompt. One to six outputs per call. Reference: https://www.spicyapi.com/docs/api/edit-image.mdGET /v1/images: List the images your account has generated. Reference: https://www.spicyapi.com/docs/api/list-images.mdGET /v1/styles: The art styles forstyleon image and video generation. No key needed. Reference: https://www.spicyapi.com/docs/api/list-styles.mdGET /v1/images/actions: The ready-made explicit stills foractiononspicy-image-action-1. No key needed. Reference: https://www.spicyapi.com/docs/api/list-image-actions.md
Moderation
POST /v1/moderations: Run a prompt or conversation through the same screen the generation endpoints use, without generating. Reference: https://www.spicyapi.com/docs/api/moderations.md
Characters
POST /v1/characters: Save a reusable identity from your own generated images, to get the same person again across images, edits and video. Reference: https://www.spicyapi.com/docs/api/create-character.mdGET /v1/characters: Your saved characters, newest first. Reference: https://www.spicyapi.com/docs/api/list-characters.mdDELETE /v1/characters/{id}: Remove a saved character. Its reference sheet stays in your image library. Reference: https://www.spicyapi.com/docs/api/delete-character.md
Audio
POST /v1/audio/speech: Turn text into speech. Returns a durable audio URL you can pass straight to video generation asaudio_url, or streams the audio as it is synthesized. Reference: https://www.spicyapi.com/docs/api/create-speech.mdPOST /v1/audio/transcriptions: Turn speech into text: voice notes, clips and call recordings, with the spoken language detected. Reference: https://www.spicyapi.com/docs/api/create-transcription.mdPOST /v1/voices: Design a voice from a description, or clone one from one of your clips with the speaker's consent. Custom voices are spoken byspicy-voice-1-custom. Reference: https://www.spicyapi.com/docs/api/create-voice.mdGET /v1/voices: Every preset voice with the models that speak it, then your designed and cloned voices, newest first. Reference: https://www.spicyapi.com/docs/api/list-voices.mdDELETE /v1/voices/{id}: Delete one of your designed or cloned voices. Reference: https://www.spicyapi.com/docs/api/delete-voice.mdPOST /v1/audio: Upload a short audio clip for lip-sync and audio-guided video. The clip is transcribed and screened before it can be used. Reference: https://www.spicyapi.com/docs/api/upload-audio.mdGET /v1/audio: Your audio clips, uploaded and generated, newest first. Reference: https://www.spicyapi.com/docs/api/list-audio.mdDELETE /v1/audio/{id}: Forget a clip so it can no longer be used as a video input. Reference: https://www.spicyapi.com/docs/api/delete-audio.md
Video
POST /v1/videos/generations: Start a video generation. Returns a task to poll. Reference: https://www.spicyapi.com/docs/api/create-video.mdGET /v1/videos/tasks/{id}: Check a video task and collect the finished clip. Reference: https://www.spicyapi.com/docs/api/get-video-task.mdPOST /v1/videos/edits: Edit a clip you generated from a prompt (outfit, style, lighting, props), or transfer a clip's motion onto the person in one of your images. Returns a task to poll. Reference: https://www.spicyapi.com/docs/api/create-video-edit.mdGET /v1/videos/motions: The curated motion clips for motion transfer onspicy-animate-1. No key needed. Reference: https://www.spicyapi.com/docs/api/list-motions.mdGET /v1/videos/actions: The ready-made sex acts foractiononspicy-motion-3andspicy-motion-3-fast. No key needed. Reference: https://www.spicyapi.com/docs/api/list-actions.md
Chat
POST /v1/chat/completions: OpenAI-compatible uncensored chat and role-play, with streaming. Reference: https://www.spicyapi.com/docs/api/chat-completions.md
Embeddings
POST /v1/embeddings: Turn text, or images from your own library, into vectors for memory, search, recommendations and near-duplicate detection. OpenAI-compatible. Reference: https://www.spicyapi.com/docs/api/create-embeddings.md
Live calls
POST /v1/realtime/sessions: Start a live voice call with a companion: returns a one-time WebSocket URL your client (a browser too) connects to, without ever seeing your API key. Reference: https://www.spicyapi.com/docs/api/create-realtime-session.mdGET /v1/realtime: The call itself: a WebSocket (wss://) opened with the session ticket. Send microphone audio, receive the persona's voice and transcripts as JSON events. Reference: https://www.spicyapi.com/docs/api/realtime-websocket.md
Input images: no uploads, by design
Image editing and image-to-video accept only URLs returned by /v1/images/generations or /v1/images/edits on the same account (GET /v1/images lists them). Uploads and third-party URLs are rejected before any charge. This keeps real-person and minor imagery out of the pipeline and is a legal-compliance requirement, not a limitation we plan to lift.
Video is asynchronous
POST /v1/videos/generations returns 202 with a task id and the price already debited. Poll GET /v1/videos/tasks/{id} every 10 to 15 seconds until status is succeeded (then output.video_url) or failed (refunded). Or set a webhook URL in the dashboard: we POST {"event": "task.succeeded" | "task.failed", "data": {...}} with an X-SpicyAPI-Signature header (hex HMAC-SHA256 of the raw body with your signing secret).
Video inputs beyond the first frame
last_frame_url: a closing frame (one of your images) the clip travels to; needsimage_url.spicy-motion-2,spicy-motion-3,spicy-motion-3-fast.reference_image_urls: up to 10 of your images guiding identity without fixing the first shot; exclusive withimage_url.spicy-character-video-1,spicy-motion-3,spicy-motion-3-fast.video_url: theoutput.video_urlof one of your own finished tasks. Onspicy-motion-2it continues the clip in place ofimage_url(2 to 10 s in;durationis the whole output, billed in full). Onspicy-motion-3andspicy-motion-3-fastit is a reference clip to extend or edit (15 s max; clip plusdurationwithin 30 s).duration: -1onspicy-motion-3andspicy-motion-3-fast: the model picks the length. Debited at the maximum, refunded to the rendered length; the task'scost_usdsettles andusage.output_secondsreports it.aspect_ratio(text-to-video only): 16:9, 9:16, 1:1, 4:3, 3:4 onspicy-video-1(default 16:9); plus 21:9 onspicy-motion-3, where omitting it lets the model choose. Prompt caps per model inlimits.maxPromptChars;negative_promptis ignored onspicy-motion-3.fps: 60(default 30), any video model: the finished clip is frame-interpolated to 60 fps for smoother motion, audio kept, for 20% of the video's price on top (input and action clip seconds included). About a minute longer. Not available at 1080P withaspect_ratio21:9 or 9:21 (a 400; use 720P). If the 60 fps step fails, the 30 fps clip is delivered and the fee refunded; the task'susage.fpssays 60 or 30.- Which inputs may travel together is per model:
GET /v1/modelsreturnsinputson every video model (first_frame,last_frame,reference_images,reference_video,continuation_clip,audio.mode,aspect_ratio,smart_durationandcombinations, the valid combinations in plain sentences). An invalid combination is a 400 that quotes them.
Audio inputs
POST /v1/audio is the one place external media comes in: a WAV or MP3 of 2 to 30 seconds (15 MB), as multipart file or JSON {"url"}. It is transcribed and the transcript screened like a prompt ($0.01 flat; 422 with the usual block fields if it fails, 503 and refunded if screening is unavailable). Pass the returned url as audio_url on video. GET /v1/audio lists clips, DELETE /v1/audio/{id} forgets one.
Driving audio versus reference audio (inputs.audio.mode): on spicy-motion-2 and spicy-video-1 the clip is driving audio, the soundtrack of the output, and the mouth follows it; it works alongside a first frame where the model takes one, one clip per request. On spicy-motion-3 and spicy-motion-3-fast it is reference audio: the model generates its own sound and uses the clip for voice, tone and beat while the words come from the prompt; up to 5 clips and 15 s total in audio_urls, only with a text prompt or references, never with a first frame (there a first frame pairs only with last_frame_url). spicy-motion-1 and spicy-character-video-1 take no audio. Send audio_mode ("driving" or "reference") to assert which one you expect; a mismatch is a 400 naming a model that offers it.
Speech and voices
POST /v1/audio/speech with {"model", "input", "voice"} returns {"id": "au_...", "url", "duration_s", "voice", "cost_usd"}: a WAV (24 kHz, 16-bit mono) on the CDN, registered as one of your audio clips (GET /v1/audio lists it with source: "speech"). input is up to 600 characters; language is Auto (default), English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French or Russian. Billed per 10,000 characters of input (CJK ideographs count 2); the text is screened like a prompt first (422, never charged).
spicy-voice-1: preset voices by name (GET /v1/voiceslists them with the models that speak them).spicy-voice-1-expressive: fewer preset voices plusinstructions(up to 1,600 characters) describing emotion, pace, pitch and delivery. The debit includes the instructions and settles down to the billed count.spicy-voice-1-custom: your own voices,vc_...ids fromPOST /v1/voices. Design one with{"name", "description"}(vocal qualities only: a description that names or imitates a real person is refused withcode: "real_person"), or clone one with{"name", "audio_url", "consent": true}, whereaudio_urlis one of your clips fromPOST /v1/audio(10 MB max) andconsent: trueattests the voice is yours or the speaker gave written consent; the attestation is stored with the voice. Creation is a one-off fee;DELETE /v1/voices/{id}removes a voice.
Make a character speak: generate the line with POST /v1/audio/speech, then pass its url as audio_url on POST /v1/videos/generations. On spicy-motion-2 it is driving audio with a first frame (the mouth follows the words); on spicy-motion-3 it is reference audio for voice, tone and beat, with a text prompt or references and no first frame. Video audio must be 2 to 30 seconds long.
Expressive speech and streaming
spicy-voice-2 and spicy-voice-2-flash take up to 5,000 characters with inline tags: control tags ([whispers], [excited], [crying], [asmr] and more) change the delivery until the next tag, sound tags ([giggles], [gasp], [sighing] and more) insert a sound; tags are billed as characters. response_format wav or mp3. stream: true returns the audio bytes as they are synthesized (pcm 24 kHz 16-bit mono by default, or wav, mp3), with the clip id in x-spicyapi-audio-id; the clip is stored afterwards like any other.
Live voice calls
- Server side:
POST /v1/realtime/sessionswith{"model": "spicy-live-1", "voice", "instructions", "turn_detection"}. The persona is screened; you need a small balance to open a call. The response hasurl, a one-timewss://URL valid for 60 seconds. - Client side (a browser is fine, it never sees the key): open the WebSocket, send
input_audio_buffer.appendwith base64 PCM16 mono 16 kHz, playresponse.audio.delta(PCM16 mono 24 kHz). Transcripts arrive for both sides. Billed per turn;spicy.usagereports the running cost; calls end at 13 minutes (spicy.session_expired). Onlysession.turn_detectioncan change mid-call.
Transcription and embeddings
POST /v1/audio/transcriptions (multipart file or JSON url, optional language) returns {"text", "language", "duration_s", "cost_usd"}, billed by the second. POST /v1/embeddings is OpenAI-compatible: spicy-embed-1 for text, spicy-embed-vision-1 for your own images and text in one space.
Video edits and motion transfer
POST /v1/videos/editswithspicy-video-edit-1orspicy-cinema-1-edit:video_url(your own clip),prompt, optionalreference_image_urls; billed input plus output seconds, settled when the task finishes.POST /v1/videos/editswithspicy-animate-1:image_urlplusmotion(fromGET /v1/videos/motions, no key needed) or your ownvideo_url;qualitystandard or pro. Explicit images and clips are refused on this model; use a dressed performer.
Draft tier and the Cinema engine
spicy-motion-draft-1 is silent image-to-video at a quarter of the price: preview motion, then render the keeper on spicy-motion-2 or spicy-motion-3. spicy-cinema-1, spicy-cinema-1-image and spicy-cinema-1-character are a second video engine: 3 to 15 seconds, 480P to 1080P, native audio always on, strong multi-reference consistency.
Actions: custom sex actions
action on spicy-motion-3 and spicy-motion-3-fast takes one of 39 ready-made sex acts from GET /v1/videos/actions (no key needed): a short reference clip (the first 4 seconds of the act, 5 on two actions) that supplies camera angle, framing, pose and motion; past its end the model carries the same act on as one continuous take. image_url (or reference_image_urls, or a character; one is required) is the woman who performs it, not a first frame. prompt is optional and sets the scene (setting, outfit, lighting); without it she is in a plain light bedroom. duration defaults to 8 (action plus duration within 30 s), aspect_ratio to the action's own shape; video_url and last_frame_url are rejected. style (from GET /v1/styles) keeps a stylised woman in her style (anime, 3d, cartoon); without it the photoreal reference clip pulls her toward realistic. The action's reference clip (4 seconds on most, duration_s in the list) is billed as input seconds on top of the output seconds. Guide with examples and prices: https://www.spicyapi.com/docs/pov-video.
Errors and limits
401 bad key, 402 insufficient balance (top up at https://www.spicyapi.com/dashboard/account), 400 invalid parameters, 422 prompt blocked by moderation (never charged), 429 rate limited or model at capacity (retry after the Retry-After header; limits: 600 requests/min per account, 120 screened requests/min, 20 video tasks in progress), 502 upstream unavailable (refunded). Shape: {"error": {"message": "...", "type": "insufficient_balance"}}. Acceptable use: https://www.spicyapi.com/acceptable-use.
Sandbox
Create a key with Sandbox ticked: every endpoint returns fixture output flagged sandbox: true after the same validation and keyword screening, nothing is billed, video tasks report processing for a few seconds then succeeded with an example clip.
For AI agents
Machine-readable overview: https://www.spicyapi.com/llms.txt. OpenAPI: https://www.spicyapi.com/openapi.json. MCP endpoint: https://www.spicyapi.com/api/mcp (Claude Code: claude mcp add --transport http spicyapi https://www.spicyapi.com/api/mcp --header "Authorization: Bearer sk-spicy-..."; OAuth clients need only the URL; guide: https://www.spicyapi.com/docs/mcp). Skills: https://www.spicyapi.com/skills/index.json. A sandbox key answers everything from fixtures with nothing billed.
Full reference as one file: https://www.spicyapi.com/llms-full.txt.