Spicy API
성인 앱(18세 이상)을 위한 무검열 이미지, 이미지-투-비디오 및 채팅 생성. OpenAI 호환 REST API와 15개의 MCP 도구 제공: 모델 가격, 비용 추정, 지출별 승인을 통한 생성, 작업, 사용량, 충전 링크 및 월별 지출 한도. 프롬프트는 청구 전에 검토되며 샌드박스 키는 청구되지 않습니다.
호스팅형 MCP 서버
npx add-mcp 'https://www.spicyapi.com/api/mcp'Claude Code, Codex, Cursor 등에 설치됩니다
문서
Spicy API documentation
Spicy API is one REST API for uncensored image, image-editing, video, video-editing, speech, live voice call, transcription, embedding and chat generation, sold per generation from a prepaid US-dollar balance: no subscription, no GPUs to run. It is OpenAI-compatible (point the OpenAI SDK at https://api.spicyapi.com/v1 with a Spicy API key) and includes ready-made sex actions for video and role-play companion models. Not to be confused with spicyapi.ai, an unrelated model aggregator.
Base URL https://api.spicyapi.com. Authenticate with Authorization: Bearer sk-spicy-... (keys from https://www.spicyapi.com/dashboard/api-keys, shown once, stored hashed). Check your balance with GET /v1/account.
Quickstart
from openai import OpenAI
client = OpenAI(api_key="sk-spicy-...", base_url="https://api.spicyapi.com/v1")
# Images: the OpenAI SDK's images.generate maps onto POST /v1/images/generations
img = client.images.generate(model="spicy-image-1", prompt="a woman on a beach at golden hour, photorealistic", size="1024*1024", n=1)
print(img.data[0].url) # durable CDN URL, reusable as image_url for edits and video
# Chat: fully compatible, streaming included
stream = client.chat.completions.create(model="spicy-chat-1", messages=[{"role": "user", "content": "hey"}], stream=True)
curl https://api.spicyapi.com/v1/videos/generations \
-H "Authorization: Bearer $SPICYAPI_KEY" -H "Content-Type: application/json" \
-d '{"model": "spicy-motion-2", "prompt": "slow build, steady rhythm", "image_url": "<a URL from /v1/images/generations>", "resolution": "720P", "duration": 5}'
# 202 {"id": "sj_...", "status": "queued", "cost_usd": 1.00} then poll every 10 to 15 seconds:
curl https://api.spicyapi.com/v1/videos/tasks/sj_... -H "Authorization: Bearer $SPICYAPI_KEY"
Models
| Model | Kind | Price | Limits |
|---|---|---|---|
spicy-image-1-pro | image | $0.09 per image | sizes 10241024, 8321216, 1216832, 12801280, 14401440, 10241536, 15361024, 10801920, 19201080, 11522048, 2048*1152; up to 6 outputs per call |
spicy-image-1 | image | $0.06 per image | sizes 10241024, 8321216, 1216832, 12801280, 14401440, 10241536, 15361024, 10801920, 19201080, 11522048, 2048*1152; up to 6 outputs per call |
spicy-image-action-1 | image | $0.15 per image | up to 4 outputs per call |
spicy-image-edit-1 | image-edit | $0.15 per image | sizes 10241024, 8321216, 1216832, 12801280, 14401440, 10241536, 15361024, 10801920, 19201080, 11522048, 2048*1152; up to 6 outputs per call; needs image_url (one of your generated images) |
spicy-motion-3 | video | $0.1 per second at 480P, $0.2 per second at 720P, $0.4 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) |
spicy-motion-3-fast | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.56 per second at 1080P | resolutions 480P, 720P, 1080P; 2 to 30 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 21:9; prompt up to 20,000 characters; reference clip (video_url) up to 15 seconds; the input clip's seconds are billed too, at the same rate; image_url optional (text-to-video without it) |
spicy-cinema-1-image | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) |
spicy-motion-2 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; needs image_url (one of your generated images) |
spicy-character-video-1 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 5,000 characters; image_url optional (text-to-video without it) |
spicy-cinema-1-character | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters |
spicy-cinema-1 | video | $0.14 per second at 480P, $0.28 per second at 720P, $0.36 per second at 1080P | resolutions 480P, 720P, 1080P; 3 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4, 4:5, 5:4, 9:21, 21:9; prompt up to 5,000 characters |
spicy-video-1 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters |
spicy-motion-draft-1 | video | $0.05 per second at 720P, $0.075 per second at 1080P | resolutions 720P, 1080P; 2 to 15 seconds; prompt up to 1,500 characters; always silent; needs image_url (one of your generated images) |
spicy-motion-1 | video | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; 2 to 10 seconds; prompt up to 1,500 characters; needs image_url (one of your generated images) |
spicy-video-edit-1 | video-edit | $0.2 per second at 720P, $0.3 per second at 1080P | resolutions 720P, 1080P; text-to-video aspect ratios 16:9, 9:16, 1:1, 4:3, 3:4; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 2 to 10 second clips (output up to 10 seconds) with up to 4 reference images, aspect_ratio, keep_audio |
spicy-cinema-1-edit | video-edit | $0.28 per second at 720P, $0.48 per second at 1080P | resolutions 720P, 1080P; prompt up to 5,000 characters; the input clip's seconds are billed too, at the same rate; edits your own 3 to 30 second clips (output up to 15 seconds) with up to 5 reference images, keep_audio |
spicy-animate-1 | video-edit | $0.24 per second at standard, $0.36 per second at pro | resolutions standard, pro; motion transfer: image_url plus a motion id (GET /v1/videos/motions) or your own 2 to 30 second clip, quality standard or pro, billed per output second; needs image_url (one of your generated images) |
spicy-companion-1 | chat | $1 per 1M prompt tokens and $2.8 per 1M completion tokens (minimum $0.001 per request) | |
spicy-companion-1-flash | chat | $0.1 per 1M prompt tokens and $0.8 per 1M completion tokens (minimum $0.0005 per request) | |
spicy-chat-1 | chat | $0.8 per 1M prompt tokens and $2.4 per 1M completion tokens (minimum $0.001 per request) | |
spicy-voice-2 | speech | $0.40 per 10,000 characters of input text | input up to 5,000 characters; 2 preset voices; instructions steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside input; WAV or MP3 URL usable as audio_url on video, or stream: true for raw audio as it is synthesized |
spicy-voice-2-flash | speech | $0.30 per 10,000 characters of input text | input up to 5,000 characters; 9 preset voices; instructions steer emotion, pace and delivery; inline tags like [whispers], [excited], [giggles] inside input; WAV or MP3 URL usable as audio_url on video, or stream: true for raw audio as it is synthesized |
spicy-voice-1-expressive | speech | $0.23 per 10,000 characters of input text | input up to 600 characters; 21 preset voices; instructions steer emotion, pace and delivery; WAV output, a URL usable as audio_url on video |
spicy-voice-1-custom | speech | $0.23 per 10,000 characters of input text; creating a voice costs $0.40 (designed) or $0.02 (cloned), once | input up to 600 characters; speaks your designed and cloned voices (vc_...); WAV output, a URL usable as audio_url on video |
spicy-voice-1 | speech | $0.20 per 10,000 characters of input text | input up to 600 characters; 45 preset voices; WAV output, a URL usable as audio_url on video |
spicy-live-1 | realtime | per 1M tokens: text in $0.46, audio in $1.86, text out $1.4, audio out $3.74, billed per turn (about half a cent per minute of conversation; opening a session needs a $0.05 balance) | 28 voices, up to 13 minutes per session, POST /v1/realtime/sessions then a WebSocket |
spicy-transcribe-1 | transcription | $0.0042 per minute of audio, billed by the second (minimum $0.0001 per request) | wav, mp3, m4a, ogg, flac, webm and more, up to 10 MB and 5 minutes |
spicy-embed-1 | embedding | $0.14 per 1M tokens | |
spicy-embed-vision-1 | embedding | $0.18 per 1M text tokens and $0.06 per 1M image tokens |
Endpoints
Core
GET /v1/models: List every model available to your account. Reference: https://www.spicyapi.com/docs/api/list-models.mdGET /v1/account: Read your balance and account status. Reference: https://www.spicyapi.com/docs/api/get-account.md
Images
POST /v1/images/generations: Generate images from a text prompt. Reference: https://www.spicyapi.com/docs/api/create-image.mdPOST /v1/images/edits: Edit an image you generated earlier, guided by a prompt. One to six outputs per call. Reference: https://www.spicyapi.com/docs/api/edit-image.mdGET /v1/images: List the images your account has generated. Reference: https://www.spicyapi.com/docs/api/list-images.mdGET /v1/styles: The art styles forstyleon image and video generation. No key needed. Reference: https://www.spicyapi.com/docs/api/list-styles.mdGET /v1/images/actions: The ready-made explicit stills foractiononspicy-image-action-1. No key needed. Reference: https://www.spicyapi.com/docs/api/list-image-actions.md
Moderation
POST /v1/moderations: Run a prompt or conversation through the same screen the generation endpoints use, without generating. Reference: https://www.spicyapi.com/docs/api/moderations.md
Characters
POST /v1/characters: Save a reusable identity from your own generated images, to get the same person again across images, edits and video. Reference: https://www.spicyapi.com/docs/api/create-character.mdGET /v1/characters: Your saved characters, newest first. Reference: https://www.spicyapi.com/docs/api/list-characters.mdDELETE /v1/characters/{id}: Remove a saved character. Its reference sheet stays in your image library. Reference: https://www.spicyapi.com/docs/api/delete-character.md
Audio
POST /v1/audio/speech: Turn text into speech. Returns a durable audio URL you can pass straight to video generation asaudio_url, or streams the audio as it is synthesized. Reference: https://www.spicyapi.com/docs/api/create-speech.mdPOST /v1/audio/transcriptions: Turn speech into text: voice notes, clips and call recordings, with the spoken language detected. Reference: https://www.spicyapi.com/docs/api/create-transcription.mdPOST /v1/voices: Design a voice from a description, or clone one from one of your clips with the speaker's consent. Custom voices are spoken byspicy-voice-1-custom. Reference: https://www.spicyapi.com/docs/api/create-voice.mdGET /v1/voices: Every preset voice with the models that speak it, then your designed and cloned voices, newest first. Reference: https://www.spicyapi.com/docs/api/list-voices.mdDELETE /v1/voices/{id}: Delete one of your designed or cloned voices. Reference: https://www.spicyapi.com/docs/api/delete-voice.mdPOST /v1/audio: Upload a short audio clip for lip-sync and audio-guided video. The clip is transcribed and screened before it can be used. Reference: https://www.spicyapi.com/docs/api/upload-audio.mdGET /v1/audio: Your audio clips, uploaded and generated, newest first. Reference: https://www.spicyapi.com/docs/api/list-audio.mdDELETE /v1/audio/{id}: Forget a clip so it can no longer be used as a video input. Reference: https://www.spicyapi.com/docs/api/delete-audio.md
Video
POST /v1/videos/generations: Start a video generation. Returns a task to poll. Reference: https://www.spicyapi.com/docs/api/create-video.mdGET /v1/videos/tasks/{id}: Check a video task and collect the finished clip. Reference: https://www.spicyapi.com/docs/api/get-video-task.mdPOST /v1/videos/edits: Edit a clip you generated from a prompt (outfit, style, lighting, props), or transfer a clip's motion onto the person in one of your images. Returns a task to poll. Reference: https://www.spicyapi.com/docs/api/create-video-edit.mdGET /v1/videos/motions: The curated motion clips for motion transfer onspicy-animate-1. No key needed. Reference: https://www.spicyapi.com/docs/api/list-motions.mdGET /v1/videos/actions: The ready-made sex acts foractiononspicy-motion-3andspicy-motion-3-fast. No key needed. Reference: https://www.spicyapi.com/docs/api/list-actions.md
Chat
POST /v1/chat/completions: OpenAI-compatible uncensored chat and role-play, with streaming. Reference: https://www.spicyapi.com/docs/api/chat-completions.md
Embeddings
POST /v1/embeddings: Turn text, or images from your own library, into vectors for memory, search, recommendations and near-duplicate detection. OpenAI-compatible. Reference: https://www.spicyapi.com/docs/api/create-embeddings.md
Live calls
POST /v1/realtime/sessions: Start a live voice call with a companion: returns a one-time WebSocket URL your client (a browser too) connects to, without ever seeing your API key. Reference: https://www.spicyapi.com/docs/api/create-realtime-session.mdGET /v1/realtime: The call itself: a WebSocket (wss://) opened with the session ticket. Send microphone audio, receive the persona's voice and transcripts as JSON events. Reference: https://www.spicyapi.com/docs/api/realtime-websocket.md
Input images: no uploads, by design
Image editing and image-to-video accept only URLs returned by /v1/images/generations or /v1/images/edits on the same account (GET /v1/images lists them). Uploads and third-party URLs are rejected before any charge. This keeps real-person and minor imagery out of the pipeline and is a legal-compliance requirement, not a limitation we plan to lift.
Video is asynchronous
POST /v1/videos/generations returns 202 with a task id and the price already debited. Poll GET /v1/videos/tasks/{id} every 10 to 15 seconds until status is succeeded (then output.video_url) or failed (refunded). Or set a webhook URL in the dashboard: we POST {"event": "task.succeeded" | "task.failed", "data": {...}} with an X-SpicyAPI-Signature header (hex HMAC-SHA256 of the raw body with your signing secret).
Video inputs beyond the first frame
last_frame_url: a closing frame (one of your images) the clip travels to; needsimage_url.spicy-motion-2,spicy-motion-3,spicy-motion-3-fast.reference_image_urls: up to 10 of your images guiding identity without fixing the first shot; exclusive withimage_url.spicy-character-video-1,spicy-motion-3,spicy-motion-3-fast.video_url: theoutput.video_urlof one of your own finished tasks. Onspicy-motion-2it continues the clip in place ofimage_url(2 to 10 s in;durationis the whole output, billed in full). Onspicy-motion-3andspicy-motion-3-fastit is a reference clip to extend or edit (15 s max; clip plusdurationwithin 30 s).duration: -1onspicy-motion-3andspicy-motion-3-fast: the model picks the length. Debited at the maximum, refunded to the rendered length; the task'scost_usdsettles andusage.output_secondsreports it.aspect_ratio(text-to-video only): 16:9, 9:16, 1:1, 4:3, 3:4 onspicy-video-1(default 16:9); plus 21:9 onspicy-motion-3, where omitting it lets the model choose. Prompt caps per model inlimits.maxPromptChars;negative_promptis ignored onspicy-motion-3.fps: 60(default 30), any video model: the finished clip is frame-interpolated to 60 fps for smoother motion, audio kept, for 20% of the video's price on top (input and action clip seconds included). About a minute longer. Not available at 1080P withaspect_ratio21:9 or 9:21 (a 400; use 720P). If the 60 fps step fails, the 30 fps clip is delivered and the fee refunded; the task'susage.fpssays 60 or 30.- Which inputs may travel together is per model:
GET /v1/modelsreturnsinputson every video model (first_frame,last_frame,reference_images,reference_video,continuation_clip,audio.mode,aspect_ratio,smart_durationandcombinations, the valid combinations in plain sentences). An invalid combination is a 400 that quotes them.
Audio inputs
POST /v1/audio is the one place external media comes in: a WAV or MP3 of 2 to 30 seconds (15 MB), as multipart file or JSON {"url"}. It is transcribed and the transcript screened like a prompt ($0.01 flat; 422 with the usual block fields if it fails, 503 and refunded if screening is unavailable). Pass the returned url as audio_url on video. GET /v1/audio lists clips, DELETE /v1/audio/{id} forgets one.
Driving audio versus reference audio (inputs.audio.mode): on spicy-motion-2 and spicy-video-1 the clip is driving audio, the soundtrack of the output, and the mouth follows it; it works alongside a first frame where the model takes one, one clip per request. On spicy-motion-3 and spicy-motion-3-fast it is reference audio: the model generates its own sound and uses the clip for voice, tone and beat while the words come from the prompt; up to 5 clips and 15 s total in audio_urls, only with a text prompt or references, never with a first frame (there a first frame pairs only with last_frame_url). spicy-motion-1 and spicy-character-video-1 take no audio. Send audio_mode ("driving" or "reference") to assert which one you expect; a mismatch is a 400 naming a model that offers it.
Speech and voices
POST /v1/audio/speech with {"model", "input", "voice"} returns {"id": "au_...", "url", "duration_s", "voice", "cost_usd"}: a WAV (24 kHz, 16-bit mono) on the CDN, registered as one of your audio clips (GET /v1/audio lists it with source: "speech"). input is up to 600 characters; language is Auto (default), English, Chinese, German, Italian, Portuguese, Spanish, Japanese, Korean, French or Russian. Billed per 10,000 characters of input (CJK ideographs count 2); the text is screened like a prompt first (422, never charged).
spicy-voice-1: preset voices by name (GET /v1/voiceslists them with the models that speak them).spicy-voice-1-expressive: fewer preset voices plusinstructions(up to 1,600 characters) describing emotion, pace, pitch and delivery. The debit includes the instructions and settles down to the billed count.spicy-voice-1-custom: your own voices,vc_...ids fromPOST /v1/voices. Design one with{"name", "description"}(vocal qualities only: a description that names or imitates a real person is refused withcode: "real_person"), or clone one with{"name", "audio_url", "consent": true}, whereaudio_urlis one of your clips fromPOST /v1/audio(10 MB max) andconsent: trueattests the voice is yours or the speaker gave written consent; the attestation is stored with the voice. Creation is a one-off fee;DELETE /v1/voices/{id}removes a voice.
Make a character speak: generate the line with POST /v1/audio/speech, then pass its url as audio_url on POST /v1/videos/generations. On spicy-motion-2 it is driving audio with a first frame (the mouth follows the words); on spicy-motion-3 it is reference audio for voice, tone and beat, with a text prompt or references and no first frame. Video audio must be 2 to 30 seconds long.
Expressive speech and streaming
spicy-voice-2 and spicy-voice-2-flash take up to 5,000 characters with inline tags: control tags ([whispers], [excited], [crying], [asmr] and more) change the delivery until the next tag, sound tags ([giggles], [gasp], [sighing] and more) insert a sound; tags are billed as characters. response_format wav or mp3. stream: true returns the audio bytes as they are synthesized (pcm 24 kHz 16-bit mono by default, or wav, mp3), with the clip id in x-spicyapi-audio-id; the clip is stored afterwards like any other.
Live voice calls
- Server side:
POST /v1/realtime/sessionswith{"model": "spicy-live-1", "voice", "instructions", "turn_detection"}. The persona is screened; you need a small balance to open a call. The response hasurl, a one-timewss://URL valid for 60 seconds. - Client side (a browser is fine, it never sees the key): open the WebSocket, send
input_audio_buffer.appendwith base64 PCM16 mono 16 kHz, playresponse.audio.delta(PCM16 mono 24 kHz). Transcripts arrive for both sides. Billed per turn;spicy.usagereports the running cost; calls end at 13 minutes (spicy.session_expired). Onlysession.turn_detectioncan change mid-call.
Transcription and embeddings
POST /v1/audio/transcriptions (multipart file or JSON url, optional language) returns {"text", "language", "duration_s", "cost_usd"}, billed by the second. POST /v1/embeddings is OpenAI-compatible: spicy-embed-1 for text, spicy-embed-vision-1 for your own images and text in one space.
Video edits and motion transfer
POST /v1/videos/editswithspicy-video-edit-1orspicy-cinema-1-edit:video_url(your own clip),prompt, optionalreference_image_urls; billed input plus output seconds, settled when the task finishes.POST /v1/videos/editswithspicy-animate-1:image_urlplusmotion(fromGET /v1/videos/motions, no key needed) or your ownvideo_url;qualitystandard or pro. Explicit images and clips are refused on this model; use a dressed performer.
Draft tier and the Cinema engine
spicy-motion-draft-1 is silent image-to-video at a quarter of the price: preview motion, then render the keeper on spicy-motion-2 or spicy-motion-3. spicy-cinema-1, spicy-cinema-1-image and spicy-cinema-1-character are a second video engine: 3 to 15 seconds, 480P to 1080P, native audio always on, strong multi-reference consistency.
Actions: custom sex actions
action on spicy-motion-3 and spicy-motion-3-fast takes one of 39 ready-made sex acts from GET /v1/videos/actions (no key needed): a short reference clip (the first 4 seconds of the act, 5 on two actions) that supplies camera angle, framing, pose and motion; past its end the model carries the same act on as one continuous take. image_url (or reference_image_urls, or a character; one is required) is the woman who performs it, not a first frame. prompt is optional and sets the scene (setting, outfit, lighting); without it she is in a plain light bedroom. duration defaults to 8 (action plus duration within 30 s), aspect_ratio to the action's own shape; video_url and last_frame_url are rejected. style (from GET /v1/styles) keeps a stylised woman in her style (anime, 3d, cartoon); without it the photoreal reference clip pulls her toward realistic. The action's reference clip (4 seconds on most, duration_s in the list) is billed as input seconds on top of the output seconds. Guide with examples and prices: https://www.spicyapi.com/docs/pov-video.
Errors and limits
401 bad key, 402 insufficient balance (top up at https://www.spicyapi.com/dashboard/account), 400 invalid parameters, 422 prompt blocked by moderation (never charged), 429 rate limited or model at capacity (retry after the Retry-After header; limits: 600 requests/min per account, 120 screened requests/min, 20 video tasks in progress), 502 upstream unavailable (refunded). Shape: {"error": {"message": "...", "type": "insufficient_balance"}}. Acceptable use: https://www.spicyapi.com/acceptable-use.
Sandbox
Create a key with Sandbox ticked: every endpoint returns fixture output flagged sandbox: true after the same validation and keyword screening, nothing is billed, video tasks report processing for a few seconds then succeeded with an example clip.
For AI agents
Machine-readable overview: https://www.spicyapi.com/llms.txt. OpenAPI: https://www.spicyapi.com/openapi.json. MCP endpoint: https://www.spicyapi.com/api/mcp (Claude Code: claude mcp add --transport http spicyapi https://www.spicyapi.com/api/mcp --header "Authorization: Bearer sk-spicy-..."; OAuth clients need only the URL; guide: https://www.spicyapi.com/docs/mcp). Skills: https://www.spicyapi.com/skills/index.json. A sandbox key answers everything from fixtures with nothing billed.
Full reference as one file: https://www.spicyapi.com/llms-full.txt.