NSFW LLM APINSFW LLM API Quickstart
NSFW LLM API Quickstart
Start building your uncensored chatbot with our single-model, OpenAI-compatible API. This quickstart covers authentication, basic requests, streaming, and tool calling using the 'uncensored' model.
Authentication and Base URL
Our API uses standard OpenAI authentication. Use the base URL https://api.nsfwllmapi.com/v1 and set your API key in the Authorization header. Sign up on the Get API key page to receive your key immediately. You can regenerate your key at any time; the old key is revoked immediately, though the rate limit counter continues from where it left off. Each account supports one key, and no card is needed for the trial.
Basic Chat Completion
Send a standard chat completion request to test the uncensored model. The model ID is always uncensored. This endpoint supports up to 64,000 tokens for the combined prompt and completion. Use this for simple text generation or roleplay setups where you need consistent content tolerance without routing complexity.
curl https://api.nsfwllmapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Ensure your system prompts allow adult themes if needed, as the model is tuned to answer without refusals for lawful adult content.
Python SDK Integration
Use the official OpenAI Python SDK to interact with the NSFW LLM API. Point the client to our base URL and pass your API key. This approach works well for backend services or scripts where you need structured responses. The SDK handles JSON serialization automatically, making it easy to integrate into existing Python pipelines.
from openai import OpenAI
client = OpenAI(base_url="https://api.nsfwllmapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Remember that the model does not support embeddings or image generation, so keep your payload focused on text completion tasks.
Node SDK Integration
For JavaScript or TypeScript projects, the OpenAI Node SDK works seamlessly with our API. Configure the base URL and API key in your client instance. This is ideal for web-based chatbots or serverless functions that need to generate uncensored text responses. The SDK supports both synchronous and asynchronous patterns.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.nsfwllmapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Handle errors gracefully, especially rate limits and credit exhaustion, to maintain a smooth user experience in your application.
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns Server-Sent Events (SSE) for real-time token delivery. This reduces perceived latency for chatbot users, providing immediate feedback as tokens are generated. Streaming is supported for all standard chat completions and works with the OpenAI SDKs.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Parse the SSE stream on the client side to display tokens as they arrive. This is critical for maintaining engagement in conversational interfaces.
Limits, Errors, and Context
The API enforces a limit of 300 requests per minute per key. If you exceed this, you receive a 429 error. Ensure your account has sufficient prepaid credit; a 402 error indicates no credit is available. An 401 error means the API key is invalid or revoked. The context window is 64,000 tokens. Requests exceeding 8 MB in body size will be rejected. Use these limits to design your retry logic and token management strategies effectively.
Under the hood: specs
Use this table to decide whether the API fits your project before you buy credit.
| Spec | Value |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Model | uncensored |
| Base URL | https://api.nsfwllmapi.com/v1 |
| API key | Bearer token in the Authorization header |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Other parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Completion length | 16,000 tokens max; 2,048 if max_tokens is not set |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Max context | 64,000 tokens, input and output combined |
| JSON mode | response_format: {"type": "json_object"} |
| Parallel requests | up to 8 in parallel per key |
| Requests per minute | 300 requests per minute per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Max body | up to 8 MB per request |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Credit expiry | paid credit never expires, no subscription |
| Volume bonus | +5% from $50, +10% from $100 |
| Key management | one active key per account; a new key replaces the old one |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Account | sign in with Google or with e-mail + password |
When a request fails
Every error is JSON with a type you can switch on. You are never charged for an error.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | body over 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Does regenerating my API key reset my rate limit counter?
No. When you regenerate your key, the old key is revoked immediately, but the rate limit counter continues from where it left off. The new key starts with the existing rate limit state, not a fresh counter.
What content is blocked?
A hard content limit always applies: sexual content involving minors is blocked. All other lawful adult, fictional, and controversial topics are tolerated by the uncensored model.
Can I use this API for embeddings or image generation?
No. This API only serves the 'uncensored' text model. Endpoints for embeddings, images, audio, video, and fine-tuning are not available. It is strictly a text-completion service.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.