Unlimited AI API: Quickstart Guide
Get your API key and start sending text requests in minutes. This guide covers the base URL, authentication, and standard OpenAI-compatible endpoints for our uncensored LLM.
Base URL and Authentication
Our API is fully compatible with the OpenAI SDK format. You must set the base URL to https://api.unlimitedaiapi.com/v1. This ensures all requests route to our dedicated uncensored model, not a proxy or aggregator. Authentication uses a Bearer token in the Authorization header. You can generate your key immediately after signing up with just an email and password. The key is shown once after signup, but you can regenerate it at any time, which invalidates the old one. We do not require a phone number or credit card for the trial tier.
First Request
Start by sending a simple chat completion request. This confirms your API key works and gives you an idea of the response speed. The model ID is always uncensored. You can send multiple messages in a single request to maintain context. The model accepts standard OpenAI message formats. There are no complex routing rules or vendor-specific quirks to manage.
curl https://api.unlimitedaiapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'Check the response for the choices array containing the generated text. If you get a 401 error, verify your API key. A 402 error means your prepaid credit is exhausted.
Python SDK Integration
Using the official openai Python library makes integration trivial. Set the base_url and api_key in your client initialization. This approach works for synchronous calls, which is the standard mode for most scripts. We support standard text input and output. We do not offer embeddings or image generation, so keep your payload focused on text. The SDK handles tokenization automatically, so you do not need to count tokens manually for basic usage.
from openai import OpenAI
client = OpenAI(base_url="https://api.unlimitedaiapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)Ensure you are using a recent version of the SDK to support the base_url parameter correctly. This setup connects directly to our infrastructure without third-party routing layers.
Node SDK Usage
For Node.js environments, the official openai package works with minimal configuration. Pass your API key and the correct base URL during initialization. This allows you to leverage existing Node.js infrastructure while maintaining full control over your API calls. We support standard request-response patterns. Streaming is handled separately via Server-Sent Events. The model is an open-weight model tuned for fewer refusals, making it suitable for creative or controversial topics within legal bounds.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.unlimitedaiapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);Remember that we charge per token, so monitor your usage via the response headers or your dashboard. There are no hidden fees or monthly subscriptions.
Streaming Responses
You can enable streaming by setting the stream parameter to true in your request. The API returns Server-Sent Events (SSE) instead of a single JSON object. This reduces perceived latency for long outputs. Each chunk contains a partial delta of the response. You must accumulate these deltas to reconstruct the full text. Pricing remains based on the total input and output tokens processed, regardless of streaming. This is ideal for UIs that want to display text as it is generated.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)Handle SSE parsing carefully to avoid memory leaks if a connection drops unexpectedly. The 100,000 token context window applies to the entire conversation history, not just the current turn.
Limits, Errors, and Context
Your API key is limited to 300 requests per minute. If you exceed this, you will receive a 429 rate limit error. The maximum request body size is 8 MB. The model has a 100,000 token context window, combining both prompt and completion tokens. This is sufficient for most long-document tasks. We do not offer fine-tuning or model routing. If you need embeddings, you must use a separate service. The hard content limit blocks sexual content involving minors, but otherwise allows lawful adult and controversial topics. Errors like 401 indicate an invalid key, while 402 indicates insufficient prepaid credit.
Technical reference
One table with every limit, feature and price that applies to your key.
| Feature | Support |
|---|---|
| API format | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Base URL | https://api.unlimitedaiapi.com/v1 |
| Authentication | Bearer token in the Authorization header |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Model | uncensored |
| Streaming | Supported (stream: true), usage included at the end |
| Other parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| JSON mode | JSON object mode via response_format json_object |
| Max context | 100,000 tokens, input and output combined |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Requests per minute | 300/min per key |
| Parallel requests | 8 requests at the same time per key |
| Request size | 8 MB request body |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Credit expiry | no monthly fee; paid credit does not expire |
| Trial credit | $0.50 for 7 days, no card |
| Price | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Bonus credit | +5% from $50, +10% from $100 |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Account | sign in with Google or with e-mail + password |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Key management | one active key per account; a new key replaces the old one |
Error codes
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | over 300/min or 8 parallel — back off and retry |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Does the uncensored model work with standard OpenAI tools?
Yes, our API is fully OpenAI-compatible. You can use the standard chat-completions endpoint with any OpenAI SDK or client that supports the <code>base_url</code> configuration. We support tool/function calling and streaming via SSE.
How does pricing work for streaming responses?
Pricing is based on the total number of input and output tokens processed, regardless of whether you use streaming or standard responses. Streaming only affects latency, not the cost per token. You pay $0.25 per 1M input tokens and $1.00 per 1M output tokens.
Can I use this API for image or audio generation?
No, we only offer text generation via the <code>chat-completions</code> endpoint. We do not support embeddings, image, audio, or video generation. If you need those capabilities, you will need to integrate a different service for those specific steps.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.