Uncensored Gemma: A Developer's Integration Guide
Updated
Uncensored Gemma refers to an open-weight large language model optimized for high-fidelity, low-refusal generation, accessible via a standard chat-completions interface. This guide explains how to integrate this text-only API into your workflow, treating it as a drop-in replacement for OpenAI-compatible clients while understanding the specific token limits and pricing structures involved.
Why Choose Uncensored Gemma?
When developers seek an uncensored gemma-style experience, they are often looking for a model that avoids the safety filters common in commercial offerings. This API provides a single, high-capacity model tuned to answer without content refusals for lawful adult use, security research, or controversial topics. Unlike platforms that aggregate multiple vendors, this service focuses on one specific open-weight model, ensuring consistent behavior and predictable latency.
The model does not refuse standard creative or analytical prompts, making it suitable for generating diverse content styles. However, it is important to note that it is an independent open-weight model, not Google's Gemma, nor is it related to xAI's Grok. It runs on dedicated servers and serves text in and text out, providing a clean, predictable environment for developers who need raw model output without the overhead of multi-model routing.
API Compatibility Overview
The API is fully OpenAI-compatible, meaning you can use the official OpenAI SDKs or any other client that speaks the OpenAI protocol. You only need to update the base_url and provide your API key. The endpoint structure mirrors the standard chat completions interface, making migration from other providers trivial.
- Endpoints: Only
POST /v1/chat/completionsandGET /v1/modelsare available. - No Extra Features: There are no embeddings, image generation, audio processing, or fine-tuning endpoints.
- Base URL:
https://api.grokuncensored.cc/v1 - Model ID: Use
uncensoredas the model identifier.
This simplicity reduces integration complexity. You do not need to handle different response schemas for different features. The response format is consistent, returning text content and token usage statistics.
Setting Up Your Environment
Integration begins with obtaining an API key. Sign up via the Get API key page using either Google authentication or an email and password. The key is displayed immediately, and no phone number is required. This process is designed for speed, allowing you to start testing within minutes.
Once you have your key, configure your environment variables. For example, in Python:
Ensure your base URL points to https://api.grokuncensored.cc/v1. You can also take advantage of the trial credit, which provides $0.50 valid for 7 days, requiring no credit card. This allows you to verify connectivity and response quality before committing to a crypto top-up.
Making Your First Request
With your environment configured, you can send a request. The API accepts standard message arrays. Here is a basic example using cURL:
curl https://api.grokuncensored.cc/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'The response will contain the generated text and token usage. If you set max_tokens to a value, the API will respect that limit. If you do not set max_tokens, the default maximum output is 2,048 tokens. For longer outputs, you can increase this up to 16,000 tokens, provided the total context (prompt + completion) stays within the 64,000 token window.
Handling Responses and Tokens
Responses are returned in standard JSON format. The token usage is included in the last chunk of a streaming response, allowing you to track costs in real-time. Pricing is transparent: $0.25 per 1M input tokens and $1.00 per 1M output tokens. Errors and refusals are free, so you can experiment without worrying about wasted credits.
When using streaming, ensure your client handles Server-Sent Events (SSE) correctly. The API returns chunks of text, and the final chunk contains the usage field. This allows for precise billing reconciliation. Note that credit is prepaid, and errors do not deduct from your balance. This model is not an embedding service, so you cannot use it for vector searches directly; you must handle embedding separately if needed.
Advanced Parameters
The API supports several parameters to control generation quality. You can adjust temperature for creativity, top_p for nucleus sampling, and seed for reproducibility. Other supported parameters include stop, presence_penalty, and frequency_penalty.
Function calling is supported via the tools and tool_choice fields. This allows the model to generate structured outputs that can be parsed by your application. Additionally, JSON mode is available via response_format: {"type": "json_object"}, ensuring the output is valid JSON. This is useful for building agents or structured data pipelines. The model is an open-weight model, tuned for high-quality text generation, making it suitable for complex instruction following.
Error Handling
Errors are returned in standard HTTP status codes. Common errors include rate limits (429) or invalid API keys (401). Rate limits are set to 300 requests per minute and 8 concurrent requests per key. If you hit these limits, implement exponential backoff in your client.
The request body is limited to 8 MB. If you exceed this, the request will be rejected. Ensure your prompts are within these bounds. Since the API is text-only, errors related to media processing will not occur. For billing errors, such as insufficient credit, the API will return a 402 status. Credit never expires, so you can top up at your convenience. Mistakes like double charges can be resolved via the Support page, as credit is not automatically refunded.
Comparison with Grok Uncensored
While often discussed alongside other uncensored models like grok uncensored, this API is an independent service. It does not serve Google's Gemma, nor is it the official Grok API from xAI. The term uncensored gemma in search results often refers to the open-weight model's behavior, which this API emulates closely. However, the model ID is simply uncensored, and it is not tied to Google's architecture.
Unlike the official Grok API, which may have different pricing and limits, this API offers a flat rate of $0.25/$1.00 per million tokens. It also supports crypto payments exclusively, with no credit card required for the trial. This makes it a flexible alternative for developers who prefer crypto or want to avoid subscription lock-ins. The model is not a reseller of Grok; it is a distinct open-weight model.
Conclusion
This API provides a robust, developer-first interface for an uncensored open-weight model. By sticking to the OpenAI-compatible protocol, it integrates seamlessly into existing workflows. The pricing is transparent, and the limits are generous for most use cases. Whether you are building a chatbot, an agent, or a content generation tool, this API offers a reliable, uncensored alternative. Remember that it is a text-only API, so plan your architecture accordingly. For more details on pricing and keys, visit the respective pages on the site.