> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nugen.in/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI Responses API Compatible Endpoint

> OpenAI `responses.create`-compatible endpoint.

Accepts the same shape as `openai.responses.create` and returns a `Response` object.
Uses the LiteLLM Python SDK for provider-agnostic routing.

**Request Body:**

- `model`: Model ID (required)
- `input`: User input — a plain string or array of `{"role": ..., "content": ...}` dicts
- `instructions` (optional): System-level instructions
- `max_output_tokens` (optional): Maximum tokens to generate
- `temperature` (optional): Sampling temperature 0–2 (default: 1.0)
- `top_p` (optional): Nucleus sampling parameter
- `stream` (optional): Stream partial output as SSE (default: false)

**Example Request:**

```json
POST /api/v3/inference/responses
Headers: {"Authorization": "Bearer <api_key>"}

{
  "model": "Qwen/Qwen2.5-0.5B-Instruct",
  "input": "What is the capital of France?",
  "instructions": "You are a helpful assistant.",
  "max_output_tokens": 200
}
```

**Example Response:**

```json
{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1704123600,
  "model": "Qwen/Qwen2.5-0.5B-Instruct",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{"type": "output_text", "text": "The capital of France is Paris."}]
    }
  ],
  "usage": {"input_tokens": 18, "output_tokens": 9, "total_tokens": 27}
}
```



## OpenAPI

````yaml https://api.nugen.in/openapi-public.json post /api/v3/inference/responses
openapi: 3.1.0
info:
  title: Nugen Intelligence API
  description: >+
    Nugen Intelligence: Powering Specialized Intelligence At Scale.


    **Need an API key?** Sign up, log in to the platform and generate an API
    key.


    **Need help?** Sign up, log in to the platform and raise a support ticket.


    **What this API does.** Align models to your domain corpus and run them with

    confidence. Train-time alignment builds domain-aligned adaptors through

    proprietary methods spanning pre-training, supervised and reinforcement

    stages. Inference-time alignment quantifies uncertainty across the full

    generation trajectory, producing confidence scores and keeping every

    response on domain.


    **Workflow.** Upload documents, create or generate a benchmark, create an

    alignment project, receive an aligned model id. That one id serves chat

    completions, completions, embeddings and reranking, and its embeddings carry

    your domain's representation space, making retrieval domain aware. You
    define the success

    criteria for your domain; the platform handles the deep learning, the GPUs

    and the data preparation.


    **Authentication.** Every endpoint requires an API key sent as a Bearer
    token:

    `Authorization: Bearer <api_key>`.

  contact:
    name: Nugen Intelligence - Customer Support
    url: https://nugen.in/signup
    email: support@nugen.in
  version: 25.4.20
servers:
  - url: https://api.nugen.in
    description: Production
security: []
paths:
  /api/v3/inference/responses:
    post:
      tags:
        - Inference
      summary: OpenAI Responses API Compatible Endpoint
      description: >-
        OpenAI `responses.create`-compatible endpoint.


        Accepts the same shape as `openai.responses.create` and returns a
        `Response` object.

        Uses the LiteLLM Python SDK for provider-agnostic routing.


        **Request Body:**


        - `model`: Model ID (required)

        - `input`: User input — a plain string or array of `{"role": ...,
        "content": ...}` dicts

        - `instructions` (optional): System-level instructions

        - `max_output_tokens` (optional): Maximum tokens to generate

        - `temperature` (optional): Sampling temperature 0–2 (default: 1.0)

        - `top_p` (optional): Nucleus sampling parameter

        - `stream` (optional): Stream partial output as SSE (default: false)


        **Example Request:**


        ```json

        POST /api/v3/inference/responses

        Headers: {"Authorization": "Bearer <api_key>"}


        {
          "model": "Qwen/Qwen2.5-0.5B-Instruct",
          "input": "What is the capital of France?",
          "instructions": "You are a helpful assistant.",
          "max_output_tokens": 200
        }

        ```


        **Example Response:**


        ```json

        {
          "id": "resp_abc123",
          "object": "response",
          "created_at": 1704123600,
          "model": "Qwen/Qwen2.5-0.5B-Instruct",
          "output": [
            {
              "type": "message",
              "role": "assistant",
              "content": [{"type": "output_text", "text": "The capital of France is Paris."}]
            }
          ],
          "usage": {"input_tokens": 18, "output_tokens": 9, "total_tokens": 27}
        }

        ```
      operationId: responses_create
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ResponsesCreateRequest'
        required: true
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema: {}
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    ResponsesCreateRequest:
      properties:
        model:
          type: string
          title: Model
          description: Model ID to use for the response
          examples:
            - model_01kmqm4nrn9fw6r
        input:
          anyOf:
            - type: string
            - items:
                additionalProperties: true
                type: object
              type: array
          title: Input
          description: Input text or array of message dicts
        instructions:
          anyOf:
            - type: string
            - type: 'null'
          title: Instructions
          description: System-level instructions (maps to system message)
        max_output_tokens:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Output Tokens
          description: Maximum tokens to generate
        temperature:
          anyOf:
            - type: number
              maximum: 2
              minimum: 0
            - type: 'null'
          title: Temperature
          description: Sampling temperature
          default: 1
        top_p:
          anyOf:
            - type: number
            - type: 'null'
          title: Top P
          description: Nucleus sampling parameter
        stream:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Stream
          description: Whether to stream back partial progress
          default: false
        metadata:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Metadata
          description: Optional metadata, ignored by the model
        updated_at:
          anyOf:
            - type: string
              format: date-time
            - type: 'null'
          title: Updated At
      type: object
      required:
        - model
        - input
      title: ResponsesCreateRequest
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.