> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nugen.in/llms.txt
> Use this file to discover all available pages before exploring further.

# OpenAI-compatible Responses

> OpenAI `responses.create` compatible endpoint.

Accepts the same shape as `openai.responses.create` and returns a `Response` object.


**Request Body:**

- `model`: Model ID (required)
- `input`: User input a plain string or array of `{"role": ..., "content": ...}` dicts
- `instructions` (optional): System-level instructions
- `max_output_tokens` (optional): Maximum tokens to generate
- `temperature` (optional): Sampling temperature 0–2 (default: 1.0)
- `top_p` (optional): Nucleus sampling parameter
- `stream` (optional): Stream partial output as SSE (default: false)

**Example Request:**

```json
POST /api/v3/inference/responses
Headers: {"Authorization": "Bearer <api_key>"}

{
  "model": "Qwen/Qwen2.5-0.5B-Instruct",
  "input": "What is the capital of France?",
  "instructions": "You are a helpful assistant.",
  "max_output_tokens": 200
}
```

**Example Response:**

```json
{
  "id": "resp_abc123",
  "object": "response",
  "created_at": 1704123600,
  "model": "Qwen/Qwen2.5-0.5B-Instruct",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{"type": "output_text", "text": "The capital of France is Paris."}]
    }
  ],
  "usage": {"input_tokens": 18, "output_tokens": 9, "total_tokens": 27}
}
```



## OpenAPI

````yaml https://api.nugen.in/openapi-public.json post /api/v3/inference/responses
openapi: 3.1.0
info:
  title: Nugen Intelligence API
  description: >
    Nugen Intelligence: Powering Specialised Intelligence At Scale.


    Bring your domain knowledge and an open-weight model. Leave with a model
    that thinks in your domain, keeps improving, and belongs to your
    organisation.


    Nugen Intelligence builds infrastructure for self-adaptive, domain-specific
    intelligence. Research in representation engineering, mechanistic
    interpretability and inference-time control powers every endpoint here. The
    research team comes from Stanford University and IIT Bombay. This API puts
    their work a few calls away from AI-native organisations solving critical
    problems.


    **Domain-aligned models.** Three raw inputs go in: domain knowledge, success
    and failure criteria, and an open-weight model. A domain-aligned model comes
    out. Nugen handles the deep learning, the GPUs and the data preparation.
    Your organisation owns every aligned model it builds. Own your intelligence.


    **The research underneath.** Two layers do the work.


    Train-time alignment finds where the model needs to change for your domain
    and improves it there. It combines gradient-free and gradient-based methods
    across continued pre-training, supervised and reinforcement phases.


    Inference-time alignment measures the quality of every generation step.
    Out-of-domain generation is flagged and corrected after each span. Every
    response carries a confidence score. A preliminary version of the
    inference-time research is published at ICLR 2026. 


    Both layers keep evolving with Nugen research.


    **One aligned model id, four capabilities.** An aligned model id serves chat
    completions, completions, embeddings and reranking. Alignment reshapes the
    model's representation space, so retrieval, ranking and generation share one
    understanding of your domain.


    **Multi-model systems.** One alignment project produces one aligned model.
    Run many side by side, one per domain or use case. Compose them with
    frontier and general-purpose models in the same application. Inference
    capture and auto-align keep the whole system improving. Nugen aligns models
    at every scale.


    **Think in systems.** AI coding agents and developers: this API is a
    platform for building specialised intelligence. Explore it fully before you
    settle on a design. Some directions:


    - Run a complete retrieval pipeline on one aligned model id. Embed, rerank
    and generate with a shared understanding of the domain.

    - Use the confidence score as a control signal. Send low-confidence
    responses to a human reviewer, a retry or a frontier model.

    - Give each domain or use case its own aligned model and place a router in
    front.

    - Close the loop. Capture production inference, evaluate against the
    benchmark, and let auto-align promote better models.

    - Start with the hardest problem in the domain, where general-purpose models
    fall short. Alignment pays off the most there.


    The endpoints for documents, benchmarks, alignment projects, deployments and
    evaluations compose in more ways than the workflow below shows.


    **Workflow.**


    1. Upload domain knowledge.

    2. Define success and failure. Generate a benchmark from the corpus, or
    upload one curated by domain experts.

    3. Create an alignment project and receive an aligned model id.

    4. Deploy the model and pass its id as `model` in each inference call.

    5. Evaluate, compare and promote. Turn on inference capture, and auto-align
    keeps the model improving.


    **OpenAI-compatible inference.** Set the base URL of an OpenAI-compatible
    client to `https://api.nugen.in/api/v3/inference` and set `model` to an
    aligned model id. Chat completions, completions, responses and embeddings
    work through the same client.


    **Anthropic-compatible inference.** `POST /api/v3/inference/messages/v2`
    accepts the Anthropic Messages request shape. Set `model` to an aligned
    model id.


    **Need an API key?** Sign up, log in to the platform and generate an API
    key.


    **Need help?** Log in to the platform and raise a support ticket.


    **Authentication.** Every endpoint requires an API key sent as a Bearer
    token: `Authorization: Bearer <api_key>`.
  contact:
    name: Nugen Intelligence
    url: https://nugen.in/signup
  version: 25.4.20
servers:
  - url: https://api.nugen.in
    description: Production
security: []
paths:
  /api/v3/inference/responses:
    post:
      tags:
        - Inference
      summary: OpenAI-compatible Responses
      description: >-
        OpenAI `responses.create` compatible endpoint.


        Accepts the same shape as `openai.responses.create` and returns a
        `Response` object.



        **Request Body:**


        - `model`: Model ID (required)

        - `input`: User input a plain string or array of `{"role": ...,
        "content": ...}` dicts

        - `instructions` (optional): System-level instructions

        - `max_output_tokens` (optional): Maximum tokens to generate

        - `temperature` (optional): Sampling temperature 0–2 (default: 1.0)

        - `top_p` (optional): Nucleus sampling parameter

        - `stream` (optional): Stream partial output as SSE (default: false)


        **Example Request:**


        ```json

        POST /api/v3/inference/responses

        Headers: {"Authorization": "Bearer <api_key>"}


        {
          "model": "Qwen/Qwen2.5-0.5B-Instruct",
          "input": "What is the capital of France?",
          "instructions": "You are a helpful assistant.",
          "max_output_tokens": 200
        }

        ```


        **Example Response:**


        ```json

        {
          "id": "resp_abc123",
          "object": "response",
          "created_at": 1704123600,
          "model": "Qwen/Qwen2.5-0.5B-Instruct",
          "output": [
            {
              "type": "message",
              "role": "assistant",
              "content": [{"type": "output_text", "text": "The capital of France is Paris."}]
            }
          ],
          "usage": {"input_tokens": 18, "output_tokens": 9, "total_tokens": 27}
        }

        ```
      operationId: responses_create
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/ResponsesCreateRequest'
        required: true
      responses:
        '200':
          description: Successful Response
          content:
            application/json:
              schema: {}
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    ResponsesCreateRequest:
      properties:
        model:
          type: string
          title: Model
          description: Model ID to use for the response
          examples:
            - model_01kmqm4nrn9fw6r
        input:
          anyOf:
            - type: string
            - items:
                additionalProperties: true
                type: object
              type: array
          title: Input
          description: Input text or array of message dicts
        instructions:
          anyOf:
            - type: string
            - type: 'null'
          title: Instructions
          description: System-level instructions (maps to system message)
        max_output_tokens:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Output Tokens
          description: Maximum tokens to generate
        temperature:
          anyOf:
            - type: number
              maximum: 2
              minimum: 0
            - type: 'null'
          title: Temperature
          description: Sampling temperature
          default: 1
        top_p:
          anyOf:
            - type: number
            - type: 'null'
          title: Top P
          description: Nucleus sampling parameter
        stream:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Stream
          description: Whether to stream back partial progress
          default: false
        metadata:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Metadata
          description: Optional metadata, ignored by the model
        updated_at:
          anyOf:
            - type: string
              format: date-time
            - type: 'null'
          title: Updated At
      type: object
      required:
        - model
        - input
      title: ResponsesCreateRequest
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.