> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nugen.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Generate Embeddings

> Generate dense vector embeddings for input text.


This endpoint creates numerical vector representations (embeddings) of text using specified embedding models. Embeddings are useful for semantic search, clustering, similarity comparisons, and RAG applications.


**Request Body:**

- `model`: Model ID (required). Use your aligned model id (e.g. `model_01kmqm4nrn9fw6r`)
- `input`: Text or list of texts to embed (required) - Single string or array of strings. Cannot be empty and must not exceed max input tokens for the model
- `dimensions` (optional): Number of dimensions for output embeddings (only applicable for Resizable Matryoshka Embedding models)


**Returns:**

- `id`: Unique identifier for the response
- `object`: Object type (always `embedding`)
- `data`: List of embedding objects, each containing:
  - `object`: Object type (always `embedding`)
  - `embedding`: Vector of floats representing the text (length depends on the model)
  - `index`: Position in the input array
- `model`: Model ID used for generation
- `created`: Unix timestamp (seconds) when response was generated
- `usage`: Token usage statistics containing:
  - `total_tokens`: Total number of tokens used by the request


**Example Request (Single Text):**

```json
POST /api/v3/inference/embeddings
Headers: {"Authorization": "Bearer <api_key>"}

{
  "model": "text-embedding-xyz",
  "input": "The quick brown fox jumps over the lazy dog"
}
```


**Example Request (Multiple Texts):**

```json
POST /api/v3/inference/embeddings
Headers: {"Authorization": "Bearer <api_key>"}

{
  "model": "model_01kmqm4nrn9fw6r",
  "input": [
    "First document for semantic search",
    "Second document for clustering",
    "Third document for similarity comparison"
  ]
}
```


**Example Request (With Dimensions):**

```json
POST /api/v3/inference/embeddings
Headers: {"Authorization": "Bearer <api_key>"}

{
  "model": "model_01kmqm4nrn9fw6r",
  "input": "Sample text to embed",
  "dimensions": 512
}
```


**Example Response:**

```json
{
  "id": "embed-abc123xyz",
  "object": "embedding",
  "data": [
    {
      "object": "embedding",
      "embedding": [0.00658765, -0.008665467, -0.007653789, 0.0023, -0.0091, 0.0125],
      "index": 0
    }
  ],
  "model": "text-embedding-ada-002",
  "created": 1705329600.0,
  "usage": {
    "total_tokens": 10
  }
}
```


**Notes:**

- Input must not exceed the maximum input tokens supported by the model
- Input cannot be an empty string
- For batch processing, pass multiple texts as an array to the `input` field
- The embedding vector length varies by model - check model documentation for specifics
- Use `dimensions` parameter only with Matryoshka embedding models that support resizing



## OpenAPI

````yaml https://api.nugen.in/openapi-public.json post /api/v3/inference/embeddings
openapi: 3.1.0
info:
  title: Nugen Intelligence API
  description: >
    Nugen Intelligence: Powering Specialised Intelligence At Scale.


    Bring your domain knowledge and an open-weight model. Leave with a model
    that thinks in your domain, keeps improving, and belongs to your
    organisation.


    Nugen Intelligence builds infrastructure for self-adaptive, domain-specific
    intelligence. Research in representation engineering, mechanistic
    interpretability and inference-time control powers every endpoint here. The
    research team comes from Stanford University and IIT Bombay. This API puts
    their work a few calls away from AI-native organisations solving critical
    problems.


    **Domain-aligned models.** Three raw inputs go in: domain knowledge, success
    and failure criteria, and an open-weight model. A domain-aligned model comes
    out. Nugen handles the deep learning, the GPUs and the data preparation.
    Your organisation owns every aligned model it builds. Own your intelligence.


    **The research underneath.** Two layers do the work.


    Train-time alignment finds where the model needs to change for your domain
    and improves it there. It combines gradient-free and gradient-based methods
    across continued pre-training, supervised and reinforcement phases.


    Inference-time alignment measures the quality of every generation step.
    Out-of-domain generation is flagged and corrected after each span. Every
    response carries a confidence score. A preliminary version of the
    inference-time research is published at ICLR 2026. 


    Both layers keep evolving with Nugen research.


    **One aligned model id, four capabilities.** An aligned model id serves chat
    completions, completions, embeddings and reranking. Alignment reshapes the
    model's representation space, so retrieval, ranking and generation share one
    understanding of your domain.


    **Multi-model systems.** One alignment project produces one aligned model.
    Run many side by side, one per domain or use case. Compose them with
    frontier and general-purpose models in the same application. Inference
    capture and auto-align keep the whole system improving. Nugen aligns models
    at every scale.


    **Think in systems.** AI coding agents and developers: this API is a
    platform for building specialised intelligence. Explore it fully before you
    settle on a design. Some directions:


    - Run a complete retrieval pipeline on one aligned model id. Embed, rerank
    and generate with a shared understanding of the domain.

    - Use the confidence score as a control signal. Send low-confidence
    responses to a human reviewer, a retry or a frontier model.

    - Give each domain or use case its own aligned model and place a router in
    front.

    - Close the loop. Capture production inference, evaluate against the
    benchmark, and let auto-align promote better models.

    - Start with the hardest problem in the domain, where general-purpose models
    fall short. Alignment pays off the most there.


    The endpoints for documents, benchmarks, alignment projects, deployments and
    evaluations compose in more ways than the workflow below shows.


    **Workflow.**


    1. Upload domain knowledge.

    2. Define success and failure. Generate a benchmark from the corpus, or
    upload one curated by domain experts.

    3. Create an alignment project and receive an aligned model id.

    4. Deploy the model and pass its id as `model` in each inference call.

    5. Evaluate, compare and promote. Turn on inference capture, and auto-align
    keeps the model improving.


    **OpenAI-compatible inference.** Set the base URL of an OpenAI-compatible
    client to `https://api.nugen.in/api/v3/inference` and set `model` to an
    aligned model id. Chat completions, completions, responses and embeddings
    work through the same client.


    **Anthropic-compatible inference.** `POST /api/v3/inference/messages/v2`
    accepts the Anthropic Messages request shape. Set `model` to an aligned
    model id.


    **Need an API key?** Sign up, log in to the platform and generate an API
    key.


    **Need help?** Log in to the platform and raise a support ticket.


    **Authentication.** Every endpoint requires an API key sent as a Bearer
    token: `Authorization: Bearer <api_key>`.
  contact:
    name: Nugen Intelligence
    url: https://nugen.in/signup
  version: 25.4.20
servers:
  - url: https://api.nugen.in
    description: Production
security: []
paths:
  /api/v3/inference/embeddings:
    post:
      tags:
        - Inference
      summary: Generate Embeddings
      description: >-
        Generate dense vector embeddings for input text.



        This endpoint creates numerical vector representations (embeddings) of
        text using specified embedding models. Embeddings are useful for
        semantic search, clustering, similarity comparisons, and RAG
        applications.



        **Request Body:**


        - `model`: Model ID (required). Use your aligned model id (e.g.
        `model_01kmqm4nrn9fw6r`)

        - `input`: Text or list of texts to embed (required) - Single string or
        array of strings. Cannot be empty and must not exceed max input tokens
        for the model

        - `dimensions` (optional): Number of dimensions for output embeddings
        (only applicable for Resizable Matryoshka Embedding models)



        **Returns:**


        - `id`: Unique identifier for the response

        - `object`: Object type (always `embedding`)

        - `data`: List of embedding objects, each containing:
          - `object`: Object type (always `embedding`)
          - `embedding`: Vector of floats representing the text (length depends on the model)
          - `index`: Position in the input array
        - `model`: Model ID used for generation

        - `created`: Unix timestamp (seconds) when response was generated

        - `usage`: Token usage statistics containing:
          - `total_tokens`: Total number of tokens used by the request


        **Example Request (Single Text):**


        ```json

        POST /api/v3/inference/embeddings

        Headers: {"Authorization": "Bearer <api_key>"}


        {
          "model": "text-embedding-xyz",
          "input": "The quick brown fox jumps over the lazy dog"
        }

        ```



        **Example Request (Multiple Texts):**


        ```json

        POST /api/v3/inference/embeddings

        Headers: {"Authorization": "Bearer <api_key>"}


        {
          "model": "model_01kmqm4nrn9fw6r",
          "input": [
            "First document for semantic search",
            "Second document for clustering",
            "Third document for similarity comparison"
          ]
        }

        ```



        **Example Request (With Dimensions):**


        ```json

        POST /api/v3/inference/embeddings

        Headers: {"Authorization": "Bearer <api_key>"}


        {
          "model": "model_01kmqm4nrn9fw6r",
          "input": "Sample text to embed",
          "dimensions": 512
        }

        ```



        **Example Response:**


        ```json

        {
          "id": "embed-abc123xyz",
          "object": "embedding",
          "data": [
            {
              "object": "embedding",
              "embedding": [0.00658765, -0.008665467, -0.007653789, 0.0023, -0.0091, 0.0125],
              "index": 0
            }
          ],
          "model": "text-embedding-ada-002",
          "created": 1705329600.0,
          "usage": {
            "total_tokens": 10
          }
        }

        ```



        **Notes:**


        - Input must not exceed the maximum input tokens supported by the model

        - Input cannot be an empty string

        - For batch processing, pass multiple texts as an array to the `input`
        field

        - The embedding vector length varies by model - check model
        documentation for specifics

        - Use `dimensions` parameter only with Matryoshka embedding models that
        support resizing
      operationId: generate_text_embeddings
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/EmbeddingRequestSchema'
        required: true
      responses:
        '200':
          description: Vector embeddings
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/EmbeddingResponseSchema'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    EmbeddingRequestSchema:
      properties:
        input:
          anyOf:
            - items:
                type: string
              type: array
            - type: string
          title: Input
          description: >
            Input text to embed, encoded as a string or a list of strings. To
            embed multiple inputs in a single request, pass a list of strings.
            The input must not exceed the max input tokens for the model, cannot
            be an empty string.
          examples:
            - The quick brown fox jumped over the lazy dog
        model:
          type: string
          title: Model
          description: >-
            Use your aligned model id (e.g. model_01kmqm4nrn9fw6r). The same
            aligned id serves chat completions, completions, embeddings and
            reranking. Alignment shapes the model's representation space, so
            aligned embeddings are recommended for domain retrieval.
          examples:
            - model_01kmqm4nrn9fw6r
        dimensions:
          anyOf:
            - type: integer
            - type: 'null'
          title: Dimensions
          description: >-
            (Applicable for only Resizable Matryoshka Embedding models) The
            number of dimensions the resulting output embeddings should have.
      type: object
      required:
        - input
        - model
      title: EmbeddingRequestSchema
    EmbeddingResponseSchema:
      properties:
        id:
          anyOf:
            - type: string
            - type: 'null'
          title: Id
          description: A unique identifier of the response.
        data:
          anyOf:
            - items:
                $ref: '#/components/schemas/EmbeddingObject'
              type: array
            - type: 'null'
          title: Data
          description: The list of embeddings generated by the model.
        model:
          anyOf:
            - type: string
            - type: 'null'
          title: Model
          description: The name of the model used to generate the embedding.
        created:
          anyOf:
            - type: number
            - type: 'null'
          title: Created
          description: The Unix time in seconds when the response was generated.
        object:
          anyOf:
            - type: string
              const: embedding
            - type: 'null'
          title: Object
          description: The object type, which is always 'embedding'.
          default: embedding
        usage:
          anyOf:
            - $ref: '#/components/schemas/EmbeddingUsage'
            - type: 'null'
          description: The usage information for the request.
      type: object
      required:
        - id
        - data
        - model
        - created
        - usage
      title: EmbeddingResponseSchema
      example:
        created: 1623645497
        data:
          - embedding:
              - 0.007897645
              - -0.006754357
              - 0.007654689
            index: 0
        id: nugen-1234
        model: model_01kmqm4nrn9fw6r
        object: embedding
        usage:
          total_tokens: 25
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    EmbeddingObject:
      properties:
        index:
          anyOf:
            - type: integer
            - type: 'null'
          title: Index
          description: The index of the embedding in the list of embeddings.
        embedding:
          anyOf:
            - items:
                type: number
              type: array
            - type: 'null'
          title: Embedding
          description: The embedding vector, which is a list of floats.
        object:
          anyOf:
            - type: string
              const: embedding
            - type: 'null'
          title: Object
          description: The object type, which is always 'embedding'.
          default: embedding
      type: object
      required:
        - index
        - embedding
      title: EmbeddingObject
      example:
        embedding:
          - 0.00658765
          - -0.008665467
          - -0.007653789
        index: 0
        object: embedding
    EmbeddingUsage:
      properties:
        total_tokens:
          type: integer
          title: Total Tokens
          description: The total number of tokens used by the request.
      type: object
      required:
        - total_tokens
      title: EmbeddingUsage
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.