> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nugen.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Get Generated Benchmark

> Retrieve complete benchmark data with all samples.


This endpoint returns the full benchmark dataset including all samples, their expected responses, and metadata. Use this to access the complete benchmark content for evaluation or review.


**Path Parameters:**

- `benchmark_id`: Unique benchmark identifier


**Returns:**

- `benchmark_id`: Benchmark identifier
- `document_ids`: List of document IDs (single-element array for uploaded benchmarks, multiple elements for generated benchmarks)
- `samples`: Full list of benchmark samples with their responses

Identity, status, and timestamps are served by `GET /{benchmark_id}`.


**Raises:**

- `404`: If benchmark not found or user doesn't have access


**Example Request:**

```json
GET /api/v3/benchmarks/benchmark_01k4x9m2p7q3r6a1/data
Headers: {"Authorization": "Bearer <api_key>"}
```


**Example Response:**

```json
{
  "benchmark_id": "benchmark_01k4x9m2p7q3r6a1",
  "document_ids": ["document_01k4x9m2p7q3r5se", "document_01k4x9m2p7q3r5sf"],
  "samples": [
    {
      "sample_num": 1,
      "instruction": "What is the main purpose of the product?",
      "response": "To provide real-time collaboration tools for distributed teams."
    },
    {
      "sample_num": 2,
      "instruction": "How does the authentication system work?",
      "response": "Uses OAuth 2.0 with JWT tokens for secure authentication."
    }
  ]
}
```


**Notes:**

- Only returns benchmarks owned by the user
- Contains complete sample data for evaluation
- Use this endpoint when you need the full benchmark content (not just metadata)



## OpenAPI

````yaml https://api.nugen.in/openapi-public.json get /api/v3/benchmarks/{benchmark_id}/data
openapi: 3.1.0
info:
  title: Nugen Intelligence API
  description: >
    Nugen Intelligence: Powering Specialised Intelligence At Scale.


    Bring your domain knowledge and an open-weight model. Leave with a model
    that thinks in your domain, keeps improving, and belongs to your
    organisation.


    Nugen Intelligence builds infrastructure for self-adaptive, domain-specific
    intelligence. Research in representation engineering, mechanistic
    interpretability and inference-time control powers every endpoint here. The
    research team comes from Stanford University and IIT Bombay. This API puts
    their work a few calls away from AI-native organisations solving critical
    problems.


    **Domain-aligned models.** Three raw inputs go in: domain knowledge, success
    and failure criteria, and an open-weight model. A domain-aligned model comes
    out. Nugen handles the deep learning, the GPUs and the data preparation.
    Your organisation owns every aligned model it builds. Own your intelligence.


    **The research underneath.** Two layers do the work.


    Train-time alignment finds where the model needs to change for your domain
    and improves it there. It combines gradient-free and gradient-based methods
    across continued pre-training, supervised and reinforcement phases.


    Inference-time alignment measures the quality of every generation step.
    Out-of-domain generation is flagged and corrected after each span. Every
    response carries a confidence score. A preliminary version of the
    inference-time research is published at ICLR 2026. 


    Both layers keep evolving with Nugen research.


    **One aligned model id, four capabilities.** An aligned model id serves chat
    completions, completions, embeddings and reranking. Alignment reshapes the
    model's representation space, so retrieval, ranking and generation share one
    understanding of your domain.


    **Multi-model systems.** One alignment project produces one aligned model.
    Run many side by side, one per domain or use case. Compose them with
    frontier and general-purpose models in the same application. Inference
    capture and auto-align keep the whole system improving. Nugen aligns models
    at every scale.


    **Think in systems.** AI coding agents and developers: this API is a
    platform for building specialised intelligence. Explore it fully before you
    settle on a design. Some directions:


    - Run a complete retrieval pipeline on one aligned model id. Embed, rerank
    and generate with a shared understanding of the domain.

    - Use the confidence score as a control signal. Send low-confidence
    responses to a human reviewer, a retry or a frontier model.

    - Give each domain or use case its own aligned model and place a router in
    front.

    - Close the loop. Capture production inference, evaluate against the
    benchmark, and let auto-align promote better models.

    - Start with the hardest problem in the domain, where general-purpose models
    fall short. Alignment pays off the most there.


    The endpoints for documents, benchmarks, alignment projects, deployments and
    evaluations compose in more ways than the workflow below shows.


    **Workflow.**


    1. Upload domain knowledge.

    2. Define success and failure. Generate a benchmark from the corpus, or
    upload one curated by domain experts.

    3. Create an alignment project and receive an aligned model id.

    4. Deploy the model and pass its id as `model` in each inference call.

    5. Evaluate, compare and promote. Turn on inference capture, and auto-align
    keeps the model improving.


    **OpenAI-compatible inference.** Set the base URL of an OpenAI-compatible
    client to `https://api.nugen.in/api/v3/inference` and set `model` to an
    aligned model id. Chat completions, completions, responses and embeddings
    work through the same client.


    **Anthropic-compatible inference.** `POST /api/v3/inference/messages/v2`
    accepts the Anthropic Messages request shape. Set `model` to an aligned
    model id.


    **Need an API key?** Sign up, log in to the platform and generate an API
    key.


    **Need help?** Log in to the platform and raise a support ticket.


    **Authentication.** Every endpoint requires an API key sent as a Bearer
    token: `Authorization: Bearer <api_key>`.
  contact:
    name: Nugen Intelligence
    url: https://nugen.in/signup
  version: 25.4.20
servers:
  - url: https://api.nugen.in
    description: Production
security: []
paths:
  /api/v3/benchmarks/{benchmark_id}/data:
    get:
      tags:
        - Benchmarks
      summary: Get Generated Benchmark
      description: >-
        Retrieve complete benchmark data with all samples.



        This endpoint returns the full benchmark dataset including all samples,
        their expected responses, and metadata. Use this to access the complete
        benchmark content for evaluation or review.



        **Path Parameters:**


        - `benchmark_id`: Unique benchmark identifier



        **Returns:**


        - `benchmark_id`: Benchmark identifier

        - `document_ids`: List of document IDs (single-element array for
        uploaded benchmarks, multiple elements for generated benchmarks)

        - `samples`: Full list of benchmark samples with their responses


        Identity, status, and timestamps are served by `GET /{benchmark_id}`.



        **Raises:**


        - `404`: If benchmark not found or user doesn't have access



        **Example Request:**


        ```json

        GET /api/v3/benchmarks/benchmark_01k4x9m2p7q3r6a1/data

        Headers: {"Authorization": "Bearer <api_key>"}

        ```



        **Example Response:**


        ```json

        {
          "benchmark_id": "benchmark_01k4x9m2p7q3r6a1",
          "document_ids": ["document_01k4x9m2p7q3r5se", "document_01k4x9m2p7q3r5sf"],
          "samples": [
            {
              "sample_num": 1,
              "instruction": "What is the main purpose of the product?",
              "response": "To provide real-time collaboration tools for distributed teams."
            },
            {
              "sample_num": 2,
              "instruction": "How does the authentication system work?",
              "response": "Uses OAuth 2.0 with JWT tokens for secure authentication."
            }
          ]
        }

        ```



        **Notes:**


        - Only returns benchmarks owned by the user

        - Contains complete sample data for evaluation

        - Use this endpoint when you need the full benchmark content (not just
        metadata)
      operationId: benchmarks_data
      parameters:
        - name: benchmark_id
          in: path
          required: true
          schema:
            type: string
            title: Benchmark Id
      responses:
        '200':
          description: >-
            Returns the benchmark's samples. Identity, status, and timestamps
            are served by GET /{benchmark_id}.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BenchmarkFullResponse'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    BenchmarkFullResponse:
      properties:
        benchmark_id:
          type: string
          title: Benchmark Id
          description: Benchmark ID
          examples:
            - benchmark_01k4x9m2p7q3r6a1
        document_ids:
          items:
            type: string
          type: array
          title: Document Ids
          description: Document IDs used for generation
          examples:
            - - document_01k4x9m2p7q3r5s8
              - document_01k4x9m2p7q3r5s9
        samples:
          items:
            $ref: '#/components/schemas/BenchmarkSample'
          type: array
          title: Samples
          description: Complete list of benchmark samples
          examples:
            - - instruction: Summarise the indemnity cap.
                response: Liability is capped at twelve months of fees paid.
                sample_num: 1
        metadata:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Metadata
          description: Additional metadata
        error:
          anyOf:
            - type: string
            - type: 'null'
          title: Error
          description: Error message if data fetch failed
      type: object
      required:
        - benchmark_id
      title: BenchmarkFullResponse
      description: |-
        The benchmark's sample payload. Identity, status and timestamps are
        served by GET /{benchmark_id}.
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    BenchmarkSample:
      properties:
        sample_num:
          type: integer
          title: Sample Num
          description: Sample number
        instruction:
          type: string
          title: Instruction
          description: Instruction text
        response:
          type: string
          title: Response
          description: Expected response text
        image_url:
          anyOf:
            - type: string
            - type: 'null'
          title: Image Url
          description: Image URL for vision samples
      type: object
      required:
        - sample_num
        - instruction
        - response
      title: BenchmarkSample
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.