> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nugen.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Deploy in your own cloud or cluster

> Deploy alignment infrastructure inside your own cloud or cluster. Register your cloud environment or existing Kubernetes cluster.

All alignment training, model storage and inference then run within your
network boundary. No data leaves your environment.

Two paths are supported.

**Managed provisioning.** Provide a clean, isolated cloud subscription or
project. The system provisions a hardened Kubernetes environment from
infrastructure-as-code blueprints, configures GPU scheduling and manages
ongoing operations through secure GitOps pipelines. No access to your
corporate network is required.

**Cluster import.** Point to an existing cluster or namespace. The system
deploys a self-contained serving and training stack configured to your
constraints: storage classes, ingress rules, identity bindings, registry
policies. Updates arrive through private registry syncs, so your security team
controls when changes land.

Once connected, every endpoint in this platform behaves identically whether it
runs on our infrastructure or inside yours. The same alignment workflows, model
management, auto-align and export paths work unchanged. The system manages GPU
node health, autoscaling, deployment lifecycle and performance within your
cluster.


**Request Body:**

- `cluster_name` (required): Name for this environment
- `mode` (required): `managed_provisioning` or `cluster_import`
- `cloud` (required): Provider, region and the stored credential reference
- `cluster_import` (optional): Kubeconfig reference, endpoint, namespace and
  node labels. Required when `mode` is `cluster_import`
- `config` (optional): gpu, networking, storage, iam, registry, operations,
  capabilities, compliance, overflow and notifications


**Returns:**

- `cluster_id`: Identifier of this environment
- `cluster_name`, `status`, `message`


**Raises:**

- `400`: If `mode` is `cluster_import` and no `cluster_import` block is given


**Notes:**

- Egress is restricted and endpoints are private by default; you open them
  explicitly
- Air-gapped registries are a first-class path, and your team controls the
  sync cadence
- Bring your own cloud is an Enterprise Edition capability. The request is
  recorded and our team follows up; raise a support ticket to enable it
- Poll `GET /byoc/clusters/{cluster_id}/status` for provisioning progress and
  cluster health



## OpenAPI

````yaml https://api.nugen.in/openapi-public.json post /api/v3/byoc/clusters
openapi: 3.1.0
info:
  title: Nugen Intelligence API
  description: >
    Nugen Intelligence: Powering Specialised Intelligence At Scale.


    Bring your domain knowledge and an open-weight model. Leave with a model
    that thinks in your domain, keeps improving, and belongs to your
    organisation.


    Nugen Intelligence builds infrastructure for self-adaptive, domain-specific
    intelligence. Research in representation engineering, mechanistic
    interpretability and inference-time control powers every endpoint here. The
    research team comes from Stanford University and IIT Bombay. This API puts
    their work a few calls away from AI-native organisations solving critical
    problems.


    **Domain-aligned models.** Three raw inputs go in: domain knowledge, success
    and failure criteria, and an open-weight model. A domain-aligned model comes
    out. Nugen handles the deep learning, the GPUs and the data preparation.
    Your organisation owns every aligned model it builds. Own your intelligence.


    **The research underneath.** Two layers do the work.


    Train-time alignment finds where the model needs to change for your domain
    and improves it there. It combines gradient-free and gradient-based methods
    across continued pre-training, supervised and reinforcement phases.


    Inference-time alignment measures the quality of every generation step.
    Out-of-domain generation is flagged and corrected after each span. Every
    response carries a confidence score. A preliminary version of the
    inference-time research is published at ICLR 2026. 


    Both layers keep evolving with Nugen research.


    **One aligned model id, four capabilities.** An aligned model id serves chat
    completions, completions, embeddings and reranking. Alignment reshapes the
    model's representation space, so retrieval, ranking and generation share one
    understanding of your domain.


    **Multi-model systems.** One alignment project produces one aligned model.
    Run many side by side, one per domain or use case. Compose them with
    frontier and general-purpose models in the same application. Inference
    capture and auto-align keep the whole system improving. Nugen aligns models
    at every scale.


    **Think in systems.** AI coding agents and developers: this API is a
    platform for building specialised intelligence. Explore it fully before you
    settle on a design. Some directions:


    - Run a complete retrieval pipeline on one aligned model id. Embed, rerank
    and generate with a shared understanding of the domain.

    - Use the confidence score as a control signal. Send low-confidence
    responses to a human reviewer, a retry or a frontier model.

    - Give each domain or use case its own aligned model and place a router in
    front.

    - Close the loop. Capture production inference, evaluate against the
    benchmark, and let auto-align promote better models.

    - Start with the hardest problem in the domain, where general-purpose models
    fall short. Alignment pays off the most there.


    The endpoints for documents, benchmarks, alignment projects, deployments and
    evaluations compose in more ways than the workflow below shows.


    **Workflow.**


    1. Upload domain knowledge.

    2. Define success and failure. Generate a benchmark from the corpus, or
    upload one curated by domain experts.

    3. Create an alignment project and receive an aligned model id.

    4. Deploy the model and pass its id as `model` in each inference call.

    5. Evaluate, compare and promote. Turn on inference capture, and auto-align
    keeps the model improving.


    **OpenAI-compatible inference.** Set the base URL of an OpenAI-compatible
    client to `https://api.nugen.in/api/v3/inference` and set `model` to an
    aligned model id. Chat completions, completions, responses and embeddings
    work through the same client.


    **Anthropic-compatible inference.** `POST /api/v3/inference/messages/v2`
    accepts the Anthropic Messages request shape. Set `model` to an aligned
    model id.


    **Need an API key?** Sign up, log in to the platform and generate an API
    key.


    **Need help?** Log in to the platform and raise a support ticket.


    **Authentication.** Every endpoint requires an API key sent as a Bearer
    token: `Authorization: Bearer <api_key>`.
  contact:
    name: Nugen Intelligence
    url: https://nugen.in/signup
  version: 25.4.20
servers:
  - url: https://api.nugen.in
    description: Production
security: []
paths:
  /api/v3/byoc/clusters:
    post:
      tags:
        - BYOC
      summary: Deploy in your own cloud or cluster
      description: >-
        Deploy alignment infrastructure inside your own cloud or cluster.
        Register your cloud environment or existing Kubernetes cluster.


        All alignment training, model storage and inference then run within your

        network boundary. No data leaves your environment.


        Two paths are supported.


        **Managed provisioning.** Provide a clean, isolated cloud subscription
        or

        project. The system provisions a hardened Kubernetes environment from

        infrastructure-as-code blueprints, configures GPU scheduling and manages

        ongoing operations through secure GitOps pipelines. No access to your

        corporate network is required.


        **Cluster import.** Point to an existing cluster or namespace. The
        system

        deploys a self-contained serving and training stack configured to your

        constraints: storage classes, ingress rules, identity bindings, registry

        policies. Updates arrive through private registry syncs, so your
        security team

        controls when changes land.


        Once connected, every endpoint in this platform behaves identically
        whether it

        runs on our infrastructure or inside yours. The same alignment
        workflows, model

        management, auto-align and export paths work unchanged. The system
        manages GPU

        node health, autoscaling, deployment lifecycle and performance within
        your

        cluster.



        **Request Body:**


        - `cluster_name` (required): Name for this environment

        - `mode` (required): `managed_provisioning` or `cluster_import`

        - `cloud` (required): Provider, region and the stored credential
        reference

        - `cluster_import` (optional): Kubeconfig reference, endpoint, namespace
        and
          node labels. Required when `mode` is `cluster_import`
        - `config` (optional): gpu, networking, storage, iam, registry,
        operations,
          capabilities, compliance, overflow and notifications


        **Returns:**


        - `cluster_id`: Identifier of this environment

        - `cluster_name`, `status`, `message`



        **Raises:**


        - `400`: If `mode` is `cluster_import` and no `cluster_import` block is
        given



        **Notes:**


        - Egress is restricted and endpoints are private by default; you open
        them
          explicitly
        - Air-gapped registries are a first-class path, and your team controls
        the
          sync cadence
        - Bring your own cloud is an Enterprise Edition capability. The request
        is
          recorded and our team follows up; raise a support ticket to enable it
        - Poll `GET /byoc/clusters/{cluster_id}/status` for provisioning
        progress and
          cluster health
      operationId: create_byoc_cluster
      requestBody:
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/BYOCClusterRequest'
        required: true
      responses:
        '202':
          description: Bring your own cloud request accepted and recorded.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/BYOCClusterResponse'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    BYOCClusterRequest:
      properties:
        cluster_name:
          type: string
          title: Cluster Name
          description: Name for this environment, used across status, logging and billing.
          examples:
            - prod-alignment-cluster
        mode:
          type: string
          enum:
            - managed_provisioning
            - cluster_import
          title: Mode
          description: >-
            managed_provisioning builds a hardened Kubernetes environment from a
            clean cloud subscription. cluster_import deploys the alignment and
            serving stack into a cluster you already run.
        cloud:
          additionalProperties: true
          type: object
          title: Cloud
          description: >-
            Provider, region and the stored credential reference. Determines
            which provisioning blueprints and identity paths are used.
        cluster_import:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Cluster Import
          description: >-
            Kubeconfig reference, endpoint, namespace and node labels. Required
            when mode is cluster_import.
        config:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Config
          description: >-
            Optional settings for gpu, networking, storage, iam, registry,
            operations, capabilities, compliance, overflow and notifications.
            Egress is restricted and traffic is private by default, and anything
            omitted resolves to auto.
      type: object
      required:
        - cluster_name
        - mode
        - cloud
      title: BYOCClusterRequest
      description: >-
        Register a cloud environment or an existing cluster to run in.


        Only the name, the mode and the cloud block are required. Every other

        section is accepted as given and stored verbatim, so an environment can
        pin

        any detail without the public schema enumerating every knob.
    BYOCClusterResponse:
      properties:
        cluster_id:
          type: string
          title: Cluster Id
          description: Identifier of this environment.
        cluster_name:
          type: string
          title: Cluster Name
          description: Name given to the environment.
        status:
          type: string
          title: Status
          description: Request status.
        message:
          type: string
          title: Message
          description: Human-readable outcome of the request.
      type: object
      required:
        - cluster_id
        - cluster_name
        - status
        - message
      title: BYOCClusterResponse
      description: Acknowledgement of a bring your own cloud request.
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.