> ## Documentation Index
> Fetch the complete documentation index at: https://docs.nugen.in/llms.txt
> Use this file to discover all available pages before exploring further.

# Start Auto-Align

> Submit a target alignment score and let the system plan the run.

The system analyses your data, base model and prior train-time alignment history to build a
multi-phase alignment plan.

Phases are sequenced and executed continuously. The system may apply
gradient-free representation adjustments, supervised learning, reward-guided
optimisation, or a combination, re-planning when the score trajectory demands
it. Checkpoints are evaluated between phases, and the system revisits earlier
phases if regressions surface.

Every parameter defaults to `auto`. You can override any of them to constrain
the search, and the system performs best when given room to explore.


**Path Parameters:**

- `alignment_id`: The alignment project to auto-align. Its base model, data and
  prior train-time alignment history warm-start the planning.


**Request Body:**

- `target_score` (required): Target quality on a 0-99 scale
- `benchmark_id` (optional): Benchmark that measures the achieved score against
  the target. Defaults to the benchmark already on the project
- `method`, `representations`, `sequencing` (optional): Constrain which
  optimisation families run and in what order
- `max_compute_hours` (optional): Upper bound on GPU hours across all phases
- `notify_on_phase_transition` (optional): Notify on each phase change
- `config` (optional): Take control of any part of the run, for reinforcement
  learning or supervised (SFT) settings: loss, optimizer, LoRA adapter, reward,
  advantage, reference policy, environment, dataset, sampling, training,
  checkpointing, evaluation, rendering and logging. Pass only what you want to
  pin; anything omitted stays auto and redundant settings are discarded. Every
  field is listed in the `AutoAlignConfig` schema, with a full example


**Returns:**

- `auto_align_id`: Identifier of this request
- `alignment_id`: The project it targets
- `status`, `message`


**Raises:**

- `404`: If the alignment project is not found or does not belong to you


**Notes:**

- Auto-alignment is an Enterprise Edition capability. The request is recorded
  and our team follows up; raise a support ticket to enable it
- Once enabled, poll `GET /alignment-projects/{alignment_id}/status` for
  progress, phase transitions and intermediate scores



## OpenAPI

````yaml https://api.nugen.in/openapi-public.json post /api/v3/alignment-projects/{alignment_id}/auto-align
openapi: 3.1.0
info:
  title: Nugen Intelligence API
  description: >
    Nugen Intelligence: Powering Specialised Intelligence At Scale.


    Bring your domain knowledge and an open-weight model. Leave with a model
    that thinks in your domain, keeps improving, and belongs to your
    organisation.


    Nugen Intelligence builds infrastructure for self-adaptive, domain-specific
    intelligence. Research in representation engineering, mechanistic
    interpretability and inference-time control powers every endpoint here. The
    research team comes from Stanford University and IIT Bombay. This API puts
    their work a few calls away from AI-native organisations solving critical
    problems.


    **Domain-aligned models.** Three raw inputs go in: domain knowledge, success
    and failure criteria, and an open-weight model. A domain-aligned model comes
    out. Nugen handles the deep learning, the GPUs and the data preparation.
    Your organisation owns every aligned model it builds. Own your intelligence.


    **The research underneath.** Two layers do the work.


    Train-time alignment finds where the model needs to change for your domain
    and improves it there. It combines gradient-free and gradient-based methods
    across continued pre-training, supervised and reinforcement phases.


    Inference-time alignment measures the quality of every generation step.
    Out-of-domain generation is flagged and corrected after each span. Every
    response carries a confidence score. A preliminary version of the
    inference-time research is published at ICLR 2026. 


    Both layers keep evolving with Nugen research.


    **One aligned model id, four capabilities.** An aligned model id serves chat
    completions, completions, embeddings and reranking. Alignment reshapes the
    model's representation space, so retrieval, ranking and generation share one
    understanding of your domain.


    **Multi-model systems.** One alignment project produces one aligned model.
    Run many side by side, one per domain or use case. Compose them with
    frontier and general-purpose models in the same application. Inference
    capture and auto-align keep the whole system improving. Nugen aligns models
    at every scale.


    **Think in systems.** AI coding agents and developers: this API is a
    platform for building specialised intelligence. Explore it fully before you
    settle on a design. Some directions:


    - Run a complete retrieval pipeline on one aligned model id. Embed, rerank
    and generate with a shared understanding of the domain.

    - Use the confidence score as a control signal. Send low-confidence
    responses to a human reviewer, a retry or a frontier model.

    - Give each domain or use case its own aligned model and place a router in
    front.

    - Close the loop. Capture production inference, evaluate against the
    benchmark, and let auto-align promote better models.

    - Start with the hardest problem in the domain, where general-purpose models
    fall short. Alignment pays off the most there.


    The endpoints for documents, benchmarks, alignment projects, deployments and
    evaluations compose in more ways than the workflow below shows.


    **Workflow.**


    1. Upload domain knowledge.

    2. Define success and failure. Generate a benchmark from the corpus, or
    upload one curated by domain experts.

    3. Create an alignment project and receive an aligned model id.

    4. Deploy the model and pass its id as `model` in each inference call.

    5. Evaluate, compare and promote. Turn on inference capture, and auto-align
    keeps the model improving.


    **OpenAI-compatible inference.** Set the base URL of an OpenAI-compatible
    client to `https://api.nugen.in/api/v3/inference` and set `model` to an
    aligned model id. Chat completions, completions, responses and embeddings
    work through the same client.


    **Anthropic-compatible inference.** `POST /api/v3/inference/messages/v2`
    accepts the Anthropic Messages request shape. Set `model` to an aligned
    model id.


    **Need an API key?** Sign up, log in to the platform and generate an API
    key.


    **Need help?** Log in to the platform and raise a support ticket.


    **Authentication.** Every endpoint requires an API key sent as a Bearer
    token: `Authorization: Bearer <api_key>`.
  contact:
    name: Nugen Intelligence
    url: https://nugen.in/signup
  version: 25.4.20
servers:
  - url: https://api.nugen.in
    description: Production
security: []
paths:
  /api/v3/alignment-projects/{alignment_id}/auto-align:
    post:
      tags:
        - Model Alignment
      summary: Start Auto-Align
      description: >-
        Submit a target alignment score and let the system plan the run.


        The system analyses your data, base model and prior train-time alignment
        history to build a

        multi-phase alignment plan.


        Phases are sequenced and executed continuously. The system may apply

        gradient-free representation adjustments, supervised learning,
        reward-guided

        optimisation, or a combination, re-planning when the score trajectory
        demands

        it. Checkpoints are evaluated between phases, and the system revisits
        earlier

        phases if regressions surface.


        Every parameter defaults to `auto`. You can override any of them to
        constrain

        the search, and the system performs best when given room to explore.



        **Path Parameters:**


        - `alignment_id`: The alignment project to auto-align. Its base model,
        data and
          prior train-time alignment history warm-start the planning.


        **Request Body:**


        - `target_score` (required): Target quality on a 0-99 scale

        - `benchmark_id` (optional): Benchmark that measures the achieved score
        against
          the target. Defaults to the benchmark already on the project
        - `method`, `representations`, `sequencing` (optional): Constrain which
          optimisation families run and in what order
        - `max_compute_hours` (optional): Upper bound on GPU hours across all
        phases

        - `notify_on_phase_transition` (optional): Notify on each phase change

        - `config` (optional): Take control of any part of the run, for
        reinforcement
          learning or supervised (SFT) settings: loss, optimizer, LoRA adapter, reward,
          advantage, reference policy, environment, dataset, sampling, training,
          checkpointing, evaluation, rendering and logging. Pass only what you want to
          pin; anything omitted stays auto and redundant settings are discarded. Every
          field is listed in the `AutoAlignConfig` schema, with a full example


        **Returns:**


        - `auto_align_id`: Identifier of this request

        - `alignment_id`: The project it targets

        - `status`, `message`



        **Raises:**


        - `404`: If the alignment project is not found or does not belong to you



        **Notes:**


        - Auto-alignment is an Enterprise Edition capability. The request is
        recorded
          and our team follows up; raise a support ticket to enable it
        - Once enabled, poll `GET /alignment-projects/{alignment_id}/status` for
          progress, phase transitions and intermediate scores
      operationId: create_auto_alignment
      parameters:
        - name: alignment_id
          in: path
          required: true
          schema:
            type: string
            title: Alignment Id
      requestBody:
        required: true
        content:
          application/json:
            schema:
              $ref: '#/components/schemas/AutoAlignRequest'
      responses:
        '202':
          description: Auto-alignment request accepted and recorded.
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/AutoAlignResponse'
        '422':
          description: Validation Error
          content:
            application/json:
              schema:
                $ref: '#/components/schemas/HTTPValidationError'
      security:
        - HTTPBearer: []
components:
  schemas:
    AutoAlignRequest:
      properties:
        target_score:
          type: number
          maximum: 99
          minimum: 0
          title: Target Score
          description: >-
            Target alignment quality on a 0-99 scale. The system iterates across
            phases until this score is reached or the compute budget is
            consumed. Targets above 85 typically cause the system to explore
            multiple optimisation strategies in sequence.
          examples:
            - 88
        benchmark_id:
          anyOf:
            - type: string
            - type: 'null'
          title: Benchmark Id
          description: >-
            Benchmark used to measure the achieved score against the target.
            Defaults to the benchmark already on the alignment project.
        method:
          anyOf:
            - type: string
            - type: 'null'
          title: Method
          description: >-
            Which optimisation families the system may use. On auto it selects
            and sequences them from the score trajectory, and may begin with one
            approach, move to another, and revisit earlier ones.
          default: auto
        representations:
          anyOf:
            - type: string
            - type: 'null'
          title: Representations
          description: >-
            How internal model representations are handled before and during
            alignment. On auto the system profiles the base model and picks a
            strategy before any gradient-based phase begins.
          default: auto
        sequencing:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          title: Sequencing
          description: >-
            Explicit phase ordering. When null the system orders phases itself
            and may interleave or revisit them. When given, it follows the order
            and still controls the hyperparameters within each phase.
        max_compute_hours:
          anyOf:
            - type: number
            - type: 'null'
          title: Max Compute Hours
          description: >-
            Upper bound on GPU hours across all phases, allocated by expected
            marginal score gain. When null the run continues until the target.
        notify_on_phase_transition:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Notify On Phase Transition
          description: Notify when the system moves between alignment phases.
          default: true
        config:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignConfig'
            - type: 'null'
          description: >-
            Take control of any part of the run, for reinforcement learning or
            supervised (SFT) settings: loss, optimizer, reward, advantage,
            reference policy, environment, dataset, sampling, training, adapter,
            checkpointing, evaluation, rendering and logging. Pass only what you
            want to pin; anything omitted stays auto, and redundant settings are
            discarded.
          examples:
            - adapter:
                alpha: 32
                dropout: 0
                init_seed: 7
                mode: lora
                quantization_bits: 4
                rank: 16
                train_attn: true
                train_mlp: true
                train_unembed: true
              advantage:
                center: group_mean
                estimator: gae
                gae_lambda: 0.95
                gamma: 0.99
                normalize_std: true
                scope: per_token
                skip_zero_advantage_groups: true
              base_model: Qwen/Qwen2.5-0.5B-Instruct
              checkpointing:
                keep_last_n: 3
                mode: training
                promote_final: true
                save_every_steps: 50
              dataset:
                answer_field: responses
                file_type: jsonl
                format: chat
                prompt_field: instructions
                shuffle: true
                split: train
                streaming: true
                synthetic:
                  enabled: false
              environment:
                fail_fast: false
                max_turns: 8
                num_agents: 1
                retry_on_failure: true
                sandbox:
                  backend: local
                  enabled: false
                tools: []
                type: multi_turn
              evaluation:
                eval_batch_size: 64
                eval_every_steps: 50
                eval_group_size: 1
                eval_temperature: 0
                metrics:
                  - reward
                  - task_success
              logging:
                log_metrics:
                  - reward
                  - loss
                  - kl
                  - entropy
                  - grad_norm
                wandb_project: alignment-id
              loss:
                cispo:
                  eps_max: 6
                custom_weights_field: token_weights
                dro:
                  beta: 0.05
                entropy_coef: 0.001
                is_ratio_level: token
                kl_coef: 0.02
                ppo:
                  clip_high: 0.2
                  clip_low: 0.2
                token_reduction: sum
                type: cispo
                z_loss_coef: 0.0001
              optimizer:
                beta1: 0.9
                beta2: 0.95
                eps: 1.e-8
                grad_clip_norm: 1
                learning_rate: 0.000025
                lr_schedule: cosine
                type: adamw
                warmup_steps: 20
                weight_decay: 0
              reference_policy:
                ema_decay: 0.999
                enabled: true
                kl_target: 0.05
                source: frozen_snapshot
              rendering:
                add_special_tokens: true
                batch_unit: tokens
                max_seq_len: 65536
                packing: true
                truncation: right
              reward:
                clip:
                  - -1
                  - 1
                composite:
                  task_success: 1
                  tool_format: 0.2
                llm_judge:
                  criteria: Were the right words used to reach the correct answer?
                  judge_model: Qwen/Qwen2.5-0.5B-Instruct
                normalize: per_group_zscore
                type: llm_judge
              sampling:
                enable_thinking: true
                group_size: 8
                max_prompt_tokens: 4096
                max_seq_len: 131072
                max_tokens: 1024
                min_p: 0
                prompt_caching: true
                prompt_groups_per_step: 64
                seed: 12345
                session_affinity_key: trajectory_id
                stop:
                  - <|im_end|>
                temperature: 1
                thinking_effort: high
                top_k: 40
                top_p: 0.95
              training:
                batch_size: 256
                early_stopping:
                  enabled: true
                  metric: eval_loss
                  patience: 3
                epochs: 3
                gradient_accumulation: 1
                inner_optim_steps: 2
                max_steps: 2000
                mode: async
                rollout_training_overlap: true
                seed: 42
                steps: 500
      type: object
      required:
        - target_score
      title: AutoAlignRequest
      description: |-
        Target a score and let the system plan the alignment to reach it.

        Every field except target_score defaults to auto. The deep configuration
        blocks are accepted as given and stored verbatim, so a run can pin any
        detail without the public schema enumerating every knob.
    AutoAlignResponse:
      properties:
        auto_align_id:
          type: string
          title: Auto Align Id
          description: Identifier of the auto-alignment request.
        alignment_id:
          type: string
          title: Alignment Id
          description: Alignment project the request targets.
        status:
          type: string
          title: Status
          description: Request status.
        message:
          type: string
          title: Message
          description: Human-readable outcome of the request.
      type: object
      required:
        - auto_align_id
        - alignment_id
        - status
        - message
      title: AutoAlignResponse
      description: Acknowledgement of an auto-alignment request.
    HTTPValidationError:
      properties:
        detail:
          items:
            $ref: '#/components/schemas/ValidationError'
          type: array
          title: Detail
      type: object
      title: HTTPValidationError
    AutoAlignConfig:
      properties:
        base_model:
          anyOf:
            - type: string
            - type: 'null'
          title: Base Model
          description: Base model by Hugging Face id; defaults to the alignment project's
        loss:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignLoss'
            - type: 'null'
        optimizer:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignOptimizer'
            - type: 'null'
        reward:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignReward'
            - type: 'null'
        advantage:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignAdvantage'
            - type: 'null'
        reference_policy:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignReferencePolicy'
            - type: 'null'
        environment:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignEnvironment'
            - type: 'null'
        dataset:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignDataset'
            - type: 'null'
        sampling:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignSampling'
            - type: 'null'
        training:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignTraining'
            - type: 'null'
        adapter:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignAdapter'
            - type: 'null'
        checkpointing:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignCheckpointing'
            - type: 'null'
        evaluation:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignEvaluation'
            - type: 'null'
        rendering:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignRendering'
            - type: 'null'
        logging:
          anyOf:
            - $ref: '#/components/schemas/AutoAlignLogging'
            - type: 'null'
      additionalProperties: true
      type: object
      title: AutoAlignConfig
    ValidationError:
      properties:
        loc:
          items:
            anyOf:
              - type: string
              - type: integer
          type: array
          title: Location
        msg:
          type: string
          title: Message
        type:
          type: string
          title: Error Type
      type: object
      required:
        - loc
        - msg
        - type
      title: ValidationError
    AutoAlignLoss:
      properties:
        type:
          anyOf:
            - type: string
            - type: 'null'
          title: Type
          description: >-
            Objective: cross_entropy for SFT, or ppo, cispo, dro for
            reinforcement learning
        ppo:
          anyOf:
            - additionalProperties:
                type: number
              type: object
            - type: 'null'
          title: Ppo
          description: PPO clip bounds, clip_low and clip_high
        cispo:
          anyOf:
            - additionalProperties:
                type: number
              type: object
            - type: 'null'
          title: Cispo
          description: CISPO settings, eps_max
        dro:
          anyOf:
            - additionalProperties:
                type: number
              type: object
            - type: 'null'
          title: Dro
          description: DRO settings, beta
        kl_coef:
          anyOf:
            - type: number
            - type: 'null'
          title: Kl Coef
          description: KL penalty against the reference policy
        entropy_coef:
          anyOf:
            - type: number
            - type: 'null'
          title: Entropy Coef
          description: Entropy bonus
        token_reduction:
          anyOf:
            - type: string
            - type: 'null'
          title: Token Reduction
          description: 'How token losses reduce: sum or mean'
        is_ratio_level:
          anyOf:
            - type: string
            - type: 'null'
          title: Is Ratio Level
          description: 'Importance ratio level: token or sequence'
        custom_loss_ref:
          anyOf:
            - type: string
            - type: 'null'
          title: Custom Loss Ref
          description: Reference to a custom loss
        custom_weights_field:
          anyOf:
            - type: string
            - type: 'null'
          title: Custom Weights Field
          description: Dataset field holding per-token weights
        z_loss_coef:
          anyOf:
            - type: number
            - type: 'null'
          title: Z Loss Coef
          description: Z-loss coefficient
      additionalProperties: true
      type: object
      title: AutoAlignLoss
    AutoAlignOptimizer:
      properties:
        type:
          anyOf:
            - type: string
            - type: 'null'
          title: Type
          description: Optimizer, e.g. adamw
        learning_rate:
          anyOf:
            - type: number
            - type: 'null'
          title: Learning Rate
          description: Peak learning rate
        beta1:
          anyOf:
            - type: number
            - type: 'null'
          title: Beta1
          description: Adam beta1
        beta2:
          anyOf:
            - type: number
            - type: 'null'
          title: Beta2
          description: Adam beta2
        eps:
          anyOf:
            - type: number
            - type: 'null'
          title: Eps
          description: Adam epsilon
        weight_decay:
          anyOf:
            - type: number
            - type: 'null'
          title: Weight Decay
          description: Weight decay
        grad_clip_norm:
          anyOf:
            - type: number
            - type: 'null'
          title: Grad Clip Norm
          description: Global gradient norm clip
        lr_schedule:
          anyOf:
            - type: string
            - type: 'null'
          title: Lr Schedule
          description: Learning rate schedule, e.g. cosine or linear
        warmup_steps:
          anyOf:
            - type: integer
            - type: 'null'
          title: Warmup Steps
          description: Warmup steps
        warmup_ratio:
          anyOf:
            - type: number
            - type: 'null'
          title: Warmup Ratio
          description: Warmup as a fraction of steps, instead of warmup_steps
        lr_multiplier:
          anyOf:
            - type: number
            - type: 'null'
          title: Lr Multiplier
          description: Multiplier on the learning rate
      additionalProperties: true
      type: object
      title: AutoAlignOptimizer
    AutoAlignReward:
      properties:
        type:
          anyOf:
            - type: string
            - type: 'null'
          title: Type
          description: >-
            Reward source: llm_judge, verifier, rubric, preference_model or
            composite
        verifier_ref:
          anyOf:
            - type: string
            - type: 'null'
          title: Verifier Ref
          description: Reference to a programmatic verifier
        llm_judge:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Llm Judge
          description: 'Judge settings: judge_model and criteria'
        rubric_ref:
          anyOf:
            - type: string
            - type: 'null'
          title: Rubric Ref
          description: Reference to a scoring rubric
        preference_model_ref:
          anyOf:
            - type: string
            - type: 'null'
          title: Preference Model Ref
          description: Reference to a preference model
        composite:
          anyOf:
            - additionalProperties:
                type: number
              type: object
            - type: 'null'
          title: Composite
          description: Weights per reward component
        clip:
          anyOf:
            - items:
                type: number
              type: array
            - type: 'null'
          title: Clip
          description: Reward clip range, [low, high]
        normalize:
          anyOf:
            - type: string
            - type: 'null'
          title: Normalize
          description: Reward normalisation, e.g. per_group_zscore
      additionalProperties: true
      type: object
      title: AutoAlignReward
    AutoAlignAdvantage:
      properties:
        estimator:
          anyOf:
            - type: string
            - type: 'null'
          title: Estimator
          description: Advantage estimator, e.g. gae or group_relative
        center:
          anyOf:
            - type: string
            - type: 'null'
          title: Center
          description: Baseline used to centre advantages, e.g. group_mean
        normalize_std:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Normalize Std
          description: Divide advantages by their standard deviation
        skip_zero_advantage_groups:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Skip Zero Advantage Groups
          description: Skip groups whose advantages are all zero
        gamma:
          anyOf:
            - type: number
            - type: 'null'
          title: Gamma
          description: Discount factor
        gae_lambda:
          anyOf:
            - type: number
            - type: 'null'
          title: Gae Lambda
          description: GAE lambda
        scope:
          anyOf:
            - type: string
            - type: 'null'
          title: Scope
          description: Per token or per sequence
      additionalProperties: true
      type: object
      title: AutoAlignAdvantage
    AutoAlignReferencePolicy:
      properties:
        enabled:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Enabled
          description: Keep a reference policy for the KL penalty
        source:
          anyOf:
            - type: string
            - type: 'null'
          title: Source
          description: frozen_snapshot or ema
        snapshot_path:
          anyOf:
            - type: string
            - type: 'null'
          title: Snapshot Path
          description: Snapshot to use as the reference
        ema_decay:
          anyOf:
            - type: number
            - type: 'null'
          title: Ema Decay
          description: EMA decay when source is ema
        kl_target:
          anyOf:
            - type: number
            - type: 'null'
          title: Kl Target
          description: Target KL for adaptive control
      additionalProperties: true
      type: object
      title: AutoAlignReferencePolicy
    AutoAlignEnvironment:
      properties:
        type:
          anyOf:
            - type: string
            - type: 'null'
          title: Type
          description: single_turn or multi_turn
        env_ref:
          anyOf:
            - type: string
            - type: 'null'
          title: Env Ref
          description: Reference to an environment
        max_turns:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Turns
          description: Maximum turns per episode
        tools:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          title: Tools
          description: Tools available to the model
        num_agents:
          anyOf:
            - type: integer
            - type: 'null'
          title: Num Agents
          description: Agents per episode
        rollout_strategy:
          anyOf:
            - type: string
            - type: 'null'
          title: Rollout Strategy
          description: Rollout strategy
        sandbox:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Sandbox
          description: 'Sandbox settings: enabled and backend'
        retry_on_failure:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Retry On Failure
          description: Retry a failed rollout
        fail_fast:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Fail Fast
          description: Stop the run on the first rollout failure
      additionalProperties: true
      type: object
      title: AutoAlignEnvironment
    AutoAlignDataset:
      properties:
        name:
          anyOf:
            - type: string
            - type: 'null'
          title: Name
          description: Dataset name; defaults to the alignment project's data
        split:
          anyOf:
            - type: string
            - type: 'null'
          title: Split
          description: Split to train on
        prompt_field:
          anyOf:
            - type: string
            - type: 'null'
          title: Prompt Field
          description: Field holding the prompt
        answer_field:
          anyOf:
            - type: string
            - type: 'null'
          title: Answer Field
          description: Field holding the reference answer
        completion_field:
          anyOf:
            - type: string
            - type: 'null'
          title: Completion Field
          description: Field holding a completion
        shuffle:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Shuffle
          description: Shuffle examples
        streaming:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Streaming
          description: Stream instead of loading in memory
        format:
          anyOf:
            - type: string
            - type: 'null'
          title: Format
          description: chat or completion
        file_type:
          anyOf:
            - type: string
            - type: 'null'
          title: File Type
          description: jsonl, json or parquet
        max_examples:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Examples
          description: Cap on examples used
        synthetic:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Synthetic
          description: >-
            Synthetic training dataset: enabled, domain_description, label_list,
            num_examples
      additionalProperties: true
      type: object
      title: AutoAlignDataset
    AutoAlignSampling:
      properties:
        group_size:
          anyOf:
            - type: integer
            - type: 'null'
          title: Group Size
          description: Completions sampled per prompt
        prompt_groups_per_step:
          anyOf:
            - type: integer
            - type: 'null'
          title: Prompt Groups Per Step
          description: Prompt groups per step
        max_tokens:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Tokens
          description: Maximum generated tokens
        max_prompt_tokens:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Prompt Tokens
          description: Maximum prompt tokens
        max_seq_len:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Seq Len
          description: Maximum sequence length
        temperature:
          anyOf:
            - type: number
            - type: 'null'
          title: Temperature
          description: Sampling temperature
        top_p:
          anyOf:
            - type: number
            - type: 'null'
          title: Top P
          description: Nucleus sampling
        top_k:
          anyOf:
            - type: integer
            - type: 'null'
          title: Top K
          description: Top-k sampling
        min_p:
          anyOf:
            - type: number
            - type: 'null'
          title: Min P
          description: Min-p sampling
        stop:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          title: Stop
          description: Stop sequences
        seed:
          anyOf:
            - type: integer
            - type: 'null'
          title: Seed
          description: Sampling seed
        enable_thinking:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Enable Thinking
          description: Allow a thinking phase
        thinking_effort:
          anyOf:
            - type: string
            - type: 'null'
          title: Thinking Effort
          description: low, medium or high
        prompt_caching:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Prompt Caching
          description: Cache shared prompt prefixes
        session_affinity_key:
          anyOf:
            - type: string
            - type: 'null'
          title: Session Affinity Key
          description: Field that keeps a trajectory on one worker
      additionalProperties: true
      type: object
      title: AutoAlignSampling
    AutoAlignTraining:
      properties:
        steps:
          anyOf:
            - type: integer
            - type: 'null'
          title: Steps
          description: Optimiser steps
        batch_size:
          anyOf:
            - type: integer
            - type: 'null'
          title: Batch Size
          description: Batch size
        gradient_accumulation:
          anyOf:
            - type: integer
            - type: 'null'
          title: Gradient Accumulation
          description: Gradient accumulation steps
        inner_optim_steps:
          anyOf:
            - type: integer
            - type: 'null'
          title: Inner Optim Steps
          description: Optimiser steps per batch
        mode:
          anyOf:
            - type: string
            - type: 'null'
          title: Mode
          description: sync or async
        rollout_training_overlap:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Rollout Training Overlap
          description: Overlap rollouts with optimisation
        epochs:
          anyOf:
            - type: integer
            - type: 'null'
          title: Epochs
          description: Epochs
        max_steps:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Steps
          description: Hard cap on steps
        seed:
          anyOf:
            - type: integer
            - type: 'null'
          title: Seed
          description: Seed
        early_stopping:
          anyOf:
            - additionalProperties: true
              type: object
            - type: 'null'
          title: Early Stopping
          description: 'Early stopping: enabled, metric, patience'
      additionalProperties: true
      type: object
      title: AutoAlignTraining
    AutoAlignAdapter:
      properties:
        mode:
          anyOf:
            - type: string
            - type: 'null'
          title: Mode
          description: lora or full
        rank:
          anyOf:
            - type: integer
            - type: 'null'
          title: Rank
          description: LoRA rank
        alpha:
          anyOf:
            - type: integer
            - type: 'null'
          title: Alpha
          description: LoRA alpha
        dropout:
          anyOf:
            - type: number
            - type: 'null'
          title: Dropout
          description: LoRA dropout
        train_attn:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Train Attn
          description: Adapt attention layers
        train_mlp:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Train Mlp
          description: Adapt MLP layers
        train_unembed:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Train Unembed
          description: Adapt the unembedding layer
        init_seed:
          anyOf:
            - type: integer
            - type: 'null'
          title: Init Seed
          description: Adapter initialisation seed
        quantization_bits:
          anyOf:
            - type: integer
            - type: 'null'
          title: Quantization Bits
          description: Base model quantisation, e.g. 4
      additionalProperties: true
      type: object
      title: AutoAlignAdapter
    AutoAlignCheckpointing:
      properties:
        save_every_steps:
          anyOf:
            - type: integer
            - type: 'null'
          title: Save Every Steps
          description: Save a checkpoint every N steps
        mode:
          anyOf:
            - type: string
            - type: 'null'
          title: Mode
          description: Checkpoint mode
        keep_last_n:
          anyOf:
            - type: integer
            - type: 'null'
          title: Keep Last N
          description: Checkpoints to keep
        resume_from:
          anyOf:
            - type: string
            - type: 'null'
          title: Resume From
          description: Checkpoint to resume from
        promote_final:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Promote Final
          description: Promote the final checkpoint to the aligned model
      additionalProperties: true
      type: object
      title: AutoAlignCheckpointing
    AutoAlignEvaluation:
      properties:
        eval_dataset:
          anyOf:
            - type: string
            - type: 'null'
          title: Eval Dataset
          description: Evaluation dataset
        eval_every_steps:
          anyOf:
            - type: integer
            - type: 'null'
          title: Eval Every Steps
          description: Evaluate every N steps
        eval_group_size:
          anyOf:
            - type: integer
            - type: 'null'
          title: Eval Group Size
          description: Samples per evaluation prompt
        eval_temperature:
          anyOf:
            - type: number
            - type: 'null'
          title: Eval Temperature
          description: Evaluation temperature
        metrics:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          title: Metrics
          description: Metrics to track
        benchmark_ref:
          anyOf:
            - type: string
            - type: 'null'
          title: Benchmark Ref
          description: Benchmark to evaluate against
        eval_batch_size:
          anyOf:
            - type: integer
            - type: 'null'
          title: Eval Batch Size
          description: Evaluation batch size
      additionalProperties: true
      type: object
      title: AutoAlignEvaluation
    AutoAlignRendering:
      properties:
        renderer_name:
          anyOf:
            - type: string
            - type: 'null'
          title: Renderer Name
          description: Renderer
        chat_template:
          anyOf:
            - type: string
            - type: 'null'
          title: Chat Template
          description: Chat template override
        add_special_tokens:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Add Special Tokens
          description: Add special tokens
        max_seq_len:
          anyOf:
            - type: integer
            - type: 'null'
          title: Max Seq Len
          description: Maximum rendered length
        truncation:
          anyOf:
            - type: string
            - type: 'null'
          title: Truncation
          description: left or right
        packing:
          anyOf:
            - type: boolean
            - type: 'null'
          title: Packing
          description: Pack sequences
        batch_unit:
          anyOf:
            - type: string
            - type: 'null'
          title: Batch Unit
          description: tokens or examples
      additionalProperties: true
      type: object
      title: AutoAlignRendering
    AutoAlignLogging:
      properties:
        wandb_project:
          anyOf:
            - type: string
            - type: 'null'
          title: Wandb Project
          description: Weights and Biases project
        log_metrics:
          anyOf:
            - items:
                type: string
              type: array
            - type: 'null'
          title: Log Metrics
          description: Metrics to log
        wandb_key:
          anyOf:
            - type: string
              format: password
              writeOnly: true
            - type: 'null'
          title: Wandb Key
          description: 'Weights and Biases API key. Write only: never returned or logged'
      additionalProperties: true
      type: object
      title: AutoAlignLogging
  securitySchemes:
    HTTPBearer:
      type: http
      scheme: bearer

````

This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.