Skip to main content
XaasIO AI Factory Platform

Private AI Factory and Token Factory, Made Simple

Enterprise-grade private AI. Validated open-source stack. On your infrastructure. With XaasIO lifecycle support and production SRE.

Train, fine-tune and evaluate models, and run HPC workloads, with XaasIO AI Factory + HPC. Publish approved models as governed, token-metered APIs with XaasIO AI Token Factory. Both run on infrastructure you control: private, sovereign or fully air-gapped.


Open source by default: Kubernetes, Slurm, Kubeflow, vLLM and KServe. Upstream-aligned. Customer-operated, co-managed or fully managed by XaasIO. Your models, prompts and data stay inside your compliance boundary, not a hyperscaler’s.

Architecture

One platform, two services, six layers

From the applications and teams that consume AI down to the accelerators: how one XaasIO platform builds models on one side and serves them as governed services on the other.
ONE PLATFORM, TWO SERVICES

One AI platform.
Two AI services.


Teams do not consume AI the same way. Data scientists, ML engineers and researchers need GPU clusters, notebooks, pipelines and schedulers to build and adapt models and to run simulations. Application teams and customers need a model API that works like the public services they already use, with keys, quotas and a bill they can understand.
Build and compute

XaasIO AI Factory + HPC

  • Train and fine-tune LLMs and SLMs on your own data with distributed GPU jobs

  • Notebooks, pipelines, experiment tracking, evaluation gates and a model registry

  • Classical HPC: Slurm scheduling, MPI, scientific software and parallel file systems

  • Develop RAG applications and agents on governed enterprise
    context

  • Bare-metal GPU and CPU clusters with InfiniBand or RoCE fabric


XaasIO delivers both from one platform. XaasIO AI Factory + HPC produces validated models and runs HPC workloads. XaasIO AI Token Factory publishes approved models as governed, token-metered services. Either can run on its own; together they cover model creation through commercial delivery.
Serve and meter

XaasIO AI Token Factory

  • A private model catalog behind OpenAI-compatible APIs, with virtual keys per team or customer

  • Shared, dedicated, reserved-throughput and batch endpoints

  • Token metering, quotas, budgets, prepaid and postpaid plans, showback and chargeback

  • Local-first routing, with cloud burst to approved public providers only by policy

  • Guardrails, PII controls and an immutable audit archive; air-gapped operation

XaasIO AI Token Factory gives your applications the experience of a public token service, from the model list to the invoice, on private, sovereign or air-gapped infrastructure.
WHICH SERVICE, FOR WHICH WORKLOAD

Build in the AI Factory. Serve from the Token Factory.

Workload Service Why
Train or fine-tune an LLM or SLM on your data
AI Factory + HPC
Distributed GPU jobs, pipelines, experiment tracking and a registry with approval gates.
Run CFD, FEA, genomics, weather or EDA simulation
AI Factory + HPC
Slurm scheduling, MPI, the scientific software stack and parallel file systems.
Develop RAG applications and agents
AI Factory + HPC
Notebooks, LangGraph, governed context from XaasIO AI Lake, evaluations and trace replay.
Serve a chat, completion or embedding API to applications
AI Token Factory
OpenAI-compatible endpoints, scoped keys, quotas and token metering.
Give a regulated application a private model endpoint
AI Token Factory
Dedicated replicas, private network path, customer-managed keys and no-log mode.
Run nightly batch inference over documents
AI Token Factory
A batch service on scheduled or off-peak capacity, metered by records or tokens.
Sell AI APIs to customers under your own brand
AI Token Factory
Plans, prepaid and postpaid accounts, invoicing and reseller settlement through the XaasIO Hyperscaler Platform.
Allocate AI budgets to departments and charge back
AI Token Factory
Organization-wide token pools, department budgets, showback and chargeback.
Train and serve models on a disconnected site
AI Factory + HPC AI Token Factory
The air-gapped build: local registries, model store, identity, metering and signed offline updates for both services.
PLATFORM COMPONENTS

Everything needed to deliver AI as a cloud service

Model Factory

Notebooks, pipelines, distributed training and fine-tuning, experiment tracking, evaluation gates, model cards and a registry with approval workflow before anything is served.

Kubeflow · MLflow · Airflow and Argo Workflows · JupyterHub · Ray

HPC Scheduler

Job queues, partitions, fair-share and reservations for tightly coupled simulation and large batch training, with usage accounting per project for showback.

Slurm · OpenPBS and Flux options · Kueue and Volcano for Kubernetes batch · MPI, NCCL, UCX

Private Inference Cloud

GPU pools, model serving with continuous batching, adapter serving, canary releases, autoscaling and GPU telemetry, on NVIDIA, AMD or Intel accelerators where validated.

vLLM · KServe · Ray Serve · NVIDIA GPU Operator · DCGM

AI Gateway and Token Metering

OpenAI-compatible APIs, virtual keys, model aliases, routing and fallback, rate limits, and a durable usage pipeline that turns every request into an attributable, rated usage record.

AI gateway · OpenTelemetry · Event-driven metering · ClickHouse analytics · XaasIO BSS integration

Identity, Policy and Guardrails

Single sign-on, organization and project boundaries, role and attribute policies, tool allow-lists for agents, PII and secret detection, content guardrails and an immutable audit archive.

Keycloak and OIDC · Open Policy Agent · OpenFGA · NeMo Guardrails · Presidio · Ceph S3 archive

Management and Observability

GPU utilization, queue depth, time to first token, inter-token latency, error rates, cost per unit and margin by model, tenant and region, alongside prompt and trace analytics.

Prometheus · Grafana · OpenTelemetry · Langfuse · XaasIO MLT

ACCESS METHODS

Consume AI the way your teams already work

Applications call an API. Data scientists open a notebook. Researchers submit a job. Platform teams automate. The same identity, policy and metering apply to all of them.
Console
Console
OpenAI-compatible API
OpenAI-compatible API
SDKs and playground
SDKs and playground
JupyterHub and VS Code
JupyterHub and VS Code
Slurm CLI
Slurm CLI
Kubeflow and Kubernetes
Kubeflow and Kubernetes
GitOps and infrastructure as code
GitOps and infrastructure as code
HOW IT WORKS

One workflow. From dataset to metered API.

01

Build

Teams train or fine-tune a model in the AI Factory, or bring an approved external model. Datasets, weights and adapters live in your object storage.
02

Validate

Evaluation, safety, license, provenance and vulnerability checks gate promotion in the registry. Nothing is served without an approval record.
03

Publish

The Token Factory creates a shared, dedicated, reserved or batch endpoint, a catalog entry and, where billing applies, a SKU with plans, quotas and regions.
04

Serve and meter

Applications call the API with scoped keys. Every request is authorized, routed, served and recorded as a usage event for dashboards, showback or invoices.
Users request a model, a workspace or an API key. XaasIO coordinates identity, policy, scheduling, serving, metering and lifecycle operations.
SERVICE CLASSES

Service profiles, not cluster internals

Customer-friendly service profiles and plans, not GPU part numbers.
Throughput, latency, tokens per second and training time depend on the validated accelerators, interconnect, storage, model, precision and batching selected for each deployment. No specific tokens-per-second, time-to-first-token or training-duration figure is guaranteed on this page. Request a validated architecture brief for deployment-specific numbers.

XaasIO AI Factory + HPC profiles

GPU Training Cluster

Multi-node GPU capacity for pre-training, continued pre-training, fine-tuning and large evaluation runs, scheduled through Slurm or Kubernetes.

HPC Simulation Cluster

CPU and GPU partitions for MPI simulation, CFD, FEA, genomics, weather and EDA workloads, with the scientific software stack and parallel scratch.

AI Workbench

Notebooks, VS Code, Ray and Kubeflow workspaces for data preparation, adapter fine-tuning, RAG and agent development, and evaluation.

XaasIO AI Token Factory plans

Shared Endpoint

Pooled, multi-tenant model APIs for developers, departments and elastic application traffic, with fair queueing and per-key quotas.

Dedicated Endpoint

Reserved replicas or accelerators for one tenant, for sensitive, regulated, customized or strict-latency workloads, with a private network path.

Reserved Throughput

A contracted capacity envelope with admission control and protected priority for production applications with predictable demand.

Also in the catalog: batch inference, embeddings and reranking, speech and image services, fine-tuned and bring-your-own-model endpoints, agent token plans, organization-wide token pools, a hybrid frontier-model gateway and the air-gapped model endpoint.

ADVANCED CAPABILITIES

Train, serve and govern AI without proprietary lock-in


Distributed Training and Scheduling

Multi-node, multi-GPU training with topology-aware placement, checkpointing, fair-share queues and reservations across Slurm and Kubernetes.

Scientific Software Stack

Reproducible environments from OpenHPC, Spack and E4S, with Apptainer containers and module environments for researchers and engineers.

Local-First Hybrid Routing

Requests stay on private models by default. Cloud burst to an approved public provider happens only where policy, data classification and budget allow, and is reconciled.

Model Registry and Promotion Gates

License, provenance, evaluation and safety checks are recorded on every model version before it receives an endpoint, a SKU or production traffic.

Token Plans, Budgets and Metering

Free tiers, prepaid credits, postpaid, bundles, committed use and department budgets, backed by a deduplicated, auditable usage pipeline that feeds showback or invoicing.

Guardrails, PII and Audit Archive

Prompt and response policies, PII and secret detection, tool allow-lists for agents, retention rules and an immutable archive of prompts, decisions and usage for compliance evidence.
WORKLOAD USE CASES

One open AI platform for research, enterprise and service providers


Sovereign AI and National Research Cloud

Private model training, classical HPC and a governed model API for government, universities and national labs, with controlled access and full-stack governance

Life Sciences, Manufacturing and Semiconductor HPC

Genomics, protein modeling, CFD, FEA, digital twins and EDA on Slurm-scheduled CPU and GPU clusters with parallel file systems.

Enterprise AI Utility

One governed access plane for every application team, with department budgets, private and public routing, reduced API-key sprawl and cost attribution.

Service Provider AI Token Service

Shared model APIs, dedicated endpoints, reserved throughput and white-label plans above commodity GPU hosting, with margin visibility by model, tenant and region.

RAG Applications and Agents on Governed Knowledge

Retrieval, embeddings and agent workflows over enterprise context, with tool governance, human approvals, trace replay and evaluations.

Air-Gapped Defense and Critical Infrastructure

Locally stored models, offline registries, signed update media, local identity and audit, with no dependency on a public API.
PLANNING TOOL

Estimate monthly token consumption


Token usage is what a model API meters and what a token plan prices. Enter typical request volumes to see how many input and output tokens an application consumes in a month, and what that costs at a rate you choose. The same numbers size a shared endpoint or a reserved-throughput plan.
Token consumption estimator

Input tokens per month 720 M
Output tokens per month 240 M
Estimated monthly charge $288

Planning estimate only. Token counts depend on the model’s tokenizer, prompt templates, retrieved context and caching. The rates are placeholders you set; XaasIO does not publish list prices on this page. A Token Factory sizing engagement produces validated throughput and capacity numbers for your models.

Deployment models

Choose the architecture that fits the workload

Enterprise Private AI Factory

GPU and CPU clusters with project isolation, AI and simulation workloads, secure engineering workspaces, enterprise identity and showback.

Sovereign AI + HPC Factory

Private AI, classical HPC, a data lake and secure model training inside one boundary, with controlled access and offline update media for air-gapped sites.

Research HPC
Cloud

Multi-tenant Slurm with project quotas, a scientific software catalog, a researcher portal and secure browser-based remote access.

Service Provider AI Token Service

A self-service model catalog, customer isolation, metering and billing, GPU and endpoint offerings, and a Token Delivery Network across owned and partner sites.
AIR-GAPPED BUILD

Air-gapped AI Factory and Token Factory

For defense, critical infrastructure, regulated and sovereign sites, the complete platform runs inside the isolated boundary: identity, registries, model store, schedulers, inference, policy, telemetry, metering and documentation. There is no public model provider, no internet identity, no hosted billing service and no online vulnerability feed in the request path. Updates arrive only as signed, scanned and approved offline bundles.
Controlled external staging
01
Acquire packages, images, charts, model weights, tokenizers and security advisories
02
Scan and review vulnerability scan, SBOM, license and provenance review
03
Approve and sign a versioned offline bundle with rollback media
Controlled transfer media Removable media procedure, two-person approval, scanned on both sides
Air-gapped sovereign boundary
01
Verify and import signature check, second scan, import to the local registry and model store
02
Platform XaasIO Hyperscaler console, AI Factory + HPC and AI Token Factory control planes
03
Run local Kubernetes, Slurm, model serving and notebooks on your accelerators
04
Account local metering, showback or chargeback, audit archive and backups

Alongside: local identity and single sign-on, PKI and certificate chains, secrets, DNS and NTP, local documentation

AI FACTORY + HPC, AIR-GAPPED

Train and simulate with no outbound dependency

  • Datasets, checkpoints, weights and adapters on local Ceph S3; scratch on local Lustre, BeeGFS or NVMe

  • Mirrored operating-system, Python and scientific software repositories: OpenHPC, Spack and E4S mirrors, an Apptainer image registry and module environments

  • Local Git, pipeline and registry services for Kubeflow, MLflow and Argo Workflows; local JupyterHub and VS Code workspaces

  • GPU drivers, CUDA or ROCm and firmware references carried in the bundle and installed by bare-metal automation

  • Slurm accounting and Kubernetes usage recorded locally for project showback

AI Token Factory, air-gapped

Serve locally stored models as governed APIs

  • Air-gapped model endpoints and private enterprise endpoints for locally stored models; offline shared models and internal token pools for departments

  • The model catalog, tokenizers, license notices and evaluation records live in the local registry; nothing is fetched at request time

  • No public-provider routing: the hybrid gateway is absent, so requests cannot fall back to an external model

  • Local metering and internal chargeback with no dependency on an external billing API; usage evidence signed for audit

  • Local identity, virtual keys, policy, guardrails and PII controls; prompts and responses retained only under the local policy

Operating rules for the air-gapped build: restricted administrator roles, an offline root of trust, signed images and models, a documented removable-media procedure, local time synchronization, no public fallback and tested backup and recovery. Update lead time, signature verification, audit completeness, backup and restore success and catalog currency are the metrics XaasIO reports for the site.

BUILD, SERVE AND BILL

Model production, model consumption and billing are separate capabilities


Each capability can be bought and operated on its own. Together they form a governed hand-off: a model is evaluated and released in the AI Factory, published and metered in the Token Factory, and settled in the XaasIO Hyperscaler Platform.

XaasIO AI Factory + HPC

Produces validated models, adapters, agents and simulation results. Runs without the Token Factory when no external model service is needed.

XaasIO AI Token Factory

Publishes approved models as governed APIs and meters them. Serves models from the AI Factory or from an external, approved pipeline.

XaasIO Hyperscaler console

Included with both services: catalog, projects, identity, keys, quotas, usage dashboards and showback for AI services.

XaasIO Hyperscaler Platform

Required for plans, invoicing, payments, multi-service catalogs, tenant billing across XaasIO modules and Token Delivery Network settlement.

OPERATING MODEL

Private AI without the integration burden


XaasIO designs, deploys, validates and supports the AI Factory Platform as one production platform rather than an unsupported collection of operators, charts, drivers and scripts.

Validated Architecture

Accelerators, drivers, interconnect, storage, schedulers, runtimes, gateway, identity and metering validated as one architecture.

Controlled Lifecycle

Secure and Sovereign

Customer-controlled hardware, identity, model registries, software repositories, encryption, key management, monitoring, logging and update procedures.

Flexible Operating Model

Customer-operated, co-managed, fully managed, or a dedicated XaasIO AI SRE Pod.

AUTONOMOUS OPERATIONS

Do more with less: autonomous operations for AI


Add XaasIO MLT, SRE Pod and AI-SRE to move from reactive GPU administration toward policy-driven, assisted and increasingly automated AI operations.

GPU health and incident triage

Endpoint SLO monitoring

Capacity and token forecasting

Model rollout verification

Cost and margin analytics

Queue and scheduler tuning

Patch and upgrade readiness

Pre and post-change validation

Backup and DR verification

Human-approved automation

Advanced AI Support Packs

Slurm, OpenPBS and Flux

Scheduler design, partitions, accounting and Slurm on Kubernetes

AMD ROCm and Intel Accelerators

Validation of runtimes and models beyond NVIDIA

NVIDIA NIM and Triton

Optional vendor runtimes where licenses and workloads justify them

Hybrid Frontier-Model Gateway

Policy-routed access to approved public providers with reconciliation

Model and Data Migration

From public AI platforms, SaaS gateways and standalone MLOps stacks

Multi-Region and Token Delivery Network

Federated inference sites, local metering and partner settlement

AI FACTORY VS TOKEN FACTORY

XaasIO AI Factory + HPC or XaasIO AI Token Factory? Both. On one platform.


Dimension AI Factory + HPC AI Token Factory
Primary purpose Build, adapt, evaluate and operationalize models, agents and HPC workloads Govern, distribute, meter and commercialize production inference
Main output A validated model, adapter, agent or simulation result A governed API endpoint and an attributable usage record
Primary users Data scientists, ML engineers, researchers, data engineers Application teams, service-product teams, platform operators, FinOps
Main input Data, documents, base models, code and job definitions An approved model, a service policy and a plan
Model training and fine-tuning Core capability, including distributed training No; publishes and serves approved models and adapters
Inference serving Deployment as part of the lifecycle Core production capability with shared, dedicated, reserved and batch endpoints
API gateway and keys Optional development access Core: virtual keys, routing, policy, rate limits and service tiers
Token metering and billing Technical experiment and endpoint metrics Core, financially reconcilable, settled through the XaasIO Hyperscaler Platform
Scheduling Slurm, Kueue, Volcano and Ray for training and HPC Inference pools, endpoint capacity and admission control
Key performance indicators Experiment velocity, evaluation quality, reproducibility, job throughput Time to first token, inter-token latency, error rate, utilization, cost and margin per unit
Typical buyer Chief data and AI office, ML platform team, research organization Service provider product owner, enterprise platform leader, FinOps
Can operate independently Yes; can deploy models without external token monetization Yes; can serve externally developed models
KEY TERMS

Glossary

AI token

A model-specific unit of input or output produced by a tokenizer or generated during inference. Counts are not directly comparable across tokenizers, and a token here has nothing to do with cryptocurrency.

Time to first token (TTFT)

Elapsed time from an accepted request to the first streamed output token. The main latency figure for interactive applications.

Inter-token latency (ITL)

The time between successive streamed output tokens, usually reported as an average or a percentile.

Reserved throughput

A contracted capacity envelope, such as tokens per minute, protected by admission control on a shared or partitioned fleet.

LoRA adapter

A small set of trained weights that adapts a base model to a domain or task and can be served over shared base weights.

RAG

Retrieval-augmented generation: the model answers from documents retrieved at request time, so answers cite governed enterprise context instead of training data alone.

MCP

Model Context Protocol: an open standard through which agents call tools and data sources. XaasIO governs which MCP tools an agent may use.

Slurm

The open-source workload manager for classical HPC: job queues, partitions, fair-share scheduling and accounting for CPU and GPU clusters.

Air-gapped deployment

Operation with no path to public networks: local model registries, identity, telemetry, metering and signed update media inside the boundary.

Frequently asked questions

What is XaasIO AI Factory Platform?
An open-source-based platform for private AI on infrastructure you control. It has two services: XaasIO AI Factory + HPC, for training, fine-tuning and evaluating models and running HPC workloads, and XaasIO AI Token Factory, for publishing approved models as governed, token-metered APIs. XaasIO designs, validates, supports and can operate it as one production platform.
What is the difference between AI Factory + HPC and AI Token Factory?
The AI Factory produces models and results: data preparation, training, fine-tuning, evaluation, registry, notebooks, pipelines and HPC jobs. The Token Factory consumes approved models: it publishes them as APIs, issues keys, routes requests, enforces quotas and budgets, meters tokens and makes usage accountable. Build in one, serve from the other.
Do I need both services?
No. Each operates independently. A research organization can run the AI Factory + HPC without external model services. A service provider can run the Token Factory to serve externally developed open-weight or licensed models. Together, model promotion becomes a governed hand-off from evaluation to endpoint.
Is the API compatible with OpenAI?
Yes. XaasIO AI Token Factory exposes OpenAI-compatible chat, completion, embedding and related APIs, so most SDKs and tools work by changing the base URL and key. Compatibility gaps for provider-specific features are documented per release.
Can it replace a public token service such as OpenAI, Nebius Token Factory or SAIL?
For open-weight, enterprise-licensed and fine-tuned models, yes: applications get the same catalog, API and pay-per-token experience on private, sovereign or air-gapped infrastructure. Frontier models that are only available from a public provider can be reached through the hybrid gateway, where policy allows, with usage reconciled against the provider
Which models can it serve?
Open-weight models, enterprise-licensed models, domain-specific and fine-tuned models, and your own adapters, in text, embedding, reranking, speech, image and multimodal categories. Each model carries license, provenance and evaluation metadata, and the runtime is chosen per model and accelerator.
Which accelerators are supported?
NVIDIA GPUs are the validated default. AMD and Intel accelerators, and CPU-only inference for embeddings and small models, are supported where the chosen runtime and model are validated for the deployment.
How is token usage metered and billed?
Every request produces a usage event with input, output and cached token counts, or task-specific units for speech, image and batch services, attributed to organization, project and key. Events are deduplicated, aggregated and rated against the subscribed plan. Showback works with the included console; invoicing, payments and reseller settlement run through the XaasIO Hyperscaler Platform.
Are our prompts used to train models?
No. Production prompts and responses are not used as training data by default. Retention of prompts and responses follows the policy you set per endpoint, including a no-log mode for private endpoints.
Can it run in an air-gapped environment?
Yes. The air-gapped build runs both services entirely inside the boundary: local identity and PKI, image, chart and model registries, object storage, schedulers, inference, policy, telemetry, metering and documentation. Locally stored models are served with no public API dependency and no public fallback, and updates arrive only as signed, scanned and approved offline bundles.
Slurm or Kubernetes?
Both. Slurm remains the stronger scheduler for tightly coupled MPI simulation and large batch training. Kubernetes runs AI pipelines, notebooks, services and inference. XaasIO maps projects, quotas and accounting across the two so that users and finance see one platform.
Can an existing HPC cluster or GPU cluster be brought in?
Yes, as a scoped assessment. Existing Slurm clusters and Kubernetes GPU clusters can be integrated where the operating system, drivers, network and storage meet the validated architecture. Bare-metal automation handles new and re-provisioned nodes.
How does XaasIO compare to Rafay AI Factory and AI Token Factory?
Both take the same approach: turn GPU infrastructure into self-service, multi-tenant, metered AI services. XaasIO delivers it as an open-source-based platform that integrates with the wider XaasIO portfolio, including classical HPC with Slurm, XaasIO AI Lake for governed context, and the XaasIO Hyperscaler Platform for billing, with customer-operated, co-managed or fully managed operation.
How does XaasIO compare to NVIDIA AI Enterprise and Run:ai?
NVIDIA AI Enterprise and Run:ai are software layers licensed from NVIDIA for NVIDIA GPUs. XaasIO AI Factory Platform is built on upstream open-source components, runs on NVIDIA and, where validated, AMD and Intel accelerators, and adds tenancy, token metering and billing. NVIDIA runtimes such as NIM and Triton can be added through a support pack where a customer licenses them.
Are other XaasIO modules required?
No. The AI Factory Platform operates independently and includes the XaasIO Hyperscaler console for its own services. XaasIO AI Lake, XaasIO MLT, XaasIO Compute, XaasIO Block and Object Storage and the XaasIO Hyperscaler Platform are optional integrations; the Hyperscaler Platform is required for invoicing and multi-service tenant billing.
Is an AI token a cryptocurrency?
No. On this page an AI token is a unit of model input or output consumption, created by a model tokenizer and inference runtime. It is unrelated to cryptocurrency, blockchain assets or token issuance.
Is XaasIO AI Factory Platform SOC 2 or HIPAA compliant?
The platform is architected with controls designed to align with SOC 2, HIPAA and CMMC control families: identity, tenant isolation, encryption, audit archive, guardrails and software supply chain evidence. Confirm the current formal certification status directly with XaasIO.

OpenAI, Anthropic, NVIDIA, AMD, Intel, Kubernetes, Slurm, Kubeflow, MLflow, Ray, vLLM, KServe, Rafay, Nebius, SAIL and other product names are trademarks of their respective owners and are named for identification and comparison only. XaasIO is not affiliated with, endorsed by or sponsored by these organizations. An AI token on this page is a unit of model input or output; it is unrelated to cryptocurrency or blockchain assets.

Build what the world depends on.

Planning a cloud for government, financial services, pharmaceuticals, healthcare, energy, transport, industry, service-provider operations, or private AI? Work with XaasIO to define the architecture, security boundary, and operating model your environment requires.