Produces validated models, adapters, agents and simulation results. Runs without the Token Factory when no external model service is needed.
Private AI Factory and Token Factory, Made Simple
Enterprise-grade private AI. Validated open-source stack. On your infrastructure. With XaasIO lifecycle support and production SRE.
Train, fine-tune and evaluate models, and run HPC workloads, with XaasIO AI Factory + HPC. Publish approved models as governed, token-metered APIs with XaasIO AI Token Factory. Both run on infrastructure you control: private, sovereign or fully air-gapped.
Open source by default: Kubernetes, Slurm, Kubeflow, vLLM and KServe. Upstream-aligned. Customer-operated, co-managed or fully managed by XaasIO. Your models, prompts and data stay inside your compliance boundary, not a hyperscaler’s.
Architecture
One platform, two services, six layers
One AI platform.
Two AI services.
XaasIO AI Factory + HPC
- Train and fine-tune LLMs and SLMs on your own data with distributed GPU jobs
- Notebooks, pipelines, experiment tracking, evaluation gates and a model registry
- Classical HPC: Slurm scheduling, MPI, scientific software and parallel file systems
- Develop RAG applications and agents on governed enterprise
context
- Bare-metal GPU and CPU clusters with InfiniBand or RoCE fabric
XaasIO AI Token Factory
- A private model catalog behind OpenAI-compatible APIs, with virtual keys per team or customer
- Shared, dedicated, reserved-throughput and batch endpoints
- Token metering, quotas, budgets, prepaid and postpaid plans, showback and chargeback
- Local-first routing, with cloud burst to approved public providers only by policy
- Guardrails, PII controls and an immutable audit archive; air-gapped operation
Build in the AI Factory. Serve from the Token Factory.
| Workload | Service | Why |
|---|---|---|
| Train or fine-tune an LLM or SLM on your data | AI Factory + HPC |
Distributed GPU jobs, pipelines, experiment tracking and a registry with approval gates. |
| Run CFD, FEA, genomics, weather or EDA simulation | AI Factory + HPC |
Slurm scheduling, MPI, the scientific software stack and parallel file systems. |
| Develop RAG applications and agents | AI Factory + HPC |
Notebooks, LangGraph, governed context from XaasIO AI Lake, evaluations and trace replay. |
| Serve a chat, completion or embedding API to applications | AI Token Factory |
OpenAI-compatible endpoints, scoped keys, quotas and token metering. |
| Give a regulated application a private model endpoint | AI Token Factory |
Dedicated replicas, private network path, customer-managed keys and no-log mode. |
| Run nightly batch inference over documents | AI Token Factory |
A batch service on scheduled or off-peak capacity, metered by records or tokens. |
| Sell AI APIs to customers under your own brand | AI Token Factory |
Plans, prepaid and postpaid accounts, invoicing and reseller settlement through the XaasIO Hyperscaler Platform. |
| Allocate AI budgets to departments and charge back | AI Token Factory |
Organization-wide token pools, department budgets, showback and chargeback. |
| Train and serve models on a disconnected site |
AI Factory + HPC
AI Token Factory
|
The air-gapped build: local registries, model store, identity, metering and signed offline updates for both services. |
Everything needed to deliver AI as a cloud service
Model Factory
Notebooks, pipelines, distributed training and fine-tuning, experiment tracking, evaluation gates, model cards and a registry with approval workflow before anything is served.
Kubeflow · MLflow · Airflow and Argo Workflows · JupyterHub · Ray
HPC Scheduler
Job queues, partitions, fair-share and reservations for tightly coupled simulation and large batch training, with usage accounting per project for showback.
Slurm · OpenPBS and Flux options · Kueue and Volcano for Kubernetes batch · MPI, NCCL, UCX
Private Inference Cloud
GPU pools, model serving with continuous batching, adapter serving, canary releases, autoscaling and GPU telemetry, on NVIDIA, AMD or Intel accelerators where validated.
vLLM · KServe · Ray Serve · NVIDIA GPU Operator · DCGM
AI Gateway and Token Metering
OpenAI-compatible APIs, virtual keys, model aliases, routing and fallback, rate limits, and a durable usage pipeline that turns every request into an attributable, rated usage record.
AI gateway · OpenTelemetry · Event-driven metering · ClickHouse analytics · XaasIO BSS integration
Identity, Policy and Guardrails
Single sign-on, organization and project boundaries, role and attribute policies, tool allow-lists for agents, PII and secret detection, content guardrails and an immutable audit archive.
Keycloak and OIDC · Open Policy Agent · OpenFGA · NeMo Guardrails · Presidio · Ceph S3 archive
Management and Observability
GPU utilization, queue depth, time to first token, inter-token latency, error rates, cost per unit and margin by model, tenant and region, alongside prompt and trace analytics.
Prometheus · Grafana · OpenTelemetry · Langfuse · XaasIO MLT
Consume AI the way your teams already work
One workflow. From dataset to metered API.
Build
Validate
Publish
Serve and meter
Service profiles, not cluster internals
XaasIO AI Factory + HPC profiles
GPU Training Cluster
HPC Simulation Cluster
AI Workbench
XaasIO AI Token Factory plans
Shared Endpoint
Dedicated Endpoint
Reserved Throughput
Train, serve and govern AI without proprietary lock-in
Distributed Training and Scheduling
Scientific Software Stack
Local-First Hybrid Routing
Model Registry and Promotion Gates
Token Plans, Budgets and Metering
Guardrails, PII and Audit Archive
One open AI platform for research, enterprise and service providers
Sovereign AI and National Research Cloud
Life Sciences, Manufacturing and Semiconductor HPC
Enterprise AI Utility
Service Provider AI Token Service
RAG Applications and Agents on Governed Knowledge
Air-Gapped Defense and Critical Infrastructure
Estimate monthly token consumption
Planning estimate only. Token counts depend on the model’s tokenizer, prompt templates, retrieved context and caching. The rates are placeholders you set; XaasIO does not publish list prices on this page. A Token Factory sizing engagement produces validated throughput and capacity numbers for your models.
Choose the architecture that fits the workload
Enterprise Private AI Factory
Sovereign AI + HPC Factory
Research HPC
Cloud
Service Provider AI Token Service
Air-gapped AI Factory and Token Factory
Alongside: local identity and single sign-on, PKI and certificate chains, secrets, DNS and NTP, local documentation
Train and simulate with no outbound dependency
- Datasets, checkpoints, weights and adapters on local Ceph S3; scratch on local Lustre, BeeGFS or NVMe
- Mirrored operating-system, Python and scientific software repositories: OpenHPC, Spack and E4S mirrors, an Apptainer image registry and module environments
- Local Git, pipeline and registry services for Kubeflow, MLflow and Argo Workflows; local JupyterHub and VS Code workspaces
- GPU drivers, CUDA or ROCm and firmware references carried in the bundle and installed by bare-metal automation
- Slurm accounting and Kubernetes usage recorded locally for project showback
Serve locally stored models as governed APIs
- Air-gapped model endpoints and private enterprise endpoints for locally stored models; offline shared models and internal token pools for departments
- The model catalog, tokenizers, license notices and evaluation records live in the local registry; nothing is fetched at request time
- No public-provider routing: the hybrid gateway is absent, so requests cannot fall back to an external model
- Local metering and internal chargeback with no dependency on an external billing API; usage evidence signed for audit
- Local identity, virtual keys, policy, guardrails and PII controls; prompts and responses retained only under the local policy
Operating rules for the air-gapped build: restricted administrator roles, an offline root of trust, signed images and models, a documented removable-media procedure, local time synchronization, no public fallback and tested backup and recovery. Update lead time, signature verification, audit completeness, backup and restore success and catalog currency are the metrics XaasIO reports for the site.
Model production, model consumption and billing are separate capabilities
Each capability can be bought and operated on its own. Together they form a governed hand-off: a model is evaluated and released in the AI Factory, published and metered in the Token Factory, and settled in the XaasIO Hyperscaler Platform.
XaasIO AI Factory + HPC
XaasIO AI Token Factory
Publishes approved models as governed APIs and meters them. Serves models from the AI Factory or from an external, approved pipeline.
XaasIO Hyperscaler console
Included with both services: catalog, projects, identity, keys, quotas, usage dashboards and showback for AI services.
XaasIO Hyperscaler Platform
Required for plans, invoicing, payments, multi-service catalogs, tenant billing across XaasIO modules and Token Delivery Network settlement.
OPERATING MODEL
Private AI without the integration burden
XaasIO designs, deploys, validates and supports the AI Factory Platform as one production platform rather than an unsupported collection of operators, charts, drivers and scripts.
Validated Architecture
Controlled Lifecycle
Documented release, patching, upgrade, compatibility, rollback and maintenance procedures for the platform and every pinned open-source component.
Secure and Sovereign
Flexible Operating Model
AUTONOMOUS OPERATIONS
Do more with less: autonomous operations for AI
GPU health and incident triage
Endpoint SLO monitoring
Capacity and token forecasting
Model rollout verification
Cost and margin analytics
Queue and scheduler tuning
Patch and upgrade readiness
Pre and post-change validation
Backup and DR verification
Human-approved automation
Advanced AI Support Packs
Slurm, OpenPBS and Flux
AMD ROCm and Intel Accelerators
NVIDIA NIM and Triton
Hybrid Frontier-Model Gateway
Model and Data Migration
Multi-Region and Token Delivery Network
AI FACTORY VS TOKEN FACTORY
XaasIO AI Factory + HPC or XaasIO AI Token Factory? Both. On one platform.
| Dimension | AI Factory + HPC | AI Token Factory |
|---|---|---|
| Primary purpose | Build, adapt, evaluate and operationalize models, agents and HPC workloads | Govern, distribute, meter and commercialize production inference |
| Main output | A validated model, adapter, agent or simulation result | A governed API endpoint and an attributable usage record |
| Primary users | Data scientists, ML engineers, researchers, data engineers | Application teams, service-product teams, platform operators, FinOps |
| Main input | Data, documents, base models, code and job definitions | An approved model, a service policy and a plan |
| Model training and fine-tuning | Core capability, including distributed training | No; publishes and serves approved models and adapters |
| Inference serving | Deployment as part of the lifecycle | Core production capability with shared, dedicated, reserved and batch endpoints |
| API gateway and keys | Optional development access | Core: virtual keys, routing, policy, rate limits and service tiers |
| Token metering and billing | Technical experiment and endpoint metrics | Core, financially reconcilable, settled through the XaasIO Hyperscaler Platform |
| Scheduling | Slurm, Kueue, Volcano and Ray for training and HPC | Inference pools, endpoint capacity and admission control |
| Key performance indicators | Experiment velocity, evaluation quality, reproducibility, job throughput | Time to first token, inter-token latency, error rate, utilization, cost and margin per unit |
| Typical buyer | Chief data and AI office, ML platform team, research organization | Service provider product owner, enterprise platform leader, FinOps |
| Can operate independently | Yes; can deploy models without external token monetization | Yes; can serve externally developed models |