Data locality, customer-controlled keys, tenant isolation, local repositories, identity and controlled update procedures inside the approved security boundary.
Open Data Platform for AI, Made Simple
co-managed, or fully managed by XaasIO.
Everything Needed to Deliver Data and Knowledge as a Service
Main components
Open Lakehouse | Store raw and curated data in S3-compatible object storage as open Iceberg tables with a catalog, lineage, versioning and time travel. Query it with SQL and Spark without copying it into a proprietary warehouse. Technical reference: Ceph S3 · Apache Iceberg · Parquet · Apache Polaris · OpenMetadata · OpenLineage |
Ingestion, Streaming and Processing | Load batch files, change data capture and event streams; enrich them in flight; transform at scale; and query the lakehouse and external systems through one SQL layer. Technical reference: Kafka · Debezium · Flink · Spark · Trino · Argo Workflows |
Search, Vector Database and Knowledge Graph | Turn documents, tables and events into keyword indexes, vector embeddings and a graph of entities and relationships, so applications retrieve what is similar and understand how it is connected. Technical reference: OpenSearch · Qdrant · pgvector · Apache AGE · JanusGraph |
Context Lake and Agent Lake | Assemble identity-aware, fresh, ranked and cited context for models and agents, and manage agents, tools, prompts and memory as governed services with approvals and rollback. Technical reference: Context API · Agent Registry · Temporal Workflows · LangGraph · Model Gateway |
Identity and Governance | Separate tenants, projects and workspaces with roles, attributes, classification, retention and policy. Integrate enterprise single sign-on and keep an audit trail across data, retrieval and agent actions. Technical reference: XaasIO IAM · SSO · RBAC and ABAC · Open Policy Agent · Audit Archive |
Use the same platform through the console, SQL, notebooks, APIs or infrastructure as code.
An Open Data Platform, Plus the Layers AI Needs
Governed facts and history in open formats, on storage you own.
S3-compatible object storage is the sovereign data plane. Iceberg tables give it schemas, transactions, hidden partitioning and time travel.
Fresh data in, federated SQL out, without a proprietary engine in between.
Kafka transports events, Flink enriches them, Spark transforms at scale and Trino queries Iceberg and external systems without copying data.
Semantic and hybrid retrieval with tenant scope, provenance and re-ranking built in.
Every vector carries its source, version, tenant scope and embedding model, so results can be filtered, cited, re-embedded and deleted.
Entities, relationships, dependencies and rules, for reasoning that similarity search cannot do.
Vectors answer what is similar. The graph answers how things are related, which dependencies, rules, time or location connect them.
The minimum necessary evidence for a user, task, model and policy, delivered with provenance.
The delivery API returns evidence, the policy decision, provenance and a size budget. Temporary context expires; it does not become memory.
Agents, tools, prompts and memory managed as governed services, with the model gateway they use.
Agents are managed as services with an owner, identity, permissions, version and evaluation record. Models remain replaceable consumers.
One Lifecycle. From Source Data to Governed Context.
Govern
Ingest
Index
Contextualize
Serve and observe
Enterprise Data and Knowledge, Presented as Simple Choices
Open Table Formats and Time Travel
Iceberg tables with schema evolution, snapshots, hidden partitioning and time travel, readable by Spark, Trino and any Iceberg-compatible engine.
Technical reference: Apache Iceberg · Parquet · Apache Polaris
Federated SQL Without Copying
Query lakehouse tables, operational databases and external systems in one SQL layer, so analysts and agents reach data where it lives.
Technical reference: Trino · Connectors · Query Governance
Streaming and Change Data Capture
Capture changes from operational databases and events from applications and devices, enrich them in flight and keep lakehouse tables and context current.
Technical reference: Kafka · Debezium · Flink
Hybrid Retrieval with Citations
Keyword, vector and graph results scored together, filtered by tenant and purpose, re-ranked and returned with provenance, so every answer can be traced to its evidence.
Technical reference: OpenSearch · Qdrant · Re-ranking · Graph Expansion
Graph-Enhanced RAG and Impact Analysis
Multi-hop reasoning over dependencies, timelines and rules for root-cause analysis, impact assessment, planning and explainable answers.
Technical reference: Apache AGE · JanusGraph · Ontologies
Governed Memory and Agent Harness
Agent memory that is classified, scoped, aged and deletable, and an execution harness with tool permissions, human approval, rollback and a kill switch.
Technical reference: Memory Lifecycle · Temporal · Approval Workflow · Kill Switch
Advanced capabilities depend on the validated storage, processing, retrieval and architecture selected for each deployment. No deployment needs every listed component.
Production data platform
Open Data Without the Integration Burden
Validated Architecture
Controlled Lifecycle
Every pinned component follow a documented release, patching, and upgrade policy based on validated upstream releases.
Sovereign and Air-Gapped Ready
Flexible Operating Model
Choose customer-operated, co-managed, fully managed or dedicated SRE Pod operations.
Built for enterprise private cloud, sovereign and regulated cloud, and cloud service provider or MSP environments.
Start with the Lake. Add Private Models and More Cloud Services When You Need Them.
- XaasIO Hyperscaler console for data services
- Lakehouse, streaming, processing and SQL
- Search, vector database and knowledge graph
- Context Lake and Agent Lake with the model gateway
- Tenants, projects, workspaces, roles and single sign-on
- Private model hosting on your GPUs
- Token-metered model APIs through the AI Token Factory
- Training and fine-tuning on lakehouse data
- Hybrid routing to approved public providers by policy
- Additional XaasIO modules in the same console
- Multi-service catalogs and tenant management
- FinOps, metering, billing and payments
Modules delivered through XaasIO Hyperscaler Platform
XaasIO Compute
XaasIO Block and Object Storage
XaasIO RDS
XaasIO AI Factory + HPC
XaasIO AI Token Factory
XaasIO MLT
Data Platform Migration Support Packs
Cloudera and legacy Hadoop estates
Cloud data warehouse and lakehouse repatriation
Proprietary search and SaaS vector databases
Standalone graph databases and RAG pilots