Skip to main content
XaasIO AI Lake · Open Data Platform for AI

Open Data Platform for AI, Made Simple

Give data, analytics and AI teams one governed platform for lakehouse tables, streams, SQL, search, vectors, knowledge graphs and task-ready context. Run it in your data center, sovereign cloud or provider environment, while retaining control of your data, keys, identity and operations.
Open source by default: Ceph S3, Apache Iceberg, Spark, Trino, Kafka, OpenSearch, Qdrant, and Apache AGE. Upstream-aligned. Customer-operated,
co-managed, or fully managed by XaasIO.
Data and AI, simplified

Everything Needed to Deliver Data and Knowledge as a Service

Teams register the data they own. XaasIO AI Lake coordinates storage, cataloging, processing, indexing, retrieval, context and access behind the scenes.

Main components


Open Lakehouse

Store raw and curated data in S3-compatible object storage as open Iceberg tables with a catalog, lineage, versioning and time travel. Query it with SQL and Spark without copying it into a proprietary warehouse.

Technical reference: Ceph S3 · Apache Iceberg · Parquet · Apache Polaris · OpenMetadata · OpenLineage

Ingestion, Streaming and Processing

Load batch files, change data capture and event streams; enrich them in flight; transform at scale; and query the lakehouse and external systems through one SQL layer.

Technical reference: Kafka · Debezium · Flink · Spark · Trino · Argo Workflows

Search, Vector Database and Knowledge Graph

Turn documents, tables and events into keyword indexes, vector embeddings and a graph of entities and relationships, so applications retrieve what is similar and understand how it is connected.

Technical reference: OpenSearch · Qdrant · pgvector · Apache AGE · JanusGraph

Context Lake and Agent Lake

Assemble identity-aware, fresh, ranked and cited context for models and agents, and manage agents, tools, prompts and memory as governed services with approvals and rollback.

Technical reference: Context API · Agent Registry · Temporal Workflows · LangGraph · Model Gateway

Identity and Governance

Separate tenants, projects and workspaces with roles, attributes, classification, retention and policy. Integrate enterprise single sign-on and keep an audit trail across data, retrieval and agent actions.

Technical reference: XaasIO IAM · SSO · RBAC and ABAC · Open Policy Agent · Audit Archive

Use the same platform through the console, SQL, notebooks, APIs or infrastructure as code.

Six layers

An Open Data Platform, Plus the Layers AI Needs

The lakehouse, streaming and SQL layers give you an open alternative to proprietary data platforms. The vector database, knowledge graph, Context Lake and Agent Lake layers turn that data into governed knowledge for models and agents.

Open Lakehouse

Governed facts and history in open formats, on storage you own.

Role
The sovereign data plane: raw, curated and gold zones on S3-compatible object storage, with Iceberg tables for schema evolution, hidden partitioning and time travel.
Covers
Structured datasets, documents, media, features, embeddings, prompts, model artifacts and audit logs, each in its own zone with retention and legal-hold rules.
Catalog and lineage
An Iceberg REST catalog, a data catalog with business glossary and ownership, and lineage events for every pipeline.
Technical reference
Ceph S3 · Apache Iceberg · Parquet and Arrow · Apache Polaris · OpenMetadata · OpenLineage and Marquez
Streams and Processing

Fresh data in, federated SQL out, without a proprietary engine in between.

Role
Batch, streaming and change-data-capture ingestion; stream enrichment; large-scale transformation; and interactive SQL over the lakehouse and external systems.
Covers
Operational databases through CDC, event streams, files and API loads, scheduled pipelines and notebooks.
Performance
Spark for ETL, embeddings and feature generation; Flink for low-latency context processing; Trino for federated queries that avoid copying data.
Technical reference
Apache Kafka and Kafka Connect · Debezium · Apache Flink · Apache Spark · Trino · Argo Workflows
Vector Database

Semantic and hybrid retrieval with tenant scope, provenance and re-ranking built in.

Role
Chunk, embed and index documents, tables and events for similarity search, combined with keyword search and metadata filters, then re-ranked and cited.
Covers
Retrieval-augmented generation, enterprise search, assistants, recommendations and agent context, with tenant, project and workspace isolation on every index.
Embedding lifecycle
Each vector records its source version, embedding model, dimension and re-embedding triggers, so deletion and corrections propagate from source data to every index.
Technical reference
Qdrant · OpenSearch hybrid search · pgvector · Milvus option · Feast feature store
Knowledge Graph

Entities, relationships, dependencies and rules, for reasoning that similarity search cannot do.

Role
Represent assets, services, customers, documents, agents and decisions as entities with typed, sourced relationships, and traverse them for multi-hop questions.
Covers
Application-to-infrastructure dependencies, impact and root-cause analysis, supply chains, identity relationships, maintenance history, fraud patterns and agent planning.
Graph-enhanced retrieval
Graph expansion adds relationship paths to vector and keyword results, so answers explain how things connect and cite the path.
Technical reference
Apache AGE on PostgreSQL · JanusGraph for large graphs · RDF and ontology tooling where required
Context Lake

The minimum necessary evidence for a user, task, model and policy, delivered with provenance.

Role
Assemble identity-aware, task-relevant, fresh and policy-filtered context from SQL, search, vectors and the graph; rank, compress and cite it; expire what no longer applies.
Context families
Identity and entitlement, task and process, operational and real-time state, history and conversations, time and location, security and regulatory obligations.
Delivery
One API returns the evidence, the policy decision, provenance timestamps and a token or size budget, so nothing outside the requester’s entitlement reaches a model.
Technical reference
Context namespaces and events · Policy-filtered retrieval · Re-ranking and compression · Freshness and expiry
Agent Lake

Agents, tools, prompts and memory managed as governed services, with the model gateway they use.

Role
A registry and runtime for agents with an owner, identity, permissions, version and evaluation record; a harness with sandboxed tools, approvals, rollback and a kill switch.
Memory
Working, episodic, semantic and procedural memory with classification, scope, ageing, expiry, correction and deletion. Promotion to organizational memory is explicit and auditable.
Model access
A gateway routes to private models, sovereign endpoints and approved public providers by policy, and meters usage. Models remain replaceable consumers of the lake.
Technical reference
Agent, tool and prompt registries · Temporal durable workflows · LangGraph · Open Policy Agent · Model gateway
How it works

One Lifecycle. From Source Data to Governed Context.

01

Govern

Inventory and classify sources, assign owners, retention and authorized purpose before anything is embedded.
02

Ingest

Load batch, stream and CDC data; validate, catalog and store it in open zones; clean and transform it.
03

Index

Extract entities, create embeddings, build keyword and vector indexes and update the knowledge graph.
04

Contextualize

Filter by identity and policy, rank, compress and cite the evidence; enforce freshness and expiry.
05

Serve and observe

Deliver context to applications, models and agents; trace, evaluate, correct and delete with propagation.
Every derived asset, from a chunk or vector to a graph edge or memory, stays linked to its source, version, policy and retention state.
Advanced Capabilities

Enterprise Data and Knowledge, Presented as Simple Choices

Open Table Formats and Time Travel

Iceberg tables with schema evolution, snapshots, hidden partitioning and time travel, readable by Spark, Trino and any Iceberg-compatible engine.

Technical reference: Apache Iceberg · Parquet · Apache Polaris

Federated SQL Without Copying

Query lakehouse tables, operational databases and external systems in one SQL layer, so analysts and agents reach data where it lives.

Technical reference: Trino · Connectors · Query Governance

Streaming and Change Data Capture

Capture changes from operational databases and events from applications and devices, enrich them in flight and keep lakehouse tables and context current.

Technical reference: Kafka · Debezium · Flink

Hybrid Retrieval with Citations

Keyword, vector and graph results scored together, filtered by tenant and purpose, re-ranked and returned with provenance, so every answer can be traced to its evidence.

Technical reference: OpenSearch · Qdrant · Re-ranking · Graph Expansion

Graph-Enhanced RAG and Impact Analysis

Multi-hop reasoning over dependencies, timelines and rules for root-cause analysis, impact assessment, planning and explainable answers.

Technical reference: Apache AGE · JanusGraph · Ontologies

Governed Memory and Agent Harness

Agent memory that is classified, scoped, aged and deletable, and an execution harness with tool permissions, human approval, rollback and a kill switch.

Technical reference: Memory Lifecycle · Temporal · Approval Workflow · Kill Switch

Advanced capabilities depend on the validated storage, processing, retrieval and architecture selected for each deployment. No deployment needs every listed component.

Production data platform

Open Data Without the Integration Burden

XaasIO designs, deploys, validates and supports XaasIO AI Lake as one production platform rather than an unsupported collection of open-source projects, connectors and scripts.

Validated Architecture

Object storage, table format, catalog, processing engines, search, vector and graph stores, context services, identity and observability are tested as a complete solution.

Controlled Lifecycle

Every pinned component follow a documented release, patching, and upgrade policy based on validated upstream releases.

Sovereign and Air-Gapped Ready

Data locality, customer-controlled keys, tenant isolation, local repositories, identity and controlled update procedures inside the approved security boundary.

Flexible Operating Model

Choose customer-operated, co-managed, fully managed or dedicated SRE Pod operations.

Built for enterprise private cloud, sovereign and regulated cloud, and cloud service provider or MSP environments.

Console included, platforms optional

Start with the Lake. Add Private Models and More Cloud Services When You Need Them.

The default XaasIO AI Lake edition includes the XaasIO Hyperscaler console for data services. Hosting private models on your own GPUs is the job of the XaasIO AI Factory Platform. Delivering Compute, RDS and other XaasIO modules through the same console needs the XaasIO Hyperscaler Platform.
Included
Included with XaasIO AI Lake
  • XaasIO Hyperscaler console for data services
  • Lakehouse, streaming, processing and SQL
  • Search, vector database and knowledge graph
  • Context Lake and Agent Lake with the model gateway
  • Tenants, projects, workspaces, roles and single sign-on
Optional
Add XaasIO AI Factory Platform
  • Private model hosting on your GPUs
  • Token-metered model APIs through the AI Token Factory
  • Training and fine-tuning on lakehouse data
  • Hybrid routing to approved public providers by policy
Explore XaasIO AI Factory
Optional
Requires XaasIO Hyperscaler Platform
  • Additional XaasIO modules in the same console
  • Multi-service catalogs and tenant management
  • FinOps, metering, billing and payments
Explore XaasIO Hyperscaler Platform

Modules delivered through XaasIO Hyperscaler Platform

XaasIO Compute

OpenStack private cloud for virtual machines, networks, and persistent storage.

XaasIO Block and Object Storage

Ceph block volumes and the S3-compatible object storage the lake runs on.

XaasIO RDS

Managed PostgreSQL, MySQL, MariaDB and MongoDB, and a CDC source for the lake.

XaasIO AI Factory + HPC

Model training, fine-tuning, LLM and SLM development and HPC computing.

XaasIO AI Token Factory

Governed inference and token-metered model APIs, the model gateway’s private endpoint.

XaasIO MLT

Monitoring, logging and telemetry for the lake, its pipelines, retrieval and agents.
These modules are optional and are not required to deploy or operate XaasIO AI Lake. Adding them to the console requires the XaasIO Hyperscaler Platform.

Data Platform Migration Support Packs

Separately scoped support for assessment, migration, cutover and managed operations – available for:

Cloudera and legacy Hadoop estates

Cloud data warehouse and lakehouse repatriation

Proprietary search and SaaS vector databases

Standalone graph databases and RAG pilots

Software supply chain assurance

Every supported XaasIO MLT release includes an upstream SBOM and XaasIO release SBOM, covering components, container images, dependencies, licenses, vulnerabilities, provenance, and signed artifacts. SPDX or CycloneDX outputs support SOC 2-aligned controls and audit evidence.
Upstream SBOM
XaasIO release SBOM
Components and container images
Licenses and vulnerabilities
Provenance and signed artifacts
SPDX or CycloneDX output
SOC 2-aligned audit evidence

Existing tables, pipelines, indexes and prompts are reviewed for reuse


SPDX and CycloneDX outputs are available for every supported release

Frequently asked questions

What is XaasIO AI Lake?
An open-source-based data platform for analytics and AI, delivered and supported as one product. It combines an Iceberg lakehouse on Ceph S3 object storage, streaming and processing with Kafka, Flink, Spark and Trino, search with a vector database and a knowledge graph, and two AI-native services: the Context Lake, which assembles governed context for models and agents, and the Agent Lake, which manages agents, tools, prompts and memory. It runs in your data center, sovereign cloud or provider environment, and XaasIO designs, validates, supports and can operate it.
How is an AI Lake different from a data lake or a lakehouse?
A data lake stores raw data cheaply. A lakehouse adds open table formats, transactions, schema evolution and SQL over that storage. XaasIO AI Lake keeps both and adds what models and agents need: semantic search through a vector database, relationship knowledge through a knowledge graph, identity-aware context assembly through the Context Lake, and governed agents and memory through the Agent Lake. Every derived asset, from a chunk or vector to a graph edge or memory, stays linked to its source, version, policy and retention state.
Is XaasIO AI Lake an open alternative to Cloudera and legacy Hadoop platforms?
Yes. The platform is built from open-source components with no proprietary table format, query engine or license key: Apache Iceberg on S3-compatible storage instead of HDFS and Hive tables, Spark and Trino for processing and SQL, Kafka for streaming, and open catalogs and lineage. Migration support packs cover Cloudera and legacy Hadoop estates, and existing Spark jobs, SQL, pipelines and Kafka topics are assessed for reuse before cutover.
Which open-source components does the platform use?
Ceph S3 object storage; Apache Iceberg with Parquet and Arrow; Apache Polaris, OpenMetadata and OpenLineage for catalog, metadata and lineage; Apache Kafka, Debezium, Apache Flink, Apache Spark and Trino for ingestion, processing and SQL; OpenSearch and Qdrant for keyword, vector and hybrid search; Apache AGE on PostgreSQL, or JanusGraph for large graphs, for the knowledge graph; PostgreSQL, Valkey and Temporal for services and durable workflows; and OpenTelemetry, Prometheus and Grafana for observability. Components are pinned per release and can be reviewed before deployment.
Why does the platform include both a vector database and a knowledge graph?
They answer different questions. A vector database finds what is similar in meaning, which suits document search, retrieval-augmented generation and recommendations. A knowledge graph records how entities are connected and what depends on what, which suits impact analysis, root cause, compliance and multi-step planning. XaasIO AI Lake scores keyword, vector and graph results together, filters them by tenant, purpose and policy, and returns them with citations, so applications get both similarity and relationships from one retrieval call.
Which vector database and graph database are used?
Qdrant is the default vector database, with OpenSearch providing keyword and hybrid search and pgvector available for small deployments. Milvus is an option for very large vector estates. The knowledge graph runs on Apache AGE over PostgreSQL by default and on JanusGraph for large graphs; RDF and ontology tooling is added where semantic standards are required. The validated selection for each deployment is documented in its design.
What is the Context Lake?
A governed service that assembles the minimum necessary evidence for a specific user, task, model and policy. It draws on SQL, search, vector and graph results and on live operational state, filters by identity and entitlement, ranks and compresses the evidence, attaches provenance and freshness timestamps, applies a token or size budget and expires context that no longer applies. Applications and agents call one Context API rather than integrating each store separately.
What is the Agent Lake?
A registry and runtime that treats agents, tools, prompts and memory as governed services. Each agent has an owner, an identity, explicit tool permissions, a version and an evaluation record. Runs execute in a harness with sandboxed tools, human approval for sensitive actions, tracing, rollback and a kill switch. Memory is classified, scoped, aged, correctable and deletable, and promotion to shared organizational memory is explicit and auditable.
Do we need GPUs or the XaasIO AI Factory Platform?
Not for the lake itself. Storage, tables, SQL, streaming, search, the vector database, the knowledge graph, the Context Lake and the Agent Lake run on CPU infrastructure, and embedding and re-ranking models can run on CPU at moderate volumes or on GPUs. Hosting private large language models on your own GPUs is the job of the XaasIO AI Factory Platform, whose AI Token Factory becomes the private endpoint behind the model gateway. Where policy allows, the gateway can also route to approved public model providers.
Does our data leave the environment? Are public AI services used?
No, by default. Data, indexes, graphs, context and memory stay inside the deployment boundary, encrypted with keys you control. The model gateway routes to private models first and uses approved public providers only where an explicit policy allows it, with the request logged and metered. Prompts, context and responses are retained according to the policy you set per application.
Can XaasIO AI Lake run in an air-gapped environment?
Yes. The fully air-gapped pattern runs every layer inside the boundary: local image, chart and model registries, object storage, catalogs, processing, search, vector and graph stores, context and agent services, identity, policy and observability. Updates arrive only as signed, scanned and approved offline bundles.
Can a cloud service provider offer it as a multi-tenant managed service?
Yes. Tenants, projects and workspaces are isolated at the storage, index, graph, context and agent levels, with per-tenant keys, quotas and audit trails. Usage events are produced for every service. With the XaasIO Hyperscaler Platform, providers publish the data, retrieval, knowledge and agent services in a self-service catalog and rate, invoice and settle them alongside Compute, storage and databases.
Is the XaasIO Hyperscaler console included?
Yes. The default XaasIO AI Lake edition includes the XaasIO Hyperscaler console for the lake services: tenants, projects, workspaces, data sources, tables, indexes, graphs, context applications, agents, access and usage. Extending the same console to Compute, Block and Object Storage, RDS, AI Factory, MLT and other modules, multi-service catalogs, billing and payments requires the XaasIO Hyperscaler Platform.
Are other XaasIO modules required?
No. XaasIO AI Lake runs on any validated S3-compatible object storage and Kubernetes foundation. XaasIO Block and Object Storage, Compute, RDS, AI Factory and MLT are optional and integrate directly: Ceph provides the object tier, RDS databases become change-data-capture sources, AI Factory hosts private models and MLT observes the pipelines, retrieval and agents.
Who operates XaasIO AI Lake?
You choose: customer-operated with XaasIO support, co-managed, fully managed by XaasIO, or a dedicated XaasIO SRE Pod. Every deployment includes validated architecture, documented release and upgrade policy, SBOMs for every release and 24×7 support options.
How does XaasIO AI Lake compare with Acceldata Open Data Platform?
Both are open-source data platforms positioned as alternatives to Cloudera, built on Spark, Trino, Kafka and open table formats. XaasIO AI Lake adds the retrieval and AI-native layers described on this page, a vector database, a knowledge graph, the Context Lake and the Agent Lake, ships with the XaasIO Hyperscaler console, and is validated for sovereign, air-gapped and service-provider deployments with XaasIO operating options.

Apache, Iceberg, Spark, Trino, Kafka, Flink, AGE, Ceph, OpenSearch, Qdrant, JanusGraph, Cloudera, Acceldata, Databricks, Snowflake and other product names are trademarks of their respective owners and are named for identification and comparison only. XaasIO is not affiliated with, endorsed by or sponsored by these organizations.

Build what the world depends on.

Planning a cloud for government, financial services, pharmaceuticals, healthcare, energy, transport, industry, service-provider operations, or private AI? Work with XaasIO to define the architecture, security boundary, and operating model your environment requires.