Documentation · OrionBelt® Platform

OrionBelt® Analytics

Ontology-grounded Text-to-SQL on top of your harmonized lakehouse. Ask in plain language from the AI client you already use; every query is checked against a model of your data before it runs.

Version 2.1.0 · MCP server · Docker image or Python 3.13+ · github.com/ralforion/orionbelt-analytics

What it is, and what it is not

OrionBelt® Analytics (OBA) is an MCP server that sits between an AI client and your SQL warehouse or lakehouse. It reads your schema, writes it down as an RDF/OWL ontology with the SQL needed to reach every concept, and gives the model the tools to find the right tables, join them correctly, and run the query. Every query passes a deterministic check against that ontology first.

It is

  • An MCP server between your AI client and your database
  • A generator of an ontology of your schema, with SQL mappings
  • A schema search that finds the tables and join paths a question needs
  • A rule-based check of every query before it runs
  • Read-only: queries run inside your database, nothing is copied out

It is not

  • A BI tool: charts appear in the conversation, your dashboards stay
  • A data platform or ETL: it works on your schemas as they are
  • An LLM: you bring the client and the model
  • A governed metrics layer: that is the OrionBelt® Semantic Layer
  • A write path: DDL and DML are refused

Architecture

Clients talk to the server over MCP (Streamable HTTP). The server talks to the database over SQL.

Each connection keeps a workspace on disk: schema cache, ontology versions, the GraphRAG index and the RDF store. Reconnecting to the same database restores it, so a known database is ready without another discovery.

A typical session

“Connect to our sales database. Which customer regions grew revenue last quarter, and how do their returns compare?”

  1. Connect. The model lists the configured databases and connects to the one you named: list_databases, connect_database(database="sales"). A known workspace is restored.
  2. Discover. discover_schema reads tables, keys and views; the GraphRAG index builds in the background.
  3. Model. generate_ontology writes the ontology. suggest_semantic_names and apply_semantic_names give cryptic columns business names. Every later question is matched against those names too, so "revenue" finds net_amt.
  4. Ask. graphrag_query_context returns the relevant tables, columns and join paths in one compact answer. For a question across two fact tables, plan_composite_query proposes a fan-trap-safe shape.
  5. Check, run, chart. execute_sql_query runs OBQC, then the query, read-only. generate_chart draws the result.

The example question is deliberately a two-fact one. Revenue and returns live in different tables, and summing both across a one-to-many join inflates the totals. OBQC blocks that query and points to the fix: aggregate each fact first, then combine. If the model insists on running it anyway, a client that can ask the user does so, naming the tables, before any inflated number is shown.

Bring your own ontology

The generated ontology is a starting point. When your layer declares no keys, or you already maintain a domain model, you can supply your own.

  1. Start. Download the generated ontology with download_artifact, or start from one you have.
  2. Edit. Add joins, business names and descriptions in OrionBelt® Ontology Builder or any OWL editor.
  3. Upload. Drop the .ttl file into the chat, or place it in the server's import folder; load_my_ontology reads it.
  4. Activate. It becomes the active ontology when it maps classes and properties to the database with oba: annotations. If it does not, it is still loaded for SPARQL, the previous ontology stays in force, and the response lists what is missing. Conformance to the OBA SHACL shapes is reported as advice.
  5. Ask. Your joins now guide the SQL, and OBQC checks every query against your model.
  6. Verify. Check the joins that matter against the data with validate_relationship. The verdict (confirmed, partial, refuted, or a target that is not unique) is written onto the relationship in the ontology, and join paths show it from then on.

An uploaded ontology applies to your session only. Colleagues on the same database keep theirs, and the database itself is never changed. The most recently activated ontology wins, so generating a new one or applying business names afterwards replaces the upload. The full story, including joins inferred from column names and informational key constraints, is on Join Discovery Without Declared Keys. The annotation vocabulary is published at the oba: namespace.

Tools

OBA exposes 30 MCP tools. Each carries the standard MCP tool annotations (readOnlyHint, destructiveHint, openWorldHint), so a client or an approval layer can treat tools by what they do. None is open-world: each acts only on the configured database and the server's workspace.

GroupTools
Connection and schemalist_databases, connect_database, list_schemas, discover_schema, get_table_details, reset_cache, cleanup_workspace, cleanup_old_versions
Ontology and namesgenerate_ontology, suggest_semantic_names, apply_semantic_names, add_semantic_context, load_my_ontology, validate_relationship, download_artifact
GraphRAGgraphrag_search, graphrag_query_context, graphrag_find_join_path, reachable_from, measurable_from, plan_composite_query
Query and chartssample_table_data, execute_sql_query (with OBQC), generate_chart
SPARQL and RDFstore_ontology_in_rdf, query_sparql, add_rdf_knowledge
Semantic modelssave_semantic_model, get_semantic_model, list_semantic_models

Parameters, return values and examples are in the tools reference.

Databases

Eight SQL engines, cloud and on-premises: PostgreSQL, MySQL, Snowflake, ClickHouse, Dremio, BigQuery, DuckDB/MotherDuck and Databricks SQL. SQL is parsed per dialect, so OBQC and the row limit follow each engine's own syntax.

Join discovery works best when primary and foreign keys are declared in the catalog. Many lakehouse layers declare none; see Join Discovery Without Declared Keys for the three ways around that.

Install and connect

Docker

The image on Docker Hub includes both embedding models GraphRAG can use, English and multilingual, so it runs in networks without internet access. Workspaces live in /data. Set the two -e values explicitly, or remove them from your .env: a .env copied from the repository's template sets MCP_SERVER_HOST=localhost and OUTPUT_DIR=tmp, which would make the published port unreachable and keep workspaces outside the volume.

docker run -d --name oba -p 9000:9000 \
  --env-file .env \
  -e MCP_SERVER_HOST=0.0.0.0 \
  -e OUTPUT_DIR=/data \
  -v oba-data:/data \
  ralforion/orionbelt-analytics

From source

git clone https://github.com/ralforion/orionbelt-analytics
cd orionbelt-analytics
uv sync
cp .env.template .env   # add your database settings
uv run server.py        # http://localhost:9000/mcp

Connect a client

# Claude Code
claude mcp add --transport http orionbelt-analytics http://localhost:9000/mcp

Claude Desktop connects through mcp-remote; LangChain, OpenAI Agents SDK, CrewAI, Google ADK, Vercel AI SDK and n8n connect to the same endpoint and discover every tool. Runnable examples are in the integrations guide. OrionBelt® Chat connects to OBA and the Semantic Layer at once.

Configuration

The server reads its settings from the environment or a .env file. Database credentials live only there; no tool takes them as a parameter.

Named databases

One server can hold several databases, each under the name people use and a description in their words. The model lists them with list_databases and connects by name; if two fit, it asks. Hosts, users and tokens are never shown.

OBA_DATABASES=finance-2025,sales

DB_FINANCE_2025_TYPE=databricks
DB_FINANCE_2025_DESCRIPTION=Finance actuals 2025: GL, cost centres, budgets
DB_FINANCE_2025_DATABRICKS_CATALOG=finance
DB_FINANCE_2025_DATABRICKS_SCHEMA=gold

DB_SALES_TYPE=databricks
DB_SALES_DESCRIPTION=Orders, customers and returns
DB_SALES_DATABRICKS_CATALOG=sales
DB_SALES_DATABRICKS_SCHEMA=gold

# Shared by both connections
DATABRICKS_SERVER_HOSTNAME=adb-1234567890.12.azuredatabricks.net
DATABRICKS_HTTP_PATH=/sql/1.0/warehouses/abc123
DATABRICKS_ACCESS_TOKEN=dapi...

A setting a named database does not define falls back to the unprefixed one, so two catalogs can share one endpoint and token. With a single database configured, connect_database() needs no argument. Each database keeps its own workspace, identified by where it points and who it signs in as.

For users who ask in a different language than the schema is named in, set GRAPHRAG_EMBEDDING_MODEL=multilingual; see GraphRAG Schema Discovery.

Transport, embedding model, workspace retention, session timeouts and per-engine notes are in the configuration reference.

Security model

There is no built-in user sign-in today. Run the server inside your network, next to the data, or behind your gateway.

Reference and vocabulary

TopicWhere
GraphRAG, explainedGraphRAG Schema Discovery
Query checks and fan-trapsOBQC: Ontology-Based Query Check
Joins without keys, your own ontologyJoin Discovery Without Declared Keys
The oba: vocabularyNamespace page · oba.ttl · SHACL shapes · example
Tool parametersdocs/tools-reference.md
Configurationdocs/configuration.md
OBQC rule referencedocs/obqc.md
Fan-trap patternsdocs/fan-trap-prevention.md
Changes per releaseCHANGELOG.md

OBA pairs with the OrionBelt® Semantic Layer: what OBA discovers can be saved as an OBML model, which the Semantic Layer compiles into governed, reusable metrics. Ad-hoc questions stay with OBA; recurring ones move to the model.

OrionBelt® Analytics is licensed under the Business Source License 1.1 (BUSL-1.1) and converts to the Apache License 2.0 on 2030-03-16.

Frequently Asked Questions

Is OrionBelt Analytics a BI or dashboard tool?

No. It answers questions in the conversation and can chart each answer there. Your existing dashboards stay where they are. For governed, reusable metric definitions, use its sibling, the OrionBelt® Semantic Layer.

Does it copy data out of my warehouse?

No. Queries run inside your database with the credentials configured on the server, and only bounded result sets come back. Write and DDL statements are refused before they reach the database.

Which AI clients work with it?

Any client that speaks MCP over HTTP: Claude Desktop, Claude Code, OrionBelt® Chat, LibreChat, and agent frameworks such as LangChain, OpenAI Agents SDK, CrewAI, Google ADK, Vercel AI SDK and n8n. ChatGPT custom GPTs connect through an MCP-to-REST bridge.

Which LLM does it use?

None of its own for writing SQL. You bring the client and the model. The server runs one small local embedding model to search the schema by meaning: all-MiniLM-L6-v2 by default, or a multilingual model for questions in another language than the schema. Both ship inside the Docker image, so nothing leaves the server.

What if my lakehouse tables have no primary or foreign keys?

Joins are inferred from column naming patterns and marked with a confidence. You can also upload your own ontology that states the joins, or declare informational PK/FK constraints in the catalog. Declaring keys and then adding business meaning in an ontology gives the best result. Details: Join Discovery Without Declared Keys.

Does it have user sign-in?

Not today. Deploy it inside your network or behind your gateway. The database is reached with credentials configured on the server, and every session uses them. Sessions keep their own ontology state, which is workflow isolation rather than access control.

How is it licensed?

Under the Business Source License 1.1 (BUSL-1.1). The licensed work converts to the Apache License 2.0 on 2030-03-16. For commercial licensing, contact licensing@ralforion.com.

Ask your lakehouse, and trust the SQL

Run OrionBelt® Analytics next to your data, or talk to us about a demo on your own schemas.

OrionBelt® Analytics on GitHub Docker Hub Contact RALFORION