Inspiration
AI is already excellent at fixing SQL syntax. But in a real data team, syntax is rarely the hardest problem.
If an engineer selects a query referencing analytics.customers, a generic AI model does not automatically know:
- which dataset is authoritative,
- which fields actually exist,
- what those fields mean,
- who owns the data,
- where it came from,
- or what dashboards and ML models depend on it.
That gap inspired Cursivis DataOps.
Generic AI understands the code. Cursivis DataOps understands the organization's data behind the code.
Instead of opening another chat, copying SQL, explaining your stack, and hoping the model does not hallucinate, you simply select what you are already working on and invoke Cursivis.
What it does
Cursivis DataOps is a context-native AI agent for data engineers, analysts, and platform teams.
It combines:
- Cursivis for instant selection-based intent
- DataHub for organizational truth
- Gemini for grounded reasoning
The core workflow is:
Select SQL / error
↓
Cursivis captures context
↓
DataHub resolves the real asset
↓
Schema + ownership + descriptions + lineage
↓
Gemini reasons over verified context
↓
Grounded answer + blast radius
↓
Copy / Insert / Replace / Save Resolution
The key difference is simple:
DataHub is queried before Gemini reasons.
So Cursivis does not ask the model to guess what a dataset looks like — it gives Gemini the actual organizational metadata first.
If DataHub cannot resolve an asset, Cursivis fails visibly instead of pretending an ungrounded answer is trustworthy.
A real example
Suppose an engineer selects:
SELECT
customer_id,
lifetime_value,
tier
FROM analytics.customers;
A generic model may confidently invent a fix.
Cursivis DataOps first resolves analytics.customers through DataHub and discovers the governed schema contains:
customer_id
lifetime_value_usd
customer_tier
updated_at
It also retrieves the dataset's owner and lineage.
Gemini can now return more than corrected syntax:
Suggested Fix
Use lifetime_value_usd and customer_tier.
DataHub Evidence
Dataset: analytics.customers
Owner: data-platform
Verified fields: customer_id, lifetime_value_usd, customer_tier, updated_at
Blast Radius The dataset feeds downstream analytics and ML workloads, including revenue reporting and churn prediction.
That turns an AI suggestion into an organization-aware engineering decision.
How we built it
Cursivis DataOps is a Windows desktop application built with C#, .NET 8, and WinUI 3.
Cursivis captures the user's current selection and extracts referenced datasets from SQL, including qualified FROM and JOIN statements.
For data-aware requests, it connects to DataHub and retrieves live metadata such as:
- dataset identity
- schema and field descriptions
- ownership
- upstream lineage
- downstream lineage
- impact / blast radius
That context is bounded and passed to the Gemini Developer API, which returns structured output that Cursivis can safely render inside its result panel.
Users can then:
- Copy the result
- Insert the corrected output
- Replace the selected content
- inspect DataHub evidence
- review downstream impact
- Save Resolution to DataHub
The write-back loop is deliberately safe: saving knowledge requires explicit user confirmation, and Cursivis performs read-after-write verification before reporting success.
No silent catalog mutations.
DataHub as agent memory
One of the most important ideas behind the project is that DataHub is not used as generic chat memory.
Instead, it becomes durable organizational memory.
A useful resolution — for example, why a field was renamed or how a dataset issue was fixed — can be reviewed and saved back into DataHub.
The next engineer or agent investigating the same asset can inherit that knowledge.
This creates a loop:
Read → Reason → Act → Learn
Rather than:
Read → Chat → Forget
Making it reproducible
We did not want the demo to depend on fake metadata or screenshots.
The project includes a deterministic DataHub demo catalog:
raw.customers
↓
analytics.customers
├──→ analytics.executive_revenue
└──→ ml.churn_prediction_features
The setup seeds real:
- schemas,
- descriptions,
- ownership,
- and lineage.
After ingestion, our scripts query DataHub again and verify that the expected assets and relationships actually exist.
If they do not, setup fails.
This means the metadata shown inside Cursivis is the same metadata running inside DataHub.
Challenges we faced
Grounding without hallucination
It would have been easy to inject example JSON into Gemini and call the result "DataHub-grounded."
We deliberately avoided that.
The production path resolves metadata from the live DataHub instance and handles missing entities, GraphQL failures, timeouts, authentication issues, and malformed responses.
Making actions safe
Reading metadata is low-risk. Writing organizational knowledge is not.
We therefore made write-back an explicit, confirmed action and verify the saved result before showing success.
Turning metadata into useful context
Raw metadata alone is noisy.
We had to transform schema, descriptions, ownership, and lineage into a small, structured context package that gives Gemini enough evidence to reason without overwhelming the model.
What we learned
The biggest lesson was:
A better model does not automatically mean better organizational context.
Even a powerful model can confidently recommend a field that does not exist.
The combination works because each layer has a clear job:
Gemini provides reasoning. DataHub provides organizational truth. Cursivis provides immediate intent and action.
Together, they turn metadata from passive documentation into a runtime reasoning layer for AI agents.
What's next
SQL debugging is only the beginning.
The same architecture can support:
- pipeline failure investigation
- schema-change impact analysis
- ownership-aware incident routing
- governed code generation
- data quality debugging
- migration assistance
- organization-wide agent memory
Our vision is for Cursivis to become the context-native interface between what a user is working on right now and everything their organization already knows about it.
Selection provides intent. DataHub provides truth. Gemini reasons over it. Cursivis turns it into action.
Log in or sign up for Devpost to join the conversation.