What is Devyra?

Devyra is an AI Decision Reliability Platform designed to test critical business decisions before they impact customers or the business.

Instead of deploying a policy and discovering weaknesses afterward, Devyra creates a controlled environment where a decision policy can be discovered, stress-tested, attacked, improved, replayed, approved, and monitored.

For the prototype, we focused on one concrete business decision:

ApexCloud — Customer Refund Approval

The system evaluates a refund policy against 10,000 synthetic customer cases and adversarial abuse scenarios such as refund splitting, high-frequency refund attempts, and linked-account abuse.

Why We Built It

Many business decisions are automated, but reliability is often measured only after deployment.

We wanted to explore a different approach:

What if an organization could attack its own decision policies before trusting them?

This led to Devyra's workflow:

Connect → Discover → Test → Attack → Detect → Analyze → Fix → Replay → Approve → Monitor

The goal is not simply to produce another AI prediction. It is to provide evidence showing where a decision policy fails, why it fails, and how a safer version performs against the same scenarios.

How It Works

Devyra starts with a decision policy and evaluates it against a deterministic synthetic benchmark.

For each test case, the system records whether the decision was correct.

Reliability = Correct Decisions / Total Test Cases

For the refund scenario:

  • Policy V1: 7,600 / 10,000 = 76.0% reliability
  • Policy V2: 9,820 / 10,000 = 98.2% reliability

The system also calculates potential exposure as the sum of wrongly approved abusive refunds.

  • V1 exposure: ₹11,62,400
  • V2 exposure: ₹38,200

The V2 policy is not presented as a manually written improvement. It is generated from the failure patterns discovered during adversarial testing, introducing controls such as:

  • Rolling 30-day refund limits
  • Refund velocity controls
  • Linked-account checks
  • Human escalation for borderline cases

The improved policy is then replayed against the full 10,000-case benchmark so the user can compare the tradeoffs between V1 and V2 before approval.

Evidence, Not Just Scores

A key part of Devyra is traceability.

When a vulnerability is discovered, the system connects the result back to concrete refund cases such as:

  • RF-008421
  • RF-009104
  • RF-007238

This allows the user to move from a metric such as "policy failed" to evidence showing which case failed and why.

Human Approval

Devyra does not automatically assume that the generated policy should be deployed.

The workflow includes a human approval step where the decision-maker can review the V1/V2 comparison and approve the proposed policy.

After approval, the policy becomes active in the Devyra demo environment, where its reliability and decision activity can be monitored.

What We Learned

Building Devyra taught us that a reliable AI system is not only about model accuracy.

The surrounding system needs:

  • Reproducible evaluation
  • Adversarial testing
  • Evidence traceability
  • Policy comparison
  • Historical replay
  • Human approval
  • Continuous monitoring
  • Auditability

We also learned an important limitation of our prototype: the abuse patterns are injected by the synthetic data generator. Therefore, the strong V2 performance partly reflects how well the generated controls match those predefined patterns. A production system would require evaluation against independently collected and continuously evolving abuse data.

Challenges

The biggest challenge was turning a complex decision-intelligence concept into a workflow that a real user could understand.

We deliberately moved the complex internal mechanisms behind a simpler product experience:

Find a decision → Test it → Understand the failures → Generate a safer policy → Replay it → Approve it → Monitor it.

The result is a working prototype focused on one measurable business problem rather than a collection of disconnected AI features.

Built for the Challenge

Devyra was built as a functional prototype with a browser-based product interface, a FastAPI backend, deterministic synthetic evaluation data, adversarial test scenarios, policy generation, historical replay, monitoring, and an audit trail.

The prototype is designed to demonstrate the complete decision-reliability workflow end to end.

Built With

  • adversarial
  • analysis
  • api
  • artificial
  • css
  • data
  • decision
  • fastapi
  • html
  • intelligence
  • javascript
  • learning
  • machine
  • policy
  • python
  • rest
  • simulation
  • synthetic
  • testing
  • visualization
  • web
Share this project:

Updates

Submission history