What is Devyra?
Devyra is an AI Decision Reliability Platform designed to test critical business decisions before they impact customers or the business.
Instead of deploying a policy and discovering weaknesses afterward, Devyra creates a controlled environment where a decision policy can be discovered, stress-tested, attacked, improved, replayed, approved, and monitored.
For the prototype, we focused on one concrete business decision:
ApexCloud — Customer Refund Approval
The system evaluates a refund policy against 10,000 synthetic customer cases and adversarial abuse scenarios such as refund splitting, high-frequency refund attempts, and linked-account abuse.
Why We Built It
Many business decisions are automated, but reliability is often measured only after deployment.
We wanted to explore a different approach:
What if an organization could attack its own decision policies before trusting them?
This led to Devyra's workflow:
Connect → Discover → Test → Attack → Detect → Analyze → Fix → Replay → Approve → Monitor
The goal is not simply to produce another AI prediction. It is to provide evidence showing where a decision policy fails, why it fails, and how a safer version performs against the same scenarios.
How It Works
Devyra starts with a decision policy and evaluates it against a deterministic synthetic benchmark.
For each test case, the system records whether the decision was correct.
Reliability = Correct Decisions / Total Test Cases
For the refund scenario:
- Policy V1: 7,600 / 10,000 = 76.0% reliability
- Policy V2: 9,820 / 10,000 = 98.2% reliability
The system also calculates potential exposure as the sum of wrongly approved abusive refunds.
- V1 exposure: ₹11,62,400
- V2 exposure: ₹38,200
The V2 policy is not presented as a manually written improvement. It is generated from the failure patterns discovered during adversarial testing, introducing controls such as:
- Rolling 30-day refund limits
- Refund velocity controls
- Linked-account checks
- Human escalation for borderline cases
The improved policy is then replayed against the full 10,000-case benchmark so the user can compare the tradeoffs between V1 and V2 before approval.
Evidence, Not Just Scores
A key part of Devyra is traceability.
When a vulnerability is discovered, the system connects the result back to concrete refund cases such as:
RF-008421RF-009104RF-007238
This allows the user to move from a metric such as "policy failed" to evidence showing which case failed and why.
Human Approval
Devyra does not automatically assume that the generated policy should be deployed.
The workflow includes a human approval step where the decision-maker can review the V1/V2 comparison and approve the proposed policy.
After approval, the policy becomes active in the Devyra demo environment, where its reliability and decision activity can be monitored.
What We Learned
Building Devyra taught us that a reliable AI system is not only about model accuracy.
The surrounding system needs:
- Reproducible evaluation
- Adversarial testing
- Evidence traceability
- Policy comparison
- Historical replay
- Human approval
- Continuous monitoring
- Auditability
We also learned an important limitation of our prototype: the abuse patterns are injected by the synthetic data generator. Therefore, the strong V2 performance partly reflects how well the generated controls match those predefined patterns. A production system would require evaluation against independently collected and continuously evolving abuse data.
Challenges
The biggest challenge was turning a complex decision-intelligence concept into a workflow that a real user could understand.
We deliberately moved the complex internal mechanisms behind a simpler product experience:
Find a decision → Test it → Understand the failures → Generate a safer policy → Replay it → Approve it → Monitor it.
The result is a working prototype focused on one measurable business problem rather than a collection of disconnected AI features.
Built for the Challenge
Devyra was built as a functional prototype with a browser-based product interface, a FastAPI backend, deterministic synthetic evaluation data, adversarial test scenarios, policy generation, historical replay, monitoring, and an audit trail.
The prototype is designed to demonstrate the complete decision-reliability workflow end to end.
Built With
- adversarial
- analysis
- api
- artificial
- css
- data
- decision
- fastapi
- html
- intelligence
- javascript
- learning
- machine
- policy
- python
- rest
- simulation
- synthetic
- testing
- visualization
- web
Log in or sign up for Devpost to join the conversation.