Inspiration

A DNS incident can make a website, API, or email service disappear in seconds.

Changing a DNS record is easy. Recovering safely is not.

During an incident, an operator must answer:

  • What exactly changed?
  • Which DNS state was actually trusted?
  • Is the new destination expected?
  • What is the minimum safe rollback?
  • Who authorized the recovery?
  • Did the provider actually end up in the intended state?

Teams often answer those questions using screenshots, tickets, memory, and manual provider-console work.

We built DomainTwin AI to turn DNS recovery into a deterministic, approval-gated, AI-assisted, and independently verified workflow.

Our core idea is simple:

Treat DNS as a recoverable digital twin.

What it does

DomainTwin's primary workflow is:

Detect → Explain → Authorize → Restore → Verify → Prove

It continuously reads live DNS through the Name.com API, stores immutable DNS snapshots, and allows an operator to explicitly designate a trusted known-good state.

When live DNS diverges from that trusted state, DomainTwin calculates a deterministic diff and combines it with health evidence to create an explainable incident.

A real recovery preserved in the live demo

For our controlled Name.com Sandbox test, the trusted state was:

A www → 203.0.113.10

We intentionally introduced drift:

A www → 198.51.100.77

DomainTwin detected the change and created Incident #8.

The incident was scored:

CRITICAL — 75/100

using three explicit deterministic factors:

  • +30 ADDRESS_RECORD_CHANGED
  • +30 HTTP_HEALTH_FAILED
  • +15 UNKNOWN_DESTINATION

The score does not come from an LLM.

It is reproducible from the evidence.

Evidence-grounded AI explanation

After the deterministic incident exists, DomainTwin can pass a structured evidence bundle to its AI explanation layer.

For Incident #8, the persisted AI analysis was successfully generated and classified the affected service as MULTIPLE with MEDIUM confidence.

The AI explanation references specific evidence IDs including:

DNS-001

HEALTH-DNS

HEALTH-HTTP

HEALTH-HTTPS

RISK-001

RISK-002

RISK-003

RISK-SCORE

The AI is explicitly constrained to explain existing evidence.

It cannot create the incident, change the deterministic risk score, approve recovery, or mutate DNS.

This is intentional:

AI explains. Deterministic systems decide. Humans authorize. DomainTwin verifies.

Verified DNS recovery

From Incident #8, DomainTwin created Recovery Plan #8 containing exactly one deterministic operation:

UPDATE A www

198.51.100.77 → 203.0.113.10

Nothing was changed automatically.

An operator first inspected the exact rollback preview and explicitly approved the recovery.

DomainTwin then executed the update through the Name.com Sandbox API.

But a successful write request is not enough for DomainTwin to declare recovery.

After execution, DomainTwin independently re-read the provider, normalized the live DNS records, calculated the resulting fingerprint, and compared it against trusted Snapshot v3.

The result was:

EXPECTED = a3b35ae640...dad54313

ACTUAL = a3b35ae640...dad54313

MATCH YES

Only after that fresh provider read did Recovery Plan #8 become:

RECOVERED

The complete recovery audit trail contains:

  1. PLAN CREATED
  2. PLAN APPROVED
  3. APPROVAL ACTOR RECORDED
  4. EXECUTION ACTOR AUTHORIZED
  5. APPLY STARTED
  6. OPERATION SUCCEEDED
  7. VERIFICATION SUCCEEDED
  8. RECOVERY COMPLETED

This is the central difference between DomainTwin and a simple DNS automation script:

recovery does not end when an API returns success; recovery ends when the provider is independently verified.

Emergency domain continuity

DomainTwin also addresses a more severe failure mode:

What if the original domain cannot be restored quickly enough?

The emergency workflow is:

Search → Check → Preview → Approve → Register → Clone → Verify → Ready

This workflow also uses the Name.com API as its execution plane.

During our controlled Sandbox drill, Emergency Plan #4:

  • selected trusted Snapshot v3
  • used Name.com domain inventory
  • performed the registration workflow
  • created the emergency domain
  • cloned the known-good DNS record
  • re-read the provider
  • calculated the emergency-domain fingerprint
  • compared it with the trusted source

The emergency target reached:

READY

with:

EXPECTED = a3b35ae640...dad54313

ACTUAL = a3b35ae640...dad54313

MATCH YES

The preserved audit trail contains eight events from PLAN CREATED through EMERGENCY DOMAIN READY.

The public judge deployment keeps registration blocked, but the completed Sandbox evidence remains available for inspection.

Why Name.com is central

The Name.com API is not an accessory to DomainTwin — it is the provider execution and verification plane the product depends on.

DomainTwin uses Name.com for:

  • managed-domain discovery
  • reading DNS records
  • DNS state comparison
  • DNS creation
  • DNS updates
  • DNS deletion
  • domain search
  • availability checks
  • Sandbox registration
  • post-operation provider verification

The same provider that executes the operation is queried again afterward so DomainTwin can independently verify the resulting state.

Without the Name.com API, DomainTwin cannot perform its core recovery or emergency-continuity workflows.

How we built it

The product consists of:

  • Next.js
  • TypeScript
  • Django
  • Python
  • Docker
  • Gunicorn
  • Apache
  • SQLite for the hackathon evidence store
  • Name.com API
  • OpenAI Responses API for optional evidence-grounded incident explanation
  • Let's Encrypt for public HTTPS

DomainTwin implements:

  • immutable DNS snapshots
  • explicit known-good baselines
  • normalized DNS fingerprints
  • deterministic DNS diffing
  • health observations
  • deterministic risk scoring
  • incident correlation
  • continuous monitoring
  • evidence-grounded AI explanations
  • multi-tenant organizations
  • RBAC
  • recovery state machines
  • emergency-continuity state machines
  • approval boundaries
  • stale-state protection
  • ordered audit trails
  • provider re-verification

Public deployment

DomainTwin is live at:

https://domaintwin.notificacionesfaronodolabs.com

Public judge credentials:

Username: judge Password: devpost

The judge account has the VIEWER role.

Judges can inspect:

  • the public verified walkthrough
  • Incident #8
  • deterministic risk evidence
  • the persisted AI explanation
  • Recovery Plan #8
  • the exact rollback operation
  • the complete recovery audit trail
  • fingerprint verification
  • Emergency Plan #4
  • emergency registration and clone evidence

The public deployment deliberately blocks mutation and registration permissions.

This lets judges inspect the real control plane without giving an anonymous public account authority over DNS.

Architecture

The deployed path is:

Internet → HTTPS / Apache → Next.js → Django / Gunicorn → Name.com Sandbox

DomainTwin runs in Docker on an Ubuntu VPS.

The Next.js application is reverse-proxied through the existing Apache HTTPS server.

The Django backend remains inside the application network.

Name.com credentials are server-side only and never exposed to browser JavaScript.

A continuous monitoring worker also runs alongside the web application.

Safety by design

A recovery product should never make an incident worse.

DomainTwin therefore fails closed.

  • DNS mutations are disabled by default.
  • Production mutations require a separate explicit permission.
  • Emergency registration has an additional independent permission boundary.
  • Public-demo mutation permissions are disabled.
  • Provider writes require human approval.
  • Recovery plans expose exact operations before execution.
  • Plans are revalidated before mutation.
  • Stale state stops execution.
  • Partial failures are never reported as success.
  • AI has no provider authority.
  • Provider credentials remain server-side.
  • RECOVERED requires a fresh provider read and matching fingerprint.
  • Emergency READY requires the same independent verification principle.

The public deployment therefore shows real historical execution evidence while remaining intentionally read-only.

Challenges we faced

The hardest part was not calling an API.

The hard part was making mutation safe.

DNS may change between preview and execution.

A network request may fail after a provider already accepted an operation.

A recovery may partially complete.

Retrying an apparently failed request can be dangerous if the first request actually succeeded.

We therefore modeled recovery and emergency continuity as explicit state machines instead of individual API buttons.

Preview, approval, execution, and verification are separate boundaries.

Emergency registration additionally uses persisted idempotency to prevent an interrupted workflow from blindly registering a domain twice.

Deployment presented another challenge.

The VPS already hosted other applications and Apache already owned ports 80 and 443.

Rather than disrupting existing infrastructure, DomainTwin binds its frontend to loopback and uses the existing Apache installation as the HTTPS reverse proxy.

We deployed the hackathon application without interrupting the other services on the server.

What we learned

The strongest role for AI in infrastructure recovery is not direct control.

A safer and more useful architecture is:

trusted state + deterministic evidence + AI explanation + human authority + provider execution + independent verification

AI is most valuable when it helps an operator understand reliable evidence rather than replacing that evidence.

We also learned that the Name.com API can power considerably more than domain search.

Search, availability, registration, DNS operations, and provider verification can be combined into a real domain-continuity control plane.

Most importantly:

200 OK is not proof of recovery.

The provider's actual resulting state is.

Progress and real-world viability

The hackathon implementation is not a static prototype.

It includes:

  • a live HTTPS deployment
  • continuous monitoring
  • multi-tenant organization support
  • RBAC
  • deterministic incident generation
  • AI incident explanation
  • approval-gated recovery
  • completed Name.com Sandbox DNS recovery
  • independent fingerprint verification
  • emergency-domain registration and DNS cloning
  • ordered audit evidence
  • a public read-only judge account

The commercial thesis is straightforward:

organizations depend on domains for websites, APIs, authentication, email, and customer access, but domain recovery remains largely provider-console-driven and operationally manual.

DomainTwin could become a provider-agnostic continuity layer across registrars, DNS providers, and cloud platforms.

What's next

The next stage is productization:

  • encrypted per-customer provider credentials or an external secret manager
  • configurable continuity policies
  • notification and escalation channels
  • longer audit retention
  • PostgreSQL-backed production storage
  • organization-level approval policies
  • additional DNS and cloud providers
  • observability
  • billing and subscriptions

Our business model is still a hypothesis rather than validated traction.

What the hackathon demonstrates is the difficult technical core with a real Name.com Sandbox integration:

detect a DNS incident, explain its evidence, recover from trusted state through the provider, independently re-read that provider, and prove the intended state is actually live.

Built With

Share this project:

Updates

Submission history