-
-
CRITICAL DNS incident detected after an unexpected A-record change, with deterministic risk factors and ordered evidence.
-
Read-only recovery preview showing the exact DNS rollback operation before human approval.
-
Live name.com domain search and exact availability check for a safe emergency continuity target.
-
Verified recovery through name.com: expected and actual DNS fingerprints match after rollback.
-
Emergency domain registered and cloned in the name.com sandbox, with DNS fingerprint verification confirming MATCH YES.
Inspiration
A DNS incident can make a website, API, or email service disappear in seconds.
Changing a DNS record is easy. Recovering safely is not.
During an incident, an operator must answer:
- What exactly changed?
- Which DNS state was actually trusted?
- Is the new destination expected?
- What is the minimum safe rollback?
- Who authorized the recovery?
- Did the provider actually end up in the intended state?
Teams often answer those questions using screenshots, tickets, memory, and manual provider-console work.
We built DomainTwin AI to turn DNS recovery into a deterministic, approval-gated, AI-assisted, and independently verified workflow.
Our core idea is simple:
Treat DNS as a recoverable digital twin.
What it does
DomainTwin's primary workflow is:
Detect → Explain → Authorize → Restore → Verify → Prove
It continuously reads live DNS through the Name.com API, stores immutable DNS snapshots, and allows an operator to explicitly designate a trusted known-good state.
When live DNS diverges from that trusted state, DomainTwin calculates a deterministic diff and combines it with health evidence to create an explainable incident.
A real recovery preserved in the live demo
For our controlled Name.com Sandbox test, the trusted state was:
A www → 203.0.113.10
We intentionally introduced drift:
A www → 198.51.100.77
DomainTwin detected the change and created Incident #8.
The incident was scored:
CRITICAL — 75/100
using three explicit deterministic factors:
+30 ADDRESS_RECORD_CHANGED+30 HTTP_HEALTH_FAILED+15 UNKNOWN_DESTINATION
The score does not come from an LLM.
It is reproducible from the evidence.
Evidence-grounded AI explanation
After the deterministic incident exists, DomainTwin can pass a structured evidence bundle to its AI explanation layer.
For Incident #8, the persisted AI analysis was successfully generated and classified the affected service as MULTIPLE with MEDIUM confidence.
The AI explanation references specific evidence IDs including:
DNS-001
HEALTH-DNS
HEALTH-HTTP
HEALTH-HTTPS
RISK-001
RISK-002
RISK-003
RISK-SCORE
The AI is explicitly constrained to explain existing evidence.
It cannot create the incident, change the deterministic risk score, approve recovery, or mutate DNS.
This is intentional:
AI explains. Deterministic systems decide. Humans authorize. DomainTwin verifies.
Verified DNS recovery
From Incident #8, DomainTwin created Recovery Plan #8 containing exactly one deterministic operation:
UPDATE A www
198.51.100.77 → 203.0.113.10
Nothing was changed automatically.
An operator first inspected the exact rollback preview and explicitly approved the recovery.
DomainTwin then executed the update through the Name.com Sandbox API.
But a successful write request is not enough for DomainTwin to declare recovery.
After execution, DomainTwin independently re-read the provider, normalized the live DNS records, calculated the resulting fingerprint, and compared it against trusted Snapshot v3.
The result was:
EXPECTED = a3b35ae640...dad54313
ACTUAL = a3b35ae640...dad54313
MATCH YES
Only after that fresh provider read did Recovery Plan #8 become:
RECOVERED
The complete recovery audit trail contains:
PLAN CREATEDPLAN APPROVEDAPPROVAL ACTOR RECORDEDEXECUTION ACTOR AUTHORIZEDAPPLY STARTEDOPERATION SUCCEEDEDVERIFICATION SUCCEEDEDRECOVERY COMPLETED
This is the central difference between DomainTwin and a simple DNS automation script:
recovery does not end when an API returns success; recovery ends when the provider is independently verified.
Emergency domain continuity
DomainTwin also addresses a more severe failure mode:
What if the original domain cannot be restored quickly enough?
The emergency workflow is:
Search → Check → Preview → Approve → Register → Clone → Verify → Ready
This workflow also uses the Name.com API as its execution plane.
During our controlled Sandbox drill, Emergency Plan #4:
- selected trusted Snapshot v3
- used Name.com domain inventory
- performed the registration workflow
- created the emergency domain
- cloned the known-good DNS record
- re-read the provider
- calculated the emergency-domain fingerprint
- compared it with the trusted source
The emergency target reached:
READY
with:
EXPECTED = a3b35ae640...dad54313
ACTUAL = a3b35ae640...dad54313
MATCH YES
The preserved audit trail contains eight events from PLAN CREATED through EMERGENCY DOMAIN READY.
The public judge deployment keeps registration blocked, but the completed Sandbox evidence remains available for inspection.
Why Name.com is central
The Name.com API is not an accessory to DomainTwin — it is the provider execution and verification plane the product depends on.
DomainTwin uses Name.com for:
- managed-domain discovery
- reading DNS records
- DNS state comparison
- DNS creation
- DNS updates
- DNS deletion
- domain search
- availability checks
- Sandbox registration
- post-operation provider verification
The same provider that executes the operation is queried again afterward so DomainTwin can independently verify the resulting state.
Without the Name.com API, DomainTwin cannot perform its core recovery or emergency-continuity workflows.
How we built it
The product consists of:
- Next.js
- TypeScript
- Django
- Python
- Docker
- Gunicorn
- Apache
- SQLite for the hackathon evidence store
- Name.com API
- OpenAI Responses API for optional evidence-grounded incident explanation
- Let's Encrypt for public HTTPS
DomainTwin implements:
- immutable DNS snapshots
- explicit known-good baselines
- normalized DNS fingerprints
- deterministic DNS diffing
- health observations
- deterministic risk scoring
- incident correlation
- continuous monitoring
- evidence-grounded AI explanations
- multi-tenant organizations
- RBAC
- recovery state machines
- emergency-continuity state machines
- approval boundaries
- stale-state protection
- ordered audit trails
- provider re-verification
Public deployment
DomainTwin is live at:
https://domaintwin.notificacionesfaronodolabs.com
Public judge credentials:
Username: judge
Password: devpost
The judge account has the VIEWER role.
Judges can inspect:
- the public verified walkthrough
- Incident #8
- deterministic risk evidence
- the persisted AI explanation
- Recovery Plan #8
- the exact rollback operation
- the complete recovery audit trail
- fingerprint verification
- Emergency Plan #4
- emergency registration and clone evidence
The public deployment deliberately blocks mutation and registration permissions.
This lets judges inspect the real control plane without giving an anonymous public account authority over DNS.
Architecture
The deployed path is:
Internet → HTTPS / Apache → Next.js → Django / Gunicorn → Name.com Sandbox
DomainTwin runs in Docker on an Ubuntu VPS.
The Next.js application is reverse-proxied through the existing Apache HTTPS server.
The Django backend remains inside the application network.
Name.com credentials are server-side only and never exposed to browser JavaScript.
A continuous monitoring worker also runs alongside the web application.
Safety by design
A recovery product should never make an incident worse.
DomainTwin therefore fails closed.
- DNS mutations are disabled by default.
- Production mutations require a separate explicit permission.
- Emergency registration has an additional independent permission boundary.
- Public-demo mutation permissions are disabled.
- Provider writes require human approval.
- Recovery plans expose exact operations before execution.
- Plans are revalidated before mutation.
- Stale state stops execution.
- Partial failures are never reported as success.
- AI has no provider authority.
- Provider credentials remain server-side.
RECOVEREDrequires a fresh provider read and matching fingerprint.- Emergency
READYrequires the same independent verification principle.
The public deployment therefore shows real historical execution evidence while remaining intentionally read-only.
Challenges we faced
The hardest part was not calling an API.
The hard part was making mutation safe.
DNS may change between preview and execution.
A network request may fail after a provider already accepted an operation.
A recovery may partially complete.
Retrying an apparently failed request can be dangerous if the first request actually succeeded.
We therefore modeled recovery and emergency continuity as explicit state machines instead of individual API buttons.
Preview, approval, execution, and verification are separate boundaries.
Emergency registration additionally uses persisted idempotency to prevent an interrupted workflow from blindly registering a domain twice.
Deployment presented another challenge.
The VPS already hosted other applications and Apache already owned ports 80 and 443.
Rather than disrupting existing infrastructure, DomainTwin binds its frontend to loopback and uses the existing Apache installation as the HTTPS reverse proxy.
We deployed the hackathon application without interrupting the other services on the server.
What we learned
The strongest role for AI in infrastructure recovery is not direct control.
A safer and more useful architecture is:
trusted state + deterministic evidence + AI explanation + human authority + provider execution + independent verification
AI is most valuable when it helps an operator understand reliable evidence rather than replacing that evidence.
We also learned that the Name.com API can power considerably more than domain search.
Search, availability, registration, DNS operations, and provider verification can be combined into a real domain-continuity control plane.
Most importantly:
200 OK is not proof of recovery.
The provider's actual resulting state is.
Progress and real-world viability
The hackathon implementation is not a static prototype.
It includes:
- a live HTTPS deployment
- continuous monitoring
- multi-tenant organization support
- RBAC
- deterministic incident generation
- AI incident explanation
- approval-gated recovery
- completed Name.com Sandbox DNS recovery
- independent fingerprint verification
- emergency-domain registration and DNS cloning
- ordered audit evidence
- a public read-only judge account
The commercial thesis is straightforward:
organizations depend on domains for websites, APIs, authentication, email, and customer access, but domain recovery remains largely provider-console-driven and operationally manual.
DomainTwin could become a provider-agnostic continuity layer across registrars, DNS providers, and cloud platforms.
What's next
The next stage is productization:
- encrypted per-customer provider credentials or an external secret manager
- configurable continuity policies
- notification and escalation channels
- longer audit retention
- PostgreSQL-backed production storage
- organization-level approval policies
- additional DNS and cloud providers
- observability
- billing and subscriptions
Our business model is still a hypothesis rather than validated traction.
What the hackathon demonstrates is the difficult technical core with a real Name.com Sandbox integration:
detect a DNS incident, explain its evidence, recover from trusted state through the provider, independently re-read that provider, and prove the intended state is actually live.
Log in or sign up for Devpost to join the conversation.