Inspiration
Every Shopify merchant eventually learns the same lesson: the store doesn't stop while you sleep. A discount code stacks with another one and quietly bleeds margin overnight. A theme edit breaks the buy button. A tracking pixel silently disappears and a week of ad spend goes unmeasured. Enterprise brands solve this with a real ops team watching dashboards around the clock. The other 99% of Shopify's 5+ million merchants just wake up and find out.
We wanted to build the thing that actually closes that gap — not another alerting dashboard that emails you at 3am, but something that behaves like a real overnight employee: it notices, it reasons about whether it's safe to act, it either fixes the problem or asks first, and it proves the fix actually worked before it tells you it's done.
What it does
NightShift AI is an autonomous AI workforce embedded directly in Shopify Admin. Four specialist agents each own one category of overnight risk:
- Product Quality — catches missing images, missing ALT text, and thin product descriptions; generates ALT text automatically (fully reversible, zero risk) and proposes description rewrites for a merchant to approve.
- Checkout Specialist — detects overlapping, stackable storewide discount codes before they compound into a real pricing bug, and deactivates the duplicate.
- Theme Guardian — snapshots a store's critical theme files and detects real content drift against the last known-good baseline, explains the change in plain English via Gemini, and either performs a genuine automated restore or hands the merchant an exact guided fix.
- Tracking Specialist — notices when a previously-installed tracking script (a Meta Pixel, GA, etc.) silently disappears from the storefront and recreates it from its own snapshot.
Every finding, regardless of which specialist made it, moves through the same trust pipeline: Observe → Reason → Assess Risk → Approve (if required) → Execute → Verify → Explain → Persist. Nothing is auto-executed unless its risk classification says it's genuinely safe; everything else waits in a real Approval Center for a human decision, with the actual revenue impact and AI confidence score shown alongside it.
Every morning, Chief Ops AI — one Gemini call through Vertex AI — reads every specialist's real findings from the night and writes the single executive paragraph a merchant actually needs to read, correlating findings across agents when they're genuinely related.
How we built it
The backend is a FastAPI + Celery + PostgreSQL + Redis service, cleanly layered (domain logic has zero framework imports, so risk assessment and detection rules are unit-testable without a database or an HTTP client in sight). The frontend is a Next.js app embedded in Shopify Admin via App Bridge. Every Gemini call — both the per-specialist detection reasoning and Chief Ops AI's executive synthesis — goes through Google's GenAI SDK, with Vertex AI as the production path so usage bills through normal Google Cloud billing rather than a separate API-key balance.
We built this test-first and fact-first: 356 backend tests cover every specialist, every API endpoint, and every safety mechanism (rollback, verification, budget guards, HMAC-verified webhooks), and every number the product ever shows a merchant — a dollar figure, a confidence score, a timestamp — traces back to something Shopify's own API actually returned or a rule the code can point to. Where the data doesn't support a claim, the product says so instead of guessing.
Challenges we ran into
Migrating providers under a hard deadline. Partway through the build we needed to consolidate everything onto Gemini for this submission — including a live-corrupted demo-data bug we only found by actually running the scenario end-to-end: a "create → snapshot → delete" demo trigger wasn't atomic, so an interrupted run left orphaned duplicate records that silently made detection look broken. We found it by tracing real API responses rather than trusting the code's own success log line.
Cloud Build vs. local Docker. Our Dockerfiles built fine locally but failed instantly on Cloud Build — two BuildKit-only features (--platform=$BUILDPLATFORM and an optional-file COPY glob) that silently behave differently under Cloud Build's classic builder. Found and fixed by actually deploying, not just reading the Dockerfile.
Gemini's two separate billing worlds. The Gemini Developer API draws from a separate "Prepay" credit balance, completely distinct from a project's normal Google Cloud billing — discovered mid-recording when a real demo beat failed with RESOURCE_EXHAUSTED despite the GCP project having credit. Vertex AI bills through normal Cloud Billing instead, which is the actual production path for this reason.
Never faking a capability. The hardest discipline was refusing to demo something the product can't actually do. When we wanted a scenario where a "broken theme" gets fixed live, we had to accept that our own trust model deliberately never lets NightShift silently overwrite a theme file without a Shopify-granted exemption — so the honest demo is the guided-restore path, not a fabricated one-click autonomous fix.
Accomplishments that we're proud of
- A trust pipeline that treats "safe to fix automatically" as a real risk classification, not a marketing claim — verified with actual re-fetches from Shopify's API, never just trusting a mutation response.
- Chief Ops AI's executive briefing is a real, auditable Gemini call over real findings every single time it appears, with a structured log line proving it happened.
- Zero fabricated data anywhere in the product — revenue estimates, confidence scores, and health scores are either directly Shopify-reported or explicitly grounded in the specific finding they describe.
What we learned
Building something that has to be trusted by a merchant with real money on the line is a different discipline than building something that just works in a demo. The constraint we kept returning to — never show a number, a claim, or a fix you can't back up — made the product slower to build but is exactly what makes it defensible in front of a real merchant, not just a judge.
What's next for NightShift AI
Today NightShift covers four categories of overnight risk. The underlying pattern — detect, reason, ask only when it matters, verify, explain — generalizes to nearly every operational decision a Shopify store makes: fulfillment exceptions, customer service triage, inventory rebalancing, marketing spend pacing. There are over 5 million Shopify merchants and almost none of them can afford a real ops team watching the store around the clock. That's the gap NightShift is built to close — not another alert that emails you at 3am, but an AI employee that actually handles it.
Built With
- alembic
- celery
- cloud-build
- docker
- fastapi
- gemini
- google-cloud
- google-cloud-run
- google-genai-sdk
- graphql
- nextjs
- opentelemetry
- postgresql
- pydantic
- pytest
- python
- react
- redis
- shopify
- shopify-app-bridge
- sqlalchemy
- structlog
- tailwindcss
- typescript
- vertex-ai
Log in or sign up for Devpost to join the conversation.