The call that comes too late

A small accounting firm does not have a site reliability team. It has an office manager, a bookkeeper, and a phone. When the invoicing system stops capturing transactions at 2 a.m., nobody notices until a client calls at 9 a.m. and asks why their payroll file never showed up. By then the damage is hours old, and the first person to learn about it is the customer.

That is the whole problem in one sentence: businesses find out about failures from the people they serve, not from the software they run.

We wanted to flip that. Give every application an early-warning system that speaks up before impact, and give Zikit the machinery to do something about it without waking a developer at 3 a.m. The hard part is not catching everything. It is catching the right things. An alert that fires on every nightly backup and payroll batch is an alert nobody reads. The value lives in the signal, not the noise.

Two acronyms, one platform

The system splits cleanly in two, and the names are not decoration.

AEGIS runs locally, at each client company. It is the Automated Environmental Guard & Integrity Service. It receives telemetry from the client's web apps, stores the routine traffic on-site, watches for anomalies, and forwards the things that matter: error reports and early-warning predictions.

NEPHOS runs in the cloud, owned and operated by Zikit. It is the Networked Event Processing & High-availability Orchestration System. It takes what AEGIS sends, adds the push events from GitHub, and runs a single pipeline: notify, analyze, propose a fix, get a human yes, validate, and ship.

Local guard, central brain. The boundary between them is a signed HTTP contract and nothing else, which is what keeps the two teams and the two codebases honest.

Inside AEGIS

AEGIS architecture: client web apps send telemetry to a local ingest API, which routes INFORMATIVE data to a local database and forwards ERROR reports and predictions to NEPHOS. An Early-Warning Sentinel reads the local database; an admin frontend serves dashboards and a Gemini-powered chatbot.

AEGIS lives on the client's premises, and it has to fit whatever the client already runs. That is why its database layer speaks to SAP HANA, Microsoft SQL Server, or PostgreSQL. We are not asking a business to migrate their data to us. We are asking for a socket.

Every instrumented web app emits two classes of telemetry. INFORMATIVE events are the everyday pulse: requests served, jobs finished, usage counts. Those stay local, in the client's own database, where the admin frontend turns them into dashboards and graphs. ERROR events take a different road. The telemetry router forwards them to NEPHOS the moment they happen.

Sitting on top of the local data is the Early-Warning Sentinel, and this is where the "before the phone rings" idea actually lives. It watches the trends, learns what a normal Tuesday looks like for that specific business, and predicts trouble before it turns into an outage. A slow climb in latency. An error rate that creeps up after a deploy. A disk filling faster than the retention window assumes. Those are the whispers worth hearing.

The client also gets a frontend with a chatbot, running on Gemini. Staff can ask plain questions about their own system and get answers grounded in the local context, without reading a single log file.

The bridge

Everything AEGIS forwards crosses the public internet, so the edge is where we start taking security seriously. Cloudflare terminates TLS, filters traffic through its WAF, and absorbs DDoS attempts before they reach a single NEPHOS process.

The reports themselves are not trusted just because they arrived. Each one carries an HMAC-SHA256 signature computed over the raw body with a per-tenant secret, sent in X-Nephos-Signature. NEPHOS recomputes it and compares in constant time. A report that does not match is rejected, logged, and never enters the pipeline. GitHub pushes use the same idea through X-Hub-Signature-256. Two doors, two keys, one rule: prove you are who you say you are.

Inside NEPHOS

NEPHOS architecture: AEGIS and GitHub webhooks enter through Cloudflare into an API gateway, which publishes to RabbitMQ. Five modules (ERROR_REPORT_SENTINEL, SECURITY_AUDITOR, LLM_BRAIN, CODER_HOTFIXER, COMMUNICATIONS) consume events and share TigerData as the sole database and Gemini as the LLM.

NEPHOS is the part that does the thinking. It runs on Vultr as a set of containers, and everything moves through RabbitMQ. The gateway verifies signatures and publishes the event; the modules subscribe to what they care about. Nothing calls anything directly. That choice is deliberate. A slow LLM call cannot stall the ingest path, and a module can be restarted or scaled without the rest of the system noticing.

Five modules carry the work.

  • ERROR_REPORT_SENTINEL owns the runtime errors coming from AEGIS. It fingerprints them, deduplicates repeats, and decides what deserves attention.
  • SECURITY_AUDITOR owns the GitHub side. When code is pushed, it reads the changed files and looks for problems a human reviewer might miss at 5 p.m. on a Friday.
  • LLM_BRAIN is the reasoning core. It takes an incident, pulls the relevant source, and asks Gemini for a root cause and a proposed patch.
  • CODER_HOTFIXER takes that patch and does something with it: clone, apply, test, commit, merge.
  • COMMUNICATIONS is the human interface. It pushes the incident to Telegram and Slack and listens for the reply that says go or stop.

Underneath all of them sits TigerData, the only database in the system, holding incidents, their full event history, approvals, and the integration map from app to repository.

One failure, end to end

Full architecture: client apps feed AEGIS, which forwards errors and predictions through Cloudflare to NEPHOS on Vultr, where RabbitMQ distributes work across modules backed by TigerData and Gemini; approved hotfixes are committed to GitHub, whose Actions redeploy the application.

Follow a single incident from left to right.

A web app throws an unhandled exception. AEGIS captures it, signs the report, and sends it to NEPHOS. Cloudflare passes it through, the gateway verifies the signature, and RabbitMQ hands it to the sentinel. The sentinel recognizes the error pattern, opens an incident, and stores it in TigerData. LLM_BRAIN reads the stack trace, maps it to the source file, and asks Gemini for a diagnosis and a fix, returned as a unified diff. Communications posts the summary to the team's channels with a question attached. A developer replies YES. The hotfixer clones the repo, applies the diff, runs the project's own build and test commands, and only if those pass does it commit to a branch and merge to main. GitHub Actions picks up the merge and redeploys the app. The fix reaches production before the client ever noticed the failure.

The incident lifecycle

Every incident moves through an explicit state machine, and every transition is written to the database. There is no in-memory-only state that vanishes on a restart.

NEW → ANALYZING → AWAITING_APPROVAL → APPROVED → PATCHING → VALIDATING → COMMITTED → MERGED → RESOLVED

A NO at the approval step lands in REJECTED and closes cleanly. A failed build or test lands in FAILED with the reason attached, and the incident stays open for a person to look at. A malformed report from AEGIS never enters the flow at all: it becomes a MALFORMED_PAYLOAD incident, the raw payload is stored, and Zikit's own team gets alerted, because that is a bug in the integration, not a customer's problem.

The state machine is what makes the system auditable. If someone asks why a patch went out, the answer is a row, not a recollection.

Why this stack

We did not pick these tools to decorate a logo slide. Each one solves a problem the architecture creates.

Vultr is the home. NEPHOS is a long-running service that receives traffic from many clients and cannot afford to be down. Vultr gives us compute we control, close to the clients, with containers that deploy the same way in staging and production. One region today; more regions and a load balancer as the client list grows. The event backbone already assumes multiple workers, so scaling out means adding machines, not rewriting code.

TigerData is the memory. Everything NEPHOS knows is time-ordered. Errors arrive minute by minute, incidents move through states, approvals are timestamped. That is the shape of data a time-series database is built for. TigerData, built on TimescaleDB and speaking PostgreSQL, lets us store the raw event stream as hypertables, compress older data, and set retention policies without throwing away anything that matters. The same engine also holds the ordinary relational tables, so there is one database to run, one to back up, and one to reason about.

Gemini is the brain. Reading a stack trace, finding the offending function, and writing a patch that actually applies is a reasoning task, and it needs a model that can hold a lot of context at once. Gemini works across the files involved, produces a structured root-cause analysis, and returns a unified diff we can validate like any other patch. It also powers the audit of pushed code and the chatbot inside AEGIS. One model, several jobs, behind a provider-agnostic interface, so the LLM layer never becomes a single point of lock-in.

Safety rails

Giving an AI the ability to write to main is a serious thing, so the design assumes it will sometimes be wrong.

No patch merges without a human YES. The diff is applied and validated against the project's real build and test commands before anything is committed, so a broken suggestion fails loudly instead of reaching users. NEPHOS's own commits are tagged, and pushes carrying those tags are ignored, which stops the auditor from auditing its own fixes in a loop. Data sent to the model is minimized to what the analysis needs. And if Gemini or a messaging provider goes down, the incident is still recorded and a deterministic message still goes out. The pipeline bends; it does not break.

Two stories, one pipeline

A runtime crash shows the healing path: an error from a client app becomes a diagnosed incident, a proposed fix, a human approval, a validated merge, and an automatic deploy, with no support ticket in between.

A pushed vulnerability shows the preventive path: a developer commits code with an injection flaw, the auditor catches it on the push, the team gets a clear finding with the file and line, and approves a fix that lands the same way. Same pipeline, different trigger.

What we are building

The phone will still ring. Clients will still call. But the call should be about a feature they want, not an outage they found first.

AEGIS watches from inside every application and speaks up early. NEPHOS listens, understands, and quietly ships the fix. Vultr keeps it running, TigerData remembers everything, and Gemini does the reading no one has time for.

That is a small business with the reflexes of a platform team. That is the point.

Share this project:

Updates

Submission history