Inspiration
Traditional vulnerability management produces a long list of findings. Teams get alert fatigue, and the findings that matter get buried. A severity score alone doesn't tell a team what to fix first. We wanted to show how exploit intelligence, asset context and business impact can change that decision, and to make the reasoning visible instead of presenting another opaque score.
What it does
Eclipse CTEM walks an analyst through the five-stage CTEM lifecycle: Scope → Discover → Prioritize → Validate → Mobilize.
- Scope shows the organization's assets and their declared context: environment, criticality, internet exposure and data sensitivity.
- Discover lists three real CVEs in the order a traditional scanner would rank them, by CVSS:
- Apache HTTP Server 2.4.49 path traversal (CVE-2021-41773)
- Apache HTTP Server 2.4.50 path traversal (CVE-2021-42013)
- Apache Struts 2 Jakarta Multipart RCE (CVE-2017-5638)
- Prioritize re-ranks the findings with a CTEM priority score out of 100. The score has four categories: Threat (EPSS and CISA KEV, 30 points), Exposure (25), Business Impact (25) and Technical Severity (CVSS, 20).
- The Apache 2.4.49 finding has the lowest CVSS score (7.5) but moves from #3 to #1 with a priority of 95. It sits on an internet-facing, business-critical production asset.
- Struts (CVSS 9.8) lands at 79.6 because it is not internet-facing.
- Apache 2.4.50 (CVSS 9.8) drops to 58.6 because it runs on a low-criticality development server.
- A radial score breakdown and a CVSS-to-CTEM rank-shift chart show exactly why each finding moved.
- Validate runs one approved, non-destructive check against each finding's service in a local Docker lab.
- The Apache checks read a controlled marker file outside the web root.
- The Struts check has the server evaluate fixed arithmetic and return the result (5421) in a response header. It runs no operating-system commands.
- The dashboard shows the evidence, the time it was observed, the template used and the limits of what the check proves.
- Mobilize uses Gemini to pick from a reviewed catalog of remediation actions and produce guidance for the finding. The guidance has sections for remediation, mitigation, implementation, verification, rollback and reasoning, plus linked sources. A notification flags that the plan needs human review before anyone acts on it.
Every value carries a label: verified intelligence, declared scenario context or live lab evidence. Missing intelligence is shown as unknown, and the score is withheld instead of guessed.
How we built it
- Frontend: Next.js 16, React 19, TypeScript and Tailwind CSS v4. The charts are hand-built SVG with no chart library. Row re-ranking and scrolling are animated with the browser's Web Animations API, and the dashboard respects reduced-motion settings. It can also run on mock data, so the frontend could be built before the backend was ready.
- Backend: Python with FastAPI and Uvicorn, bound to localhost only.
- Lab: Docker Compose runs three deliberately vulnerable Vulhub containers on an isolated bridge network. Only one has a host port.
- Validation: The validate endpoint accepts only a finding ID. The backend maps that ID to one fixed Nuclei template and one lab service, confirms the container's Compose labels and network, then runs ProjectDiscovery's Nuclei in a container. Validation can't be pointed at an arbitrary target or template. A match counts as confirmed only if the expected proof appears in the response.
- Intelligence: We pull dated EPSS scores from FIRST and the full CISA KEV catalog into a cached snapshot. If a refresh fails, the previous snapshot is kept, and values that were never retrieved stay unknown.
- Scoring: A transparent, category-based prototype model (
ctem-demo-v0.1). It is explicitly not a calibrated risk model, and validation results never change the score. - Remediation: Gemini selects and orders actions from an approved catalog grounded in vendor advisories. The backend checks every selected action ID before displaying the plan, and nothing is executed automatically.
Challenges we ran into
The hard part was connecting a compelling demo to evidence we could describe honestly. We replaced synthetic comparison findings with real CVEs. We built separate safe checks for Apache and Struts. And we kept declared business context separate from what the lab actually observed. Integrating a redesigned dashboard with the three-finding backend also meant keeping each finding's different validation evidence intact: a file read for Apache, a computed header for Struts.
We also worked through Git merge conflicts between parallel frontend and backend branches, Windows-specific issues (PowerShell script policies and dev-server setup), and getting the Gemini API key into the right backend process for the live demo. Smaller details mattered too. Displaying an EPSS score of 0.99992 as "100%" would have overstated certainty, so we show it as 99.992%.
Our biggest late-stage challenge was moving from localhost to a public website. The full system runs on a teammate's laptop: the FastAPI backend, the Docker lab and the Gemini API key all live there. When we deployed the dashboard to Vercel, the hosted site couldn't reach that local backend, so it fell back to its built-in mock mode. Prioritization and the validation flow still display using mock data that mirrors the real results, but Gemini remediation can't run, because it only works through the backend. Making the site fully live means exposing the backend securely over HTTPS, allowing the site's address in the backend's cross-origin (CORS) settings, and pointing the deployed site at the backend.
Accomplishments that we're proud of
- Built an end-to-end flow from scoping through prioritization, live validation and AI-assisted remediation.
- Confirmed all three vulnerabilities in the controlled lab with fixed, approved checks. There's no arbitrary scanning, and the Struts check runs no operating-system commands.
- Made the re-ranking explainable through score breakdowns, a visual rank comparison and evidence shown next to each finding.
- Kept unknown intelligence, declared context and live validation clearly labeled, so the dashboard never claims more than it knows.
What we learned
A high CVSS score doesn't always mean a finding belongs first in the queue: in our demo, the lowest-CVSS finding was the most urgent. We also learned that evidence needs careful boundaries:
- EPSS estimates the likelihood of exploitation, not the likelihood of compromise.
- A lab check confirms only what it observed.
- An AI-generated plan needs reviewed sources and constrained actions.
Building those distinctions into the product made the demo more useful and more credible.
What's next for Eclipse CTEM
- From recommendations to action: integrations with EDR/XDR and patch or configuration-management tools, so analysts can review and approve targeted mitigation steps. Scheduled workflows could then run pre-approved actions and re-validate to confirm the fix.
- Real discovery: importing findings from vulnerability scanners and asset inventories instead of a prepared lab.
- Configurable scoring weights per organization.
- Notifications when a finding crosses a priority threshold or validation confirms an exposure.
Log in or sign up for Devpost to join the conversation.