-
-
Continuous HTTP 200 OK, confirming active EC2 workload availability
-
Live terminal split-screen demonstrating instant traffic drop to connection timeout upon threat payload execution
-
AWS Console showing detachment from the workload firewall and binding to 0-ingress/0-egress security group
-
Point-in-time forensic EBS snapshot created by Lambda with incident ID for threat identifying.
-
CloudWatch execution telemetry confirming containment, snapshotting, and notification completed in under 2 seconds.
-
Amazon SNS automated email alert delivering structured forensics data directly to security engineers.
-
Operational reset restoring the baseline workload security group and verifying return to HTTP 200 OK.
Inspiration
In modern enterprise cloud environments, defensive failure can often stem from MTTR (Mean-Time-To-Respond) latency gap. Security operation centers spend anywhere from 15 minutes to multiple hours performing manual alert triage, ticket escalations, and console navigation. During this time, an adversary can execute an SSH brute-force or credential access attack in seconds, already exfiltrating sensitive data before any mediation.
What it does
Cloud SIEM & Serverless SOAR continuously monitors cloud compute infrastructure, detects active threats, and autonomously executes a three-stage containment and forensic preservation protocol:
- Intelligent Detection: Amazon GuardDuty analyzes VPC flow logs, DNS queries, and operating system access events to identify high-severity threats.
- Real-Time Event Routing: Amazon EventBridge evaluates incoming threat telemetry against strict JSON event patterns, immediately triggering the serverless response engine without polling delays.
- Sub-Second Autonomous Containment: An AWS Lambda remediation engine (Python 3.12 / Boto3) executes three actions simultaneously:
- Zero-Trust Network Isolation: Strips the active workload security group and attaches an isolated quarantine security group (0 ingress / 0 egress rules), cutting attacker command-and-control (C2) and active sessions instantly without rebooting the instance.
- Forensic Snapshotting: Queries attached Amazon EBS storage volumes and takes immediate, point-in-time forensic snapshots tagged with
ForensicLocked = True, the incident ID, and threat type to preserve non-volatile digital evidence. - SOC Alerting: Publishes a structured incident summary payload to an encrypted Amazon SNS topic, delivering actionable threat forensics directly to on-call security engineers via email.
How we built it
- Infrastructure as Code (IaC): The complete cloud infrastructure—modular VPC, compute workloads, security groups, IAM roles, EventBridge rules, SNS topics, and GuardDuty detectors—was codified in modular HashiCorp Terraform.
- Remediation Engine: Built a lightweight Python 3.12 Lambda function leveraging
boto3to interact with EC2, EBS, and SNS APIs with strict idempotency and zero external dependencies. - Local DevSecOps & Unit Testing: Developed locally in an OrbStack Linux environment on macOS using VS Code Remote-SSH. Function logic was validated offline using
pytestandmototo mock AWS service interactions before deploying to real infrastructure. - Static Security Analysis: Enforced CIS AWS Foundations Benchmarks across our Terraform modules using Checkov, resolving 20 baseline security failures to achieve 49 passed checks and 0 failures.
- CI/CD Automation: Implemented GitHub Actions to run automated Python unit tests,
terraform fmt -check,terraform validate, and static Checkov scans on every push.
Challenges we ran into
- EventBridge Payload Authorization: While simulating high-severity GuardDuty findings via the AWS CLI, we encountered
NotAuthorizedForSourceException. We discovered that EventBridge strictly forbids user-supplied payloads using the reservedaws.namespace (such asaws.guardduty). We engineered a deterministic test harness dispatching valid schema envelopes to validate the SOAR engine with exact millisecond precision. - Free Tier Lambda Concurrency Constraints: Checkov rule
CKV_AWS_115flags Lambda functions lacking reserved concurrency. However, applying reserved concurrency limits on our AWS Free Tier account caused deployment errors due to account-level concurrency pool depletion. We addressed this by omitting the hard limit and documenting an explicit architectural skip comment explaining the trade-off. - Web Server Permissions Under SELinux: During split-screen traffic testing, Apache returned
403 Forbiddenover port 80 despite correct network routing. We leveraged AWS Systems Manager (SSM) Run Command to diagnose Amazon Linux 2023 filesystem contexts and executerestoreconwithout needing an insecure, open SSH port 22.
Accomplishments that we're proud of
- Sub-Second MTTR: Achieved an end-to-end incident containment time of < 500 ms (measured via CloudWatch Lambda execution telemetry), representing a 99.9% reduction in remediation time compared to manual SOC triage.
- Flawless Checkov Audit: Completed an enterprise-grade security audit with 49 Passed checks, 0 Failed checks, and 15 documented architectural skips.
- Zero-Ingress Management Plane: Built a workload with zero open inbound ports (no port 22 SSH) by managing instances purely through AWS Systems Manager (SSM) Session Manager.
- Zero Cloud Cost Architecture: Deployed a production-patterned SIEM/SOAR system entirely within AWS Free Tier limits, avoiding costly managed NAT gateways and third-party SaaS licenses.
What we learned
- Event-Driven vs. Polling Architectures: EventBridge decoupled rule matching provides near-instantaneous execution compared to traditional log scrapers or cron-based polling agents.
- Boto3 Error Handling & Idempotency: Designing incident response scripts requires defensive programming—handling non-standard volume attachments, missing instance metadata, and ensuring failed snapshot attempts do not block network isolation.
- Shift-Left Security: Catching misconfigurations (such as unencrypted EBS volumes or over-permissive IAM roles) using Checkov in local development prevents costly cloud remediation cycles post-deployment.
What's next for Cloud SIEM & Serverless SOAR: Event-Driven Incident Response
- Volatile Memory Dump Automation: Integrate AWS Systems Manager Run Command to capture volatile RAM state (using LiME / AVML) prior to instance shutdown or isolation.
- Automated Reverse Proxy / Honeypot Redirection: Instead of dropping traffic completely, dynamically reroute malicious IP connections to a containerized honeynet to collect adversary TTPs (Tactics, Techniques, and Procedures).
- Multi-Account AWS Organizations Fan-Out: Scale the architecture into a centralized Security Operations account that consumes GuardDuty findings across member accounts using AWS Organizations and EventBridge cross-account bus routing.
Built With
- amazon-ebs
- amazon-ec2
- amazon-sns
- amazon-web-services
- aws-lambda
- boto3
- checkov
- cloudwatch
- devsecops
- eventbridge
- github-actions
- guardduty
- incident-response
- moto
- orbstack
- pytest
- python
- soar
- systems-manager
- terraform
Log in or sign up for Devpost to join the conversation.