Systema

Learn system design by operating systems, not memorizing diagrams.

Systema is an interactive systems-design laboratory built around a smart-warehouse digital twin. Learners make architecture decisions, operate the resulting system under deterministic load and failure, inspect real metrics, receive evidence-bound AI coaching, revise the design, and compare the outcome.

The product implements the complete MVP 1, MVP 2, and MVP 3 learning journey in a local-first application that does not require Docker, cloud infrastructure, or an AI API key.

Contents

Product experience

The core learning loop is:

Choose architecture
        ↓
Run a deterministic scenario
        ↓
Observe events and operational metrics
        ↓
Receive evidence-bound coaching
        ↓
Revise the architecture
        ↓
Rerun the same seed and compare
        ↓
Review mastery evidence
        ↓
Start an adaptive follow-up scenario

A learner can:

  1. open Systema without creating an account;
  2. choose Reliable Order Processing or Peak-Season Surge;
  3. select a beginner, intermediate, or advanced coaching level;
  4. answer a diagnostic question about retry ambiguity;
  5. configure queueing, workers, retries, idempotency, reservation, circuit breaking, and dead-letter handling;
  6. run a seeded simulation;
  7. watch persisted events replay through an animated architecture diagram;
  8. inspect completed and failed orders, duplicates, overselling, latency, throughput, recovery, queue depth, and complexity;
  9. receive structured coaching grounded in those results;
  10. apply an evidence-based revision and rerun the same seed;
  11. inspect before-versus-after results and an eight-dimension mastery report;
  12. download the report or begin the generated adaptive scenario.

Every visible demo action is functional. There are no placeholder buttons in the golden flow.

IncludAI K-12 Inclusive Learning Portal

IncludAI extends Systema into a production-quality, AI-powered K-12 inclusive education ecosystem designed for neurodiverse learners (ADHD, Autism, Dyslexia, Dyscalculia, Auditory Processing, Executive Function).

Key Features:

  1. Digital Learning Twin: A continuously evolving student profile mapping grade levels, math/reading scores, learning velocity, and neuro-accessibility configurations.
  2. Multi-Agent AI Router: A backend router that directs user queries to specialized agents:
    • Tutor Agent: Socratic concept helper.
    • Teacher Copilot: Lesson structure, worksheet, and MCQ quiz generator.
    • Parent Coach: Progress summarizes and homework suggested activities.
    • Wellbeing Agent: Fatigue levels tracker and break reminders helper.
  3. Explainable AI (XAI): Interactive SHAP-like local feature contribution bar charts explaining exactly why a pathway or risk prediction was calculated.
  4. Neurodiverse Adaptive UI Toolbar: Real-time styling modifiers:
    • OpenDyslexic Font & Letter Spacing: Clean custom letter styling and line spacing rules.
    • ADHD Focus Shield: Visually dims background page contents to isolate focus on the active task card.
    • Low Sensory Theme / Sepia: Reduces bright visual distractions for autistic and sensory-sensitive students.

Implemented challenges

Challenge 1: Reliable Order Processing

This challenge demonstrates retry ambiguity, overload, partial failure, consistency boundaries, and recovery.

Property Value
Scenario ID reliability-traffic-spike-v1
Deterministic seed 4242
Duration 120 simulated seconds
Baseline traffic 40 orders/minute
Peak traffic 190 orders/minute from second 30 to 80
Injected failure Fulfilment unavailable from second 45 to 53
Throughput target At least 100 orders/minute
Correctness target Zero duplicates and zero oversold units

The golden flow begins with:

  • queue disabled;
  • two workers;
  • three retries;
  • idempotency disabled;
  • inventory reservation disabled;
  • circuit breaker disabled;
  • dead-letter queue disabled.

The deterministic baseline currently produces:

Metric Baseline
Orders received 212
Orders completed 133
Failed orders 79
Duplicate fulfilments 12
Oversold units 8
Average latency 293 ms
Throughput 66.5 orders/minute
Error rate 37.3%

The recommended revision enables queueing, idempotency, inventory reservation, circuit breaking, dead-letter handling, and four workers. With the same seed it currently produces:

Metric Revised
Orders received 212
Orders completed 212
Failed orders 0
Duplicate fulfilments 0
Oversold units 0
Average latency 2,948 ms
Throughput 106 orders/minute
Error rate 0%

The higher latency is intentional evidence: buffering protects work but makes queue wait visible. Systema teaches that reliability mechanisms introduce costs rather than presenting a universally “correct” architecture.

Challenge 2: Peak-Season Surge

This challenge tests whether the learner can transfer their reasoning to a different workload.

Property Value
Scenario ID peak-season-surge-v1
Adaptive seed 7331
Duration 120 simulated seconds
Baseline traffic 60 orders/minute
Peak traffic 240 orders/minute from second 25 to 85
Injected service outage None
Throughput target At least 140 orders/minute
Primary focus Queue growth, worker scaling, and saturation

Worker processing is capped at the equivalent of eight workers to model a database-contention ceiling. Adding workers therefore stops providing unlimited linear throughput.

Architecture controls

Control Behavior represented by the simulator
Event queue Accepts and buffers work when downstream processing cannot keep up
Worker count Changes per-second processing capacity
Retries Creates recovery attempts and possible duplicate delivery
Idempotency keys Suppresses duplicate fulfilment side effects
Inventory reservation Makes the stock check and decrement atomic
Circuit breaker Emits fail-fast evidence during fulfilment failure
Dead-letter queue Retains undeliverable direct-call work for later processing
Retry backoff Validated architecture input retained for retry-policy experiments

The simulator reports operational complexity so that adding resilience mechanisms is never presented as free.

Metrics and evidence

The Go simulation engine derives:

  • orders received;
  • orders accepted;
  • orders completed;
  • failed orders, calculated as received minus completed;
  • duplicate fulfilments;
  • oversold inventory units;
  • average latency;
  • p95 latency;
  • throughput per minute;
  • error rate;
  • recovery time;
  • maximum queue depth;
  • operational complexity.

It also emits ordered domain events such as:

  • ORDER_RECEIVED;
  • ORDER_COMPLETED;
  • ORDER_REJECTED;
  • SERVICE_UNAVAILABLE;
  • CIRCUIT_OPEN;
  • MESSAGE_DEAD_LETTERED;
  • DUPLICATE_DELIVERY;
  • DUPLICATE_SUPPRESSED;
  • DUPLICATE_FULFILMENT;
  • INVENTORY_OVERSOLD.

Java persists the complete result and the exact architecture used for that run. The browser receives a Server-Sent Event replay from Java, so the animated view and text timeline are derived from persisted evidence.

AI coaching

The complete default AI mode is deterministic and offline. It requires no provider credentials.

The Python learning engine:

  • validates requests and responses with Pydantic;
  • treats learner text as untrusted data;
  • detects retry-without-idempotency misconceptions;
  • detects inventory-consistency misconceptions;
  • recognizes queue, idempotency, and reservation strengths;
  • includes actual simulation values in feedback;
  • asks one remediation question;
  • changes guidance depth for beginner, intermediate, and advanced learners;
  • proposes eight structured rubric dimensions;
  • generates a fixed-seed adaptive follow-up.

Example feedback from the flawed golden design:

Retries attempted recovery, but because idempotency was disabled, 12 operations were processed twice.

Python is advisory. It cannot update sessions, alter simulation metrics, or award final mastery independently.

Mastery and adaptation

After the revised run, Java creates a transparent 32-point mastery report:

Dimension Maximum Evidence authority
Requirement identification 4 Structured language evidence
Reliability reasoning 4 Persisted design and run metrics
Consistency reasoning 4 Reservation, idempotency, and correctness metrics
Scalability reasoning 4 Workers and before/after throughput
Failure handling 4 Retry, queue, circuit breaker, and DLQ choices
Observability 4 Learner use of measured evidence
Trade-off explanation 4 Learner explanation of benefit and cost
Evidence-based revision 4 Same-session before/after comparison

Each dimension includes:

  • a score;
  • the evidence used;
  • a targeted recommendation.

The UI exposes the same information through a mastery map and an evaluator ledger. The report is downloadable as Markdown.

The weakest evidence selects the next scenario focus. The generated Peak-Season Surge uses seed 7331 and is launchable as a new valid learning session.

MVP 2 and MVP 3 feature status

MVP 2

  • [x] live simulation event stream;
  • [x] architecture animation;
  • [x] mastery scoring;
  • [x] learner-level adaptation;
  • [x] multiple misconception types;
  • [x] controlled failure injection;
  • [x] presenter mode;
  • [x] Playwright end-to-end tests;
  • [x] responsive desktop, tablet, and mobile UI;
  • [x] Devpost documentation;
  • [x] two-minute demo script.

MVP 3

  • [x] runnable adaptive follow-up scenario;
  • [x] richer eight-dimension rubric;
  • [x] visual mastery map;
  • [x] detailed AI evidence;
  • [x] circuit breaker and dead-letter queue controls;
  • [x] correlation IDs;
  • [x] service health diagnostics;
  • [x] evaluator view;
  • [x] downloadable learning report;
  • [x] second challenge module.

System architecture

┌────────────────────────────────────────────────────────────┐
│ Untrusted browser                                          │
│ Next.js 16 + React 19 + TypeScript                  :3001 │
└───────────────────────────┬────────────────────────────────┘
                            │ JSON API + SSE
                            │ X-Correlation-ID
                            │ Idempotency-Key
                            ▼
┌────────────────────────────────────────────────────────────┐
│ Authoritative learning boundary                            │
│ Java 21 + Spring Boot                              :18080 │
│                                                            │
│ Session state machine · validation · persistence           │
│ idempotency · audit · mastery · reports · SSE ownership    │
└───────────────┬──────────────────────┬─────────────────────┘
                │                      │
                ▼                      ▼
┌──────────────────────────┐  ┌──────────────────────────────┐
│ Go simulation engine     │  │ Python AI learning engine    │
│ deterministic events     │  │ FastAPI + Pydantic           │
│ metrics + direct SSE     │  │ feedback + rubric + adapt    │
│                  :18090 │  │                      :18000 │
└──────────────────────────┘  └──────────────────────────────┘
                │
                │ Java persists authoritative evidence
                ▼
         ┌──────────────┐
         │ SQLite       │
         │ Flyway V1–V4 │
         └──────────────┘

Ownership rules

  • TypeScript owns interaction and presentation state.
  • Java owns learner identity, session state, persistence, idempotency, audit evidence, mastery, report generation, and stream authorization.
  • Go owns simulation events and operational metrics.
  • Python owns structured coaching, misconception classification, language evidence, and adaptive-scenario proposals.
  • Bash owns repeatable local setup, lifecycle, reset, health, and validation.

See ARCHITECTURE.md for system-context, container, component, state-machine, sequence, entity-relationship, deployment, trust-boundary, and failure-mode diagrams.

Session state machine

CREATED
  → DIAGNOSTIC_COMPLETED
  → CHALLENGE_STARTED
  → DESIGN_SUBMITTED
  → SIMULATION_COMPLETED
  → FEEDBACK_READY
  → REVISION_SUBMITTED
  → SIMULATION_COMPLETED
  → FEEDBACK_READY
  → FINAL_ASSESSMENT_COMPLETED

Java rejects invalid transitions with 409 INVALID_SESSION_STATE. A failed Go or Python dependency call does not falsely advance the session.

Persistence

SQLite stores:

  • guest learners;
  • learning sessions and current state;
  • the current submitted architecture;
  • simulation results and per-run architectures;
  • tutor interactions and engine provenance;
  • mastery reports;
  • adaptive scenarios;
  • idempotency responses;
  • correlation-linked audit events.

The default database is .systema/systema.db. ./scripts/demo-reset.sh preserves the previous database as a timestamped backup before starting a clean demo.

Repository structure

systema/
├── apps/
│   └── web/                         Next.js learner interface
│       ├── app/                     page, layout, and responsive CSS
│       ├── components/              lab, navigation, diagram, metrics
│       ├── lib/                     typed Java API client
│       └── tests/e2e/               Playwright browser flow
├── services/
│   ├── learning-orchestrator-java/  state and persistence authority
│   ├── simulation-engine-go/        deterministic event simulator
│   └── ai-learning-engine-python/   offline coaching and adaptation
├── packages/
│   ├── contracts/                   shared JSON Schemas
│   └── design-tokens/               shared visual tokens
├── scripts/                         setup, lifecycle, health, reset, tests
├── tests/
│   ├── contracts/                   schema validation
│   └── e2e/                         multi-service acceptance scripts
├── docs/                             API, AI, education, testing, judge guide
├── ARCHITECTURE.md                  detailed system architecture
├── DEMO_SCRIPT.md                   two-minute presentation script
├── DEVPOST_SUBMISSION.md            hackathon submission copy
└── SECURITY.md                      security posture and production gaps

Prerequisites

The setup script checks for:

Tool Required version or role
Node.js 20 or newer
npm web dependency installation
Go 1.23 or newer
Python 3.12 or newer
uv Python environment and dependency management
Java JDK 21
curl health and acceptance checks
unzip local Gradle bootstrap
Google Chrome Playwright browser acceptance only

Gradle is bootstrapped under .systema/tools; a global Gradle installation is not required.

Docker, Redis, Kafka, Kubernetes, and a cloud database are not required.

Installation and startup

Fresh clone

cd systema
cp .env.example .env
./scripts/setup.sh
./scripts/dev.sh

Open http://localhost:3001.

Zentro can continue running independently at http://localhost:3000.

Stop services

./scripts/stop.sh

Reset the demonstration

./scripts/demo-reset.sh

The reset script stops Systema, moves the existing SQLite database to a timestamped backup, and starts a fresh instance.

Verify health

./scripts/health.sh

Expected output:

✓ simulation
✓ ai
✓ orchestrator
✓ web

Logs and PID files are stored under .systema/logs and .systema/pids.

Configuration and ports

Defaults are defined in .env.example.

Variable Default Purpose
AI_MODE offline selects the complete deterministic coach
DEMO_MODE true marks the local demonstration mode
WEB_PORT 3001 Next.js learner UI
ORCHESTRATOR_PORT 18080 Java public API
AI_PORT 18000 Python internal API
SIMULATION_PORT 18090 Go internal API
WEB_ORIGIN http://localhost:3001 allowed browser origin
SIMULATION_BASE_URL http://localhost:18090 Java-to-Go URL
AI_BASE_URL http://localhost:18000 Java-to-Python URL
NEXT_PUBLIC_ORCHESTRATOR_URL http://localhost:18080 browser-to-Java URL
SYSTEMA_DB_PATH .systema/systema.db SQLite file location
OPENAI_API_KEY empty reserved for a future live provider adapter

Shell-provided values override .env, which allows the test suites to start isolated copies on non-default ports.

API overview

All orchestration endpoints are versioned under /api/v1.

Method Endpoint Responsibility
POST /guest-learners create a guest learner
GET /challenges return challenge definitions
POST /sessions create an idempotent learning session
GET /sessions/{id} retrieve session state
POST /sessions/{id}/diagnostic complete the diagnostic step
POST /sessions/{id}/designs submit initial or revised design
POST /sessions/{id}/simulations execute the trusted scenario
GET /sessions/{id}/simulations/{runId}/stream replay persisted SSE events
POST /sessions/{id}/feedback generate and persist coaching
POST /sessions/{id}/final-assessment generate mastery and adaptation
GET /sessions/{id}/mastery retrieve the mastery report
GET /sessions/{id}/report download the Markdown report
GET /diagnostics inspect dependencies and AI provenance

Mutation requests propagate X-Correlation-ID. Replay-sensitive mutations use Idempotency-Key. Errors share this structure:

{
  "code": "INVALID_SESSION_STATE",
  "message": "Cannot transition from CREATED to SIMULATION_COMPLETED.",
  "correlationId": "00000000-0000-4000-8000-000000000000",
  "details": {}
}

Java owns the workload definition. The browser submits the selected architecture and seed, but cannot remove the trusted challenge failure or lower the assessment traffic.

See docs/api-contracts.md for request examples and internal service contracts.

Testing and acceptance

Service and contract tests

./scripts/test-all.sh

This runs:

  • Go simulation unit tests;
  • Python API and learning-engine tests;
  • Java state-machine tests;
  • TypeScript/Vitest component tests;
  • TypeScript compilation;
  • shared JSON Schema validation.

Multi-service vertical slice

./tests/e2e/vertical-slice.sh

This starts all services on isolated ports and verifies:

  • the full Reliable Order Processing flow five consecutive times;
  • same-seed before/after comparison;
  • real metric-grounded feedback;
  • persisted eight-dimension mastery;
  • Markdown report generation;
  • SSE replay;
  • all service diagnostics;
  • a seed-7331 Peak-Season Surge run.

Browser acceptance

./tests/e2e/playwright.sh

Playwright drives local headless Chrome through:

  • the full design, run, coaching, revision, comparison, mastery, evaluator, and diagnostics journey;
  • mobile navigation at 390 × 844;
  • tablet layout at 1024 × 900;
  • horizontal-overflow assertions.

Complete submission gate

./scripts/validate-submission.sh

The submission gate combines every service test, schema validation, the optimized Next.js production build, five-flow vertical acceptance, and Playwright.

Verified acceptance results

The current implementation passes:

  • seven Go simulation behavior tests;
  • five Python API/learning tests;
  • Java state-transition tests;
  • two TypeScript component tests;
  • an optimized Next.js production build;
  • five consecutive golden learning flows;
  • SSE replay and dependency diagnostics;
  • the second challenge;
  • three Playwright desktop/tablet/mobile scenarios.

The Python suite currently prints one non-failing upstream Starlette/httpx deprecation warning. It does not affect product behavior or acceptance.

Two-minute demo

  1. Open port 3001 and select Reliable Order Processing.
  2. Choose Beginner, answer the diagnostic, and enter the lab.
  3. Run the flawed two-worker direct-call design at seed 4242.
  4. Show rejected work, duplicate fulfilments, overselling, and low throughput.
  5. Read the tutor’s metric-grounded retry/idempotency explanation.
  6. Click Apply revision and rerun seed 4242.
  7. Show zero duplicates, zero overselling, higher throughput, and higher queue wait.
  8. Open the mastery map, evaluator ledger, and health diagnostics.
  9. Download the learning report.
  10. Start the seed-7331 adaptive Peak-Season Surge.

The presenter overlay guides these steps inside the workspace. The timed script is in DEMO_SCRIPT.md.

Responsive and accessibility behavior

  • persistent desktop navigation and collapsible mobile navigation;
  • mobile, tablet, and desktop workspace layouts;
  • semantic buttons and navigation landmarks;
  • accessible names for architecture and health views;
  • visible keyboard focus;
  • reduced-motion support;
  • text event timeline as an alternative to animation;
  • no document-level horizontal overflow at tested breakpoints.

Operations and troubleshooting

A port is already in use

Change the relevant values in .env. Keep the URL pairs aligned:

  • WEB_PORT with WEB_ORIGIN;
  • ORCHESTRATOR_PORT with NEXT_PUBLIC_ORCHESTRATOR_URL;
  • SIMULATION_PORT with SIMULATION_BASE_URL;
  • AI_PORT with AI_BASE_URL.

Then restart:

./scripts/stop.sh
./scripts/dev.sh

A service does not become ready

Run:

./scripts/health.sh

Inspect:

.systema/logs/web.log
.systema/logs/orchestrator.log
.systema/logs/simulation.log
.systema/logs/ai.log

The Java build cannot find Gradle

Run ./scripts/setup.sh. The wrapper downloads its distribution into the project-local .systema/tools directory.

Python dependencies are missing

Verify Python 3.12 and uv, then run:

cd services/ai-learning-engine-python
uv sync --dev

Playwright cannot launch a browser

The test configuration uses the installed Google Chrome channel. Install Chrome or adjust apps/web/playwright.config.ts to an available Playwright browser.

Start from clean evidence

Use:

./scripts/demo-reset.sh

The previous database is preserved rather than deleted.

Security

Implemented local-demo controls include:

  • request validation and bounded payloads;
  • restricted browser CORS;
  • parameterized JDBC access;
  • server-owned scenario definitions;
  • operation-scoped idempotency records;
  • correlation IDs across service boundaries;
  • session/run ownership checks for SSE;
  • Java-only authoritative persistence;
  • no execution of learner code;
  • no browser exposure of provider secrets;
  • visible AI-engine provenance.

See SECURITY.md and docs/threat-model.md for the complete trust model.

Scope and limitations

Systema is a local guest-learning product and hackathon submission. It does not currently provide:

  • authentication or authorization;
  • multi-tenant isolation;
  • classroom or teacher administration;
  • payments or certificates;
  • a live LLM provider adapter;
  • TLS for localhost traffic;
  • a production network database;
  • distributed rate limiting;
  • Kubernetes or cloud deployment;
  • a native mobile application.

The product demonstrates observable improvement within a learning session. It does not claim scientifically proven educational effectiveness. A controlled study is required to measure retention, transfer, and misconception correction.

Additional documentation

License

Systema is released under the MIT License.

Built With

Share this project:

Updates

Submission history