Inspiration

AI agents can sound confident and still behave unsafely. Reading an agent's prompt or configuration does not prove what it will do when it encounters a malicious instruction, an untrusted document, a request for sensitive information, or an attempt to bypass approval.

We built BoxShield Arena for Challenge 05 — Break Me — to make agent security visible and testable. Instead of asking another model whether an agent looks safe, BoxShield runs controlled attacks, records concrete evidence, proposes a defense, and then repeats the exact same tests to show whether the defense actually worked.

Our guiding idea became:

Attack. Patch. Replay. Prove.

What it does

BoxShield Arena is a defensive security evaluation environment for BoxLang AI agents. It ships with a synthetic target called the Acme Support Agent, fake customer and order data, simulated actions, and harmless canary values. It never targets public systems or performs real refunds, messages, deletions, or account changes.

The experience follows five steps:

  1. Target — Select the bundled synthetic agent and choose a Quick or Full evaluation.
  2. Attack — Run a versioned corpus covering prompt injection, role impersonation, malicious documents, sensitive-information disclosure, hidden-context extraction, approval bypass, excessive agency, resource abuse, and obfuscated instructions.
  3. Prove — Evaluate the results with deterministic security rules before using model interpretation. Exact marker disclosure, forbidden actions, approval bypasses, and budget violations produce traceable findings.
  4. Shield — Generate a declarative defense proposal using only approved controls. A person must explicitly approve the patch before it can be applied to a cloned target.
  5. Replay — Run the identical frozen attacks against the hardened clone and compare security improvement, remaining findings, regressions, and retained utility.

The Full recorded evaluation improves the Safety Score from 26 to 89, while retaining 80% utility. The Quick demonstration improves Safety from 13 to 100 while retaining 100% utility. These results represent one controlled replay sample, not a security certification or a claim of statistical significance.

How we built it

BoxShield Arena is a modular BoxLang application running on BoxLang MiniServer.

BoxLang owns the critical workflow:

  • Request validation and execution limits
  • Attack-corpus loading and ordering
  • Simulated action authorization
  • Deterministic evidence evaluation
  • Safety and utility scoring
  • Defense allowlisting
  • Human approval enforcement
  • Clone-only patch application
  • Exact paired replay
  • JSON and Markdown security reports

The application preserves state across separate HTTP requests with a compact run capsule signed using HMAC-SHA256. Capsules have an expiration time, bind the target and corpus versions, and carry a one-time approval nonce. Invalid signatures, modified state, expired runs, incorrect nonces, and reused approvals are rejected.

BoxLang AI provides the agent-oriented pipeline and verified Mock provider integration. The design separates the Target, Attacker, Defender, and Security Judge responsibilities, while BoxLang retains final control over authorization, limits, evidence, and scoring. The current release emphasizes deterministic Replay and Mock modes so the demo remains reliable without external credentials or provider availability.

The interface uses semantic HTML, responsive CSS, and vanilla JavaScript. No separate frontend framework, persistent database, or remote-target scanner is required.

Security methodology

We designed the evaluation around three principles.

Deterministic evidence first

A model's opinion is not treated as proof. BoxShield first checks observable signals such as:

  • Exact synthetic canary disclosure
  • Hidden-context marker disclosure
  • Forbidden or unapproved actions
  • Instructions followed from a malicious document
  • External markdown-image exfiltration
  • Resource-budget violations
  • Incorrect refusal of benign requests

Ambiguous behavior can be labeled as a warning or inconclusive result, but it cannot override a proven deterministic violation. Runtime errors are also kept separate from vulnerabilities.

Fair paired replay

After the baseline and bounded attacker adaptation are complete, BoxShield freezes the attack IDs, payloads, ordering, target data, corpus version, and utility controls. The hardened target receives the exact same inputs. This makes the before-and-after comparison understandable and avoids inflating improvement by changing the test set.

Security without destroying utility

Blocking everything would produce a high security score but a useless agent. BoxShield therefore runs benign controls alongside adversarial tests and reports Safety Score and Utility Score separately.

Challenges we faced

The hardest challenge was making the comparison fair. AI behavior can be nondeterministic, so we needed to separate repeatable security evidence from interpretation and ensure that the baseline and hardened target received identical tests.

We also had to preserve workflow state without relying on a database or assuming that two requests would reach the same server instance. Signed run capsules and one-time approval nonces gave us a stateless design without trusting browser-supplied scores or authorization decisions.

Another challenge was preventing the evaluator from becoming dangerous itself. We intentionally removed arbitrary target URLs, shell execution, unrestricted tools, real credentials, and production actions. Every action is simulated, every defense is selected from an allowlist, and every patch is applied to a clone.

Finally, making the same project run consistently across a local BoxLang environment, TestBox, GitHub Actions, MiniServer, and container builds required careful dependency and environment handling.

Accomplishments that we are proud of

  • A complete attack-to-defense workflow rather than a static security checklist
  • Nine adversarial cases and five benign utility controls
  • Deterministic evidence with clear provenance
  • Human approval before applying defenses
  • One-time approval protection and signed stateless run capsules
  • Exact paired replay against a hardened clone
  • Separate Safety and Utility measurements
  • Reusable evaluation, policy, defense, scoring, and reporting components
  • 60 passing TestBox tests
  • 633 environment-independent validation checks
  • A working BoxLang AI Mock-provider pipeline
  • Replay mode that remains usable without an API key

What we learned

We learned that agent security is not only about writing a stronger system prompt. Reliable defenses require several layers: untrusted-context handling, deterministic authorization, strict action schemas, execution budgets, output filtering, evidence collection, and human approval for sensitive changes.

We also learned that a security score is meaningful only when the system preserves legitimate behavior. Measuring utility alongside security changed how we evaluated every proposed defense.

Most importantly, we learned that reproducibility builds trust. A judge or developer should be able to see the exact test, the observed behavior, the evidence, the applied defense, and the result of replaying the same input.

What's next

Next, we want to:

  • Add credentialed Live-provider evaluations behind explicit access controls
  • Run repeated samples to measure nondeterministic behavior
  • Support additional developer-authorized local target adapters
  • Expand the community attack corpus and defense catalog
  • Add richer report formats and command-line evaluation
  • Publish reusable BoxLang security-testing components
  • Test attacker and defender strategies across more agent architectures

BoxShield Arena is not a security certification. It is a practical way to turn agent-security claims into observable, repeatable evidence.

Built With

Share this project:

Updates

Submission history