-
-
Assessment input: learning outcome, assignment prompt, and rubric
-
Construct model identifies the learner-originated evidence to protect
-
Construct bypass detected after the AI stress test
-
Smallest repair proposed for the vulnerable assessment
-
Exact re-attack verifies whether the repair actually closes the bypass
-
Repair fails when the same exploit still succeeds: STILL VULNERABLE
-
Not every assessment contains a construct bypass
-
No construct bypass detected — no repair required
Inspiration
Generative AI has added a new layer of work for teachers and course designers. When I design an assessment today, it is no longer enough to ask whether the learning outcome, assignment, and rubric are aligned. We also have to think about what AI can now do for the student, and whether a strong submission still proves that the student actually demonstrated the intended skill. Doing that well takes time. A teacher may need to look at the learning outcome, identify what evidence should come from the learner, imagine different ways AI could complete the task, judge whether those outputs could still score highly, revise the assessment, and then test the revised version again. Construct Guardian grew from the idea that an AI agent could take on much of this repeated work. Instead of asking the teacher to manually red-team every assessment, the agent does the stress-testing cycle for them: it models the evidence, attacks the task, looks for a construct bypass, suggests the smallest repair, and then re-runs the same attack to see whether the repair really worked. The question behind the project is simple: Can AI earn the grade without proving the learning? The goal is not to detect or punish AI use. It is to reduce the burden on educators and help them design assessments where the grade still reflects the learning that was actually demonstrated.
What it does
Construct Guardian starts with three things from the educator: a learning outcome, an assignment prompt, and a rubric. It first looks at what the assessment is actually trying to measure and what evidence should come from the learner. Then it tries different ways AI could complete the assessment. If AI can produce work that would score well while bypassing the evidence the assessment was meant to capture, Construct Guardian identifies a Construct Bypass and shows the educator what was bypassed. The agent then suggests a small, targeted change rather than redesigning the whole assessment. But it does not assume that the repair worked. It runs the same successful attack again against the revised assessment. If the same attack still works, the assessment remains: STILL_VULNERABLE If the attack can no longer complete an important learner-originated requirement, the bypass can be closed. And sometimes there is no bypass in the first place. In those cases, Construct Guardian simply reports: No construct bypass detected and does not suggest an unnecessary repair. This became an important part of the project for us: Construct Guardian does not assume every assessment is vulnerable, and it does not assume every repair works. It tests both.
How we built it
This became one of the most important parts of Construct Guardian. When the agent suggests a repair, we do not simply assume the revised assessment is better. The system keeps the original attack identity and runs the same attack again. If the attack can also satisfy the new requirement, the repair has not solved the problem. If it still crosses the configured thresholds, the assessment remains vulnerable. If the repair introduces evidence that the same exploit genuinely cannot produce, then the bypass can be closed. For us, this was important because a repair should not be considered successful just because it sounds stronger. It should survive the attack that exposed the weakness.
Challenges we ran into
A lot of the difficulty came from making the workflow reliable rather than simply making an LLM produce an answer. I ran into Bedrock runtime issues, IAM permissions, streaming compatibility, output-token limits, timeouts, and structured-output retries. We also found logic problems that were more important than the infrastructure issues. At one point, the trace showed that an attack had completed a new repair requirement, while the final explanation said the attack had been blocked by that same requirement. That was clearly inconsistent. I changed the logic so that a human-only block can only be used when the repair is actually marked as protected and the exact re-attack really fails to complete it. If the attack completes the repair and still crosses the thresholds, the assessment must remain STILL_VULNERABLE. That behavior is now covered by regression tests.
Accomplishments that we're proud of
One thing we are especially happy with is that Construct Guardian does not give the same answer every time. In different test cases, it can conclude that: there is no construct bypass, a bypass exists, a repair successfully closes it, or a repair does not work and the assessment remains vulnerable. That made the system feel much more useful than a tool that simply warns teachers that “AI is a risk.”I am also proud that repairs are tested rather than trusted, and that the exact successful attack is preserved during re-attack. The live application has also successfully run through Strands + Amazon Bedrock, with deterministic fallback available when a provider call cannot complete in time.
What we learned
The biggest lesson from building Construct Guardian was that the most useful question is not: Did the student use AI? A better question is: What evidence still needs to come from the learner for this assessment to remain valid? We also learned that agents are most useful here when they take repetitive work away from the educator rather than trying to replace the educator. The agent can generate attacks, inspect evidence, test repairs, and repeat the process. The teacher still decides what counts as meaningful evidence of learning.
What's next for Construct Guardian
The next step is to make Construct Guardian useful across more assessments and courses. I would like to explore LMS integration, batch assessment testing, discipline-specific attack strategies, reusable assessment templates, and configurable evidence policies. We also want to test the approach with real educators and quality teams across different disciplines. The longer-term goal is not to make education “AI-proof.” It is to help educators adapt assessment design without having to manually anticipate every new way AI might change how students complete their work. The agent does the repetitive adversarial work; the educator keeps the academic judgment. Red-team the assessment, not the student.
Built With
- agents
- amazon
- amazon-web-services
- anthropic
- bedrock
- claude
- iam
- next.js
- node.js
- react
- sdk
- strands
- typescript
- zod
Log in or sign up for Devpost to join the conversation.