TutorTrust (aka "Gemma Says No") just leveled up. We didn't just ask Gemma "hey, will you write my essay?" and call it a day — we tried to break it. We disguised cheating requests as roleplay, sob stories, "just the first part, I swear," and fake professor permission slips (spoiler: some worked). We put its scores on trial by hand-labeling responses ourselves and catching the AI judge grading its own homework a little too generously. We translated prompts into Spanish and Swahili to see if the guardrails survive outside English. We even tested whether Gemma acts extra well-behaved the moment it thinks it's being watched — very "straightens up when the teacher walks by" energy. Base model vs. instruction-tuned, four judge models, two languages, one very suspicious AI — the scorecard is in, and it's spicier than expected.
Log in or sign up for Devpost to join the conversation.