Hardening the model-to-MATLAB boundary
After submitting the demo, I added a reproducible adversarial benchmark for the strict review contract. It runs one checked GPT-5.6 control plus 14 hostile mutations covering executable-action injection, candidate/model spoofing, field smuggling, invalid score types and ranges, contradictory verdicts, oversized findings, and malformed JSON.
Result: 15/15 checks passed; every adversarial output failed closed. The benchmark now runs in the release gate. PR #43 passed both Quality checks, and the main-branch Quality run is green.
Log in or sign up for Devpost to join the conversation.