-
-
Landing Page — Platform overview seen by first-time visitors before registration
-
Authentication — Role-based login and registration for students and instructors
-
Problem Management — Instructor panel for creating, editing, and ordering assignment problems
-
Student Monitoring Dashboard — Whole-class overview of submission statistics and integrity flags
-
Per-Student Analytics — In-depth report of an individual student's scores and behavioral metrics
-
Submission Review — Full code history, test case results, and AI chat log for each submission
-
Problem List — Student view of available assignments with completion status indicators
-
Workspace + AI Mentor — Integrated code editor, real-time terminal, and Socratic AI hint panel
-
Learning Progress — Personal student dashboard for tracking completion and overall performance
Inspiration
Academic integrity in programming education has become a measurable crisis. According to the 2025 HEPI and Kortext Student Generative AI Survey, 92% of undergraduate students in the UK reported using AI tools in their coursework, and 88% used them specifically for graded work. A 2026 CEPR study tracking 26,811 students over 30 months found that students who used generative AI for homework saw an average grade increase of 18%, while their scores on closed-book independent exams dropped by 18 to 24%. The same study estimated that approximately 80% of this learning decline was attributable to cognitive offloading, the habit of delegating thinking to an external tool until the student never develops the skill independently.
Programming assignments are uniquely exposed to this pattern. Unlike essays, which carry traces of a writer's individual reasoning and voice, code has no such fingerprint. A generative AI model can produce syntactically correct, fully functional Python code for a standard undergraduate exercise in seconds. A student can copy it, run it against the test suite, receive a passing grade, and submit, without any conventional detection system flagging it, because AI-generated code produces structurally unique output on every generation, making text-comparison tools ineffective.
We mapped out the available tools and found a consistent gap. Platforms such as HackerRank and Codio offer plagiarism detection through code similarity comparison, but none of the widely used platforms combine AI dependency tracking, a pedagogically constrained AI mentor, and process-based assessment in a single environment. ThinkCode was built to fill that specific gap.
What It Does
ThinkCode is a self-hostable web-based programming education platform. It does not attempt to ban AI use. Instead, it changes what AI use looks like inside the platform and makes the quality of a student's learning process visible to both the student and the instructor.
Socratic AI Mentor with a 5-Level Constraint System
The built-in AI Mentor is designed around one rule: it will not write executable code for the student under any circumstances. Students select a help level before each interaction, and the system prompt enforces strict behavioral boundaries at inference time.
| Level | What the Mentor Does |
|---|---|
| Level 0 | Moral support and encouragement only. No technical guidance. |
| Level 1 | Redirects the student toward the relevant concept through questions. |
| Level 2 | Narrows down which part of the code contains the problem. |
| Level 3 | Provides a logical walkthrough in pseudocode only, no real code. |
| Level 4 | Offers full architectural guidance in plain language. Writing code is still forbidden. |
The model runs entirely on-premise via Ollama using qwen2.5-coder:3b. No student data is sent to any external server at any point.
AI Dependency Tracker
The platform passively monitors behavioral signals throughout each session and computes an AI Dependency Score (AIDep) using the following formula:
AIDep = [(S_ins x 0.3) + (S_cpx x 0.2) + (S_freq x 0.3) + (S_itr x 0.2) + P_bin] x 100
Where S_ins captures sudden large code insertions exceeding 100 characters in under one second, S_cpx flags unexpected use of advanced syntax inconsistent with beginner-level work, S_freq measures AI query frequency normalized per hour, S_itr detects anomalously low trial-and-error counts, and P_bin is a binary paste penalty triggered when a student pastes a block exceeding 50 characters from an external source. Detection is purely behavioral and requires no screenshots, keyloggers, or device access.
Plagiarism Detection
Rather than comparing raw text, submissions are tokenized into logical units and compared using the Jaccard Similarity Coefficient:
J(A, B) = |A intersect B| / |A union B|
This approach is resistant to cosmetic edits such as variable renaming, which routinely bypass conventional text-comparison tools. The combined Copy Score is:
CopyScore = max(max(J) x 100, P_paste)
P_paste = min(N_paste x 45, 100)
Process-Based Assessment
Every session produces a Process Score that rewards the quality of the learning journey, not only whether the final submission passes:
ProcessScore = 60 + min(N_cmp x 2, 30) + min(N_err x 1.5, 20) - (AIDep x 0.3) - (CopyScore x 0.3)
N_cmp rewards independent compilation attempts up to a maximum of 30 additional points. N_err rewards errors encountered and resolved up to a maximum of 20 additional points. AIDep and CopyScore apply deductions proportional to detected shortcut behavior.
Sandboxed Code Execution
All student code runs in an isolated child process confined to a uniquely named temporary directory. Every execution is automatically terminated after 5 seconds if not completed, preventing infinite loops and resource exhaustion. Output streams back to the browser in real time via a persistent WebSocket connection.
Instructor Analytics Dashboard
Instructors can view compilation counts, error rates, AI query frequency, Copy Score, AIDep, and Process Score for every student across every assignment, alongside sequential code history snapshots showing how each submission evolved.
How We Built It
The full stack uses only open-source components.
| Layer | Technology | License |
|---|---|---|
| Backend | Node.js + Express.js | MIT |
| Database | SQLite via better-sqlite3 | MIT |
| Real-time communication | ws (WebSocket library) | MIT |
| Frontend | Vanilla HTML5, CSS3, JavaScript | - |
| Code editor | CodeMirror | MIT |
| AI runtime | Ollama | MIT |
| AI model | qwen2.5-coder:3b by Alibaba Cloud | Apache 2.0 |
| Authentication | JSON Web Token (jsonwebtoken) | MIT |
The AI Mentor constraint system is implemented through a structured system prompt injected at inference time before every student interaction. The prompt defines the mentor's role, prohibits it from producing executable code in any framing, and sets level-specific behavioral boundaries. The model receives no persistent memory between sessions, so constraint enforcement is consistent regardless of conversation history.
Behavioral monitoring runs server-side. The client sends editor events, paste events, and AI query logs to the server, which computes AIDep and CopyScore continuously and stores snapshots to SQLite. Code execution happens in a spawned child process using Node's built-in child_process module with a hard 5-second kill signal.
Challenges We Ran Into
Making the AI constraint robust against adversarial prompting. Students who want to extract code can be creative. Early versions of the system prompt were bypassed by framing requests as hypotheticals, debugging requests, or translation tasks. We iterated through multiple prompt versions, testing each against a set of adversarial phrasings, until the constraint held consistently across all tested inputs.
Designing behavioral detection without device access. The accuracy of AIDep depends on distinguishing normal fast typing from paste events and AI output dumps. Calibrating the thresholds required careful analysis of what realistic beginner coding sessions look like in terms of insertion speed, session duration, and trial-and-error frequency, to avoid flagging legitimate behavior as suspicious.
Keeping the sandbox stable under adversarial submissions. Students submitting code that intentionally or accidentally contains infinite loops, excessive memory allocation, or malformed system calls had to be isolated completely from the main server process. The 5-second forced termination and per-execution temporary directory approach required careful testing to ensure the server remained stable regardless of what was submitted.
Designing an experiment that could isolate variables. To produce credible evidence that the platform works as intended, the three-session experimental design had to hold conditions constant except for the specific variable being tested. Controlling for session assignment, problem difficulty, and participant group consistency across 1,050 sessions required structured planning before any data collection began.
Accomplishments That We Are Proud Of
Experimental evidence supporting the core design claims.
We ran a controlled three-session experiment with 1,050 coding sessions conducted across 350 students, each completing three sessions under different experimental conditions, to measure the effect of monitoring and AI Mentor presence on student behavior and outcomes.
In Session 1, which was unmonitored with no AI Mentor, the pass rate was 84.29% and the mean AI Dependency Score was 0.749. The short mean session duration of 9.86 minutes alongside this high AIDep score is consistent with students relying heavily on external AI assistance to produce passing submissions quickly.
In Session 2, which introduced monitoring without any AI Mentor, the pass rate dropped to 31.71% and mean AIDep fell to 0.082. This confirms that a substantial portion of Session 1 passes were dependent on external assistance rather than independent understanding. Without support, students who had to work independently struggled significantly.
In Session 3, which combined monitoring with the Socratic AI Mentor, the pass rate recovered to 66.57% with a mean AIDep of 0.149. Mean session duration was 22.47 minutes, between the 9.86 minutes of Session 1 and the 36.25 minutes of Session 2. Students spent enough time to work through problems genuinely, but the structured mentor made the process more efficient than working without any guidance at all. The 0.149 AIDep in Session 3 reflects legitimate interaction with the platform's own mentor rather than external AI exploitation.
These results show that the Socratic Mentor design achieves its intended purpose: it helps students succeed through guided independent work rather than by providing answers directly.
Privacy-preserving integrity detection.
The AIDep system produces actionable behavioral data without requiring any form of device surveillance. This is a deliberate design constraint, not a technical limitation.
Full data sovereignty for institutions.
Because the AI model runs on-premise via Ollama, institutions can deploy ThinkCode without transmitting any student data to a third-party service. This is directly relevant for institutions operating under data protection regulations such as GDPR or FERPA.
What We Learned
The most significant insight from building ThinkCode is that monitoring without support is not a solution. Session 2 of the experiment demonstrated this directly: removing external AI access without providing structured alternatives caused the pass rate to collapse to 31.71%. Students who had relied on external assistance were left without the skills to proceed independently. A platform that only enforces restrictions without scaffolding genuine learning replaces one problem with another.
We also learned that the quality of the AI constraint is as important as the constraint itself. A rule that says "do not write code" is only as strong as the prompt engineering that enforces it. Adversarial testing of the system prompt was essential, not optional.
Finally, designing the Process Score formula required us to be explicit about what we believe learning looks like in a coding context. Rewarding compilation attempts and errors resolved is a deliberate pedagogical stance: struggle, iteration, and correction are the signals of genuine learning, not smooth first-attempt success.
What's Next for ThinkCode: Learn to Code, Not Copy
Multi-language sandbox support. Expanding code execution beyond Python to include Java, C++, and JavaScript, which are the most commonly taught languages in undergraduate curricula.
LMS integration. API connectors for Moodle and Google Classroom to allow ThinkCode to plug into existing institutional workflows without requiring a separate login or gradebook.
Adaptive difficulty engine. Dynamically adjusting problem difficulty based on each student's verified performance history, so students who have demonstrated competence are challenged appropriately rather than being allowed to coast.
Student peer-review module. Structured collaborative debugging with Process Score applied to peer interactions, extending the integrity framework to collaborative work.
AI model fine-tuning. Training the mentor model on curated pedagogical datasets to improve the quality and consistency of Socratic responses across a wider range of problem types and student phrasings.
Exportable analytics reports. PDF and CSV exports of per-student and per-class behavioral data for institutional reporting and accreditation documentation.
Responsive mobile interface. Adapting the workspace for smaller screens to improve accessibility for students in environments where a desktop computer is not always available.
Built With
- academic-integrity
- artificial-intelligence
- codemirror
- css3
- edtech
- education-technology
- express.js
- html5
- jaccard-similarity
- javascript
- jwt
- machine-learning
- natural-language-processing
- node.js
- ollama
- python
- qwen2.5-coder
- socratic-method
- sqlite
- websocket
Log in or sign up for Devpost to join the conversation.