Inspiration

The rise of generative AI tools like ChatGPT has quietly created a crisis in programming education. A CEPR study tracking over 26,000 students found that while AI use boosted homework scores by 18%, unassisted exam performance dropped by 20% — clear evidence of "cognitive offloading," where students outsource the actual thinking to a machine and walk away with nothing learned. A HEPI survey also found 92% of UK undergraduates now use generative AI in coursework, and nearly 7,000 proven AI-related academic misconduct cases were recorded in the UK alone in 2023–24.

Programming assignments are especially vulnerable to this. Unlike an essay, code has no personal "voice" — an AI can produce a perfectly working solution in seconds, and no conventional plagiarism checker can tell the difference. We asked ourselves: what if a platform could look past the final code and actually reward the process of learning to solve a problem, instead of just the output? That question became ThinkCode.

What it does

ThinkCode is a self-hosted, web-based programming education platform that replaces the traditional "does it compile / pass tests" grading model with a process-oriented scoring system. Instead of just checking output, it:

  • Provides a Socratic AI Mentor with 5 constraint levels — from moral support only, up to a full architecture walkthrough — that is deliberately forbidden from ever writing working code for the student
  • Passively tracks behavioral signals (typing bursts, paste events, compile/error frequency) to compute an AI Dependency Score
  • Detects logic-level plagiarism between submissions using the Jaccard similarity coefficient on tokenized code, independent of variable names or formatting
  • Calculates a Process Score that rewards genuine struggle (failed compiles, resolved errors) and penalizes shortcuts (AI over-reliance, copy-paste)
  • Runs all student code inside an isolated sandbox (child process, 5-second timeout, real-time WebSocket streaming) so nothing malicious or runaway can affect the server
  • Gives instructors a real-time analytics dashboard showing exactly how each student arrived at their solution — not just whether it worked

How we built it

  • Backend: Node.js + Express.js, handling authentication (JWT), routing, and orchestrating sandboxed code execution
  • Database: SQLite via better-sqlite3 for lightweight, file-based persistence with no external DB server required
  • Real-time layer: WebSockets (ws) to stream stdout/stderr from the sandbox back to the browser as code runs
  • Frontend: Vanilla HTML5/CSS3/JavaScript with CodeMirror as the in-browser code editor
  • AI layer: Ollama running qwen2.5-coder:3b entirely on-premise, so no student code or chat data ever leaves the server — important for institutions with data-privacy requirements
  • Core algorithms: custom formulas for AI Dependency Score, Copy Score (Jaccard-based), and Process Score, computed server-side from behavioral telemetry rather than any single "gotcha" signal

We designed the scoring formulas on paper before writing a single line of code, since the entire product's credibility depends on those numbers being fair, transparent, and explainable to instructors rather than a black box.

Challenges we ran into

  • Designing an AI mentor that helps without giving answers. Tuning the prompt-level constraints for each of the 5 hint levels took a lot of iteration — too strict felt useless, too loose and it just wrote the code anyway.
  • Detecting AI dependency without surveillance. We deliberately avoided screen capture or invasive keystroke logging, which meant getting creative with behavioral proxies like insertion size, typing burst patterns, and paste-length thresholds instead.
  • Keeping the sandbox safe and fast. Balancing execution timeout, process isolation, and real-time output streaming over WebSocket required careful process lifecycle management to avoid zombie processes or server slowdown.
  • Running AI fully local. Getting qwen2.5-coder:3b to respond quickly enough on modest hardware while staying inside the Socratic constraints was a constant tuning process.
  • Defining a fair scoring formula. Balancing the weights in AIDep, CopyScore, and ProcessScore so the numbers actually reflected genuine effort — and not just penalize students who happen to type fast — took several rounds of testing against real coding sessions.

Accomplishments that we're proud of

  • Built a fully working end-to-end MVP — from sandboxed code execution to AI mentoring to analytics — in a short development window
  • Designed and implemented three original, mathematically grounded scoring systems (AI Dependency Score, Copy Score, Process Score) rather than relying on any existing off-the-shelf detection tool
  • Got a local LLM (Qwen2.5-Coder) working reliably as a constrained Socratic tutor, without sending a single byte of student data to an external API
  • Built a real-time, WebSocket-driven code execution sandbox that is both safe and responsive enough for a genuinely usable coding experience
  • Created a platform that, as far as we found, is the only tool combining AI dependency detection, leveled Socratic mentoring, and process-based grading in one self-hosted package

What we learned

We learned that fighting AI misuse in education isn't about blocking AI — it's about redesigning what gets measured. Once grading rewards the process instead of just the final output, the incentive to cheat naturally weakens. We also learned a great deal about running LLMs locally for privacy-sensitive use cases, and how much careful prompt engineering matters when the AI's job is to withhold answers rather than hand them over. On the engineering side, we deepened our understanding of process isolation, WebSocket lifecycle management, and designing scoring systems that need to be both mathematically sound and intuitively fair to end users.

What's next for ThinkCode – Socratic AI Mentor for Coding Education

  • Multi-language sandbox support (Java, C++, JavaScript)
  • LMS integration (Moodle / Google Classroom API)
  • Student peer-review and collaborative debugging module
  • Adaptive difficulty engine based on historical performance
  • Mobile-responsive interface for broader accessibility
  • Exportable analytics reports (PDF/CSV) for institutional use
  • Fine-tuning the AI model on domain-specific pedagogical datasets

Built With

Share this project:

Updates

Submission history