Live demo: [https://www.bowsala.app/]
| Demo login | |
|---|---|
| username : admin | |
| Password : admin123456 | |
| Role | Institute Admin — full access to students, attendance, payments, and AI grading |
Inspiration
In Iraq, the exam is still a piece of paper. A private institute runs weekly assessments across hundreds of students, and every paper is marked by hand — so results reach students a week or two late, long after the feedback would have changed how they study.
The tools that fix this elsewhere don't work here. Gradescope and similar platforms are Latin-script, English-first, left-to-right, and priced for enterprises. Arabic handwriting cursive, four positional letter forms per letter, optional diacritics is one of the hardest OCR problems in the field, and the commercial grading market has ignored it.
We already ran the rest of that workflow. Bowsala is an Arabic-first institute management platform used by real Iraqi institutes for student records, attendance, payments, and reporting. Grading was the last manual step in an otherwise digital pipeline.
What it does
Bowsala turns a stack of handwritten Arabic exam papers into scored, explained, recorded results — without removing the teacher.
- Teacher photographs answer sheets from a phone and uploads them with the model answer and rubric.
- Agents identify each paper, match it to the right student in the existing database, read the handwritten Arabic, and evaluate against the rubric.
- Each question gets a mark with partial credit, plus a short Arabic justification.
- Low-confidence answers illegible writing, crossed-out work, off-rubric responses are flagged for human review instead of guessed.
- The teacher approves or overrides. Approved results write into the student record and notify parents.
Each question's score is a rubric-weighted sum rather than one holistic guess:
$$ s_q = \sum_{i=1}^{n} w_i \cdot c_i, \quad \sum_{i=1}^{n} w_i = 1, \quad c_i \in [0,1] $$
A question escalates to a human when confidence \(\kappa_q\) falls below the threshold \(\tau\). That decomposition is what makes a mark explainable — when a teacher asks why two marks were lost, the system points at the criterion.
How we built it
Grading is an orchestrated workflow, not one model call. An orchestrator delegates to specialists:
| Agent | Job |
|---|---|
| Ingestion | Deskews and normalizes phone photos, detects answer regions, rejects unusable pages |
| Identity | Matches each paper to a student record; holds unresolved papers rather than guessing |
| Evaluation | Reads the handwriting end-to-end with Gemini's multimodal capability no separate OCR stage — and scores semantically against the rubric |
| Audit | Second pass checking scoring consistency across papers and defensibility against the rubric |
| Feedback | Writes the student-facing explanation in Arabic |
| Reporting | Commits approved grades, updates the performance profile, notifies parents |
Agents get real tools: student lookup, rubric retrieval, gradebook write, historical performance query, notification dispatch. The write path is gated no agent commits a grade to a permanent record before a teacher approves it.
The platform is RTL-native throughout, with cost per graded paper treated as a first-class engineering constraint rather than an afterthought.
Challenges we ran into
- Arabic handwriting. No mature baseline to build on. We designed the pipeline to stay robust to imperfect reading rather than assuming a clean transcript.
- Crossed-out work. Scratched-out answers are the fastest way to make a vision model confidently read text the student explicitly rejected. Handled explicitly.
- Partial credit. The hard case isn't a correct or blank answer it's a half-right one. Matching how experienced teachers actually mark took repeated iteration against real papers.
- Real input quality. Institutes have phones and fluorescent lighting, not scanners. Our early pipeline worked on clean scans and collapsed on what teachers actually produced.
- Trust. A teacher who doesn't trust the output re-grades everything by hand, which makes the tool worse than useless.
- Cost and latency. Multi-agent costs more per paper. A 400-paper batch forced hard choices about which agents run on every question versus only flagged ones.
Accomplishments that we're proud of
- End-to-end on real Arabic handwriting: phone photo → scored, justified, recorded result, no manual transcription anywhere.
- It ships inside a product with real users, not a prototype looking for a use case.
- Genuinely Arabic-first grading at a price an Iraqi institute can pay a gap nobody had filled.
- The failure mode is "ask the teacher," never "quietly assign the wrong mark." Non-negotiable for something touching student records.
What we learned
- Decomposition beats one big prompt. Our first version was a single well-engineered prompt. When it was wrong we couldn't tell which part of its reasoning failed. Splitting into narrow agents improved accuracy, but the bigger win was diagnosability.
- A second agent reviewing the first isn't redundant. The audit pass caught inconsistency — the same answer marked differently on two papers — that no amount of prompt engineering on the scorer eliminated.
- Knowing when to stop is a feature. A flagged answer costs a teacher fifteen seconds. A wrong grade committed silently costs their trust in the whole product.
- Teachers want the reasoning, not the number. A bare mark has to be verified from scratch, which saves nothing. A mark with a reason can be approved at a glance.
- Grading is a workflow, not a model. The model reads the paper. Everything else — matching, approval, gradebook writes, appeals — is the actual product.
What's next for Bowsala
- [ ] Baccalaureate prep institutes — the highest-volume, highest-stakes grading market in Iraq, overlapping our existing customers
- [ ] Handwriting adaptation on a labeled corpus from real graded papers
- [ ] Offline-first grading that queues uploads and syncs when connectivity returns
- [ ] Exam authoring agents — question banks and rubrics aligned to the national curriculum
- [ ] A curriculum analytics agent surfacing cohort-level weak topics as teaching recommendations
- [ ] Structured appeal handling with a full audit trail
- [ ] MENA expansion — same problem, same script, same gap
Log in or sign up for Devpost to join the conversation.