Inspiration
Workplace language learners can understand a phrase in class yet still freeze when a real phone rings. We wanted to bridge that gap with one short, low-pressure rehearsal grounded in language the learner had actually been taught—not a generic speaking test and not an automated CEFR score.
What it does
After a lesson is delivered, its two or three stored target phrases automatically become the locked aims for a Phone Call Practice task. A trainer reviews the scenario and support, explicitly approves the task, and assigns it. The learner sees the exact language expected, a masked destination, and a clear consent step before starting one bounded call.
CALL-E then places the real phone call. When it completes, the application returns a speaker-labelled transcript and phrase-level evidence showing whether each lesson phrase was used independently, as a meaning-preserving variant, after prompting, or was not observed. The result is deliberately narrow: achieved, needs practice, or no evidence, plus one correction and one next step. It cannot change CEFR level, mastery, readiness, points, badges, or progression.
How we built it
The implementation adds a distinct phone-call-practice lifecycle to the existing CEFR lesson workflow. The trusted backend freezes lesson provenance, validates one operator-controlled E.164 destination, requires consent and fresh one-call approval, and calls CALL-E with a stable idempotency key. Provider IDs are persisted before status polling. Phone numbers are masked in ordinary records, credentials remain server-side, and ambiguous outcomes fail closed instead of redialling.
The live demonstration shows the complete path: lesson-derived trainer brief, authentic CALL-E provider evidence and audio, the real conversation, and the transcript-grounded result returned to the application. A synthetic no-call path remains available for safe repeatable testing.
Challenges we ran into
The hardest work was not simply making a call. We had to protect against duplicate paid calls after timeouts, reconcile provider responses without trusting a completion flag as learning evidence, preserve privacy, and keep a tiny formative task completely separate from formal CEFR assessment. Real SIP audio routing also exposed practical device and codec issues before the final successful run.
Accomplishments that we're proud of
- One consented, end-to-end CALL-E call completed with authentic audio and transcript evidence.
- A 55-second result with 14 normalized speaker-labelled turns.
- Phrase evidence fails closed: one expected phrase was correctly withheld because its required wording was absent from the learner transcript.
- No automatic redial, no browser-selected destination, and no score or progression side effects.
- A reusable public Agent Skill with consent, masking, preview, idempotency, and safety guidance.
What we learned
Phone agents need product-level safeguards around the provider API. Idempotency, explicit approval, transcript grounding, bounded evidence, and honest failure states matter as much as the conversation prompt. In education, one authentic interaction can be valuable evidence without becoming a high-stakes assessment.
What's next
Next steps are an isolated judging environment, clearer provider retention controls, more locales and workplace scenarios, and a tightly controlled pilot with trainers and consenting learners.
Built With
- call-e
- docker
- github
- javascript
- playwright
- postgresql
- python
- react
- rest-api
Log in or sign up for Devpost to join the conversation.