The problem

Foreign residents in Japan often receive important phone calls about deliveries, utilities, housing, schools, appointments, or payments.

A call gives the listener little time to understand each sentence. The listener must also choose a safe response while the caller waits. A missed date, price, or commitment can cause a real problem.

What AlphaMind does

AlphaMind listens through the computer microphone and converts Japanese speech to text. It explains the caller's latest message in Simplified Chinese. It then presents two relevant Japanese reply choices.

The application warns the user before a reply can confirm a date, price, quantity, or promise. The user chooses the reply. AlphaMind never speaks or accepts a commitment for the user.

Why this matters

Phone calls remain common for essential services in Japan. They are difficult for language learners because there are no gestures, subtitles, or extra time.

AlphaMind gives the listener three things at the decision point:

  • The meaning of the caller's latest message.
  • Two replies that address the actual request.
  • A warning when the reply can create a commitment.

One App, two Agents

AlphaMind is one Python application with two independent Agents:

  1. The Japanese Call Listener owns the LiveKit audio session, voice activity detection, local Whisper speech recognition, speaker routing, final-turn ordering, and result delivery.
  2. The Japanese Reply Agent owns the judgment-heavy reasoning task. It explains the caller, creates distinct reply choices, and identifies facts that need confirmation.

A typed handoff connects the Agents. The Listener sends accepted caller turns and relevant caller context to the Reply Agent. Completed turns remain in capture order while reasoning is busy. Each Agent has one responsibility and a separate trace.

How we use Strands Agents

The Japanese Reply Agent uses Strands for the complete specialist reasoning task. It is not a free-text prompt wrapper.

Each reasoning attempt creates a Strands Agent with four components:

  1. A Vifu model adapter connects the App-private OpenAI-compatible Provider to the Strands model interface.
  2. The system prompt defines the listener perspective, context rules, and limits on unknown facts.
  3. A structured Strands @tool named card is the only output path.
  4. An AfterToolsEvent hook ends the Agent turn after the card is published.

The normal card tool requires one Chinese explanation, exactly two Japanese reply choices, their Chinese meanings, and an optional confirmation message. Personal-information requests use a narrower tool schema. The application supplies safe reply choices and does not let the model invent names, dates, amounts, or numbers.

Pydantic validates the tool arguments and final Assist Card. Deterministic semantic rules reject duplicate replies, copied caller requests, reversed speaker roles, incorrect dates, invented numbers, and unsafe commitments. The result sink also rejects missing or repeated tool calls.

If a card fails validation, the next Strands attempt receives the rejected card and exact errors. The retry uses the same Provider. If the second card remains unsafe, AlphaMind returns a fixed clarification card instead of showing the model output.

Trace metadata records the Strands framework, selected Provider, transcription Provider, and attempt number. The test suite covers the Agent loop, tool contract, typed handoff, validation, and bounded retry behavior.

Key features

  • Local Japanese speech recognition with multilingual Whisper.
  • Local speaker identification for the listener and callers.
  • LiveKit audio, VAD, transcript ordering, and result delivery.
  • Strands reasoning with structured tool output.
  • Two listener-side Japanese reply choices for each accepted caller turn.
  • Simplified Chinese meanings for the caller message and each reply.
  • Confirmation warnings for dates, prices, quantities, and promises.
  • Duplicate and non-speech transcript filtering.
  • Ordered processing while the reasoning Agent is busy.
  • Local traces for both Agent turns.

Run and test

The public repository contains the source code, Apache 2.0 license, architecture diagram, locked dependencies, model setup, and test instructions:

https://github.com/chenyanming/alphamind

Run the deterministic two-Agent demo:

uv sync --frozen
uv run --frozen python main.py demo

Run the interactive microphone mode:

uv run --frozen python main.py

Run the 75-test local suite:

uv run --frozen python -m unittest discover -s tests -v

Current scope

The current interface explains calls in Simplified Chinese. It listens through the computer microphone and does not connect directly to the telephone network.

The voice pipeline runs in the local process. Reasoning uses a user-selected OpenAI-compatible Provider. The project does not include a hosted live demo or AgentCore deployment.

Built With

  • livekit-agents
  • pydantic
  • python
  • strands-agents
  • vifu
  • whisper.cpp
Share this project:

Updates

Submission history