Inspiration
Most AI assistants wait for you to speak first. We wanted to explore something more fundamental to human interaction: shared attention.
Before children can carry on conversations, they learn to coordinate attention with others. This is called joint attention—when two people focus on the same object or event.
One form, Responding to Joint Attention (RJA), happens when someone notices another person's gaze, head movement, or gesture and follows it toward the same target.
Research has connected joint attention with social communication and language development, and differences in joint-attention behavior are common among autistic children.
That led us to ask:
Could a physical AI companion initiate joint attention, recognize whether someone follows its attention, and respond in a way that makes practicing this interaction engaging?
That became Ottis.
Ottis is not a treatment for autism, and our prototype has not been clinically tested. Instead, we are exploring whether embodied AI could provide another way to practice a specific social-communication skill: Responding to Joint Attention.
Why Joint Attention Matters
Joint attention helps connect another person's communication with the world around us.
For example, when a parent looks toward a dog and says, "Look, a dog," the child must notice the parent's cue, shift attention toward the dog, and connect the word with the object.
Some autistic children may be less likely to follow another person's gaze or gestures consistently. Research has therefore explored joint-attention interventions as one way of supporting social communication.
Ottis is not designed to force eye contact. The goal is functional:
Can the user recognize that another agent is directing attention somewhere and determine what it is referring to?
This is why Ottis focuses on attention shifts toward real objects rather than simply measuring whether someone looks at the robot.
Why a Robot?
A robot can repeatedly provide clear and controlled attention cues while adapting based on the user's response.
Instead of:
robot moves → prerecorded response
Ottis creates a closed loop:
cue → observe → interpret → respond → try again
For example, Ottis could begin with a subtle eye movement. If the user does not respond, it could increase the cue using a head turn or verbal prompt.
This creates the possibility of adaptive prompting:
Eyes → Eyes + Head → Verbal Cue
Because each cue can be controlled, a robot could also help researchers study which types of movement are most effective for communicating attention.
The Research Behind Our Hypothesis
Our idea was inspired by research in joint attention and human-robot interaction.
Murza et al. found support for explicitly teaching joint-attention skills to young autistic children.
Bottema-Beutel found strong relationships between joint attention—particularly RJA—and language development in autistic children.
Kumazaki et al. studied a humanoid robot with movable eyes and found improvements in joint-attention performance during robotic interaction.
We also studied demonstrations of joint-attention interactions to better understand the timing between an attention cue, a person's response, and feedback.
Warren et al. developed a robotic system that delivered joint-attention prompts, detected children's responses, and adjusted its prompting based on their behavior.
So et al. later compared robot-based and human-based joint-attention interventions.
The research is still developing, however. A robot attracting someone's attention does not necessarily mean that person is improving at joint attention.
That distinction shaped our hypothesis:
If an embodied AI repeatedly initiates joint-attention cues, recognizes whether the user follows them, and provides immediate adaptive feedback, it may create an engaging environment for practicing Responding to Joint Attention.
Ottis is our prototype of that closed loop.
What it does
Ottis can see objects, direct its attention toward them, observe whether a person follows its attention, and react.
The interaction works like this:
- Ottis sees an object using computer vision.
- Ottis chooses a target and physically moves its eyes or head toward it.
- The person responds, while Ottis tracks facial landmarks, head orientation, and approximate gaze.
- Ottis detects joint attention by comparing the person's attention with the object it intentionally selected.
- Ottis reacts using movement and natural speech.
Eventually, Ottis could adapt the strength of its cues based on the user's responses.
How we built it
Ottis combines computer vision, conversational AI, speech, embedded systems, and robotics.
We use YOLO and OpenCV to detect objects in the environment.
MediaPipe Face Landmarker provides facial and iris landmarks, which we combine with head orientation to estimate coarse attention direction.
We are not trying to determine the exact pixel someone is looking at. Instead, Ottis asks:
Did the person's attention shift generally toward my target?
Ottis then compares that estimate with the object it intentionally selected.
An LLM connects visual context with conversation, while ElevenLabs provides natural speech.
A Arduino receives movement commands and controls Ottis's motors and servos.
We separated vision, speech, conversation, behavior, and robot control into independent modules so that one slow component does not freeze the entire interaction.
Challenges
Gaze estimation
Detecting a face is easy compared with determining whether someone actually followed a robot's attention.
Lighting, glasses, camera position, head rotation, and small iris movements can all affect estimation, so we focused on coarse attention direction rather than exact gaze.
Turning vision into movement
Ottis does not just need to detect a green screen. It needs to physically look at the green screen.
That requires translating:
camera coordinates → physical direction → servo angles
Building the full loop
Computer vision often ends with a bounding box.
Ottis has to continue:
camera → object → target → movement → human response → detection → reaction
Connecting that entire loop was one of the hardest parts of the project.
Different systems, different speeds
Vision runs continuously, LLM calls depend on network latency, speech waits for audio, and servos require immediate commands.
Our modular architecture allowed these components to work without constantly blocking one another.
Accomplishments
The biggest thing we're proud of is that Ottis isn't just an LLM placed inside a robot.
Its physical embodiment is part of the interaction.
Ottis chooses something in the real world, physically directs its attention toward it, observes another person's response, and uses that response
Log in or sign up for Devpost to join the conversation.