Inspiration

Just last week, NASA released a list of 11 critical challenges it’s working to solve before humans travel deeper into space, with a major one prompted towards keeping astronauts healthy when Earth is no longer in direct communication range. As missions go farther and last longer, astronauts may have to investigate medical problems themselves before Houston can respond. That’s why we built Iris. Iris is an autonomous medical investigation system that combines astronaut health measurements and spacecraft telemetry with NASA’s Open Science Data Repository, or OSDR, which contains years of spaceflight health and biological research. When an astronaut reports a symptom, Grok doesn’t just generate a response. We used principles of retrieval augmented generation (RAG) to retrieve relevant NASA evidence, compare the astronaut’s measurements against their personal baseline, and determine what information it needs to investigate next.

What it does

Iris is designed around a simple idea: a medical anomaly should start an investigation, not immediately produce a diagnosis.

An astronaut can approach the system and report something as simple as, "I've been feeling dizzy and I have a headache." The astronaut interacts with Iris through an ESP32 voice interface, while Grok Voice handles the voice input and output.

From there, Iris investigates across multiple layers of mission data:

  • Astronaut health: heart rate, blood pressure, temperature, strength, food intake, personal statistics, and individual baselines
  • Spacecraft telemetry: oxygen levels, CO₂, and water purity
  • Space environment: radiation levels and solar activity
  • Crew context: health data and logs from other astronauts
  • Historical evidence: NASA OSDR spaceflight health and biological research

Instead of comparing an astronaut only against population-level "normal" values, Iris first asks whether something is unusual for that specific astronaut.

For example, if an astronaut normally has a resting heart rate around 62 BPM and their current rate is 84 BPM, Iris treats that deviation as a signal for further investigation. It does not immediately turn that into a diagnosis.

The investigation follows a loop:

observe → compare → identify missing evidence → collect evidence → reevaluate

Iris can then look for relationships across the mission. If an astronaut’s health has been changing while CO₂ levels have also increased, Iris can surface that relationship and investigate whether historical spaceflight evidence supports or weakens the hypothesis.

The goal is not for Iris to simply say, "CO₂ caused this." Instead, it explains what changed, what relationships it found, what historical evidence is relevant, what potential hazards exist, and what risks could emerge if the conditions continue.

How we built it

Iris is built with Next.js, TypeScript, Vercel AI SDK, Grok APIs, and an ESP32.

The ESP32 acts as the physical voice interface. It connects over Wi-Fi and handles audio input and output, while Grok Voice processes the astronaut's speech and provides spoken responses.

The backend contains a simulated mission state with astronaut health measurements, spacecraft telemetry, external space conditions, and crew logs. These values let us model changing conditions during the demo without requiring a full hardware sensor setup.

We also loaded NASA OSDR data into our backend and built a RAG layer around it. Rather than relying on a live connection to Earth, the historical research is available locally, which better represents the connectivity constraints of deep-space missions.

Grok acts as the investigation engine. It receives the astronaut's report, accesses the relevant mission variables, compares measurements against the astronaut's personal baseline, retrieves relevant historical evidence, and reasons about what information should be examined next.

The final investigation is designed to communicate:

  • What changed from the astronaut's baseline
  • Which environmental or health variables are related
  • What historical NASA evidence is relevant
  • Whether that evidence supports or weakens a hypothesis
  • Potential hazards
  • Possible future health risks if conditions persist
  • The uncertainty behind the conclusion
  • Citations for supporting evidence

Challenges we ran into

One of the biggest challenges was figuring out how to make historical spaceflight data useful without pretending that a historical astronaut is a direct prediction of the current astronaut.

OSDR contains data from many different missions, experiments, populations, and environments. A useful system therefore needs to retrieve evidence based on the type of health event and surrounding conditions, rather than simply finding an astronaut who looks similar.

We also had to design the system so that an abnormal measurement does not automatically become a diagnosis. A deviation from baseline should lead to another question: what evidence would be useful next?

Another challenge was combining completely different data types. Astronaut physiology, spacecraft telemetry, radiation data, crew logs, and historical research all describe different parts of the same mission. Iris needs to reason across those sources while keeping track of what is actually known versus what is only a hypothesis.

Finally, we wanted the system to work as a voice-first experience. An astronaut should be able to describe a problem naturally and hear Iris explain its investigation instead of interacting with a traditional dashboard.

Accomplishments that we're proud of

We built a system that turns a simple voice-reported symptom into a multi-source investigation.

Instead of just answering a medical question, Iris can connect:

astronaut → personal baseline → spacecraft → space environment → crew context → NASA research

We are especially proud of implementing RAG-like features rather than simply adding citations to an LLM response. Historical NASA research becomes evidence that Iris can use to evaluate what is happening during the current mission.

We also built the system around explainability. Iris is designed to show how it got from an observation to a hypothesis, what evidence supports that hypothesis, and where uncertainty remains.

What we learned

We learned that AI for healthcare needs much more than a model that can generate convincing answers.

Context matters. Personal baselines matter. Environmental conditions matter. Historical evidence matters. Most importantly, the system needs to understand the difference between an observation, a correlation, and a causal conclusion.

We also learned how difficult it is to work with real scientific datasets. NASA's research is incredibly valuable, but it comes from different missions and experimental contexts, so retrieval and interpretation are just as important as the model itself.

Most importantly, we realized that deep-space healthcare is not simply about putting an AI doctor on a spacecraft. It is about giving astronauts an autonomous investigation layer that can gather evidence and reason about a changing environment when communication with Earth is limited.

What's next for Iris

The next step is replacing our simulated mission data with real physiological sensors and expanding the number of live spacecraft and environmental data sources.

We also want to improve longitudinal personal baselines, detect shared health events across crew members, incorporate more NASA datasets, and use stronger statistical methods for identifying relationships between physiological changes and environmental conditions.

Long term, we see Iris becoming an autonomous medical investigation layer for deep-space missions, helping crews gather and interpret evidence when immediate support from Earth is not possible.

Built With

Share this project:

Updates

Submission history