Inspiration
Cancer researchers know where to look: GEO, ClinicalTrials.gov, cBioPortal, Dockstore, GDC, CELLxGENE, and many other catalogs. Yet a practical question remains difficult to answer: Which publicly available datasets, tools, and other research outputs could help advance my research? Web search mixes research papers, resource portals, and irrelevant results. Language models without external retrieval can invent plausible GEO accessions, clinical trial identifiers, or DOIs, while static model knowledge quickly falls behind new releases. We built Cancer Output Atlas to provide a better answer: a living catalog that returns traceable public resources and abstains when it cannot find a reliable match.
What it does
Cancer Output Atlas is an agent-managed, living graph of public cancer research outputs, updated daily. A researcher enters a research goal, and the agent helps them discover existing public research outputs they can explore for potential reuse. Results include links to source resources and are grouped into seven categories aligned with the NCI Office of Data Sharing Impact Prize’s definition of cancer research outputs (https://www.nih.gov/challenges/nci-office-data-sharing-impact-prize) : data, software, tools, methods and protocols, models, clinical trial results, and biospecimens.
- Live demo: https://cancer-output-atlas-ulao4pneza-uc.a.run.app
- Source (Apache-2.0): https://github.com/qxiong888/cancer-output-atlas
- Demo video: https://youtu.be/OsEG8Q8w7Oo
How we built it
- Build the graph through daily harvesting. Scheduled batch jobs collect public metadata from approved sources into a prebuilt graph. New records enter only through harvesting; searches never trigger ingestion.
- Retrieve strictly from existing records. The system parses each research goal into structured criteria, ranks existing graph nodes, and groups matches by output type. If no reliable match is available, it abstains.
- Put the agent on the live request path. Every GET /api/find request creates an agent using the official Strands Agents SDK. The agent must call the find_public_outputs tool, whose returned payload becomes the user-facing answer. The language model does not generate or rewrite the match list.
- Deliver a simple experience on Cloud Run. The interface has one goal box, one search button, and results organized by output type. The September 3, 2026 snapshot contains 9,935 public resource records across seven output types. Daily harvesting continues to evolve the catalog; this figure represents a dated snapshot.
Challenges we ran into
- Explaining why wording changes results. Matching considers multiple facets of a research goal, so searches within the same disease area can return different totals. We document live checks—for example, a breast cancer RNA-seq query returned 416 matches, including 39 data records—and verify that nonsensical goals produce no matches. The live app and demo video show this behavior in context.
- Documenting a catalog that keeps changing. Daily harvesting expands coverage, while submission materials need stable, verifiable numbers. We anchor our reported size to the September 3 snapshot.
- Verifying the live agent integration during a submission freeze. We needed to keep the Cloud Run deployment stable while ensuring that every search invoked an official Strands agent and its required tool.
Accomplishments that we're proud of
- A public Cloud Run demo with strict retrieval from the catalog.
- An official Strands agent on the live search path, exposed through agent=strands in search responses.
- A dated snapshot of 9,935 public resource records spanning seven output types, supported by daily harvesting.
- A clear product contract: abstain when reliable matches are unavailable, never invent identifiers, and leave controlled data at its source.
- An Apache-2.0 open-source release with a README, architecture diagram, and public demo video.
What we learned
- Abstention builds trust. Weak matches and invented identifiers waste researchers’ time. Returning no result is a useful outcome when the catalog cannot support an answer.
- Separating harvesting from search makes retrieval auditable. Every returned match can be traced to an existing graph record.
- Architecture claims should be verifiable in the running product. Putting Strands on the live search path lets reviewers inspect the integration directly.
- Public metadata defines the product’s boundaries. The interface, safety messaging, and retrieval behavior must consistently reflect that scope.
What's next for Cancer Output Atlas
We plan to extend the same strict retrieval approach to additional disease areas, starting with autoimmune research and then cardiovascular research, while continuing to use public metadata only. Future options include exploring Amazon Bedrock AgentCore and publishing a build story on builder.aws.com with “Agents for Humans” in the title. These are potential follow-on projects; neither is required for the current submission.
Built With
- cbioportal
- cellxgene
- clinicaltrials.gov
- dockstore
- europe-pmc
- gdc
- gemini
- geo
- google-cloud
- google-cloud-run
- hugging-face
- human-cell-atlas
- python
- scperturb
- strands-agents-sdk
- tcia
Log in or sign up for Devpost to join the conversation.