Demo Available!

Inspiration

We've all felt the frustration of academic search. You type a query into Google Scholar and get a flat ranked list with no sense of how ideas relate to each other. We wanted to ask: what if you could explore the entire landscape of human scientific knowledge the way you explore a galaxy? What if semantically similar papers clustered together visually, and you could fly between fields of study the way you'd travel between star systems?

The universe metaphor felt right. Knowledge is vast, scattered, and full of unexpected connections, just like space. That became search-comète: a 3D interactive galaxy of 54,000 research papers, navigable in real time.


What it does

search-comète is a semantic search engine visualised as a living 3D galaxy. Each star is a research paper. Stars cluster by field: machine learning glows violet, biology teal, physics amber. The closer two stars are, the more semantically similar their papers.

Key features:

  • Semantic search powered by Elasticsearch: type a query and matching papers light up across the galaxy, with the camera flying to the top result
  • 54,190 papers across 15 scientific disciplines, embedded with all-MiniLM-L6-v2 and reduced to 3D with UMAP
  • Time Travel: a year slider that dims papers published after a selected year, letting you watch entire fields of knowledge emerge over time
  • Comets: periodic AI-selected papers fly across the scene; clicking one pulls you to a cross-disciplinary highlight
  • Wormhole: a dark/light mode toggle with supernova and dwarf planet collapse animations
  • Ethereal sound design: procedural Web Audio ambient music and sound effects for navigation, star discovery, and mode transitions
  • Detail panel: click any star to see the abstract, authors, citation count, semantic relevance score, kNN similar papers, and the raw Elasticsearch query
  • Meteor of the Day: a daily highlighted paper surfaced from the corpus, showing the most cited, newest, or a random discovery each session
  • Star Tags: click any paper's detail panel to pin it, turning it gold so you can find it again later as you explore the galaxy

How we built it

Data pipeline

We built a fetcher targeting OpenAlex (250M+ papers, no API key required). 300 search queries across 15 disciplines fetched around 200 papers each. Papers were deduplicated, embedded using sentence-transformers (all-MiniLM-L6-v2, 384 dimensions, L2-normalised), then compressed to 3D coordinates with UMAP. The resulting stars.json and Elasticsearch index were generated locally and deployed to Elastic Cloud.

Backend

A FastAPI app serving routes at both / and /api/ for compatibility. Elasticsearch queries use hybrid BM25 + field-boosted text search for /search, paginated search_after for /stars, and kNN dense vector lookup for similar papers. Deployed to Railway with a minimal Docker image (~200MB).

Frontend

A Three.js WebGL galaxy renderer with instanced meshes (three size tiers by citation count), UMAP-positioned star clusters, additive blending nebulae, and per-cluster edge graphs. Custom modules handle the renderer, search, detail panel, comets, sound, and theming. Built with Vite, deployed to Vercel.

Infrastructure

  • Elastic Cloud (GCP us-central1) for Elasticsearch 9.3.1
  • Railway for the FastAPI backend
  • Vercel for the Vite frontend
  • Vercel rewrites proxy /api/* to Railway, keeping API keys server-side

Challenges we ran into

Bulk indexing timeouts: sending 54,000 dense vector documents to Elastic Cloud kept dropping the connection. Using parallel_bulk with 4 threads overwhelmed it. Fixed by switching to streaming_bulk with chunk_size=100, 60s timeouts, and exponential backoff retries.

UMAP coordinates vs scene positions: the raw UMAP x/y/z values did not match Three.js scene coordinates after the galaxy remapping pass, so every flyTo was flying to the wrong place. Fixed by building a _remappedPos lookup populated during _buildStars().

Relative imports on Railway: the FastAPI app used from .models import ... which broke when Docker ran uvicorn main:app directly. Fixed by converting to absolute imports.

7GB Docker image: pip freeze from the dev environment included torch, sentence-transformers, playwright, and dozens of pipeline-only packages. Stripped requirements.txt down to the 7 packages the backend actually needs.


Accomplishments that we're proud of

  • A performant 3D renderer handling 54,000 instanced objects at 60fps
  • The entire frontend works off stars.json with a local TF-IDF fallback if the backend is down
  • A full ambient music system and UI sound effects built entirely from Web Audio oscillators, with no external audio files
  • The comet system: AI paper selection, real-time 3D animation, camera tracking, and screen-space visible trails at any zoom level
  • Full production deployment across three cloud providers in a single hackathon

What we learned

  • UMAP is genuinely powerful for making high-dimensional semantic spaces navigable. The clusters it produces feel intuitive without any manual labelling.
  • The Web Audio API is capable enough to build a full procedural ambient score with no samples or libraries.
  • Three.js instanced meshes are the right tool for large point clouds but require careful attention to frustum culling and draw usage flags.
  • Railway and Vercel together make a solid zero-DevOps production stack for hackathon projects that need a real backend.

What's next for search-comète

  • True ELSER semantic search: replace BM25 with Elastic's ELSER sparse vector model for genuinely semantic retrieval
  • Black hole trash bin: drag papers into a black hole to remove them from your galaxy, with a restore panel for recently deleted papers
  • Big Bang animation: replay the history of science, watching papers appear in chronological order from a single point
  • Personalised comets: embed the user's search history and pick comets using cosine similarity against their demonstrated interests
  • Collaborative galaxies: shared sessions where multiple users explore the same galaxy simultaneously
  • Sub-cluster tagging: add custom tags within disciplines to create finer groupings, for example tagging papers within physics as electromagnetism or special relativity, forming their own mini-clusters like solar systems within a galaxy

Built With

Share this project:

Updates

posted an update —

more features to be added/already added: Choose colours for topics/type of planet ect Make pinned planets ect into neutron stars (with small pulsar- if time make it either flash or rotate around)/stars? Add hyperspace animation when traversing data files Add sound when traversing files See creation history animated (start with big bang animation) ->obsidian can do this Time travel mode (data is always updating, we can have like a progression of the data timeline) We can have comets that fly in the background and when you click them they pull you to a key paper that would interest you based on the papers that you have indexed. ->AI selects this based (vector embed the previous indexed papers of the user and map onto the dimensions of the research papers, then choose the research paper closest to the embedded vector - closest match. Use the user indexed papers for each galaxy to make an embedded vector; can have separate close matches based on subject. Meteor can have a daily fact ->can enable in settings Supernova explosion for light mode Dwarf planet dying for dark mode (maybe) Add tags to files for even more personalised categorising eg. in physics we can make a tag called electromagnetism, have another tag called special relativity ect ->maybe even make these stuff in the same tag into its own cluster (like a solar system)

Log in or sign up for Devpost to join the conversation.

posted an update —

Use black hole to remove data (planets) ->drag them in (will prompt a are you sure message) ->tap black hole to see recently deleted files to restore files? Can select after x days it will actually permanently delete

Log in or sign up for Devpost to join the conversation.

Submission history