NASA’s Astronomy Picture of the Day (APOD) archive contains more than 30 years of entries and 9,700+ images covering galaxies, planets, nebulae, stars, and other astronomical phenomena. However, the archive is primarily organized by date, and traditional keyword searches struggle when users describe what they are looking for without using the exact words in an entry.

I built AstroSearch to make the archive searchable by meaning rather than just keywords. Instead of requiring users to know the exact terminology used in an APOD description, AstroSearch can understand the semantic meaning of a query and retrieve conceptually related astronomy images.

For example, a search for “planet with rings” can retrieve an image of Uranus even when the query does not exactly match the wording of the original caption.

How It Works

AstroSearch uses a custom Transformer encoder built from scratch in PyTorch to convert text into numerical vector representations called embeddings.

The APOD titles and explanations are first processed using a custom BPE tokenizer trained on the APOD corpus. The Transformer encoder then learns meaningful representations through contrastive learning. During training, each APOD title and its corresponding explanation form a positive pair, while other entries in the same batch act as negative examples. Using an InfoNCE loss, the model learns to bring related text representations closer together in vector space while pushing unrelated ones apart.

After training, embeddings for the APOD archive are precomputed. When a user enters a query, AstroSearch encodes the query into an embedding and compares it against the archive using cosine similarity, returning the most semantically relevant results.

To demonstrate the difference between semantic and traditional search, AstroSearch also includes a TF-IDF keyword-search baseline. Users can switch between the two approaches and compare their results using the same query.

Main Features Semantic search: Finds APOD entries based on conceptual meaning rather than exact keyword matches. Custom Transformer encoder: A bidirectional Transformer encoder built from scratch in PyTorch specifically for representation learning. Custom BPE tokenizer: A 5,000-token vocabulary trained on the APOD corpus. Contrastive learning: Uses title–explanation pairs and in-batch negatives with an InfoNCE objective to learn semantic representations. Semantic vs. keyword comparison: A built-in TF-IDF baseline lets users directly compare traditional keyword retrieval with semantic search. Large astronomy archive: Searches across 9,700+ NASA APOD entries spanning more than 30 years. Live web application: A React frontend connected to a FastAPI backend allows users to interact with the search system directly.

Technology Stack

Machine Learning: Python, PyTorch, NumPy, scikit-learn Backend: FastAPI Frontend: React, Vite NLP: Custom BPE tokenizer, Transformer encoder, contrastive learning, InfoNCE loss Search: Embedding-based cosine similarity and TF-IDF Data: NASA Astronomy Picture of the Day archive

Intended Audience

AstroSearch is designed for astronomy enthusiasts, students, educators, researchers, and anyone curious about exploring NASA’s APOD archive. It provides a more intuitive way to discover astronomy content by allowing users to search using natural-language concepts rather than having to know specific keywords, titles, or dates.

Ultimately, AstroSearch demonstrates how a custom machine-learning retrieval system can make a large scientific archive more accessible through semantic search.

Share this project:

Updates

Submission history