🔭 COSMIC Data Fusion

🌌 Inspiration

The universe does not arrive neatly organized into one dataset.

Astronomical observations are collected by different surveys, telescopes, and instruments, each capturing different properties of celestial objects. A single object may have information about its position from one catalog, photometric properties from another, and additional observational data from another.

We wanted to explore a simple question:

What happens when we stop looking at astronomical datasets independently and start combining their information?

That question led us to COSMIC Data Fusion, a machine-learning pipeline that brings together heterogeneous astronomical data and uses the resulting fused representation to classify celestial objects.

Our goal was to build more than a classifier. We wanted to explore how data fusion can transform fragmented astronomical observations into a unified representation that can support scientific analysis.


🔬 What We Built

COSMIC Data Fusion is an astronomical data-fusion and machine-learning pipeline designed to integrate, process, and classify celestial objects across multiple astronomical catalogs.

Our benchmark combines cross-matched data from:

  • SDSS DR17
  • 2MASS
  • Gaia DR3
  • A built-in astronomical catalog

We evaluated the pipeline on a standardized benchmark of 1,000 celestial targets across five primary astrophysical classes:

  • ⭐ Star
  • 🌌 Galaxy
  • 🌀 Nebula
  • ✨ Cluster
  • 🔭 AGN / QSO

The pipeline takes heterogeneous astronomical observations, preprocesses and fuses their features, and feeds the resulting representation into a classification model.


⚙️ How We Built It

1. Data Collection & Cross-Matching

We gathered astronomical information from multiple catalogs and combined observations belonging to corresponding celestial targets.

The challenge was not simply collecting more data, but making information from different astronomical sources usable together.

2. Data Preprocessing

The datasets contained differences in structure, scale, feature availability, and representation.

We therefore performed preprocessing and feature preparation before constructing the unified dataset used by the machine-learning pipeline.

3. Data Fusion

The processed features from the different astronomical sources were combined into a unified representation.

This fusion step is the heart of COSMIC. Instead of asking a model to learn from one isolated observation, we provide it with complementary information from multiple astronomical datasets.

4. Classification

The fused representation was used to classify each target into one of five astrophysical categories:

Star | Galaxy | Nebula | Cluster | AGN/QSO

5. Evaluation

We evaluated the resulting predictions on 1,000 benchmark targets using accuracy, precision, recall, F1-score, and a confusion matrix.


📊 Results

The model correctly classified 898 out of 1,000 targets, giving an overall accuracy of:

89.8% Accuracy

Metric Result
Overall Accuracy 89.8%
Macro Precision 88.5%
Macro Recall 87.9%
Macro F1-Score 88.2%
Weighted F1-Score 90.1%

The class-level results were:

Class Precision Recall F1-Score
Star 88.6% 93.5% 0.910
Galaxy 84.3% 86.7% 0.855
Nebula 86.5% 80.0% 0.831
Cluster 89.7% 86.9% 0.883
AGN/QSO 93.6% 92.2% 0.929

These results also show why evaluating only overall accuracy is not enough. The confusion matrix allowed us to investigate which astronomical classes were being confused and where the model struggled.


🧩 Confusion Matrix

The confusion matrix below shows the predictions made across all five astrophysical classes.

Rows represent the actual class, while columns represent the predicted class.

Actual \ Predicted Star Galaxy Nebula Cluster AGN/QSO
Star 187 4 0 3 6
Galaxy 5 156 1 2 16
Nebula 3 1 64 8 4
Cluster 6 2 9 113 0
AGN/QSO 10 22 0 0 378

The diagonal represents correct classifications. The model's strongest diagonal counts were observed for AGN/QSO, Stars, and Galaxies, while Nebulae showed comparatively more confusion with Clusters.

🔍 What the Errors Tell Us

The interesting part of a confusion matrix is not only what the model got right, but where it got confused.

Galaxy ↔ AGN/QSO

This was one of the most prominent sources of confusion. Some AGN/QSO targets were classified as galaxies and some galaxies were classified as AGN/QSOs.

A possible astrophysical explanation is that AGN host-galaxy light can contribute to the observed signal, producing objects with characteristics that overlap between extended galaxies and point-like AGN/QSO sources.

Nebula ↔ Cluster

Nebulae and clusters also showed noticeable overlap. Diffuse emission regions and stellar associations can produce similar observational characteristics, making the distinction more challenging using the available features.

Star ↔ AGN/QSO

Some AGN/QSOs were classified as stars and vice versa. Unresolved quasars can appear point-like in optical observations, creating similarities with stellar sources when high-resolution or spectroscopic information is unavailable.

This analysis helped us understand that classification errors were not simply random failures. Some occurred between classes with genuinely overlapping observational characteristics.


🧠 What We Learned

One of the biggest lessons from COSMIC was that the hardest part of an ML project is often not the model itself.

Working with scientific data introduced challenges that do not appear in clean textbook datasets.

We learned about:

  • Cross-matching heterogeneous astronomical catalogs
  • Scientific data preprocessing
  • Feature preparation and fusion
  • Machine-learning classification
  • Class-level evaluation
  • Confusion-matrix interpretation
  • Precision, recall, and F1-score
  • Understanding model errors through an astrophysical lens

The confusion matrix was particularly valuable because it allowed us to move beyond a single accuracy number and investigate why certain classes were being confused.


🚧 Challenges We Faced

Heterogeneous Data

Different astronomical catalogs use different structures, measurements, and feature sets. Combining them into a consistent representation required careful preprocessing.

Data Quality

Scientific datasets can contain missing values, inconsistent measurements, and differences in scale. Preparing the data for machine learning became an important part of the pipeline.

Class Overlap

Some astronomical objects naturally exhibit overlapping observational characteristics.

For example, an unresolved AGN/QSO can resemble a star in imaging data, while galaxies containing active nuclei can overlap with AGN/QSO characteristics.

Evaluating Beyond Accuracy

Because the classes had different numbers of examples, accuracy alone did not tell the complete story.

We therefore examined precision, recall, F1-score, and the confusion matrix to understand the model's behavior across individual classes.


🏆 Accomplishments

We built an end-to-end pipeline that takes heterogeneous astronomical observations and transforms them into a unified machine-learning workflow.

Our key accomplishments include:

  • Integrating information from SDSS DR17, 2MASS, Gaia DR3, and a built-in catalog
  • Creating a standardized benchmark containing 1,000 celestial targets
  • Supporting classification across five astrophysical classes
  • Achieving 89.8% overall classification accuracy
  • Achieving a 0.882 macro F1-score
  • Using confusion-matrix analysis to investigate astrophysically meaningful classification errors

Most importantly, we moved from simply asking "How accurate is our model?" to asking:

"What can the model's mistakes tell us about the astronomical data?"


🚀 What's Next

COSMIC Data Fusion can be extended far beyond the current benchmark.

Future versions could explore:

  • Larger astronomical surveys
  • Additional wavelengths and observational features
  • More sophisticated feature-fusion techniques
  • Deep-learning architectures
  • Anomaly and outlier detection
  • Celestial object characterization
  • Cross-survey astronomical discovery
  • More detailed astrophysical classifications

With larger and richer datasets, we hope to investigate whether data fusion can help uncover patterns that remain difficult to identify when astronomical observations are analyzed independently.


🌠 Final Takeaway

COSMIC Data Fusion explores a simple idea with a very large dataset behind it: the universe contains more information than any single catalog can capture.

By combining observations from multiple astronomical sources and applying machine learning, we built a pipeline capable of classifying celestial objects while also giving us a way to investigate where and why the model makes mistakes.

898 correct classifications out of 1,000 targets gave us the number. The confusion matrix gave us the story behind it.

Built With

Share this project:

Updates