We will be undergoing planned maintenance on Oct 7th 6:00AM UTC / Oct 7th 2:00AM ET

Signal

Inspiration

There is an enormous amount of data available about the world, but most of it is difficult to understand without spending hours cleaning, analyzing, and visualizing it.

A dataset might contain thousands or millions of rows, but a person looking at it usually wants to know something much simpler: what is changing, why is it changing, and what should I pay attention to?

I wanted to build something that could take a large, messy dataset and turn it into an interactive story that people can actually understand.

What it does

Signal is an AI-powered data exploration platform that turns raw datasets into interactive explanations.

A user can upload a dataset or connect a public data source. Signal automatically analyzes the data, identifies interesting patterns and relationships, and creates visualizations around the most important findings.

Instead of asking the user to figure out what questions to ask, Signal proactively searches for things that are worth investigating.

For example, given a dataset about cities, Signal might identify that one city's population is growing unusually quickly, discover which factors are correlated with that growth, and create an interactive visualization showing how the trend developed over time.

Signal can:

  • Automatically profile a dataset
  • Find unusual trends and outliers
  • Detect correlations between variables
  • Generate interactive charts
  • Explain statistical findings in plain language
  • Let users ask follow-up questions about the data
  • Show the underlying data and calculations behind each finding

The main idea is to turn data analysis from a blank canvas into an exploration where the most interesting discoveries are surfaced automatically.

How I built it

I built Signal as a web application with an interactive data visualization interface.

When a dataset is uploaded, the application first analyzes its structure, including column types, missing values, distributions, and relationships between variables.

A Python analysis pipeline then runs statistical tests and searches for trends, correlations, clusters, and unusual observations.

The results are passed to an AI model that determines which findings are interesting enough to surface and generates explanations based on the actual analysis.

The main workflow is:

  1. Upload or select a dataset.
  2. Automatically profile the data.
  3. Search for statistically interesting patterns.
  4. Rank the most significant findings.
  5. Generate interactive visualizations.
  6. Explain each finding using the underlying data.
  7. Ask follow-up questions and explore the dataset further.

I wanted to make sure the AI could not simply invent an interesting-sounding conclusion. Each generated insight is connected to an actual analysis result and visualization so the user can inspect the evidence.

Challenges I ran into

One of the biggest challenges was deciding what makes a pattern interesting.

A dataset can contain hundreds of statistically significant relationships that are not actually useful or meaningful. Signal therefore needs to consider more than just statistical significance when deciding which discoveries to show.

Another challenge was making the AI work with arbitrary datasets. Different datasets can have completely different structures, units, and meanings, so the analysis pipeline needs to adapt rather than assuming a fixed schema.

I also had to think carefully about how to present uncertainty. A correlation does not necessarily mean that one variable causes another, and an unusual observation is not automatically an error.

The goal was to make Signal feel intelligent without hiding the actual analysis behind an AI-generated explanation.

What I learned

I learned that data analysis is often less about calculating statistics and more about deciding which questions are worth asking.

Modern tools make it relatively easy to create a chart once you know what you want to investigate. The difficult part is finding the interesting questions in the first place.

I also learned that AI can be much more useful when it is connected to deterministic tools. Instead of letting a language model make conclusions from raw numbers on its own, Signal uses Python for the actual analysis and AI for interpreting and communicating the results.

This combination makes it possible to explore large amounts of data while still keeping the results grounded in the underlying calculations.

What's next for Signal

I would like to expand Signal into a more autonomous data research system.

Instead of analyzing one dataset at a time, future versions could combine multiple public datasets and automatically discover relationships between them.

For example, Signal could combine economic, environmental, geographic, and demographic datasets to investigate why certain trends occur in different regions.

I would also like to add an agent that can formulate its own hypotheses, test them against the data, reject weak explanations, and continue investigating promising ones.

Another feature I would like to build is collaborative exploration. Multiple users could investigate the same dataset, save discoveries, and build a shared visual story around what they find.

My goal is to make large datasets feel less like spreadsheets and more like something people can explore and discover.

Share this project:

Updates

Submission history