About the Project

Inspiration

News spreads faster than ever, but determining whether a claim is trustworthy is still difficult. Most fake-news systems reduce this problem to a simple Fake/Real classification, while LLM-based approaches can produce convincing answers without showing how they reached them.

We wanted to build something more transparent.

NewsPortal combines machine learning, prompt engineering, information retrieval, and external fact-checking to create an evidence-aware news verification pipeline.

What We Built

NewsPortal is an AI-powered news discovery and verification platform. Its core verification system follows a multi-stage pipeline:

News Article → ML Ensemble + Claim Extraction → Evidence Retrieval → LLM Evidence Analysis → Score Fusion → AI Explanation

Instead of asking an LLM to directly decide whether an article is fake, different components perform specialized tasks and their signals are combined to produce the final result.

Machine Learning

The ML layer uses dataset-specific ensembles trained and evaluated on WELFake, LIAR, and ISOT.

  • WELFake: XGBoost, LightGBM, Random Forest, SVM, SGD, Logistic Regression, and Naive Bayes
  • LIAR: XGBoost, Random Forest, SVM, SGD, Logistic Regression, and Naive Bayes
  • ISOT: DistilBERT, XGBoost, LightGBM, Random Forest, and Logistic Regression

Classical models use fitted TF-IDF representations, while DistilBERT processes raw text independently for the ISOT pipeline.

The individual predictions are normalized into a common Real/Fake representation and combined to produce an ML consensus score. This gives the verification system an independent machine-learning signal rather than making the LLM responsible for the entire decision.

Prompt Engineering

Prompt engineering is a core part of NewsPortal's verification pipeline.

We use Groq with Llama 3.1 8B through three focused LLM stages:

  1. Claim Extraction
    The article is converted into a specific, verifiable claim that can be investigated independently.

  2. Evidence Analysis
    Search results for the claim are analyzed and classified as:

    • SUPPORT
    • CONTRADICT
    • UNRELATED
  3. Reason Generation
    Once the numerical verification result has been calculated, the LLM generates concise reasoning explaining the result to the user.

This separation follows a simple principle:

Extract → Verify → Explain

The LLM is therefore used where language understanding is valuable, rather than being treated as a black-box truth detector.

Evidence Retrieval & Fact-Checking

After claim extraction, NewsPortal searches the web for supporting or contradicting evidence.

The retrieved evidence is analyzed by the LLM and converted into a web-evidence score. The same claim is also checked against Google Fact Check Tools to identify previously published professional fact-checks.

This gives the system multiple independent sources of information:

ML prediction + live web evidence + professional fact-checking

Score Fusion

The final verification score combines the ML prediction with the web evidence.

First, the raw web score is converted into a normalized fake-news score:

$$ \mathrm{web_fake} = 1 - \frac{\mathrm{web_raw}}{100} $$

We then calculate how decisive the web evidence is:

$$ \mathrm{neutrality} = \left|\mathrm{web_fake} - 0.5\right| \times 2 $$

The influence of the web evidence is calculated as:

$$ \mathrm{web_weight} = 0.7 \times \mathrm{neutrality} $$

Finally, the ML and web signals are combined:

$$ \mathrm{score} = \mathrm{web_weight} \times \mathrm{web_fake} + \left(1-\mathrm{web_weight}\right) \times \mathrm{ml_score} $$

This makes the system evidence-aware. Strong, decisive web evidence can significantly influence the prediction, while weak or ambiguous evidence does not completely override the ML signal.

When a recognized professional fact-check is available, it can provide an authoritative result for the claim.

Why This Approach Is Different

The goal is not simply to produce another Fake/Real classifier. The goal is to make the decision traceable.

A single LLM prompt may produce a verdict, but NewsPortal can expose:

  • ML model predictions
  • Final verification confidence
  • Extracted claim
  • Live evidence
  • Professional fact-check information
  • Web evidence influence
  • AI-generated reasoning

This makes the verification process easier to inspect and understand instead of hiding the decision behind a single model output.

How We Built It

NewsPortal uses a modular architecture:

  • React 19 + Vite for the frontend
  • Node.js + Express for the main API and application logic
  • MongoDB for users, news, preferences, and application data
  • FastAPI + Python for the ML and verification microservice
  • Scikit-learn for classical ML models and TF-IDF
  • PyTorch + Hugging Face for DistilBERT
  • Groq / Llama 3.1 8B for prompt-engineered LLM tasks
  • Serper for live web search
  • NewsData.io + Google Trends for news discovery and trends

The Node.js backend acts as the main application layer, while the FastAPI service isolates the Python ML environment. The two communicate through REST APIs, allowing the ML pipeline to remain independent from the main application.

What We Learned

Building NewsPortal taught us that combining multiple specialized components can be more useful than relying on a single powerful model.

We gained practical experience with ensemble learning, NLP, transformer models, prompt engineering, information retrieval, LLM integration, API design, microservices, and model deployment.

One of our biggest learnings was the importance of giving LLMs specific responsibilities and structured outputs. Claim extraction, evidence analysis, and explanation are easier to control and evaluate when they are separated into distinct steps.

Challenges

One major challenge was integrating models trained on different datasets while keeping their outputs consistent. Each dataset required its own preprocessing, model pipeline, and label normalization.

We also faced dependency and deployment challenges while serving multiple ML models through FastAPI.

Another challenge was handling web evidence. Search results can be incomplete, irrelevant, or contradictory, so the system needed a way to prevent weak evidence from overwhelming the ML prediction. This led to the neutrality-aware score-fusion mechanism.

Finally, designing prompts that consistently produced useful claims and structured evidence classifications required iterative prompt refinement.

Final Result

NewsPortal goes beyond simply answering “Is this news fake?”

It attempts to answer:

“What is the claim, what does our ML system predict, what evidence supports or contradicts it, and why did we reach this result?”

The result is an evidence-aware verification system that combines machine learning and prompt engineering to make misinformation analysis more transparent, explainable, and useful.

Share this project:

Updates

Submission history