-
NewsPortal combines trusted news discovery with AI-powered verification to help users identify misleading information.
-
HomePage shows Personalized news aggregated from trusted sources, with trends and recommendations for relevant discovery.
-
AI verification analyzes articles using ML ensembles, prompt-engineered claim extraction, live web evidence, and fact-checking.
-
Transparent verification results combine ML predictions, evidence, confidence scores, and AI-generated explanations.
-
Modular architecture connecting React, Node.js, FastAPI ML services, MongoDB, web search, fact-checking, and Groq.
About the Project
Inspiration
News spreads faster than ever, but determining whether a claim is trustworthy is still difficult. Most fake-news systems reduce this problem to a simple Fake/Real classification, while LLM-based approaches can produce convincing answers without showing how they reached them.
We wanted to build something more transparent.
NewsPortal combines machine learning, prompt engineering, information retrieval, and external fact-checking to create an evidence-aware news verification pipeline.
What We Built
NewsPortal is an AI-powered news discovery and verification platform. Its core verification system follows a multi-stage pipeline:
News Article → ML Ensemble + Claim Extraction → Evidence Retrieval → LLM Evidence Analysis → Score Fusion → AI Explanation
Instead of asking an LLM to directly decide whether an article is fake, different components perform specialized tasks and their signals are combined to produce the final result.
Machine Learning
The ML layer uses dataset-specific ensembles trained and evaluated on WELFake, LIAR, and ISOT.
- WELFake: XGBoost, LightGBM, Random Forest, SVM, SGD, Logistic Regression, and Naive Bayes
- LIAR: XGBoost, Random Forest, SVM, SGD, Logistic Regression, and Naive Bayes
- ISOT: DistilBERT, XGBoost, LightGBM, Random Forest, and Logistic Regression
Classical models use fitted TF-IDF representations, while DistilBERT processes raw text independently for the ISOT pipeline.
The individual predictions are normalized into a common Real/Fake representation and combined to produce an ML consensus score. This gives the verification system an independent machine-learning signal rather than making the LLM responsible for the entire decision.
Prompt Engineering
Prompt engineering is a core part of NewsPortal's verification pipeline.
We use Groq with Llama 3.1 8B through three focused LLM stages:
Claim Extraction
The article is converted into a specific, verifiable claim that can be investigated independently.Evidence Analysis
Search results for the claim are analyzed and classified as:SUPPORTCONTRADICTUNRELATED
Reason Generation
Once the numerical verification result has been calculated, the LLM generates concise reasoning explaining the result to the user.
This separation follows a simple principle:
Extract → Verify → Explain
The LLM is therefore used where language understanding is valuable, rather than being treated as a black-box truth detector.
Evidence Retrieval & Fact-Checking
After claim extraction, NewsPortal searches the web for supporting or contradicting evidence.
The retrieved evidence is analyzed by the LLM and converted into a web-evidence score. The same claim is also checked against Google Fact Check Tools to identify previously published professional fact-checks.
This gives the system multiple independent sources of information:
ML prediction + live web evidence + professional fact-checking
Score Fusion
The final verification score combines the ML prediction with the web evidence.
First, the raw web score is converted into a normalized fake-news score:
$$ \mathrm{web_fake} = 1 - \frac{\mathrm{web_raw}}{100} $$
We then calculate how decisive the web evidence is:
$$ \mathrm{neutrality} = \left|\mathrm{web_fake} - 0.5\right| \times 2 $$
The influence of the web evidence is calculated as:
$$ \mathrm{web_weight} = 0.7 \times \mathrm{neutrality} $$
Finally, the ML and web signals are combined:
$$ \mathrm{score} = \mathrm{web_weight} \times \mathrm{web_fake} + \left(1-\mathrm{web_weight}\right) \times \mathrm{ml_score} $$
This makes the system evidence-aware. Strong, decisive web evidence can significantly influence the prediction, while weak or ambiguous evidence does not completely override the ML signal.
When a recognized professional fact-check is available, it can provide an authoritative result for the claim.
Why This Approach Is Different
The goal is not simply to produce another Fake/Real classifier. The goal is to make the decision traceable.
A single LLM prompt may produce a verdict, but NewsPortal can expose:
- ML model predictions
- Final verification confidence
- Extracted claim
- Live evidence
- Professional fact-check information
- Web evidence influence
- AI-generated reasoning
This makes the verification process easier to inspect and understand instead of hiding the decision behind a single model output.
How We Built It
NewsPortal uses a modular architecture:
- React 19 + Vite for the frontend
- Node.js + Express for the main API and application logic
- MongoDB for users, news, preferences, and application data
- FastAPI + Python for the ML and verification microservice
- Scikit-learn for classical ML models and TF-IDF
- PyTorch + Hugging Face for DistilBERT
- Groq / Llama 3.1 8B for prompt-engineered LLM tasks
- Serper for live web search
- NewsData.io + Google Trends for news discovery and trends
The Node.js backend acts as the main application layer, while the FastAPI service isolates the Python ML environment. The two communicate through REST APIs, allowing the ML pipeline to remain independent from the main application.
What We Learned
Building NewsPortal taught us that combining multiple specialized components can be more useful than relying on a single powerful model.
We gained practical experience with ensemble learning, NLP, transformer models, prompt engineering, information retrieval, LLM integration, API design, microservices, and model deployment.
One of our biggest learnings was the importance of giving LLMs specific responsibilities and structured outputs. Claim extraction, evidence analysis, and explanation are easier to control and evaluate when they are separated into distinct steps.
Challenges
One major challenge was integrating models trained on different datasets while keeping their outputs consistent. Each dataset required its own preprocessing, model pipeline, and label normalization.
We also faced dependency and deployment challenges while serving multiple ML models through FastAPI.
Another challenge was handling web evidence. Search results can be incomplete, irrelevant, or contradictory, so the system needed a way to prevent weak evidence from overwhelming the ML prediction. This led to the neutrality-aware score-fusion mechanism.
Finally, designing prompts that consistently produced useful claims and structured evidence classifications required iterative prompt refinement.
Final Result
NewsPortal goes beyond simply answering “Is this news fake?”
It attempts to answer:
“What is the claim, what does our ML system predict, what evidence supports or contradicts it, and why did we reach this result?”
The result is an evidence-aware verification system that combines machine learning and prompt engineering to make misinformation analysis more transparent, explainable, and useful.
Log in or sign up for Devpost to join the conversation.