-
AI-powered dashboard for extracting microplastic degradation data from research papers.
-
Designed to work with Grok, OpenAI, Ollama, Gemini, and other compatible AI providers.
-
python-based backend connecting PDF processing with AI-powered scientific data extraction.
-
Modular extraction pipeline supporting multiple AI providers and structured outputs.
-
System showing the workflow research PDFs to AI extraction, validation, structured datasets, and future ML-based degradation prediction.
Inspiration
Microplastic research is growing rapidly, but a major challenge is that experimental data is scattered across hundreds of research papers. Important information such as polymer type, particle size, treatment conditions, degradation time, catalyst concentration, temperature, pH, and degradation percentage is usually buried inside paragraphs, tables, figures, and supplementary information.
As a chemistry student working on microplastic degradation research, I wanted to solve this problem from the data side: Can AI turn scattered experimental information in research papers into a structured, traceable dataset that researchers can actually use?
This idea led to Microplastic Degradation Data Extractor AI — a tool designed to transform research papers into structured experimental records that can eventually support machine-learning models for predicting and optimizing microplastic degradation.
What it does
The project extracts experimental information from microplastic research papers and converts it into structured data.
It focuses on variables such as:
- Polymer type — PE, PP, PS, PET, PVC, PA/Nylon
- Polymer form and particle size
- Degradation/treatment method
- Catalyst or reagent
- Catalyst loading
- Temperature and pH
- Treatment duration
- Initial and final mass
- Degradation/removal endpoints
- Measurement methods
- Experimental evidence
- Page, table, and figure references
The system supports processing multiple research-paper PDFs and produces structured JSON and CSV datasets.
A key principle is:
If the paper does not report a value, the system should not invent one.
The goal is not simply to extract numbers, but to preserve the scientific context and provenance behind those numbers.
How I built it
I built the project as a modular Python application with a Streamlit interface.
The basic pipeline is:
Research PDF → Text extraction → AI extraction → Structured schema → Normalization → Validation → JSON/CSV dataset
The main technologies include:
- Python for the core application
- Streamlit for the interactive interface
- Pydantic for structured scientific data schemas
- PyMuPDF for extracting text from PDFs
- Pandas for dataset handling
- OpenAI, Ollama/Llama, and Grok API as AI extraction approaches
- JSON/CSV for research-data output
I initially experimented with OpenAI for structured extraction. API costs became a limitation for processing many research papers, so I explored Ollama and local Llama models as an alternative. Ollama removed API costs but was slower on local hardware. I then explored the Grok API as a faster cloud-based option.
This evolution helped me design the application around a provider-independent extraction layer rather than tying the project to a single AI model.
Challenges I ran into
One of the biggest challenges was that scientific papers are not written as clean datasets.
The same variable can appear in different formats across papers. For example, one paper may report degradation as percentage mass loss, another as particle-size reduction, another as TOC mineralization, and another through spectroscopic evidence.
I also faced challenges with:
- Extracting information from long research PDFs
- Separating multiple experiments within the same paper
- Preserving the relationship between conditions and results
- Handling missing experimental values
- Normalizing different polymer names and terminology
- Validating extracted numerical values
- Managing AI API costs
- Running local LLMs efficiently
- Making the system work with multiple PDFs
- Maintaining evidence and provenance for extracted values
A particularly important scientific challenge is distinguishing degradation, removal, fragmentation, and adsorption. A particle disappearing from a measurement does not necessarily mean that the polymer was chemically degraded.
Accomplishments that I'm proud of
I built a working foundation for converting unstructured microplastic research literature into structured experimental data.
Some parts I am particularly proud of are:
- Building a multi-PDF extraction workflow
- Creating a structured Pydantic schema for experimental data
- Adding normalization and validation
- Supporting multiple AI providers
- Designing the system around evidence and provenance
- Creating JSON and CSV outputs suitable for further analysis
- Connecting a chemistry research problem with AI, NLP, and data science
More importantly, the project moves toward a larger research goal: building a high-quality experimental dataset that can eventually be used to study and predict microplastic degradation.
What I learned
This project taught me that applying AI to scientific research is much more than sending a PDF to an LLM and asking for a summary.
I learned about:
- Scientific information extraction
- Prompt engineering for structured outputs
- Pydantic data validation
- PDF text processing
- Data normalization
- Research-data provenance
- LLM API integration
- Local LLM deployment with Ollama
- Dataset design
- The importance of missing-data handling
- The difference between scientific endpoints
- Designing software around interchangeable AI providers
I also learned that data quality is just as important as model performance. A powerful ML model cannot compensate for incorrectly extracted or poorly defined experimental data.
What's next for Microplastic Degradation Data Extractor AI
The current project is the foundation for a much larger research pipeline.
My next goals are:
- Build a dataset from 50+ primary microplastic degradation papers.
- Add stronger paper metadata such as DOI, authors, year, and
paper_id. - Improve extraction from tables, figures, and scanned PDFs.
- Add better evidence and page-level provenance.
- Introduce automated extraction-quality checks.
- Build a standardized microplastic degradation database.
- Analyze relationships between polymer properties, treatment conditions, and degradation.
- Develop machine-learning models to predict degradation performance.
- Compare different degradation technologies using standardized experimental variables.
- Eventually explore optimization models for selecting degradation conditions.
The long-term vision is to move from:
Research Papers → Structured Data → Scientific Dataset → Machine Learning → Degradation Prediction & Optimization
and use AI not just to read the literature faster, but to help turn fragmented scientific knowledge into reproducible, machine-readable research data.
Built With
- ai
- chemistry
- csv
- data-extraction
- data-science
- environmental-science
- git
- github
- grok-api
- json
- llama-3.1
- llm
- machine-learning
- microplastics
- natural-language-processing
- ollama
- openai
- pandas
- pdf-processing
- pydantic
- pymupdf
- python
- research
- scientific-data
- streamlit


Log in or sign up for Devpost to join the conversation.