Enhanced Investor Document Verification System: SERVER Burners
Inspiration
The world of venture capital is flooded with pitch decks, and verifying the credibility of these documents is often a manual and error-prone process. We were inspired by the power of AI to automate due diligence, reduce human error, and uncover hidden risks. Tools like large language models and Perplexity Sonar presented a unique opportunity to create a system that could act as a “digital analyst team,” making investment decisions faster, smarter, and more transparent.
What it does
Our system automates the entire investor document verification pipeline:
- Accepts PDF/DOCX pitch decks via a web-based dashboard.
- Parses and extracts relevant data using NLP.
- Engages multiple specialized AI agents (financial, legal, environmental, competitor analysis, claim validation).
- Aggregates their findings into a coherent, traceable report.
- Outputs insights through visual dashboards and downloadable reports in PDF, JSON, and HTML formats.
- Tracks every step with audit logs for compliance.
How we built it
- Backend (FastAPI + Celery):
- High-performance async APIs for responsiveness.
- SQLAlchemy ORM + Pydantic models for robust data handling.
- Celery + Redis for background processing.
- Frontend (Streamlit):
- Real-time status updates, file uploads, and visual report interface.
- Score visualization and downloadable reports.
- Document Parsing & NLP:
- Used
PyMuPDF,pdfplumber,python-docxfor ingestion. - Applied
spaCyand HuggingFace transformers for NER and fact extraction.
- External Integration:
- Integrated Perplexity Sonar API for real-time web research.
- Implemented rate-limiting, retries, and fallbacks for reliability.
- DevOps:
- Local and Docker-based deployment support.
- Auto-setup scripts and documentation for contributors.
Challenges we ran into
- Async Orchestration: Coordinating agents and background tasks without race conditions.
- API Reliability: Handling Perplexity API rate limits and downtime with retries and graceful fallbacks.
- Cross-Platform Dependencies: Issues with
libmagicand Redis on Windows required extra setup steps. - Serialization: Ensuring nested models from multiple agents were JSON-serializable.
- Schema Syncing: Rapid iteration meant constant database and model updates.
- Frontend Feedback: Communicating real-time task progress in Streamlit while handling long processing times.
Accomplishments that we're proud of
- End-to-End Automation: Complete pipeline from pitch upload to insight report.
- Multi-Agent System: Specialized agents contribute domain-specific analysis.
- Explainability: Traceable decisions backed by full audit logs and detailed outputs.
- Security & Compliance: GDPR features and robust validation built-in.
- Scalable Design: Easily extensible to new agents and deployment environments.
What we learned
- How to orchestrate multiple AI agents and combine their outputs effectively.
- Async programming and background processing with FastAPI + Celery.
- Robust integration of external APIs with retry and fallback logic.
- Importance of clean data modeling and serialization with SQLAlchemy and Pydantic.
- Best practices for compliance, including audit trails and user data safety.
What's next for SERVER Burners
- Agent Expansion: Add more agent types (e.g., fraud detection, market sentiment).
- LLM Optimization: Fine-tune smaller models for local inference to reduce API dependency.
- Collaboration Features: Enable investors to comment, annotate, and share findings securely.
- Enterprise Features: Role-based access, multi-user dashboards, and advanced analytics.
- AI Feedback Loop: Train models using feedback from real investment decisions to improve accuracy.
Log in or sign up for Devpost to join the conversation.