Inspiration
I wasted hundreds of hours pulling founder names from Wikipedia, scraping tech stacks, hunting for emails across websites, checking if companies were hiring. Every lead research tool was either $200/month or gave you dead emails and generic data.
I asked: What if you could know everything about a company in 3 seconds?
That became LeadScan. A tool that actually works for finding real leads - from Stripe to your local dentist.
What it does
For salespeople: Paste a domain. Get back founder names, founding year, tech stack, revenue estimates, hiring signals, emails, and a personalized cold email you can actually send.
For developers: Open API. No auth. Hit the endpoints directly to build on top of LeadScan.
For teams: Batch scan 20 domains, compare competitors, export to CSV, share reports with your team.
How we built it
Started with a simple scraper, but realized it only worked for tech startups. Pivoted to build a system that works for everyone:
Wikipedia integration (depth-aware wikitext parsing - those templates are evil) Sub-page scraping for small businesses (contact, about, team pages in parallel) Schema.org + JSON-LD parsing for structured data (phones, hours, addresses) Groq AI for filling gaps (founder names, revenue estimates, key products) Caching system (30-day TTL, so repeat requests are instant) Rate limiting + validation (protection against abuse) Built on Next.js 15, deployed on Vercel, open source.
Challenges we ran into
Wikipedia wikitext was a nightmare. {{Unbulleted list|a|b|c}} templates were leaking raw into the UI. Spent way too long building a depth-aware parser that actually works.
Small businesses hide their data. Startups have /pricing pages and job postings. Dentists? Phone number buried in microdata, address in a contact form. Had to scrape 5 pages per domain in parallel.
Phone extraction is weirdly hard. Emails are everywhere. Phone numbers? Contact forms, microdata, sometimes random footer text. Regex + microdata parsing required.
AI hallucination is real. Groq's good but sometimes guesses at revenue or founder names. We score lower when AI fills in data so users know to verify.
Accomplishments that we're proud of
Built a scraper that works equally well for Stripe and your local coffee shop. Designed a scoring algorithm that actually predicts "will this person respond to cold email?" Created AI outreach that doesn't sound like a bot (actually personalized per company). Made the whole thing free and open source so people can run it themselves. Shipped production-grade features: caching, rate limiting, validation, error handling.
What we learned
Small businesses are different. You can't use startup heuristics. You need schema.org, sub-page scraping, regex patterns for phones. Build once for startups, then multiply the effort for everyone else.
Optional AI > broken AI. When the Groq key is missing, fall back to rule-based detection. Users get value either way. Don't make AI a blocker.
Caching changes everything. First request takes 5 seconds. Second request takes 50ms. That difference is huge for UX.
Real validation saves headaches. Validate domains upfront. Reject IPs, localhost, blacklisted test domains. Clear error messages. People will actually use your tool if it feels polished.
What's next for Leadscan
Bulk enrichment API - give us 1000 domains, we'll rank them all by conviction score CRM integrations - Salesforce, HubSpot plugins to enrich your contacts Real-time signals - job posting changes, funding announcements, new hires Paid tiers - free tier covers most use cases; paid for priority scraping + bulk operations
Log in or sign up for Devpost to join the conversation.