Inspiration

In July 2026, Bloomberg and an independent researcher caught a consumer shopping app claiming affiliate commissions on sales it hadn't driven. The mechanism, confirmed across more than fifty retail sites: the app's code opened a background tab during checkout and overwrote the referral credential of whichever party had actually earned it — including shoppers who'd arrived through a completely different, legitimate affiliate link. The affiliate network suspended the app within days. The company called it a bug from a prior release and fixed it within a day of being contacted.

This is a real, named category — cookie-stuffing, attribution fraud, litigated before against PayPal's Honey. It is not a consent-banner violation, and a static code review would not have caught it, because the behavior is conditionally triggered, minified, and only exists at runtime.

That gap is what we wanted to close, and the incident above isn't an outlier. A 2026 academic study [1] analyzed 1,000 Android apps and nearly 87 million log entries and found that roughly two-thirds of them collect or log data their own privacy policy never mentions. Usually that's not malice — the team writing the policy and the team writing the logging code rarely talk to each other — but the outcome is the same undisclosed data flow, just at a much larger scale than one shopping app.

We're also skeptical of compliance and security tools that quietly overclaim what they actually detected, so we set a harder bar for ourselves: every capability that couldn't be made fully reliable would have to disclose exactly where it falls short, in the same report, every time. That discipline ended up shaping most of the interesting engineering in this build.

Right now, companies find out their product doesn't match what it discloses from a reporter, a regulator, or a plaintiff's lawyer — never from their own engineering or compliance process. Verascope exists so that discovery happens internally, before it happens publicly.

What it does

You give Verascope a target, one of three ways: a quick demo against a controlled example, your own repository built and run live in a sandbox, or a live URL you attest you're authorized to test. Four agents run against it:

  • Agent 1 — Code health & team risk. Dependency install, test suite results, CI/CD coverage, and authorship concentration across the most-changed files — all reported as data points, not verdicts.
  • Agent 2 — Security, licensing & AI-vendor exposure. Committed secrets, copyleft dependencies, hardcoded model calls with no fallback layer.
  • Agent 3 — Runtime behavior & disclosure (the marquee capability). Reads your actual privacy policy and extracts every claim about data sharing and third parties, each cited to its exact source sentence. Then runs your product for real — real browser, real network traffic — and diffs what it observes against what was disclosed. Every request lands in one of three buckets: conforms, undisclosed, or contradicted.
  • Agent 4 — Report synthesis. One severity-tiered report, with coverage details — which mode ran, which flows were tested — placed right after the executive summary, because it changes how everything below it should be read.

The demo storefront proves the mechanism: a scan surfaces nine cited findings, anchored by the planted attribution-override pattern firing on an unclicked background request. But a controlled demo only proves the pattern, not the product, so we ran the full pipeline against a live URL we actually own and control, start to finish: three reported flows, six cited findings, a recorded consent timestamp, and zero credential values retained anywhere in the report.

One rule holds across all four agents: no finding survives without a citation — an exact file and line, or an exact command and its real output. That check runs in application code, not just a prompt instruction, so a model that gets sloppy can't slip a finding through. And Verascope never tells you that you violated a specific law. It tells you what it observed and flags it for legal and compliance review — determining lawful basis is a human job.

How we built it

The core stack:

Layer Technology Role
App Next.js 14 (App Router), TypeScript, Tailwind, shadcn/ui Single deployable — frontend and backend in one runtime
Reasoning Google Gen AI SDK (gemini-3.5-flash) Powers all four agents' analysis, behind a citation-grounding gate that rejects any finding that can't be traced to real evidence
Sandbox execution Vercel Sandbox Isolated, ephemeral execution for repo builds and runtime browser traces
Data Postgres via Supabase Scan history and state
Deployment Vercel Hosting for the single deployable

The interesting engineering lives in Agent 3, and most of it is about knowing when to stop trusting the browser.

The original plan differentiated browser posture by mode — plain headless Chromium for demo_app and repo_build, a real installed Chrome channel for user_url to close the cheapest automation tells. That plan didn't survive contact with Vercel's actual runtime, and Challenges below covers why. What shipped instead: every mode launches the same plain, non-hardened Chromium from a pre-built, verified snapshot, so no sandbox spends its runtime installing a browser. stealthPosture reports as none on every scan, stated plainly rather than implied otherwise.

Reading the privacy policy has its own honest limit. Extracting third-party-sharing claims from policy text is grounded in published NLP research: precision runs in the 90%+ range when the extractor names an entity, but recall is meaningfully lower, because policies routinely bury the real answer in vague collective language like "our partners." A vague claim gets extracted and flagged as vague — never quietly filled in with a guess.

Accepting a user-supplied URL and having a server fetch it is the textbook SSRF shape, so before any external host is contacted we resolve the hostname and reject anything that lands in a private, loopback, link-local, or cloud-metadata range — then re-check immediately before the actual fetch, closing the DNS-rebinding gap between the first check and the request. The guard's own test suite runs 34 assertions across IPv4, IPv6, DNS-failure, and rebinding cases before we trusted it with a real scan.

How we used Codex

Every feature in Verascope — the SSRF guard, both static agents, the full five-stage runtime pipeline, the report UI — was built with Codex from Phase 1 through Phase 4, working almost entirely in its goal-tracking mode: set a goal, let it work autonomously, and when it hits a real blocker instead of a trivial one, it stops and asks rather than guessing past it.

Where Codex accelerated the workflow. The clearest example is the SSRF guard: we set the goal test-first — reject every unsafe IPv4, IPv6, mapped-IP, malformed, and DNS-failure case, plus DNS-rebinding, with no cached approval on repeat calls — and Codex built the full 34-assertion suite against that goal in one sitting, the kind of exhaustive edge-case coverage that's slow and tedious to write by hand. The same pattern held for the Gemini-call resilience layer: once flaky upstream 503/429/500/504 responses showed up during a live run, we set the goal — don't fail a scan on a transient provider hiccup — and Codex implemented the retry-with-backoff logic and its contract tests in the same pass.

Where we made the key product, engineering, and design decisions. Codex executes toward a goal; the goals themselves, and the calls about how to handle a blocker, were ours. Three concrete examples: when the original plan to run static-agent model calls through the OpenAI Agents SDK stopped being the right fit, we made the call to swap to the Google Gen AI SDK for Gemini rather than force it. When Vercel's sandbox runtime turned out not to support the differentiated Chrome-channel browser posture we'd planned for user_url, we made the call to cut that scope three days before the deadline rather than keep chasing a browser-packaging problem — and to disclose stealthPosture: none honestly instead of quietly shipping a lesser version without saying so. And when a required Gemini API key wasn't available mid-build, we made the call to defer that test rather than let Codex substitute synthetic findings just to keep the goal moving.

How GPT-5.6 and Codex contributed to the final result. Codex — running on GPT-5.6 — wrote effectively all of the application code across every phase: the SSRF guard, both static agents, the full runtime pipeline, sandbox lifecycle management, and the report UI. Verascope's own deployed agents reason with Gemini rather than GPT-5.6 directly, because the product's static and runtime analysis needed Gemini specifically — but GPT-5.6 was the model behind every decision Codex executed in building that system, from the first scaffolded file to the last passing contract test.

Challenges we ran into

Building the runtime diff engine was the easy part. Aiming it at the right thing was much harder. An independent assessment of v2 found that Agent 3's primary test was pointed at the wrong pattern entirely — checking for consent-banner violations, when the incident that motivated the whole product was attribution-override fraud, a completely different mechanism. Worse, the test had a structural ceiling no amount of tuning could fix: no matter how good it got, it could never run against anything except the demo app we'd built for ourselves. Our own PRD described a real-target input in the user journey that was never actually wired to the backend, the frontend, or the technical spec — all three quietly hardcoded the runtime target shut.

v3 fixes both problems as one change, not two. The general diff engine stays the umbrella — every observed request still sorts into conforms, undisclosed, or contradicted, so the tool catches whatever gap exists, not only the one pattern we happened to read about — but attribution-override is now sharpened into the highest-severity check inside that system, because it reproduces the exact fraud mechanism instead of a proxy for it. v2 had actually gotten the precision backwards: treating any unscripted request as suspect sounds thorough, but it mostly flags ordinary analytics and manufactures false urgency. And there are now three real, user-chosen ways to reach a target instead of one fixed demo, so the fix that matters is provable, not just described.

Two smaller limits are worth naming honestly rather than glossing over. The browser posture is the bigger one: we'd planned a real installed Chrome channel for user_url specifically, to close the cheapest automation tells, but Vercel's sandbox runtime ships neither Chrome nor Chromium on Node 22 or Node 24, and Playwright's Chrome-channel installer refused to provision one itself — Vercel's sandbox identifies as Amazon Linux, and that installer only supports Ubuntu and Debian. With the deadline close, we cut the differentiation instead of chasing it further: every mode ships the same plain Chromium now, and the report states stealthPosture: none outright rather than implying coverage the build doesn't have. And the CNAME-cloaking check that flags disguised third-party trackers catches only the classic version of that evasion by design — it doesn't detect generic-cloud or direct server-side tracking, and the report never implies otherwise.

Accomplishments that we're proud of

No silent target substitution. If user_url mode fails to resolve, the run is marked skipped with a specific reason — never quietly swapped for the demo app, because that would produce a misleading report dressed up as a resilient one.

Graceful degradation. A repo that can't self-build in the sandbox — which is most real production repos, missing a database or API keys a bare checkout doesn't have — still returns a complete, honest, static-only report instead of failing the entire scan. Proved against a real public repository, not a synthetic fixture: vercel/next-learn came back repo_not_runnable, and the report still carried its cited static findings.

No stealth evasion against non-consenting targets. That line existed in every draft of this product, and it survived every rewrite.

What we learned

The biggest lesson was that a runtime-testing product is only as real as its weakest wiring. It's easy to write a user journey that describes testing a real target; it's a different thing to actually route that choice through the backend, the sandbox, and the report. We also learned that shipping the honest, lesser version of a feature beats chasing the originally scoped one against a deadline: when Vercel's runtime turned out not to support the browser posture we'd planned, disclosing stealthPosture: none plainly earned more trust than a half-finished workaround would have. A check that names its own blind spot is worth more than one that pretends not to have any.

What's next for Verascope

  • All six analysis categories instead of four, with per-jurisdiction consent-law logic
  • Private repos via secure upload or scoped OAuth, instead of public GitHub only
  • A CI-integrated pass/fail gate, so a release can't ship a new undisclosed data flow without someone seeing it first
  • Full multi-tenant auth, so this stops being a single-demo-user tool and becomes something a compliance team actually runs on a schedule

Sources

[1] Chen Z., Ahir L. J., Suleiman A., Yao K., Tang Y., Shang W., & Hou D. (2026). Do Privacy Policies Match with the Logs? An Empirical Study of Privacy Disclosure in Android Application Logs. In Proceedings of the 30th International Conference on Evaluation and Assessment in Software Engineering (EASE 2026). ACM. arXiv:2604.18552. https://ece.uwaterloo.ca/~wshang/pubs/log_policy_ease_2026.pdf

Built With

Share this project:

Updates

Submission history