Inspiration

A 90-year-old, $1 million mathematics problem became a warning for every organisation using AI.

On 8 September 2026, OpenAI announced a claimed solution to the Navier–Stokes Millennium Prize Problem, one of seven problems for which the Clay Mathematics Institute allocated a $1 million prize [1], [2]. OpenAI reported that an internal system involving about 10,000 concurrent agents produced the result after an 88-hour run [2]. The mathematical claim was extraordinary. The priority controversy that followed raised an equally urgent question for the technology community: what happens to valuable unpublished work after someone enters it into an AI tool?

Mathematicians Tristan Buckmaster and Levent Alpöge had spent much of the previous year developing related results while using several AI systems, including Codex. Buckmaster stated that drafts from their project had been entered into Codex and questioned whether their work could have influenced OpenAI’s system [3]. OpenAI denied that its researchers or agents accessed their work and later reported that an investigation found Buckmaster’s recent Codex prompts could not have influenced the model [2].

Whichever account ultimately prevails, the controversy demonstrates the cost of discovering the question after sensitive material has crossed into an external system. At that point, the creators must rely on provider policies, technical assurances and retrospective investigations.

This risk extends far beyond mathematics. A developer may paste proprietary source code to debug an error. A researcher may upload an unpublished result for feedback. A founder may ask an AI assistant to refine a confidential product strategy. OWASP identifies confidential business data, proprietary algorithms, credentials and legal documents as sensitive information at risk in large language model applications, and recommends sanitising data before it enters a model [4]. NIST likewise identifies leakage of sensitive data and exposure of trade secrets among the risks created or intensified by generative AI [5].

Traditional pattern scanning can recognise a password or card number. A broad classifier can decide that text sounds financial, legal or technical. Neither signal alone answers the question that matters to an organisation:

Is this our confidential information?

That is why we built ndAI, short for NDA + AI. It creates a local checkpoint before submission, recognises material belonging to the organisation, and applies the organisation’s own rules. The safest time to protect an idea is before someone presses Send.

What it does

ndAI checks what an employee types, pastes or uploads through a supported workflow before it is submitted to ChatGPT, Claude or Gemini. The browser extension sends the content to a FastAPI service running on the employee’s machine. The service returns one of four decisions: Allow, Warn, Rewrite or Block.

Each sentence, file row or text segment is evaluated separately so that one sensitive line is not diluted by a long, harmless prompt.

The decision pipeline has five stages:

  1. Secrets scan: Pattern and entropy checks identify credentials, passwords and personal data. A live credential is always blocked.
  2. Company material detection: Each segment is compared with internal documents and public material covering the same subjects. A match records the source document, section and sensitivity tier: public, internal, confidential or restricted.
  3. Cumulative exposure detection: ndAI records which sections of an internal document a person or their team has exposed during the previous 14 days. It can therefore catch a document disclosed gradually across several prompts or several members of a configured team.
  4. Readable policy: Rules in policy.yaml combine the matched sensitivity, destination and employee role.
  5. Rewrite and re-check: ndAI edits only the flagged segments, then runs the result through the full detector again. It strengthens the rewrite once if necessary and blocks the request if confidential meaning remains.

Our demonstration policy is deliberately easy to inspect:

Content Company model Approved vendor Consumer AI
Public or internal Allow Allow Allow
Confidential Allow Warn Rewrite
Restricted Allow Rewrite Block

For example, Dana can send Gemini a vague sentence about evaluating an acquisition. If Dana later sends a detail about the seller’s lawsuit, the two prompts may together cover enough sections of a restricted acquisition memo to trigger a block. The same rule can apply when another configured team member sends the second fragment.

When rewriting is possible, ndAI removes the confidential proposition while preserving the safe intent of the request. Masking names alone is insufficient because “Company A will acquire Company B” still reveals an acquisition. When the sensitive information is itself the question, such as “check this confidential calculation,” ndAI blocks rather than producing a misleading rewrite.

The service also scans supported text and structured files, including CSV, Markdown, text, XLSX and Parquet, by splitting them into reviewable segments. Its dashboard shows decisions, matched document references, hashes and redacted previews. The audit trail does not retain the original prompt text.

How we built it

We built ndAI as three connected components:

  • A Chrome Manifest V3 extension written in TypeScript intercepts supported typing, pasting and upload actions. It displays the decision before submission and opens a review flow for files.
  • A local FastAPI service runs the detection pipeline at 127.0.0.1. It handles secret scanning, document comparison, exposure history, policy decisions, calibration and rewriting.
  • A React and Vite dashboard receives a live event stream and shows recent decisions and detector health.

The central technical problem was matching meaning rather than exact wording. We use sentence-transformers to create semantic sentence representations based on the Sentence-BERT approach, which enables efficient similarity comparisons between sentences [6]. Each incoming segment is compared against two corpora:

  • The internal corpus contains documents labelled public, internal, confidential or restricted.
  • The public corpus contains public writing about the same industries and subjects.

The public comparison is essential. A sentence counts as company material only when the evidence points more strongly toward an internal source. Topic similarity without a matching document can produce a warning, but it cannot produce a block by itself.

The context engine stores document and section references in SQLite. It maintains separate exposure histories for each user and configured team, while team membership comes from policy.yaml rather than being inferred from prompt text. Audit records contain the decision, references, a hash and a redacted preview.

For rewrites and local answers, we run the instruction-tuned Qwen2.5 3B model through Ollama [7]. The model receives only the flagged segments and an action such as removing a detail, generalising a sentence or dropping it. After each rewrite, the detector makes the final decision.

We also implemented calibration by document type. It can adjust detection thresholds using decision history, while roles, clearance and destination rules remain under explicit policy control.

Challenges we ran into

Distinguishing ownership from subject matter: Our first risk was building a detector that merely learned that mergers, research and source code are sensitive topics. Such a system would flag public reporting and normal industry discussion. The internal-versus-public comparison turned detection into a provenance decision: does this resemble the organisation’s document, or does it only concern the same topic?

Detecting leaks spread across messages: A single prompt may expose too little to cross a threshold, while several prompts can reconstruct the substance of a document. We needed to represent documents as sections, record partial exposure without storing raw prompts, and combine evidence across a person or explicitly configured team.

Rewriting without preserving the secret: A fluent paraphrase can still disclose the same fact. We therefore made rewriting iterative and required every rewritten result to pass through the detector again. If the protected fact is necessary to perform the task, the correct decision is Block.

Creating an audit trail without creating another liability: Security teams need a reason for each decision, but a database full of employee prompts would itself be sensitive. We kept document references, hashes, decisions and redacted previews while excluding original prompt text.

Working inside changing AI websites: Browser integrations can break whenever a host site changes its interface. We constrained the extension to supported typing, pasting and upload paths and kept the security logic behind a stable local API.

Scoping an honest prototype: Our current extension demonstrates the text and structured-file workflows. Robust PDF and DOCX extraction, returning edited copies of files, the dashboard context graph and the local-answer experience inside the extension remain roadmap items.

Accomplishments that we're proud of

We completed a working path from browser interception to local detection, policy evaluation, rewriting, re-checking, blocking and dashboard reporting. The system can attribute a semantic match to a specific document section, distinguish internal material from public discussion, and detect exposure accumulated across messages.

We also built an evaluation suite comparing ndAI with two simpler baselines on the same 56 labelled examples from a synthetic company corpus [8]:

Detector Leaks detected
Pattern scanner 8.8%
Generic category classifier 17.6%
ndAI 73.5%

Additional prototype results were:

  • A 5.9% false alarm rate on public material about the same subjects.
  • Every constructed cumulative-leak sequence was detected, with no false triggers across 16 harmless or public context prompts.
  • 80% of prompts routed to rewriting passed ndAI’s re-check and could proceed.
  • The core detection path took approximately 25 to 60 milliseconds per check in our prototype environment. Local model rewriting is a separate, slower path.

These figures are early evidence rather than production accuracy. The evaluation set is small, synthetic and designed by our team. We include the qualification beside the results because a security product should be as transparent about its evidence as it is about its decisions.

What we learned

Our main lesson was that confidentiality is relational. A sentence is not confidential merely because it sounds technical or financial. It becomes sensitive when it reveals material that belongs to a particular organisation and is not already public. That shifted our approach from broad topic classification toward source comparison and provenance.

We learned that the unit of analysis matters. Checking an entire prompt at once can dilute a sensitive sentence, so we scan sentence by sentence or row by row. At the same time, checking every prompt in isolation misses disclosures that accumulate. Effective detection therefore needs both fine-grained segmentation and controlled historical context.

We learned to treat rewriting as part of the security boundary. Good prose is not proof of safe prose. The rewritten content must be checked by the same detector, and some tasks cannot be sanitised without destroying their purpose.

We also learned that false alarms need difficult negative examples. Public documents should discuss the same subjects as the internal corpus. Otherwise, a high score may only prove that the detector recognises a topic.

Finally, readable policy matters. Security teams should be able to understand why a confidential document is allowed in a company model, warned for an approved vendor and rewritten for a consumer tool. Calibration may tune thresholds, but it should never silently change roles or clearance.

What's next for ndAI

  1. Expand and independently review the evaluation: Grow beyond 200 labelled examples, keep roughly one third as public material on overlapping topics, and test with pilot organisations.
  2. Connect real company sources: Ingest documents incrementally from shared drives, Confluence and code repositories while preserving sensitivity labels and access controls.
  3. Add an approval workflow: Let employees request an exception and route blocked requests to an authorised reviewer with a clear explanation.
  4. Cover more AI entry points: Add a local proxy for IDE assistants, command-line tools and other applications outside the browser.
  5. Complete file handling: Implement robust backend extraction for PDF and DOCX and return edited copies of reviewed files.
  6. Improve security visibility: Add the contextual exposure graph and clearer views of calibration and user history to the current dashboard.
  7. Offer a local alternative: Surface the existing local-answer endpoint in the extension so a blocked question can still be answered without leaving the device. Our goal is to let people benefit from AI while keeping company knowledge under company control.

References

[1] Clay Mathematics Institute, “The Millennium Prize Problems.” [Online]. Available: https://www.claymath.org/millennium-problems/. [Accessed: Sep. 13, 2026].

[2] OpenAI, “On the Navier–Stokes Millennium Prize Problem,” Sep. 8, 2026. [Online]. Available: https://openai.com/index/navier-stokes-solution/. [Accessed: Sep. 13, 2026].

[3] T. Buckmaster, “Today, Levent Alpöge and I have made public three results,” public statement, Sep. 8, 2026. [Online]. Available: https://cims.nyu.edu/~tristanb/statement.pdf. [Accessed: Sep. 13, 2026].

[4] OWASP Foundation, “LLM02:2025 Sensitive Information Disclosure,” OWASP GenAI Security Project, 2025. [Online]. Available: https://genai.owasp.org/llmrisk/llm022025-sensitive-information-disclosure/. [Accessed: Sep. 13, 2026].

[5] C. Autio et al., Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, National Institute of Standards and Technology, Gaithersburg, MD, USA, Jul. 2024, doi: 10.6028/NIST.AI.600-1.

[6] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks,” in Proc. 2019 Conf. Empirical Methods in Natural Language Processing and 9th Int. Joint Conf. Natural Language Processing, Hong Kong, China, 2019, pp. 3982–3992, doi: 10.18653/v1/D19-1410.

[7] A. Yang et al., “Qwen2.5 Technical Report,” arXiv:2412.15115, Dec. 2024, doi: 10.48550/arXiv.2412.15115.

[8] ndAI Team, “ndAI,” GitHub repository, Sep. 2026. [Online]. Available: https://github.com/massive-attack-team/ndAI. [Accessed: Sep. 13, 2026].

Built With

Share this project:

Updates

Submission history