Inspiration

Inspiration

Large language models are increasingly used for coding, debugging, research, and everyday work. However, users often paste information into public AI tools without realizing that their prompts may contain API keys, passwords, personal information, internal business details, or proprietary code.

This inspired us to build ChronoGuard, a privacy-focused tool that helps users identify potentially sensitive information before they share it with a public LLM.

The core idea is simple: detect the risk first, explain it, and provide a safer version of the prompt.

What We Built

ChronoGuard uses a hybrid detection approach combining deterministic security rules with a local machine-learning classifier.

The rule-based scanner detects concrete patterns such as:

  • API keys and access tokens
  • Passwords and credentials
  • Email addresses and phone numbers
  • IP addresses
  • Database connection strings
  • Internal and confidential information

Alongside this, we built a lightweight ML classifier using TF-IDF and Logistic Regression. The classifier analyzes contextual language that may indicate sensitive or proprietary information even when there is no obvious secret pattern.

The two signals work together:

User Input ↓ Rule-Based Detection ↓ Local ML Context Classification ↓ Risk Assessment ↓ SAFE / SUSPICIOUS / CRITICAL ↓ Automatic Sanitization ↓ Safe Version of the Prompt

ChronoGuard produces a heuristic privacy risk score from 0–100 and automatically replaces detected sensitive values with clear placeholders such as "[REDACTED_API_KEY_1]".

How We Built It

The application was built as a local-first web application using:

  • Next.js + TypeScript for the web interface
  • Python + Flask for the ML service
  • scikit-learn for TF-IDF and Logistic Regression
  • Regex-based detection for structured secrets and PII
  • Tailwind CSS for the interface

The ML classifier was trained using a synthetic dataset containing both sensitive and benign contextual examples. We evaluated it using 5-fold cross-validation; after refining the contextual training examples, the classifier reached 88.9% synthetic cross-validation accuracy.

The ML service runs locally on "127.0.0.1", and the application is designed so that the core rule-based scanner continues working even if the ML service is unavailable.

What We Learned

One of the biggest lessons was that privacy detection cannot rely on a single technique.

Regex is excellent for structured information such as API keys and email addresses, but it cannot reliably understand context. On the other hand, a lightweight ML classifier can recognize contextual sensitivity but should not be trusted as the only security mechanism.

This led us to the hybrid architecture:

«Rules provide deterministic detection; ML provides contextual intelligence.»

We also learned that synthetic ML datasets require careful construction. Adding contextual sensitive examples improved the model's recognition of proprietary and confidential language, while adding benign examples containing words such as "internal" and "private" helped prevent the classifier from treating every occurrence of those words as sensitive.

Challenges

One of the main challenges was balancing security detection with false positives.

For example, the word "confidential" can appear in a legitimate technical question such as "confidential computing." A simple keyword-based rule may flag this even though the context is not actually sensitive.

Another challenge was making the ML component useful without making the application dependent on an external AI API. We therefore kept the classifier lightweight and local, allowing ChronoGuard to operate without sending user prompts to a third-party AI service.

We also designed the application with an offline fallback: if the local ML service is unavailable, the deterministic scanner continues to detect concrete sensitive patterns.

Why It Matters

ChronoGuard is designed around a simple principle:

The safest sensitive prompt is one that is identified and sanitized before it leaves the user's environment.

Rather than attempting to replace public LLMs, ChronoGuard acts as a privacy layer between users and those tools.

The current implementation is a functional local web application and a foundation for future extensions such as browser integration, desktop protection, and additional detection models.

Limitations

ChronoGuard is a prototype and should not be considered a complete enterprise security product. Its risk score is heuristic rather than a formally calibrated probability of data leakage, and the ML classifier was trained and evaluated using synthetic data.

For this reason, the system is intended to assist users in identifying potential privacy risks, not to guarantee that every sensitive piece of information will be detected.

Built With

Share this project:

Updates

Submission history