Inspiration
While building an AI-assisted tool to automate job application forms and cover letters (jobagent), I strictly instructed the AI assistant to store my personal contact details in a local, .gitignore-protected JSON file.
During an automated test generation session, the coding assistant silently wrote unit test fixtures containing my actual name, private email, phone number, and physical address—hardcoded directly as string literals inside a Python test file.
I only caught it by pure chance during a quick manual spot-check before pushing upstream. It was a wake-up call: in the age of autonomous coding agents generating hundreds of lines of diffs across dozens of files, developers cannot reliably catch subtle personal data leaks line-by-line.
Around the same time, I read about non-autoregressive Small Language Models (SLMs) and Jev/Julia-1 from SupersonicLabs, specifically fine-tuned for confidential data classification. I brainstormed with AI on the best real-world application, immediately remembered my near-miss PII leak, and realized this was the perfect hackathon project: an automated, local Git pre-commit guardrail. To make it run natively and efficiently on local developer CPUs without external servers, I chose Julia-1.
What it does
Latch is an intelligent, privacy-preserving Git pre-commit guardrail that stops sensitive PII (names, phone numbers, addresses, personal emails) and live API credentials before code ever enters Git history or reaches a remote repository.
Key features:
- Pre-Commit Interception: Automatically inspects staged additions before
git commitcompletes. If confidential data is detected, the commit is blocked (fail-closed). - Binary Search Dissection: Uses token-aware binary dissection to pinpoint the exact offending line window (e.g., lines 1–14) rather than rejecting the whole file blindly.
- Zero-Cloud Local Inference: Runs 100% on-device using a lightweight 144.3M parameter Julia-1 SLM via a local daemon (
127.0.0.1:5138), meaning your source code and secrets are never sent to third-party cloud APIs. - Mutual Authenticated IPC: Restricts daemon access using ephemeral cryptographic tokens and loopback validation, isolating daemon state per repository.
- Developer Observability: Includes built-in real-time throughput metrics, rolling percentile latency tracking, and a Prometheus
/metricsexporter compatible with Grafana.
How we built it
- Core Engine & Architecture: Built in Python with strict clean-code architecture, zero external web server frameworks (using Python's native
http.server), and separation of configuration (config.json) from ephemeral secrets (.latch/daemon.token). - Model Inference: Implemented
JuliaEngineusing PyTorch on CPU threads to load and evaluate the 144M Julia-1 model weights on structured diff payloads. - Diff Parsing & Batching: Developed a custom
DiffParserthat handles multi-file staged diffs, token-aware chunking with line overlaps,# latch:ignoreinline pragmas, and glob allowlists. - Daemon Lifecycle & Client: Created a persistent background daemon that keeps model weights warm in memory to eliminate cold-start overhead, falling back cleanly to in-process execution if needed.
- Testing & Verification: Built a comprehensive 91-test automated unit suite covering fail-closed behavior, prompt-injection resilience, token authorization, gitignore boundary resolution, and realistic multi-branch merge simulation demos.
Challenges we ran into
- Intra-File Structural Dilution: We discovered that wrapping small secrets inside large functions, heavy imports, and nested dictionary scaffolding can dilute the token share of the secret, reducing the SLM's confidence. We tackled this by engineering surgical chunking and prompt structuring.
- Path-Token Model Bias: Learned that diff headers containing tokens like
tests/orfixtures/can bias small models toward assuming content is non-sensitive mock data. We addressed this through allowlist isolation, ensuring external repositories never inherit test exemptions. - Cross-Platform Background Daemons: Reliable detached background process spawning and loopback socket management across Windows and Linux required custom process group isolation and authenticated shutdown protocols.
- Git Simulation Cleanliness: Creating realistic multi-file PR merge demos without corrupting the host developer's Git reflog or leaving unreferenced blobs required building a disposable temporary sandbox clone architecture.
Accomplishments that we're proud of
- True 0.0% False Negative Rate: Detected 100% of benchmark leaks across test records and credentials without missing a single leak.
- 100% Prompt Injection Resilience: Adversarial attempts to bypass pre-commit inspection via comments (e.g.
system override,return clean) were neutralized. - Honest Engineering: Fully removed scripted terminal simulations in favor of authentic, reproducible live Git CLI demonstrations running real hook enforcement against the warm local model.
- Sub-5-second Local Evaluation: Achieved real-time neural inference on standard laptop CPUs without needing expensive GPU cloud infrastructure.
What we learned
- Small, task-specific language models (144M parameters) can achieve enterprise-grade classification accuracy on focused domain tasks while running entirely on commodity developer hardware.
- Pre-commit developer tooling must prioritize zero-dependency, fail-closed design—if a security tool fails silently or requires complex cloud setups, developers will simply bypass it.
What's next for Latch
- AI Agent Middleware & Hook: Integrating Latch as a guardrail specifically for autonomous coding agents (Claude Code, Cursor, Copilot Workspace, Aider). As agents increasingly write and commit code autonomously, Latch acts as the local firewall preventing agent hallucinations from committing private developer data.
- Server-Side CI Enforcement: Providing a containerized GitHub Action and pre-receive hook to establish a defense-in-depth policy that catches leaks even if a developer bypasses local hooks via
--no-verify. - IDE Diagnostic Provider: Creating a lightweight VS Code extension that queries the local Latch daemon on file save, highlighting leaked PII with red squiggly diagnostics before the developer even opens the Git panel.
Built With
- ai
- bash
- cli
- cybersecurity
- devtools
- git
- git-hooks
- huggingface
- ipc
- julia-1
- machine-learning
- natural-language-processing
- pii-detection
- powershell
- privacy
- pytest
- python
- pytorch
- rest-api
- small-language-models
- transformers
Log in or sign up for Devpost to join the conversation.