Inspiration

This started as something else entirely. Originally, I was building a symbol system called Click Clack, modeled on DNA. It was a sorta goof around project I would play with whenever I finished all my other work. There was 4 main symbols, ╬, ╫, ┼, ╪, and each glyph connected at the top and sides so the written text formed one continuous strand.

It looked pretty neat when a whole screen was filled with it, but at the end of the day it was mostly a novelty. Here is an example:

╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪ ╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼╪╫╫╪╬┼╬╬╫┼╪┼

I spent a lot of time playing around with it but never really found a place for it. A True Purpose so to speak. I thought maybe it could become some sort of AI communication protocol that was lightweight, fast, and able to mutate to keep security top notch. That sent me down a pretty deep rabbit hole looking at how AI agents actually communicate.

MCP was the main one that caught my attention, and the more I looked at it the more I noticed this huge hole in security. It seemed like a wild wild west where agents were just willy nilly sending messages around while users trusted them with highly sensitive data. There was no strong way to prove which agent sent what, where it went, or what that agent was actually allowed to do. MCP already handled moving messages between everything, but there still wasn't much proving who sent them or what they were actually allowed to do.

That is where 7h3 Protocol was born. What I eventually shipped is clean, easy to use, and requires minimal to no installation depending on how you use it. Every message carries a signature, every capability comes from a parent grant with a token, and every receipt carries a hash from the previous receipt. If somebody changes a link or removes one, the strand fails verification.

That was the idea hiding inside ClickClack the whole time. It kinda just fell into the right spot and gave me that Ah-haa moment I had been waiting for. Funny how things work out.

What it does

7h3 Protocol is a deterministic authorization layer for agent tool calls. This submission extends it to WebMCP and gives browser based agents a real security boundary instead of another system trying to guess whether a request looks suspicious.

Google Chromes agent security model is mostly probabilistic. It uses classifiers, prompt spotlighting, filters, and other models that try to recognize something malicious before the agent acts. What gets me is that OpenAI straight up admits you cannot trust a tool just because of its name or because it claims it only reads information. Fair enough, that part does make sense, but then the answer is pretty much to let each website rely on whatever security it already has.

That is where the whole thing falls apart. Existing website security was built around humans clicking buttons and filling out forms, not AI agents acting on their behalf across multiple tools. If a model can be persuaded past a security boundary, then it was never really a boundary. It was pretty much a suggestion.

With 7h3 Protocol, the refusal isn't a judgement call. The AI doesn't have to decide whether it should behave or if something feels suspicious. The signature either works or it doesn't, and the action is either inside the allowed scope or it isn't. There is no prompt that produces a valid signature, no matter how cleverly somebody writes it or how much Unicode nonsense they stuff into it.

The demo uses a business ledger with real invoices and "real" money. Obviously it isn't real money but it simulates real money. I didn't want to put up a screen that simply says “attack blocked” because anyone can fake that. I wanted people to be able to check the result themselves, change things, and watch verification fail instead of just having to believe me. So there is 2 main demos.

/compare

/compare runs one compromised agent through four hostile actions against two identical copies of the same books. One copy is guarded by 7h3 Protocol, while the other is connected straight to its handlers like a lot of agent apps are right now. It is the same agent, the same books, and the same attacks on both sides.

On the unguarded side, 4 out of 4 attacks work and $2,750 gets moved. On the guarded side, 0 out of 4 work and nothing moves. The unguarded side also creates 0 receipts while the guarded side creates 4, so the money gets moved and theres no real record showing how it happened either. Its like getting robbed and the security camera decided to take the day off too.

/verify

/verify shows the real grant, receipt chain, and signed manifest inside an editable box. You dont have to take my word for anything because you can start changing the evidence yourself. Rewrite a refusal into an approval and it reports bad-signature.

You can also delete one of the receipts and every receipt left behind will still verify on its own because those individual receipts weren't changed. Only the chain notices that one was removed. That part matters more than it may seem at first because a signature proves somebody didn't change a receipt, while the chain proves somebody didn't quietly delete one and hope nobody noticed.


So the second example starts as one compromised agent tries four of the same type of attacks against two identical copies of the books. One is connected directly to its handlers like most apps are, and the other is protected by 7h3 Protocol.

Same agent, same tools, same attacks. The only difference is the authorization layer.

Click Run the attack on both. When it finishes, say:

Four out of four attacks worked on the unguarded side. Zero out of four worked on the guarded side.

The agent will try to pay a $1,850 invoice, wire $900 to an offshore account, delete an invoice, and export the data of four customers. The guarded side blocks all four and tells you exactly why. The payment is way over the $50 limit, and the agent never had permission to do the other three things at all.

On the unguarded side, the balance gets drained out from $2,317.50 to $47.50. An invoice gets deleted all together and four customers have their data exposed. The guarded side remains unphased.

The unguarded side loses the money and exposes the data. But the craziest part is it writes zero receipts. There is no record any of it ever happened. The guarded side blocks all four and records every refusal.

Click the button once. If you run it twice, the numbers change because the first attack already paid and deleted records. Reload the page before each demo run through.


The 3rd demo is dealers choice. Verify the results yourself. Skim through the technical side.


There is 19 WebMCP tools across 4 routes, including the landing page, so an agent has something it can call wherever it ends up landing.

How we built it

The main piece is a canonical JSON envelope signed with Ed25519. Sorting the keys alphabetically sounds like a boring little detail and it might be haha. What it does is it makes sure every SDK signs the same exact bytes. TypeScript, Python, Rust and Go all have to agree with each other or the cross-language tests fail.

The wire version is 7h3/0.1 and its locked in. I also added X25519 with ChaCha20-Poly1305 for end to end encryption. ML-DSA handles post-quantum signatures and BLS12-381 is there when a grant needs more than one approver. Getting all of that working across HTTP, WebSocket, gRPC, queues and webhooks was a bit of a headache to say the least. I got it workin and working well.

The part I really like is that it runs through Web Crypto with zero runtime dependencies. It can sit natively on Cloudflare Workers without needing a mountain of packages just to stay alive. The WebMCP part is fairly small because its really just the gate sitting on top. A tool call comes in and it checks the signature, what the agent has permission to do, and where that permission came from.

Delegation

Delegation is one part I really want to point out because I actually got it wrong the first time.

Say if I give an agent permission to spend $50. That agent can pass the grant to another agent, but it cant suddenly turn the limit into $500. Passing it around doesn't reset anything either. The lowest limit found anywhere in the chain is still the one that counts.

My first version only looked at the final token though. So an agent farther down the chain could pretty much write itself a bigger limit and the system would ignore the $50 limit that came before it. I caught it during testing and just sat there for a second like, well thats definitely not good.

That isnt delegation. The original limit follows the grant through the entire chain, no matter how many agents pass it along.

The corrected rule is:

\text{cap}_{\text{eff}}(f) = \min \{\, \text{cap}_t(f) \;:\; t \in \text{chain}, \; f \in \text{caps}(t) \,\}

The little condition under the minimum is easy to overlook, but it matters. If a token doesn't mention a certain field, that doesn't mean the limit disappears. It just means that token has nothing new to add for that field. If nobody in the chain sets a ceiling at all, the tools own limit takes over. Basically, silence doesn't equal permission. So if the system cant clearly find permission for something, it just blocks it.

Challenges we ran into

I went through my own code five different times trying to break it and ended up finding 13 real defects. Some were regular bugs, but a few made me stop and stare at the screen because the system wasn't doing what I thought it was doing at all. Three of them really stuck with me, mostly because everything looked fine right up until it wasn't.

The one that scared me

The queue verifier was returning the wrong field. The signature covered envelope.body.content, but after verification the application was handed another field called payload. That payload had never been signed.

I tested the whole thing from beginning to end. The content that actually got signed was:

{"job":"reindex","amount":10}

The message passed verification, but this is what came back out:

{"job":"DROP TABLE users","amount":1000000000}

So the signature wasn't broken. It did exactly what it was supposed to do and verified the signed content. The problem was that after all that work, the application grabbed a completely different field. It was basically like installing a giant steel lock on the front door while leaving the window wide open right next to it.

What really got me was finding the same bug in TypeScript, Python and Rust. I had repeated the exact same mistake three times in three different languages without noticing it. The cryptography was working perfectly and the system was still unsafe. That one honestly made my stomach drop a little.

The one that taught me the most

I somehow managed to publish two packages that were completely broken. Not broken in some rare little edge case either. They failed with ERR_MODULE_NOT_FOUND as soon as somebody tried to import them, yet every unit test was passing like nothing was wrong.

The problem was Vitest accepts extensionless ESM specifiers while plain Node doesn't. My tests were checking the source code inside the testing environment, but they weren't checking the package a real person would download and install. So technically the tests passed, they were just testing the wrong damn thing.

Adding the missing .js extensions fixed the packages, but I didn't want to slap a bandage on it and move along. Now every multi-file package gets packed into a tarball, installed, and imported with plain Node. I also tested each check against its own broken build to make sure it actually fails when its supposed to. A test you've never seen fail might just be a pretty green checkmark sitting there lying to your face.

The one hiding in plain sight

Another problem was sitting right in the README the entire time. I claimed the audit log was “signed and chained,” but it was only signed. Each receipt had its own signature, so changing one would break verification, but someone could delete a receipt completely and every receipt left behind would still pass.

Nothing remaining had been altered, so nothing looked wrong. The documentation described what I meant to build, and after reading it so many times my brain just filled in the missing part. It wasn't until I read the documentation like I was trying to prove it wrong that the problem finally stood out.

That is why receipts now carry prevHash. If a receipt is deleted, the others can still pass their own signature checks, but the chain catches the missing link. Signing the receipt proves it wasn't changed. Chaining the receipts proves one wasn't quietly removed.

I also wasted almost a full day on a Rust fix that had been correct the entire time. A stale incremental build artifact made it look broken, then I ran cargo clean and suddenly everything passed. Later, a bug I had fixed a month earlier came crawling back during a documentation rewrite. Apparently bugs can sneak back in through the paperwork too.

What we learned

The biggest thing I learned is that probabilistic defenses and actual security boundaries are not the same thing. A classifier that catches prompt injection 99 percent of the time sounds great until you remember the attacker can keep trying and only needs it to fail once. Its still useful, but its a filter. It isnt a wall.

A signature check isnt trying to guess whether something looks suspicious. It checks whether the request has real authority behind it. The agent either has the signature and permission or it doesn't. If that system is broken, it tends to break in the same repeatable way, which means a test can actually catch it.

I also learned to test the thing people install, not just the code sitting in front of me. Two packages went out broken while the test suite stayed green because my source worked and my tests worked, but the finished package didn't. The bug was sitting in that annoying little space between what I built and what the user actually received.

Writing the threat model down ended up helping a lot more than I expected too. Some of my worst bugs were found where the documentation promised something stronger than the code delivered. I kept reading the README like it matched the code because it described what I meant to build. It took comparing them directly before I noticed the code wasn't actually doing all of it. Reading those claims like an attacker instead of the person who wrote them is what exposed the difference.

The other big lesson was to ship the counterfactual. Just showing a screen that says “attack blocked” doesn't prove much because anybody can connect a button to a red denied message. I needed to run the exact same attack against an unguarded copy so people could see what the refusal actually prevented.

Its the same agent, the same books, and the same four actions. The unguarded side moves $2,750 and writes no receipts. The guarded side moves nothing and records every refusal. Seeing both happen side by side makes the difference pretty hard to ignore.

What’s next

Im not done with it yet. The next step is native ML-DSA grants for WebMCP so browser tool calls can use post-quantum signatures from end to end. I also want threshold-signed grants for higher value tools, where more than one person or key has to approve something before it can happen.

A shared receipt format is another thing I want to build. An agent should be able to carry a receipt between different sites and prove what it was authorized to do somewhere else. Not just show up saying “trust me bro” and expect every system to go along with it.

I would also like to see signed tool manifests become a normal WebMCP standard. Signing the tool surface when it gets deployed isn't some huge impossible job, but it would make an entire class of quiet attacks loud and obvious. Sites shouldn't have to adopt my whole library just to get that basic protection.

Every package is available today on npm, PyPI and crates.io. There are 776 tests and zero runtime dependencies. So thats where the project is now. WebMCP lets the agent use the tools, and 7h3 keeps track of who gave it permission, how far that permission goes, and what actually happened.

TLDR: Think of a wax seal on an envelope. If the seal is broken, you don't accept the letter.

Share this project:

Updates

Submission history