Inspiration
Someone is at home, the doorbell never rings, and a card comes through the door saying they were out. The retailer's position is that the tracking says delivered, so the matter is closed. Ofcom asked 4,058 UK adults about this in 2025: 20% had a delivery person who did not knock loudly enough or ring the bell, and 16% got a "you were not in" card while they were at home. 68% had some delivery issue, the same share as 2023 and 2024.
The thing that made us build it was a smaller number. We coded 17,851 r/amazonprime posts by pattern match; 691 of them were "marked delivered but never received". Eleven of those 691 mentioned camera footage of any brand. People who own the exact evidence that would settle the argument are almost never using it, and Citizens Advice found one in three consumers with a delivery problem takes no action at all because they do not think it will help.
What it does
Paste a tracking number, or photograph the carrier's own tracking page. Doorstep Receipt reconciles the claimed delivery window against the Ring event history for that window and produces one page with a stamped verdict, the events found inside the window or the empty set shown as an empty set, and a plain sentence ready for a retailer's claim form. It never says a driver lied.
How we built it
Python and FastAPI, with the Ring client written against the documented request shapes and exercised by fixture payloads in the real JSON:API envelope; the demo replays HMAC-signed webhooks through the same signature-verification code a live Ring webhook would hit.
The part we did not expect to have to build is the coverage engine, and it is the one that makes the verdict honest. An empty event log means either nobody came or nobody was watching, and Ring publishes no offline webhook, so there is no event to key on. Without it the app would print "not consistent with camera" over an hour when the doorbell was unplugged. status_poller.py samples GET /v1/devices/{id}/status on a schedule, coverage.py treats only the span between two online samples less than thirty minutes apart as covered, and a "not consistent" verdict over a window containing any gap is downgraded to indeterminate with the gap printed on the page. Every existing reconciliation test had to gain coverage samples, which was the right outcome: a test without them was testing the wrong rule.
The second thing is a language guard, twelve rules across two categories: dishonesty, theft, fraud, fault, false claims, asserted omission, asserted absence of a person, named-actor verdicts, faked scans and overclaims of proof or certainty. reconcile.py calls it on its own sentences and the assembled document is checked as a whole before it renders, because the one artefact this product exists to produce is a page somebody attaches to a claim.
Amazon Bedrock does the one thing nothing else could. Carrier tracking APIs authenticate a shipper account, not a recipient, and the inbox route needs Gmail's restricted scopes plus an annual third-party security assessment. So TrackingPageAdapter sends the PNG bytes of a tracking page to Converse in an image content block with a forced toolConfig, which means the model answers by filling the record_tracking_page schema rather than writing prose somebody then parses. It worked first time on all four sample pages, UPS, USPS, Royal Mail and Amazon Logistics, at high confidence in 4 to 6 seconds, normalising a USPS number printed as 9400 1118 9922 3197 4284 90 without being asked and returning "Friday, September 20, 2026 at 2:11 P.M." verbatim beside the parsed value, because reformatting a third party's assertion without showing the original reads as tampering even when it is not. Three things guard it: carriers.py knows the published number formats and status vocabularies and outranks the model where they disagree, with the disagreement printed on the reading as a warning; nothing is stored until a person accepts it on a review screen with every field editable; and a low-confidence read refuses to prefill the time and status at all, because a wrong time shifts the window and produces a confident finding about the wrong hour.
The screen and the document it produces are deliberately two different objects. The screen follows a vendored Montserrat and IBM Plex Mono, an indigo masthead and a gold strapband, so the app itself reads as software you are using right now. The filed record that actually leaves the app and gets attached to a claim is a separate stylesheet, plain black ink on white paper, no colour, a system serif stack rather than a vendored one, because it is going to be printed and photocopied and read by someone whose job is to reject things, and a decorative choice on that page is a small claim that it was made by a marketing department rather than a witness. The dispute pack carries a manifest.sha256 written in the format GNU coreutils actually reads, so the sha256sum -c line printed at the bottom of the document is a command a claims handler can paste and run, not one that only looks like it would work. 135 tests pass with no live Ring device.
Challenges we ran into
There is no Ring simulator to test any of this against: real testing needs an actual device, registered in the US, tied to an active Protect subscription, and pulling the clip behind an event needs the paid continuous-recording tier or the API answers with 416 TIMESTAMP_NOT_FOUND. The problem that matters more for a reconciliation product specifically doesn't show up as an error at all: event history only starts from the moment consent was granted, with no way to backfill, so a claim about a delivery from before that moment can never be checked against anything, only reported as a window the app was never watching. The webhook signature scheme has the same gap in miniature: the reference gives one sentence and no worked example, nothing showing how the bytes are actually assembled before signing, so what is implemented here is written down in code as a stated assumption rather than left for someone to guess at.
The cost of that design lands on the wrong person. Citizens Advice found a third of people with a delivery problem already take no action, so asking them to type a tracking number is asking the least likely person to act to do one more thing. The only answer is to make it take under fifteen seconds, which is what the photographed-page path is for.
Bedrock also has three id shapes with three different failure modes. A bare id returns a ValidationException about on-demand throughput, which sends you to the wrong document when the fix is the us. inference profile.
The word "verbatim" turned out to be a promise the code almost did not keep. claimed_at_text is the delivery time exactly as the carrier's page printed it, quoted inside a class="verbatim" span on the filed record, the strongest typographic claim the document makes, and it was one of two fields the model's tool call returned that never went through the guard at all — a string presented to the reader as quoted, that the model had in fact written freely. Fixed by routing every field the tool call returns through the same _apply_guard, not the two fields that were obviously prose.
A second gap was worse, because it meant the guard itself had never been shown a sentence trying to beat it. The only assertion in the original suite ran a clean page through is_clean() and would still have passed with the rule list emptied — proof the guard is correct in isolation, never proof it is applied against a hostile input. Feeding the twelve regexes real evasions found two: the driver liеd about it with a Cyrillic е, and the driver didn't deliver it with the curly apostrophe a word processor inserts by default, both read clean. The fix folds lookalike Cyrillic and Greek letters and normalises punctuation before matching, one character for one so a finding's span still points at the original text, pinned by a parametrised test naming each evasion it closes.
Accomplishments that we're proud of
The verdict can be indeterminate, and the page says why, with the coverage gap printed. A product that could only say yes or no would be more impressive and would be wrong.
What we learned
An empty event log is not evidence of absence unless you can also show the camera was watching. We nearly shipped a verdict engine that could not tell an empty doorstep from an unplugged doorbell, and no Ring event exists that would have told it apart.
What's next
Wiring the Ring client to a live device, and a carrier adapter for anyone who publishes a recipient-side API.
Built With
- amazon-bedrock
- boto3
- fastapi
- python
- ring-api
- uvicorn
Log in or sign up for Devpost to join the conversation.