Inspiration
For about thirty years, websites have been built on one quiet assumption: someone is looking at them.
That assumption is starting to break. More and more often, the visitor who arrives at a site isn't a person. It's an AI agent that a person sent on a small errand: find the venue's address, compare two prices, register for the conference. The agent shows up, does its best, and leaves. If it fails, nobody watches it fail. The person who sent it just hears "I couldn't find that."
We started from one sentence, and it became our tagline: Your next visitor is not human.
What caught our attention wasn't that agents are coming. Everyone knows that. It was how differently they read. A person glances at a page and gets it all at once. An agent reads the same page three ways:
- The screenshot: the page as it looks. It's rich, but slow and costly to process, so agents use it last.
- The HTML: the page as it was written, before anything is painted on screen.
- The accessibility tree: the browser's plain-language summary of every control, along the lines of "this is a button, it's called Register, and you can press it." Screen readers use this same layer to describe a page to blind visitors.
Most of the time the three views agree. Where they don't, sites quietly break. Take a "button" made from a plain box with some code attached. It looks like a button and works when you click it with a mouse, but in the accessibility tree it's nothing at all. It's a painting of a button. Magritte might have captioned it Ceci n'est pas un bouton.
A person never notices. An agent gets lost.
Designers already have a routine for checking their work: they test for speed, for search ranking, and for accessibility. Nobody tests whether an agent can find its way around. We wanted to add that to the routine.
Why "Pytheas"
Pytheas of Massalia was a Greek navigator. Around 325 BC he sailed north past the edge of the known world, around Britain and toward a place he called Thule, and wrote down what he found: a sun that wouldn't set, a sea that seemed to freeze, and a link between the moon and the tides. Many people back home didn't believe him.
That seemed like the right name. The web has a new kind of traveler now, and someone needs to draw the map and write down honestly what's out there.
What we learned
1. Agents read what you wrote, not what you meant
People are forgiving readers. Our eyes fill in gaps: we see a blue rounded rectangle and think "button." An agent working from structure doesn't do that. Building Pytheas taught us that the agent's view is, in a sense, the most honest view of a website. It shows the distance between what a designer meant and what they actually said.
2. The curb-cut effect is real
When cities cut ramps into curbs for wheelchair users, it turned out almost everyone used them: parents with strollers, travelers with suitcases, kids on bikes. The same thing happens here. Every fix Pytheas suggests also helps a person using a screen reader, a keyboard, or a slow connection. That includes a real button in place of a painted one, a label tied to its field, and a link that doesn't hide behind a hover. We came to believe that "agent-ready" and "human-ready" are mostly the same thing seen from two sides. That idea became the last line of our demo: the same fixes that help an agent help a person.
3. A sitemap lists addresses. A map explains why you'd go there
A traditional sitemap tells a search engine where the pages are. It doesn't say what a page is for, what questions it answers, or how a visitor would actually reach it by clicking. An agent needs the second kind of map, which is closer to a guidebook than a phone book. That's what the agent sitemap Pytheas generates tries to be.
4. Measure what you actually care about
Our Agent Navigability Score is out of 100 and has six parts:
$$\text{Score} = \underset{\text{task success}}{35 \cdot \frac{\text{tasks finished}}{\text{tasks tried}}} + \underset{\text{structure}}{20 \cdot \frac{\text{well-built controls}}{\text{all controls}}} + \underset{\text{visual}}{15 \cdot \frac{\text{clearly visible controls}}{\text{all controls}}} + \underset{\text{path}}{10 \cdot \text{efficiency}} + \underset{\text{discover}}{10 \cdot \frac{\text{pages within 3 clicks}}{\text{all pages}}} + \underset{\text{tooling}}{T}$$
Path efficiency compares the shortest possible route with the route the agent actually took, averaged over every errand it finished:
$$\text{efficiency} = \frac{1}{n} \overset{n}{\underset{i=1}{\sum}} \frac{\text{shortest path of errand } i}{\text{steps taken on errand } i}$$
An agent that needs six clicks for a two-click errand earns a third of the path points.
We ran into Goodhart's law early: when a measure becomes a target, it stops being a good measure. Ten points come from "agent tooling," meaning whether the site publishes the files that guide agents. Pytheas writes those files for you, so installing them raises your score by definition, whether or not anything actually got easier. So we decided the real evidence would always be two plain numbers, shown on their own: did the agent finish the task, and how many steps did it take? Everything else is supporting detail.
5. A report holds two kinds of truth, so label them
Some things Pytheas knows by rule. Is this control covered by another element? Is this field missing its label? Our eleven checks give the same answer every time. Other things need judgment. Is this page easy for an agent? Did the visit go well? For those we ask a model, and models, like people, don't give the exact same answer twice. We learned to say that plainly ("every scan is judged fresh, so the score can move a few points") instead of presenting a judgment as a measurement. When no model is available, Pytheas says the score hasn't been written. It doesn't make one up. An empty box is more honest than a confident guess.
6. An agent that can click can also do harm
When Pytheas sends an agent through a site to test it, that agent is acting on a real, live website. So we drew a hard line: it never submits a purchase or a registration. If its plan includes something that looks like pay, check out, or sign up, it stops and explains why. The same rule applies to the forms we annotate for agents: we never mark a registration or a purchase to submit itself. The person stays the one who says yes.
How we built it
The shape of it. You paste a web address. Pytheas opens the site in a real browser and visits up to twenty pages, following the navigation the way a visitor would. It saves all three views of every page. It runs eleven checks on every button, link, and field. A model then reads the three views and scores the site. When there's an errand to try, Pytheas also plans a short visit and actually clicks through it. Then the report opens: the score, a map of the site, an Easy or Hard mark on each page, and a list of fixes with the corrected code. One button, Copy all as a prompt, turns the whole list into instructions a coding assistant can apply in one pass.
A few principles shaped how we got there:
- We wrote before we built. Our first commit wasn't code. It was a concept document covering the problem, the research behind it (Google's web.dev guidance on agent-friendly websites), the score, the eleven checks, the screens, and a list of what to cut if time ran short. Every step ended with a "done when" sentence, so we always knew what finished meant.
- We built with agents, so we drew them a map first. Much of Pytheas was built alongside AI coding agents working in parallel. One built a deliberately broken demo site, one the crawler, one the scoring, one the map and report, one the fixes and exports, and one the shared plumbing. To keep them from colliding, we wrote a single shared guide. It listed the decisions already made, who owned which files, the exact shape of the data passed between them, and checkpoints that had to pass before the next wave of work began. Partway through, we noticed the irony: we had written an agent sitemap for our own agents. It worked for exactly the reason Pytheas says it should. Agents do far better when someone tells them where things are and what each place is for. We were testing our own thesis on ourselves before the product could scan a single page.
- We made the wait part of the show. A scan takes about half a minute. Instead of a spinner, the map draws itself, one page at a time, left to right from the home page. The wait turns into the first look at the result.
- We built Pytheas to pass its own test. It has real buttons, labels tied to their fields, and navigation you can click. Color is never the only signal: every problem also gets an icon and a word. A tool that lectures people about legibility has to be legible itself.
- We showed the real thing. Our demo video is the app scanning a real website, serai.world, with nothing mocked. The only editing is cutting out waits and zooming in on details.
Under the hood: Next.js, React, and Tailwind for the interface. Headless Chrome for the crawl. xAI's Grok for the judging and the click visits. A live stream of events so the map can draw as pages arrive.
Challenges we faced
- A staged demo vs. the real world. Our first plan centered on "Example Summit," a fake conference website with eight flaws we planted on purpose: a Register button that was really a plain box, a menu that appeared only on hover, an invisible layer covering a control, and more. It was a great test suite. It showed that our checks caught every planted flaw and that a clean page came back with nothing. But a tool that only shines on a site built to fail doesn't prove much. Late in the weekend we made a hard call: we removed the staged site and the stored replays, and now Pytheas scans real sites only. The site in our video scores 89, "Agents can use this site," which is far less dramatic than a red failing grade. We decided a true 89 was worth more than a staged failure.
- The same site, a different score. A model's judgment wobbles. Run the same scan twice and the score can shift by a few points. We couldn't remove that, so we built around it. The rule-based checks became the steady backbone, and the judged parts are labeled as judgment. Our own test suite had the same problem in miniature: one test compared a record that included a randomly generated ID, and it failed about one run in eight. Randomness turned up at every level of the project.
- We ran into the same costs agents do. Screenshots are the expensive view, which is why agents use them last. Our scorer hit the same wall: passing every page's screenshot back and forth was slow and costly. We cut scoring down to one model call per scan, or two when a click visit is planned. In the second call, the agent plans its whole route once and Pytheas walks it locally, so the pages don't have to be sent again at every step. We didn't learn how agents read from a paper. We learned it from our own bill.
- Finding a face. The design changed more than anything else. It started as a literary journal on warm paper, became a navigator's log on a night sea, then an editorial illustration, then a navy-and-cyan network, and finally warm white with a quiet navy network in the background. The hardest lesson was restraint. The map is the illustration, and the rest of the interface should stay quiet.
- Trust and secrets. For a while, the scan page asked visitors to paste in their own model key. We removed that: Pytheas uses one key kept on the server, and the browser never sees it. A website shouldn't ask strangers for their secrets, which is exactly the kind of problem we'd flag on someone else's site.
- Scope. Twenty planned steps, one weekend. The concept document had a cut list from the first hour, saying what to protect and what to drop first. Some of our favorite ideas are still past the edge of our map: testing with several kinds of agents, checking the score on every deploy, and scoring a design file before any code exists.
What's next
Pytheas the navigator didn't sail north to conquer anything. He went to see what was there and write it down truthfully, even when people didn't want to believe it.
That's what we want for the web as it gains its second audience. We want a way to see your site the way your next visitor will, an honest account of where that visitor gets lost, and a clear path to fixing it, for agents and for everyone else.
Get your site ready for its next visitor.
Built With
- claude
- css
- cursor
- grok
- grokbot
- html
- javascript
- typescript
Log in or sign up for Devpost to join the conversation.