Inspiration

Public TestFlight programs are surprisingly difficult to discover. There are many interesting beta apps, but the links are scattered across repositories, forums, social posts, and individual developer pages. For users, finding something relevant often means manually scanning long lists and opening links one by one.

At the same time, independent developers need testers, but getting a public TestFlight link in front of the right people is difficult.

FlightDeck started from a simple idea: what if a TestFlight directory could work not only for people browsing manually, but also for AI agents helping users discover software?

WebMCP made that possible. Instead of forcing an agent to inspect a page, understand its layout, and simulate clicks, the website can expose structured tools that describe exactly what the agent can do.

That creates value on both sides: users can ask an AI assistant to find interesting beta apps, while developers gain another discovery channel for their public TestFlight programs.

What it does

FlightDeck is a live directory of public TestFlight programs across iOS, iPadOS, macOS, tvOS, and visionOS.

People can use the site normally to:

  • search by app name
  • filter by platform
  • filter by availability
  • see whether a beta is accepting testers, full, closed, or removed
  • open the real TestFlight join link

AI agents can access the same catalog through six WebMCP tools:

  • search_testflight_apps
  • get_testflight_app
  • list_testflight_platforms
  • get_catalog_stats
  • filter_visible_catalog
  • prepare_testflight_join

This allows an agent to handle requests such as:

“Find macOS betas that are accepting testers.”

“Show me available visionOS projects.”

“Find this app and give me its TestFlight link.”

One important part of the project is that the agent and human interfaces are connected. For example, when an agent uses filter_visible_catalog, the page itself updates to show the same filtered results. The user can immediately see what the agent found.

FlightDeck currently tracks roughly 1,000 real public TestFlight programs and automatically refreshes its data.

How we built it

FlightDeck is intentionally lightweight.

The frontend is built with plain HTML, CSS, and JavaScript and is hosted on GitHub Pages.

The TestFlight catalog is synchronized from the open-source pluwen/awesome-testflight-link dataset. A GitHub Actions workflow periodically downloads the latest data, transforms it into the local catalog format, validates the generated files, and deploys the updated site.

WebMCP is implemented using imperative tool registration with:

document.modelContext.registerTool()

Each tool has a defined input schema, structured output, and explicit validation.

The six tools share the same underlying catalog used by the visible interface, so the human-facing website and the agent-facing tools never operate on separate datasets.

We also added strict argument validation. Invalid platform names, unsupported availability values, invalid limits, missing IDs, and unexpected fields return structured INVALID_ARGUMENT errors rather than being silently accepted.

The WebMCP implementation was tested in Google Chrome with WebMCP enabled and with the WebMCP Model Context Tool Inspector.

Challenges we ran into

The biggest challenge was understanding what makes a WebMCP integration genuinely useful rather than simply exposing arbitrary functions.

The first version of the project accidentally registered seven tools instead of six. A declarative WebMCP tool was being generated automatically from HTML attributes on the search form in addition to the six imperative tools.

That duplicate tool could change the page state, but its WebMCP call did not complete correctly. It was a particularly useful failure because the site appeared to work visually while the agent-facing contract was broken.

We removed the declarative duplicate and kept a single, explicit implementation for each capability.

Another challenge was validation. Early versions were too permissive. For example, an invalid limit could be silently clamped and an unknown platform could simply return zero results. That is acceptable for some user interfaces, but poor behavior for an agent tool because it hides errors from the caller.

We changed the tools to reject invalid arguments explicitly and return structured errors.

Keeping UI state and agent state synchronized was another important design problem. We wanted agents to do more than query hidden data. Some actions should also be visible to the person using the page, so tools such as filter_visible_catalog update the actual interface.

Accomplishments that we're proud of

The most important accomplishment is that FlightDeck is not a mock demo. It operates on real public TestFlight data.

The published site currently contains roughly 1,000 beta programs and automatically refreshes from the upstream source.

We are also proud that the site remains fully useful without AI. WebMCP is an additional capability rather than a requirement. A person can visit FlightDeck and use it like a normal website, while an agent can interact with the same application through structured tools.

The final implementation exposes exactly six WebMCP tools, with no duplicate declarative tools and no console errors.

We also built automated tests that verify tool registration, schemas, argument validation, error handling, and protection against unintended UI mutations on invalid requests.

The full project is open source, automatically deployed, and accessible through a public live URL.

What we learned

The biggest lesson was that making a website agent-friendly is not the same thing as giving an agent access to the DOM.

A good agent interface needs clear capabilities, predictable contracts, explicit validation, and structured results.

WebMCP also changes how we think about website UX. Traditionally, the website is the interface. With WebMCP, the website can have two interfaces at the same time: one designed for people and another designed for agents.

They do not have to be independent. In FlightDeck, the most interesting interactions happen when both interfaces cooperate.

We also learned that strict validation matters much more for agent-facing tools than it may initially appear. Silent normalization can make an agent believe a request succeeded when it actually changed meaning.

Finally, we learned that relatively simple web applications can become significantly more useful to AI systems without requiring a backend, an external API, or a large framework.

What's next for FlightDeck

The next step is improving discovery beyond exact app-name search.

We want agents to be able to help users find projects based on intent, such as:

“Find experimental productivity apps.”

“Show me unusual macOS utilities.”

“Find new visionOS projects worth testing.”

That will require richer metadata such as categories, developer information, descriptions, and potentially App Store metadata where available.

Another direction is better support for developers who are looking for testers. FlightDeck could provide a simple submission workflow for developers to add or update their public beta programs.

We also want to improve catalog freshness and make availability changes easier to detect.

Longer term, FlightDeck could become a two-sided discovery layer between people looking for interesting experimental software and developers looking for engaged testers, with AI agents acting as the bridge between them.

Built With

Share this project:

Updates

Submission history