Inspiration

The hidden threat in agents. Agents quietly guess at work that was never meant to be uncertain - dates, randomness, math, lookups - and that guessing hides inside natural-language prompts where no test can see it. It inflates cost and latency, and it poisons evals: a "failed" agent often failed on flaky plumbing, not on its reasoning. This is the threat agents smuggle into production, and it is fundamentally a testing problem - you cannot test what you cannot separate.

Two of Alan Turing's ideas, put to work by coding agents, answer it:

  • The Bombe cracked Enigma by eliminating every setting that logic ruled out, leaving only the truly uncertain few. A coding agent does the same to your agent - it separates the mechanical steps from genuine reasoning in the agent's trajectory, extracting the deterministic work (e.g. randomness for synthetic test data, dates, math) so the model handles only real judgment. What's left is the only thing worth testing.
  • The Universal Machine becomes any machine once given a program. Here, skills are the programs: an open framework where AI-experts add their own skills for new patterns, and the agent is composed into a deterministic Flow that feeds the model.

The guiding principle: only one node may guess. Everything else is plain code you can read and verify.

That separation is where the quality value lands - before the first test case is ever written. The moment TuringFlow pulls the deterministic work out, the agent is already cheaper, more predictable and auditable; every path through the resulting Flow then becomes a test case automatically, so the agent arrives with its regression suite already attached.

What it does

TuringFlow TestShield runs a coding agent over your UiPath agents and reshapes them around a single rule - only one node is allowed to be probabilistic. It works in three phases (ANALYZE → BUILD → VERIFY):

  1. Analyze · The Bombe - Harvest: the coding agent captures each agent's trajectory in one of two ways - it downloads the execution trace logs from Test Cloud, and/or injects an additional output argument into the agent so it emits its trajectory tasks as structured output. Classify: customizable skills then rate every trajectory task on a deterministic-vs-probabilistic scale and route each one by precedence - a task that is deterministic is lifted out into the preceding front script (or an earlier deterministic step in the trajectory); only genuinely probabilistic work stays in the prompt. Clean up: finally the system and user prompts are rewritten free of the deterministic tasks now handled in code (e.g. random seed generation), so the model is left with judgment alone.
  2. Build · Universal Machine - Builds a front script (seed, Orchestrator config, REST tokens), rewrites the agent prompt, and composes the agent into a UiPath Flow: deterministic nodes feeding a single model node. Crucially, any decision point found in the agent's trajectory is externalized too - lifted out of the model and expressed as explicit branching in the UiPath Flow - so the control flow is deterministic structure you can see, not something the model guesses. Ships it via CLI to Orchestrator.
  3. Verify · Path Coverage - A coding agent challenges the Flow by walking every possible token-flow path, minting a test case for each - a full regression suite. To actually walk each path, it derives the input arguments that drive every branch, so each test case feeds the Flow exactly the args needed to reach its token flow - turning "every path" into a real, executable coverage set. The cases are written straight into UiPath Test Manager - with requirements and the Turing findings attached - so it produces test cases before any requirement or test exists.

Live example: Choosing the day of the week needs today's date - but an LLM has no clock, so it can only guess. TuringFlow found that dependency in the prompt, extracted "get current time (ms)" into a deterministic UiPath script, and injected it ahead of the agent - spotted, extracted, and injected automatically - so the result is reliable and repeatable.

Architecture: on-prem analysis, cloud execution. The coding agent runs locally on Claude Code (Opus 4.8) on an AI-Expert on-prem device; the rebuilt Flow runs in Orchestrator, with requirements, test cases, test sets and execution logs landing in UiPath Test Manager.

How we built it

  • UiPath Test Cloud
  • Claude Code: Plugins, Subagents, MCP-servers, Skills, Commands, Hooks, scheduled tasks, etc. - running locally on Opus 4.8
  • UiPath CLI - ships the rebuilt Flow to Orchestrator
  • UiPath Skills - the customizable, extensible analysis programs (randomness, dates, math, stubs)
  • UiPath Coding Agents
  • UiPath Flow - the resulting deterministic-scripts → agent → end graph
  • UiPath Agent Builder
  • UiPath Data Fabric & Test Manager - requirements, test cases, test sets, execution logs
  • UiPath Orchestrator - assets (config & credentials) and storage buffers (Context Grounding & RAG)

Why UiPath Flow, not Maestro. Flow's decisive advantage is that a coding agent can author and adjust the workflow itself through Skills - a .flow is something the agent reads, writes and rewires as code. That's what makes TuringFlow possible: the agent can mechanically lift a deterministic step out of another agent and drop a script node in front of it. Maestro can't be reshaped by a coding agent this way.

Example - random number generation. Ask an agent to "pick a random business partner" and it will spend a model call faking randomness inside its reasoning. TuringFlow instead extracts that into a small script node that generates the random number deterministically (from a seed) and feeds the result into the agent. Why pulling it outside the model wins:

  • Fewer tokens / cheaper - the random draw costs zero tokens instead of a model round-trip.
  • Faster / lower latency - a script runs instantly; no LLM call sits on the critical path.
  • Less testing - a seeded script is deterministic, so its output needs no test at all; only the one remaining probabilistic node has to be verified.
  • Reproducible & auditable - same seed, same draw, so runs and evals are repeatable instead of flaky, and the logic is plain code you can read.

Challenges we ran into

  • Finding hidden determinism in prompts. Deterministic dependencies (like "needs today's date") are buried in natural-language prompts and agent trajectories. Reliably detecting them - and proving a step is mechanical rather than a judgment call - is the core hard problem.
  • Cleanly extracting and re-wiring. Pulling a deterministic task out of an agent's own prompt and injecting an equivalent UiPath script ahead of it, without changing the agent's intended behavior, has to be automatic and repeatable.
  • Covering every path. Walking every token-flow path through the rebuilt Flow to mint a full regression suite - and doing it before any requirement or test exists - meant generating tests from agents dropped into agentic tests.
  • On-prem / cloud split. Keeping analysis local (on the AI-Expert device with Claude Code) while execution runs in the cloud (Orchestrator + Test Cloud) required a clean ingest → deploy-back loop.

Accomplishments that we're proud of

  • A working, automatic extraction: TuringFlow found a real deterministic dependency in an agent's prompt, extracted it into a UiPath script, and injected it ahead of the agent - running live in UiPath.
  • Tests for free - every path through the Flow becomes a test case, generated before anyone writes one, and written straight into Test Manager with requirements and findings attached.
  • An agent architecture where only one node may guess - everything else is plain, auditable code.
  • An extensible coding-agent framework, local on Claude Code (Opus 4.8) and deployed on UiPath.

What we learned

  • Cheaper & faster: deterministic steps run as code, not tokens - lower cost, lower latency, fewer surprises.
  • Predictable & auditable: containing the probabilistic work to a single node makes the rest of the agent plain code you can read and verify.
  • Evals you can trust: when hallucination stays contained, results measure the agent's reasoning - not flaky plumbing.
  • Coding agents can do more than write code - they can re-architect other agents, separating the deterministic from the probabilistic and generating the tests to prove it.

What's next for TuringFlow TestShield

  • Grow the open skills framework - let AI-experts contribute their own skills for new deterministic patterns (randomness, dates, math, stubs, and beyond), since the Universal Machine becomes any machine once given a program.
  • Broaden path-coverage and regression generation across more agent and Flow shapes.
  • Deepen the Test Manager / Data Fabric integration so requirements, test cases and Turing findings stay continuously in sync as agents evolve.

Built With

  • claudecode
  • codingagent
  • flow
  • meastro
  • testcloud
  • testmanager
  • uipathcli
  • uipathcodingagent
  • uipathflow
  • uipathmaestro
  • uipathtestcloud
+ 8 more
Share this project:

Updates

Submission history